Analysis
Cycling power meters and FTP explained, watt by watt
What a cycling power meter physically measures, why normalised power exists, what FTP really approximates, and why critical power is better founded.
By CricketTaken EditorialPublished Analysis18 min read
A heart rate monitor tells you what your body is doing about the effort. A power meter tells you what the effort is. That distinction sounds academic and it is the whole reason cycling training was rebuilt around the device: heart rate is a response, subject to sleep, caffeine, heat, stress and the previous three days, whereas a watt is a watt on any morning of the year.
Nothing else in endurance sport gives you a clean measurement of work done rather than effort felt. Running has pace, which is corrupted by wind and hills. Swimming has time per length, which is corrupted by turns. Cycling has a number that would read identically if you were pedalling on the moon, and the consequences of that go a long way beyond training.
This is what the device actually measures, what the numbers derived from it mean, and where they quietly stop telling the truth.
What is being measured, and where
A power meter is a strain gauge with a radio attached.
The physics is short. Power is the rate at which work is done, one watt being one joule per second. On a bicycle, the work is rotational, so power is torque multiplied by angular velocity. Torque is how hard you are twisting the cranks. Angular velocity is how fast they are going round, which is cadence in different units. Multiply the two and you have watts.
Torque is the hard part. You cannot measure a twist directly, so the meter measures the tiny deflection that the twist causes in a piece of metal it has been bonded to. Strain gauges change electrical resistance when they are stretched or compressed by a very small amount. Bond several of them to a crank arm, a spider, a pedal spindle or a hub shell, measure how their resistance changes as you push, and you can infer the force. Add a way of knowing where in the pedal stroke you are and how quickly you are getting there, and the maths falls out.
Where the gauges sit determines what the number includes.
Pedal-based meters read the force at the point where the shoe meets the machine, which is as close to the rider as it is possible to get. They read both legs directly if there is a unit on each side. They also live in the most abused location on a bicycle and are the easiest to swap between bikes.
Crank arm and spider units read slightly downstream, at the point where the leg's effort has already been fed into the drivetrain. A spider unit sits between the crank and the chainrings and captures both legs together, which is why the good ones are trusted.
Bottom bracket and axle designs read the twist in the spindle itself.
Hub-based meters read at the far end of the chain. Because the chain, jockey wheels and bearings all consume a little of what you produced, a hub reads a fraction lower than a crank on the same effort. Neither is wrong. They are measuring at different points in a system that loses energy along the way, and the difference is real rather than an error.
Then there are the estimating devices, which are a different animal. A trainer that calculates power from a known resistance curve and a wheel speed, or a head unit that derives it from speed, weight and gradient, is modelling power rather than measuring it. The model can be decent. It is not the same thing, and it degrades quickly when the assumptions break, which they do in wind and on rough surfaces.
Two practical points follow from all of this, and both catch people out.
Single-sided meters double one leg. A unit that only reads the left crank measures half your output and multiplies by two. If your legs are asymmetric, and most are to some degree, the number is systematically off, and the asymmetry can change with fatigue and with position on the bike.
The zero offset is not optional. Strain gauges drift with temperature. A meter that was calibrated in a warm room and then ridden in the cold will read differently, which is why every unit has a zero-offset routine and why doing it before a hard session is worth the fifteen seconds. A ride that produced surprisingly good numbers on a cold morning may have produced surprisingly good numbers because the metal was cold.
Manufacturers publish an accuracy figure. It is measured under laboratory conditions, on a steady load, at a controlled temperature. Real riding is none of those, and the useful discipline is to treat your own meter as internally consistent and to distrust comparisons between different riders' devices.
- The leg applies forceThe crank, spider, pedal spindle or hub shell deflects by an amount far too small to see, and by an amount proportional to the force applied.
- Strain gauges read the deflectionTheir electrical resistance changes with the strain. That change, corrected for temperature, is the raw measurement the whole system rests on.
- Torque is calculatedThe deflection is converted to a twisting force using the known geometry and stiffness of the component it is bonded to.
- Torque is multiplied by angular velocityCadence, expressed in radians per second, turns torque into watts. This is the only genuine calculation of power in the chain.
- The value is broadcastUsually once a second over ANT+ or Bluetooth, to a head unit that has done none of the measuring and all of the displaying.
- The head unit derives everything elseAverages, normalised power, intensity, training load and zone time are all computed from that one-second stream after the fact.
Every derived metric in cycling, from normalised power to training load, is arithmetic performed on the output of step four.
A watt is a rate, which is why averages lie
Power is instantaneous. The number on the screen is what you are producing right now, and it changes constantly, because pedalling is not a smooth application of force. Within a single revolution the torque rises and falls twice. Between revolutions it varies with gradient, gear and attention.
That instantaneous quality makes the raw stream nearly unreadable, so everything gets averaged. And averaging a rate over a period during which the rate was wildly variable throws away exactly the information you wanted.
Consider a ride average. It includes every second you spent freewheeling down a descent producing nothing. It includes the traffic lights if the recording did not pause, and it excludes them if it did, which means two riders who did identical work can report very different averages depending on a setting. A three-hour ride in the hills with a stated average of 180 W might contain forty minutes of zero, which means the pedalling average was considerably higher, and neither number describes what the legs went through.
Now consider two one-hour efforts. Both average 250 W. One was held perfectly steady on a turbo trainer. The other was a chaotic group ride: five minutes hard, five minutes soft, over and over. The averages are identical. The rides are not remotely comparable, and any rider who has done both knows which one hurt.
That gap is the reason the next number exists.
Normalised power, and the fourth-power trick
Normalised power is an attempt to answer the question: what steady effort would have cost my body the same as this variable one?
The calculation has four steps and is worth knowing precisely, because it explains the behaviour that confuses people.
Take a rolling thirty-second average of the power stream. Raise every one of those averaged values to the fourth power. Take the mean of the results. Then take the fourth root of that mean.
Two design choices are doing all the work.
The thirty-second window exists because the body does not respond instantly. A two-second spike does not have two seconds' worth of physiological consequence; the cardiovascular and metabolic response takes a while to arrive and a while to subside. Smoothing at roughly that timescale is a rough model of how the effort is actually experienced.
The fourth power exists because the cost of producing power rises much faster than the power does. Doubling the wattage does not double the difficulty. Raising to the fourth power builds that asymmetry into the arithmetic: a value twice as large contributes sixteen times as much to the mean before the root is taken. Hard efforts dominate the result, soft efforts almost vanish, and the fourth root at the end brings the whole thing back to something expressed in watts.
Here is the arithmetic on an invented ride, with round numbers so it can be checked by hand.
A rider does one hour alternating five minutes at 350 W and five minutes at 150 W, six times each. The average is exactly 250 W. Raise the two values to the fourth power, average them, take the fourth root, and the result is a little under 297 W. A second rider holds 250 W dead steady for the same hour. Same average. Normalised power of 250 W.
Forty-seven watts of difference between two rides that a spreadsheet would call identical. That is the number normalised power was built to recover, and once you have it, two more follow immediately.
Intensity factor is normalised power divided by threshold power. An hour spent at your threshold has an intensity factor of one. Anything above one for a long duration is either a race or a mistake.
Training stress score combines duration and intensity into a single figure, scaled so that one hour at threshold scores one hundred. That scaling is the whole point: it gives a common currency for comparing a short brutal session with a long easy one, and it is the basis of every training load chart you have ever seen.
All three inherit one dependency. They are calculated from a threshold number, and the threshold number is an estimate.
What FTP actually is, and the honest sentence about it
Functional threshold power is usually defined as the highest power a rider can sustain in a quasi-steady state for approximately an hour.
Notice how much hedging is in that sentence. It is there for a reason.
The concept the definition is reaching for is a genuine physiological boundary. Below a certain intensity, the body reaches a stable state: lactate production and clearance balance, oxygen uptake plateaus, and the effort is sustainable for a long time. Above it, nothing stabilises. Lactate accumulates, oxygen uptake drifts upwards, and the clock starts running on a failure that will arrive whether you want it to or not. That boundary is real, it is measurable in a laboratory, and it goes by several names depending on the protocol used to find it. The relationship between power output and what is happening in the blood is the thing FTP is trying to stand in for.
FTP is a field proxy for that boundary. It is not the boundary.
This matters more than it sounds, because the proxy and the thing it approximates come apart in individuals. For some riders the hour-power figure sits very close to their true metabolic threshold. For others it sits meaningfully above or below, depending on how well they hold a steady state, how much anaerobic capacity they have to spend, and how good they are at suffering in a controlled way for an hour. A rider with a large anaerobic reserve can drag their measured hour power above their sustainable metabolic rate, which produces zones that are all slightly too hard and a training plan that quietly digs a hole.
The honest formulation is this. FTP is a useful, repeatable, self-consistent number that anchors a training system. It is a physiological threshold only by approximation, and the size of the approximation varies from rider to rider and cannot be known without laboratory testing.
Anyone who tells you their FTP as though it were a biological constant has skipped that paragraph.
The tests, and what each one is really measuring
Almost nobody does the actual hour. It is unpleasant, it requires a course or a trainer and a clear head, and it is difficult to pace properly. So the number is nearly always estimated.
The twenty-minute test. Warm up properly, including at least one hard opener to clear the legs, then ride twenty minutes as hard as you can hold. Take about ninety-five per cent of the average. The five per cent haircut is a correction for the fact that twenty minutes is a shorter and therefore faster effort than an hour.
That correction is a population average applied to an individual. Riders with a big anaerobic contribution go proportionally harder over twenty minutes than over sixty, so the standard multiplier flatters them. Riders who are pure diesel lose less over the longer duration and are penalised by it. The multiplier is right on average and wrong for most specific people.
Ramp tests. Power increases in fixed steps until the rider cannot continue, and threshold is estimated as a fraction of the best minute. These are shorter, easier to administer and considerably less horrible, and they are correspondingly more sensitive to the rider's anaerobic capacity, because the last two minutes of a ramp are almost entirely anaerobic. A rider with a big sprint gets a flattering ramp result and then finds their threshold sessions unmanageable.
Split tests. Two efforts of different lengths with a recovery between them, used to fit a curve rather than to take a single number. Slightly more work, considerably more information, and the direct ancestor of the critical power approach below.
Modelled estimates from ride data. Software that watches your best efforts over many durations and infers a threshold without any test at all. This works surprisingly well for riders who race or do hard group rides, because they regularly produce genuinely maximal efforts. It works badly for riders who train alone at moderate intensity, because the model never sees a maximum and has nothing to fit.
The practical rule cuts through all of it. The absolute value matters less than you think, and the consistency matters more. A threshold estimate that is five per cent too high is a nuisance. A threshold estimate obtained by a different protocol every three months is useless, because you cannot tell improvement from measurement noise. Pick a test, do it in the same conditions, and compare it only against itself.
Zones are built on the estimate, which is the weak link
The seven-zone model that most training plans use is defined as bands of percentages either side of threshold power, and it is a genuinely good piece of design. Each band corresponds to a rough physiological intention: recovery, aerobic base, tempo, threshold work, maximal oxygen uptake, anaerobic capacity, and neuromuscular sprint work.
Every one of those boundaries is a percentage of a number that was estimated by a twenty-minute test on a Tuesday. If the estimate is high by five per cent, every zone is high by five per cent, and the rider's easy days are tempo, their tempo is threshold, and they arrive at their hard session already tired. This is the most common self-inflicted injury in structured training and it looks exactly like a lack of talent.
The zones also have a subtler flaw: they are drawn as sharp lines across something that is a smooth curve. There is no metabolic event at 76 per cent of threshold. The boundaries are administratively convenient, not physiologically discrete, and a rider who agonises over whether a ride averaged high zone two or low zone three is arguing about a line somebody drew for tidiness.
Watts per kilogram, and why climbing is a different sport
On a climb, most of the power a rider produces goes into lifting the combined mass of rider and bicycle against gravity. Gravity does not care how big the engine is in absolute terms. It cares about the ratio of power to mass.
That is why the number quoted about climbers is watts per kilogram, and why it makes the arithmetic of professional cycling brutal. A big engine in a big body climbs no better than a small engine in a small body.
Two invented riders, chosen to make the point cleanly.
- 300Rider A, threshold power in watts
- 4Rider A, watts per kilogram at 75kg
- 260Rider B, threshold power in watts
- 4.19Rider B, watts per kilogram at 62kg
Constructed figures. Rider A produces forty more watts at threshold and still climbs slower, because the mass in the denominator decides the ratio.
Rider A produces forty watts more. On a sustained steep climb, Rider B goes away, because 4.19 beats 4.00 and the extra forty watts are being spent carrying an extra thirteen kilogrammes uphill.
Now put the same two riders on a flat road against the clock and the answer reverses. On the flat, the dominant resistance is aerodynamic, and aerodynamic drag scales with frontal area and the square of speed rather than with mass. A larger rider has more frontal area, but not proportionally more: doubling a rider's mass does not double the hole they punch in the air. Absolute watts win, and Rider A rides away. The same logic explains why aerodynamic position matters so much more than weight on flat terrain, and why a good time trial is decided by numbers that have nothing to do with the ones that decide a mountain stage.
The full treatment of the ratio, including why it is a much less useful figure over short durations than over long ones, deserves its own explanation of watts per kilogram. The short version: quoting a rider's power-to-weight without saying for how long and on what terrain is quoting nothing at all.
Critical power, which is better founded and less popular
There is a model that does the job FTP does, with a firmer theoretical basis, and it has been in the physiology literature for decades.
Critical power treats a rider as having two things: a sustainable rate of aerobic energy production, which is critical power itself, and a finite reservoir of work available above that rate, called W prime and measured in joules rather than watts.
The relationship is simple and produces a curve rather than a point. Time to exhaustion at any given power above critical power is the size of the reservoir divided by how far above critical power you are riding. Ride two hundred watts above and the reservoir empties fast. Ride twenty watts above and it takes a long time. Ride below critical power and the reservoir refills, at a rate that itself depends on how far below you are.
The advantages over a single threshold number are considerable.
It predicts durations rather than describing one. Given the two parameters, you can estimate how long a rider can hold any power, which is exactly the question a racer actually has.
It separates two qualities that FTP smears together. Two riders with the same threshold can have very different reservoirs, and the one with the bigger reservoir can attack repeatedly while the other can only ride steadily. In a road race that is the difference between winning and finishing. It is also the number that makes team tactics legible: a domestique burning their reservoir on the front is spending a quantity that can be measured and will not come back quickly.
It models recovery. The reservoir refilling below critical power is a description of what happens when a rider sits in the bunch after a turn, and it is the only common model that has anything at all to say about repeated efforts.
The disadvantages are why it has not displaced FTP. It requires several maximal efforts of different durations to fit properly, and maximal efforts are expensive in freshness. The curve fits well over a middle band of durations and less well at the extremes, where sprint efforts and very long efforts both misbehave. And the parameters are not constants: critical power drops as a rider becomes glycogen-depleted, dehydrated or simply tired, which means the model describes a fresh rider more accurately than a rider four hours into a race.
That last limitation is where the interesting work is now happening.
Durability, and the thing none of these numbers measure
A threshold figure describes a rested rider. Races are not decided by rested riders.
The quality that decides them is the ability to still produce power after several hours of work, and it is increasingly treated as a separate parameter rather than as a consequence of the others. Two riders with the same fresh threshold can differ enormously in what is left after three hours of hard riding, and the one who holds more of it wins races that the fresh numbers say they should lose.
The field measurement is straightforward in concept: produce a maximal effort of a given length when fresh, accumulate a large amount of work at a moderate intensity, then repeat the same effort and compare. The decline is the measurement. In the laboratory the same idea is applied to threshold and to the whole power-duration curve, showing how both shift downwards as work accumulates.
It is not yet a standardised metric. The protocols in the published studies were built around professional riders doing amounts of work that most amateurs cannot fit into a day, there is no agreement on the right test, and the numbers are not comparable between methods. What it does is name something that every experienced rider already knows and that no threshold number captures: your fourth hour is a different rider from your first.
For the racer this is the entire game. Positioning, sheltering, eating and the drafting economics of riding in a bunch are all about arriving at the decisive moment with more of your first-hour self intact than the person next to you.
The limits, stated plainly
A power meter measures one thing extremely well and is silent on everything else. The silences are worth listing, because most misuse of the device comes from forgetting them.
Power says nothing about cost. Two hundred and fifty watts in a cool headwind and two hundred and fifty watts in still heat at the end of a long day are the same number and are not the same experience. Core temperature, hydration and glycogen state all change what a given wattage costs, and the meter cannot see any of them.
Power says nothing about fatigue. The number is an output. It reports what you produced, not what producing it did to you, and not what you have left. This is why heart rate has not become obsolete: the divergence between the two, where heart rate drifts upward while power stays flat, is a fatigue signal that neither instrument gives on its own.
Power says nothing about how you feel. Perceived exertion is not a primitive substitute for real data. It is a second channel that integrates everything the meter cannot see, and riders who abandon it entirely in favour of the screen tend to ride through the days when they should have gone home.
Power says nothing about speed. Speed is the outcome of power against resistance, and resistance is set by aerodynamics, gradient, road surface, tyres and wind. A rider who improves their threshold by five per cent and their position not at all may go slower on a windy day than they did last year, and the number will not explain why.
The measurement itself has a floor. Drift, temperature, single-sided estimation and installation all put noise around the reading. A three-watt improvement is not an improvement. It is measurement.
Why indoor and outdoor numbers disagree
Almost every rider who trains both indoors and outdoors ends up with two different threshold figures, and concludes that one of their devices is broken. Usually neither is.
Indoors there is no wind, no cooling and no freewheeling. Heat builds up, which raises heart rate and lowers the power a rider can hold for a long effort, sometimes considerably. Against that, an indoor effort is completely uninterrupted: no corners, no junctions, no descents, no moments where the legs get four seconds off. The two effects pull in opposite directions and the balance is individual, which is why some riders test higher indoors and some lower.
There is also a measurement difference. A smart trainer and a crank meter are reading at different points in the drivetrain, and if the trainer is estimating rather than measuring, its curve is a model of a generic bicycle rather than a reading from yours. Riders who use both should decide which device is the reference and stop asking the other one to agree.
The sensible arrangement is to hold two threshold figures, one for each environment, and to use each for the sessions done in that environment. That sounds like an admission of failure and it is simply an accurate description of the situation: the sustainable power of a rider in a hot still room and the same rider in moving air are different quantities, and forcing them into one number loses information rather than gaining it.
Pacing, which is the use most riders undervalue
Training gets all the attention, and the device's sharpest application is pacing.
An even-paced effort is close to the fastest way to cover a fixed distance against mostly aerodynamic resistance, and human beings are terrible at riding evenly. Left to feel, almost everyone goes out too hard, because the first ten minutes of a hard effort feel manageable in a way the last ten do not. A power meter makes the ceiling visible before the damage is done, and it is the difference between a good time trial and a bad one far more often than fitness is.
On a climb the logic inverts slightly. Gradient varies, and holding constant power over changing gradient means going faster on the shallow sections and slower on the steep ones, which is not quite optimal but is close enough and is enormously better than the alternative of matching the effort to the gradient and blowing up on the first steep ramp.
In a road race the meter is nearly useless in the moment and invaluable afterwards. Racing is reactive, and a rider who is looking at a screen when the attack goes has already lost. What the file tells you the next morning is which efforts cost what, how many hard minutes you had in you, and whether you were dropped because you lacked the fitness or because you spent everything in the first hour fighting for position.
Using the thing honestly
The device is worth having, and the way to get value out of it is unglamorous.
Zero it before every hard session, and do it in the conditions you are about to ride in rather than in the kitchen.
Test the same way every time, and treat the test result as a comparison against your own previous tests rather than as a fact about your body. If the number goes up and the sessions built on it feel harder rather than easier, the number is wrong.
Watch the power-duration curve rather than the single threshold figure. Your best five seconds, best minute, best five minutes and best twenty minutes tell a story that one number cannot, and the shape of the curve tells you what kind of rider you are and which part of it is not improving.
Compare normalised power against average power on your rides. A large gap means a variable ride, which is either good racing practice or bad pacing, and knowing which is a matter of what you were trying to do.
And keep the perceived effort. Write down whether it felt easy or grim. The meter cannot tell you that you were coming down with something, that the heat was worse than the thermometer suggested, or that you should have taken the day off, and the riders who last longest in the sport are the ones who kept listening after they bought the instrument.
More on the physics and the economics of the sport is collected in the cycling archive, where the recurring theme is that the number on the screen is the easy part and everything behind it is not.
Common questions
What does a cycling power meter actually measure?
It measures how much a small piece of metal in the drivetrain bends under load. Strain gauges convert that bending into a torque figure, the unit measures how fast the cranks or wheel are turning, and multiplying torque by angular velocity gives power in watts. Everything else on the screen is arithmetic performed on that one measurement.
What is FTP in cycling?
Functional threshold power is an estimate of the highest power a rider can hold in a steady state for a long effort, conventionally described as around an hour. It is used as the anchor for training zones and for pacing. It is an estimate of a physiological threshold rather than a direct measurement of one, which is the part most explanations skip.
How do you test your FTP?
The common field protocol is a maximal twenty-minute effort after a proper warm-up, with the average power multiplied by about 0.95 to allow for the fact that twenty minutes is shorter than an hour. Ramp tests take a fraction of the best minute of an incrementally increasing effort instead. Both are proxies, and both can be gamed by pacing, so the useful comparison is the same test repeated in the same conditions.
Why is normalised power higher than average power?
Because normalised power weights hard efforts far more heavily than easy ones, using a rolling thirty-second average raised to the fourth power. A ride with surges and recoveries produces a much higher normalised figure than a steady ride of the same average, which reflects how much harder the variable ride actually was. On a perfectly steady ride the two numbers converge.
Is watts per kilogram more important than raw watts?
On a climb, yes, because most of the power is going into lifting body and bike against gravity, and mass is in the denominator. On the flat and in a time trial, raw watts matter far more, because the resistance is aerodynamic and a bigger rider does not pay a proportional penalty. Any judgement about a rider that quotes one number without the terrain is incomplete.
Filed under Cycling·cycling · training · power meter · ftp · sports science