Analysis
NFL fourth down analytics: why coaches punted anyway
How expected points, win probability and break-even conversion rates value a fourth down, why coaches ignored the answer for two decades, and where it breaks.
By CricketTaken EditorialPublished Analysis17 min read
Fourth and two at the opponent's 38, early in the second quarter, scores level. The punt team runs on. In the stands, half the crowd groans. On the broadcast, a graphic appears saying the model wanted a fourth-down attempt, and the coach has already turned away.
This is the cleanest case in sport of a settled analytical answer meeting institutional resistance. The maths behind NFL fourth down analytics was worked out, published and replicated a long time ago, and for most of two decades almost nobody acted on it. The interesting question was never whether the numbers were right. It was why an industry full of intelligent, competitive people kept doing the thing the numbers said was worse.
The answer has less to do with statistics than with who gets blamed.
What NFL fourth down analytics actually measure
The whole field rests on one idea, and it is not a complicated one: give every situation on a football field a value in points, and you can compare things that otherwise cannot be compared.
Expected points is that value. Take every occasion in a large historical sample when a team has faced first and ten at its own 25, and look at the next scoring event. Sometimes the offence drives and scores a touchdown, sometimes it kicks a field goal, sometimes it punts and the opposing team scores next. Average the outcomes, counting the opponent's points as negative, and you get a number: the net points a team in that exact situation can expect. Do that for every combination of down, distance and yard line and you have a surface covering the whole field.
The trick is what this makes possible. A punt is not a scoring play, so in raw terms it is worth nothing. In expected-points terms it is worth a great deal, because it takes the ball from a situation with a low value for you and gives your opponent one with a low value for them, which is a positive change from your side of the ledger. Once every option ends in a situation, and every situation has a price, a punt and a conversion attempt are finally denominated in the same unit.
Expected points added is the same idea applied to a single play: the value of the situation after it, minus the value of the situation before it. A twelve-yard gain on second and eight is worth more than a twelve-yard gain on third and twenty, because one produced a new set of downs and the other did not. That is the logic behind the statistic that now sits under most football analysis, and fourth-down decisions are the place it produces the sharpest conclusions.
Two features of the surface matter for what follows.
It is steeply non-linear near the goal lines. A first and ten at your own 1 is worth substantially less than nothing, because a safety is live and any turnover is catastrophic. A first and ten at the opponent's 5 is worth close to the value of a touchdown discounted by the chance of not scoring one. Between the twenty-yard lines it is much flatter, which is why field position arguments carry more weight at the ends than in the middle.
And it already contains the opponent's response. The value of punting is not the value of your own next possession, it is the negative of what you have handed the other team. Coaches who describe a punt as playing for field position are correct; the model simply insists on pricing that field position rather than treating it as self-evidently worth having.
- 4Downs to gain ten yards
- 4Options available on fourth down
- 20Yard line after a punt touchback
- 40Seconds on the play clock to decide
Rule-defined figures from the NFL playing rules, not model outputs.
The four options, and the two that get forgotten
Every fourth down offers four choices, and most discussion collapses them into two.
Go for it. The offence runs a play and either gains the distance or does not. Success means a new set of downs at a spot slightly ahead of the current one. Failure means the opponent takes over at approximately the line of scrimmage, which is the part coaches feel in their stomach and models treat as just another situation with a price.
Attempt a field goal. Worth the probability of making it multiplied by three, adjusted for what happens after each outcome. A make gives the opponent a kickoff. A miss, outside the ten-yard line, gives the opponent the ball at the spot of the kick, which is worse than a punt and is the reason long attempts are penalised twice over: they are less likely to succeed and more costly when they fail.
Punt. Worth the field position handed over, net of the return, and heavily dependent on where the kick is taken from. This is where the option quietly collapses. A punt from your own 30 gains an enormous amount of ground. A punt from the opponent's 38 gains hardly any, because a touchback puts the ball on the 20 and even a perfectly placed kick pins the opponent only a little deeper. The marginal value of punting falls towards nothing as you approach the opposing red zone, which is exactly the zone where coaches punted most stubbornly.
And the fourth, which almost nobody prices: deliberately taking a penalty or a safety. A delay of game to give the punter room, a false start to move a field goal back out of a tricky hash, a snap through the end zone to concede two points and take a free kick from your own 20 with a lead and a clock running. These are real options with real values, they show up a handful of times a season, and they are almost never in the model output on a broadcast graphic.
The rule for choosing among them is not "compare going for it with punting". It is compare going for it with the best of the alternatives, and which alternative is best changes by yard line. Inside the opponent's 35, the field goal is usually the relevant comparison. Beyond your own 40, it is the punt. Somewhere in between there is a band where all three are close, and that band is where the arguments happen.
The break-even rate is the whole calculation in one number
Here is the piece that turns a model into something a coach can hold in his hand.
Going for it produces one of two outcomes. Call the value of converting S, the value of failing F, and the value of the best alternative A. Going for it is worth taking whenever
p × S + (1 − p) × F is greater than A
where p is the probability of converting. Solve for the p at which the two sides are equal and you get the break-even conversion rate:
p = (A − F) ÷ (S − F)
That is the entire theory. Everything else is estimating the three values and estimating p.
A worked example makes it concrete. The numbers below are invented for the arithmetic and are not model outputs, historical rates or estimates of anything real.
Fourth and two at the opponent's 38, scores level, second quarter. Suppose in our constructed model that converting leaves the offence with first and ten at the 36, worth +2.6 points. Suppose failing hands the opponent the ball at the 38, worth −1.4 to the offence. Suppose a punt from there, netting very little, leaves the offence at −0.2, and a long field goal attempt, weighing the chance of a make against the field position conceded on a miss, is worth +0.9.
The best alternative is the field goal at +0.9, not the punt. So:
p = (0.9 − (−1.4)) ÷ (2.6 − (−1.4)) = 2.3 ÷ 4.0 = 57.5%
The offence needs better than a 57.5% chance of gaining two yards to justify the attempt. If the model's conversion estimate for this offence against this defence is 55%, the correct call in this constructed example is to kick, and every graphic saying otherwise is wrong.
Now move the same play back seven yards, to the opponent's 45. The field goal is a 62-yard attempt and effectively off the table, so the best alternative reverts to the punt at −0.2:
p = (−0.2 − (−1.4)) ÷ (2.6 − (−1.4)) = 1.2 ÷ 4.0 = 30%
The distance to gain has not changed. The offence has not changed. The break-even has halved, purely because one of the alternatives disappeared. That is the single most useful thing to understand about fourth-down decisions: the recommendation is usually driven by the collapse of the alternatives rather than by any change in how likely you are to convert.
Why the break-even barely moves and the conversion rate falls off a cliff
Run the same arithmetic across different distances at a fixed spot and something counterintuitive appears. The break-even rate hardly changes.
The reason is visible in the formula. F is the same whatever the distance, because a failed attempt hands the ball over at roughly the same spot either way. A is the same, because the punt and the field goal do not care how many yards you needed. Only S moves, and it moves in the direction that makes going for it look better: converting on fourth and seven leaves you further downfield than converting on fourth and one, which raises the denominator and lowers the break-even.
So why is nobody going for it on fourth and seven? Because the other half of the comparison, the probability of actually converting, falls much faster than the break-even does.
The two lines cross somewhere between two and four yards in that constructed model. To the left of the crossing you go, to the right you kick, and the position of the crossing is what changes when the field position changes, the score changes, or the quality of the offence changes.
This reframes the whole argument. Coaches were not wrong to think fourth and seven is a bad idea. They were wrong about fourth and one and fourth and two in the middle of the field, where the punt is nearly worthless and short-yardage conversion is a genuine strength for most professional offences. The systematic error was concentrated in a fairly narrow band, and it was large within it.
Where win probability disagrees with expected points
Points are not the objective. Winning is. Most of the time maximising points maximises the chance of winning, and where that stops being true, expected points gives the wrong answer with total confidence.
A win probability model does the same job in a different currency. It estimates, from score margin, time remaining, down, distance, field position, timeouts and a measure of the two teams' relative strength, the probability that a team in this exact state goes on to win. Every fourth-down option can then be priced in win probability instead of points, and the recommendation is whichever option has the highest one.
They disagree in predictable places.
Late, when points are lumpy. Trailing by five with four minutes left, a field goal is close to worthless, because it does not change the fact that you need a touchdown. Expected points loves it. Win probability does not. The same logic runs the other way when trailing by two: three points is the whole game.
At the end of a half. Expected points is built on the next scoring event, which quietly assumes there is time for one. With eight seconds left there is not, and a model that has not been told about the clock will happily value a first down that can never be used.
When possessions are running out. Late in a close game the number of drives remaining becomes a hard integer rather than a smooth average. Punting to preserve field position is worth more when the opponent has four possessions left and much less when they have one, because in the second case you are trading a scoring chance for territory you will never get to use. Anyone who has watched the endgame that a two-minute drill is designed for has seen this in practice, whether or not it was labelled.
When the two teams are mismatched. A heavy underdog should prefer variance. Two options with the same expected points can have very different distributions of outcomes, and the option with the fatter tail is worth more to the team that is losing on average. Win probability models capture some of this through the pregame strength input; expected points, by construction, captures none of it.
The practical rule most analysts use is unglamorous. In the first three quarters of a close game, expected points is the cleaner and more stable measure. In the fourth quarter, and at the end of the second, win probability is the one that matters. Where they agree, which is most of the time, the recommendation is safe. Where they diverge, the divergence itself is the information.
- Read the exact game stateDown, distance measured properly rather than as the announced yardage, yard line, score margin, time, timeouts, whether the game is indoors, and a pregame measure of how good the two teams are.
- Estimate the probability of convertingA model trained on historical plays in similar situations, adjusted for the strength of the offence, the quarterback and the defence. This is the number a coach thinks he knows better than the model, and sometimes he does.
- Price the two outcomes of going for itCompute the value of the situation after a conversion and the value of the situation after a failure, in points or in win probability, and weight them by the conversion estimate.
- Price the field goalMultiply the make probability by the value of a successful kick, and the miss probability by the value of handing the opponent the ball at the spot of the attempt rather than at the end zone.
- Price the puntModel the net punt, including the chance of a touchback, the chance of a return, and the resulting position, then take the negative of what the opponent's situation is worth.
- Take the largest number, and report the gapThe recommendation is whichever option scores highest. The size of the gap is the part that gets cropped out of broadcast graphics and is the only part that tells you whether the decision was close.
The sequence used by public open-source implementations and by the league's own public tool. The order is fixed; the modelling difficulty is concentrated in the first and last steps.
Why coaches punted anyway, and why it was rational for them
The economics of this were set out clearly in 2006, in a paper by David Romer asking whether professional football teams behave like profit-maximising firms. The finding was that coaches punt and kick far more than the numbers justify, and that the deviation is systematic rather than random. The more durable contribution was the explanation, because it is an explanation about incentives rather than about mathematics.
A coach is not maximising points. A coach is maximising a mixture of points and continued employment, and those two things are not aligned on fourth down.
- Two options with almost the same expected valueThe model says going for it is worth a little more. In points, the gap is small. In terms of how the two decisions will be discussed afterwards, the gap is enormous.
- The punt fails invisiblyPunting costs a fraction of a point on average. When the opponent then drives eighty yards and scores, nobody says the punt caused it. The cost is spread across the next fifteen plays and is attributed to the defence.
- The failed attempt fails visiblyA stopped fourth-down play is a single frame. It has a timestamp, a play call and one person who chose it. If the opponent scores from that field position, the causal chain is short enough to fit in a headline.
- The criticism is therefore asymmetricDeviating from convention and failing is punished far more heavily than following convention and losing slowly. The expected career cost of the aggressive choice exceeded its expected competitive gain.
- The gain was real but slow to accumulateCorrecting fourth-down decisions is worth a fraction of a win per season, which is meaningful across a career and invisible across a month. A coach whose job is judged over a much shorter horizon than that is not being irrational by ignoring it.
- The equilibrium only broke when the convention movedOnce the conventional choice became the aggressive one, the asymmetry reversed. Punting on fourth and one at midfield is now the deviation that gets questioned, and the same incentive that produced excessive punting now produces the opposite.
A description of the incentive structure, not of any individual decision. It applies with less force now than it did, largely because the conventional choice has become the visible one.
That last step is the important one, and it explains the timing. The mathematics did not improve much between the mid-2000s and the point at which behaviour changed. What changed was the social cost of each option.
There were mechanical reasons for the delay as well, and they deserve credit rather than mockery. A head coach in 2006 had roughly twenty-five seconds to make the decision, no way to compute anything, and no trusted source of numbers. Nobody was going to overturn a lifetime of coaching instinct on the basis of a table they had not seen. The decision was also unavailable in practice: without a prepared answer, the default is whatever the special teams coordinator is already lining up.
How public models changed behaviour more than private ones did
Teams have had internal analysts for a long time. What actually moved the league was the models being public.
Three things happened, roughly in sequence. Play-by-play data became freely available and parseable, so anyone could rebuild an expected-points surface. Open-source implementations appeared, letting anyone compute a fourth-down recommendation for any situation with a couple of lines of code and see exactly which assumptions produced it. And publishers started grading every decision, every week, by name.
The effect of that third step was out of proportion to its sophistication. Once a newspaper is running a bot that says a coach should have gone for it, and once a broadcast is putting a recommendation on screen before the punt team has finished jogging out, the punt stops being invisible. The cost of the conservative choice moved from zero to something, and the whole equilibrium described above depended on it being zero.
Coaches responded the way anyone does to a changed incentive: they hired people to prepare answers in advance. The modern version of this is a laminated card, produced during the week, showing for each yard line and each distance whether the recommendation is go, kick or punt at the current score, with the boundaries already computed. The decision no longer has to be made in twenty-five seconds; it has to be looked up. That is a far smaller cognitive task and a far smaller act of courage.
It also changed play-calling further back in the sequence, in a way the models themselves did not anticipate. A coordinator who knows the offence will go on fourth and two calls third and eight differently, because a checkdown that gains six yards is now a good outcome rather than a wasted down. Third down became a setup down. That second-order gain is real, it is worth something, and no fourth-down model prices it, because the models were trained on a league in which nobody behaved that way. The same interaction shows up in short-yardage design, where the option concepts built to put a defender in conflict are more valuable when there are effectively four downs to work with rather than three.
The honest limits of the models
Anyone quoting a fourth-down recommendation should be able to state its weaknesses. There are five that matter.
The team-strength adjustment is coarse. Models adjust for how good the offence and defence are, but with a small number of parameters estimated over a season. They do not know that the left guard is playing hurt, that this defensive front has stopped every short-yardage run all afternoon, or that the quarterback cannot execute a sneak. A coach's private information here is real, and is the strongest legitimate reason to override a number.
The conversion estimate hides a selection problem. Historical fourth-down attempts were not a random sample. For years they were disproportionately made by teams that were losing, late, out of alternatives, against defences that knew exactly what was coming. Estimating conversion probability from that sample understates what a team with a full playbook and a normal game state would achieve, and the models correct for it imperfectly.
Win probability is noisy where it matters most. The tails of a win probability model are estimated from the fewest observations, and the model is sensitive to the pregame strength input it was handed. Two reputable models can differ by several percentage points on the same late-game state, which is often larger than the gap they are being used to adjudicate.
A recommendation is a mean, not a distribution. Two options with identical expected value can carry very different risk. That is not always an error to be corrected; sometimes the variance is exactly what a team should want, and sometimes it is exactly what it should avoid, and a single number cannot express either.
Everything is measured in yards that were themselves estimated. The distance to gain is set by officials placing a ball by eye, and the difference between a genuine fourth and one and a generous fourth and one is a large swing in conversion probability. The models have improved on this by measuring distance from tracking data rather than from the announced yardage, which is a quiet but substantial upgrade.
A sixth problem sits underneath all of those, and it is the one least often admitted. There is no single model. Different implementations use different training windows, different era adjustments, different ways of handling overtime and different sources for team strength, and they will not always agree on which option is best. When a broadcast shows one recommendation as though it were the answer, it is showing the output of one estimator among several, without an error bar. That is a reasonable thing to put on a screen and an unreasonable thing to treat as a verdict, and the distinction is lost almost every time.
Where NFL fourth down analytics have been pushed too far
The correction has overshot in identifiable places, and saying so is not a defence of punting.
Treating a tiny edge as an instruction. A recommendation worth a fraction of a percentage point of win probability is not a recommendation, it is a coin flip with a decoration. Broadcast graphics almost never show the size of the advantage, which turns a probabilistic judgement into a binary verdict and makes every coach who disagrees look stupid. The gap is the most important number on the screen and it is the one that gets cropped.
Ignoring who is on the field. A league-average conversion estimate applied to a team that is demonstrably poor in short yardage produces a recommendation the team cannot execute. The model says the average offence should go. It does not say this offence should go.
Very deep in your own territory, early. The expected-points surface is steepest near your own goal line, which means the model's errors there are also largest. Recommendations to go for it on fourth and short from inside your own fifteen in the first quarter rest on a part of the surface with the least data and the most extreme values.
Two-point conversion charts applied without context. The same family of models produces two-point recommendations, and they are more sensitive to the score state than fourth-down calls are. A chart built for the fourth quarter is not a chart for the first, and the rules and situations governing the two-point attempt reward reading the specific state rather than the general table.
Assuming the historical baseline still holds. Models are trained on the past, and the past being modelled had different rules. Changes to kickoffs alter where drives start, which shifts the entire field-position surface underneath every punt calculation, and the kickoff has been rewritten more than once in recent years. A punt value computed from a decade of data that included a different kickoff regime is measuring a game that no longer exists. Anyone reading efficiency numbers alongside these decisions runs into the same vintage problem, which is one reason opponent-adjusted team ratings get re-baselined so often.
What has changed recently
Two developments are worth knowing, because both bear directly on the arithmetic above.
The first is the argument over the pushed quarterback sneak. A proposal to ban it was voted on by owners in May 2025 and fell short: twenty-two votes in favour against ten opposed, where twenty-four are needed to pass a rule change. No ban was among the proposals submitted for 2026. That matters to fourth-down analysis far more than it looks, because a team with a genuinely reliable short-yardage play does not have a league-average conversion probability on fourth and one. It has a much better one, and every break-even comparison shifts with it. A rule change here would not change the model. It would change one of the model's inputs, for some teams more than others.
The second is that the recommendation has become part of the broadcast rather than a post-match argument. The league's own public decision tool and the open-source implementations now feed graphics that appear before the play, which completes the reversal described earlier: the numbers are no longer a critique delivered afterwards but a prediction made in front of everyone, in real time, against which the coach's choice is immediately compared.
What to actually check on the next fourth down you watch
Five questions, in order, will tell you more than any graphic.
How far is it really? Not the announced distance. Whether the ball is a foot short of the marker or two yards short is the largest single input into the conversion estimate, and it is the one the television scoreboard rounds away.
What is the best alternative worth, not the default one? Between roughly the opponent's 40 and the opponent's 35 is the dead zone where the punt gains almost nothing and the field goal is long. That is where the break-even collapses and where a punt is hardest to defend.
Is this a points question or a wins question? Before the fourth quarter, expected points. Late, win probability. If someone is quoting expected points with two minutes left and a two-score deficit, they are using the wrong tool.
Is this offence the average offence? The recommendation is generic until someone adjusts it. A team that cannot convert short yardage should not behave like a team that can, and a team that can should go more often than the average table says.
How big is the edge? Under about a percentage point of win probability, the model is expressing indifference, not issuing an order. Above five, the disagreement is real and worth an argument.
Get those five straight and fourth down stops being a culture war between traditionalists and spreadsheets. It becomes what it always was: a comparison between four priced options, three of which are usually easy and one of which is usually wrong. The rest of the tactical and financial machinery around it is covered across the American football section, and none of it changes the underlying point, which is that the most expensive decisions in the sport are the ones nobody was ever blamed for.
Common questions
What is the break-even conversion rate on fourth down?
It is the probability of converting at which going for it and the best alternative are worth exactly the same. Below that number you should kick or punt, above it you should go, and the whole of fourth-down analysis is a way of estimating both sides of that comparison. It is calculated by taking the value of the best alternative, subtracting the value of a failed attempt, and dividing by the gap between success and failure.
Why do NFL coaches punt when the numbers say to go for it?
The clearest explanation is that the punishment for a visible failure is larger than the reward for an invisible gain. A failed fourth-down attempt is a single identifiable decision that everyone saw, while a punt that quietly cost a fraction of a point is attributed to nobody. A coach maximising job security rather than points will punt more often than a coach maximising points, and both are behaving rationally within their own incentives.
What is the difference between expected points and win probability?
Expected points values a situation in points, averaged over what has historically happened from that down, distance and field position. Win probability values the same situation in the only currency that ultimately counts, the chance of winning the game. They usually agree, and they diverge at the end of halves and when the score margin makes points non-linear, such as when a field goal cannot close a two-score gap.
Do fourth-down models account for how good the offence is?
Modern models adjust for team and quarterback strength, but the adjustment is coarse compared with what a coach knows. The model does not see which lineman is playing through an injury, how the defensive front has handled short yardage all afternoon, or that the backup quarterback cannot run the sneak. That gap is the strongest legitimate reason to override a recommendation, and it is not a reason to override every recommendation.
Has going for it on fourth down become too popular?
In places, yes. The models produce recommendations with a size attached, and a great many are so close that the difference is smaller than the model's own error. Treating a recommendation worth a fraction of a percentage point of win probability as an instruction is a misuse of the tool, and so is applying a league-average conversion estimate to a team that is demonstrably worse than average in short yardage.
Filed under American Football·nfl · analytics · tactics · coaching · statistics