Analysis
Expected points added in football, and why yards lie
What expected points added measures in football: how an expected points model is trained on next scores, how EPA is computed, and how to read the figure.
By CricketTaken EditorialPublished Analysis21 min read
Third and three, and the running back gains four. First down, drive continues, everybody moves on.
Third and eight, and the running back gains four. Fourth down, the punt team runs on, drive over.
Same four yards. Opposite afternoons. Any statistic that records those two plays identically has told you something false about a game of football, and the yardage column does exactly that, several hundred times a season, without ever appearing to be wrong. That is the problem expected points added in football was built to fix, and understanding EPA is mostly a matter of understanding why the fix had to be this drastic.
Yards are not a bad unit because they are imprecise. They are a bad unit because they are the wrong quantity. Football is not played to accumulate distance, it is played to reach scoring positions and to deny them, and the relationship between distance covered and scoring position reached depends entirely on where you started and how many attempts you have left. A yard on first and 10 does not do the same work as a yard on fourth and one. Nothing that counts yards can know that.
Four yards, four completely different plays
Start with the four-yard gain and put it in four different situations on the same patch of grass, the offence's own 40-yard line.
On first and 10, four yards is fine. It sets up second and six, which is a comfortable state, and the offence keeps two attempts to find the remaining six.
On second and 12, after a holding penalty, four yards is fine as well, in the sense that third and eight is better than third and 12.
On third and three, four yards is excellent. It resets the whole apparatus: new first down, three fresh attempts, ten yards of runway.
On third and eight, four yards is a failure. The offence surrenders the ball. The four yards it gained change nothing except the punter's starting point.
Read the gaps rather than the bars. Three of those four plays raised the value of the offence's position and the fourth lowered it, and the one that lowered it lowered it further than any of the others raised it. Yardage says four, four, four, four. The football says something closer to a small gain, a small gain, a large gain and a loss.
That is the entire idea. Everything after this is engineering.
What an expected points model actually is
An expected points model answers one question: from this exact situation, how many net points will this team eventually score before the next score of the half, on average, across every comparable situation in a large historical archive?
The construction is less exotic than the name suggests. Take a very large set of historical plays, and for each one, look forward to whatever the next score in that half turned out to be. There are only a handful of possible answers. The offence in possession scored a touchdown, or kicked a field goal, or was tackled for a safety. The other team did one of those three things. Or the half ended with nobody scoring at all.
Each of those outcomes gets a signed point value, positive when the team in possession is the one that scores and negative when it is not. A touchdown carries its full seven, or a shade under it if the model prices the conversion attempt separately. A field goal carries three. A safety carries two. No score carries zero.
Now group the plays by state. Down, distance to the sticks, yard line, and usually time remaining in the half. Within each group, the model estimates the probability of each of those next-score outcomes and multiplies through by their point values. Add them up and the result is the expected points of that state.
So a first and 10 at midfield worth 2.0 expected points does not mean two points are about to be scored. It means that across all the comparable first-and-10s at midfield in the archive, the next score in the half was worth an average net two points to the offence, once you have counted the times they drove for a touchdown, the times they kicked, the times they punted and eventually got the ball back and scored anyway, and the times they threw an interception that was returned for seven the other way.
The modern implementations do this with a classifier that outputs all the next-score probabilities at once rather than a lookup table, which lets the estimate vary smoothly with field position and distance instead of jumping between buckets. They also carry a few state variables the original work did not: seconds left in the half, whether the offence has a timeout, occasionally the score margin, occasionally home advantage. The extra variables change the values. They do not change the concept.
The word doing the most work in that description is "average". Expected points is not a prediction about your drive. It is a statement about the population of drives that historically began in situations like yours. The model has no idea that your left tackle is injured, that the opposing defensive front has been unblockable all afternoon, or that it started raining ten minutes ago. It is a base rate, and base rates are useful precisely because they ignore the things you think you know.
Where the idea came from, and why it took forty years to arrive
The core insight is older than most people assume. Virgil Carter, a quarterback with a graduate degree in operations research, worked with Robert Machol in the early 1970s on a study asking what a first down at each point on the field was actually worth in points, using play-by-play records from a single season. Their answer was a curve: value rises steadily as you approach the opponent's end zone, and it goes negative in your own territory close to your own goal line.
That paper contained the whole framework. It had no down or distance dimension, no clock, and a sample small enough that the curve had to be smoothed by hand, and none of that matters. The move that counted was deciding to value a state rather than count an event.
It sat mostly unused for decades, for the dull reason that nobody had the data. Valuing states requires a complete, machine-readable, play-by-play record of tens of thousands of plays, and until that existed there was nothing to fit a model to. The statistic did not wait on a mathematical breakthrough. It waited on a database.
When play-by-play archives did arrive, expected points models arrived with them almost immediately, and within a few years they were driving the arguments about strategy that eventually changed how teams treat fourth down. That is the usual pattern in sports analytics. The idea is old, the data is new, and the idea gets credited to whoever had the data.
Why a first down at your own goal line is worth less than nothing
The most counter-intuitive property of an expected points curve is that part of it lives below zero, and it is worth sitting with, because it is where the model most clearly beats intuition.
First and 10 at your own two-yard line is a negative state. Not a low-value state. A negative one, meaning the model expects the next points scored in that half to be scored by the other team.
The reasons compound. An offence backed up against its own end zone is running a restricted playbook, because a sack or a fumble in that space is a safety and a botched throw over the middle is an interception with a very short field behind it. It needs to travel a long way before a field goal is even conceivable. And the single most likely way that possession ends is a punt from deep in its own territory, which hands the opponent a shorter field than average, which raises the opponent's chance of scoring next above the offence's own.
Sum those and the expectation flips sign. Possession is not worth having, at that spot, in expectation.
The values on this curve are constructed for this article and are not the output of any published model. The shape is the point: value rises roughly with proximity to the opponent's end zone, rises faster inside field goal range, and crosses zero somewhere near the offence's own goal line. Where exactly it crosses differs between models and between eras.
Show the numbers
| Item | Expected points on 1st and 10 |
|---|---|
| Own 1 | -0.4 |
| Own 10 | 0.3 |
| Own 25 | 1.1 |
| Own 40 | 1.6 |
| Midfield | 2.1 |
| Opp 40 | 2.7 |
| Opp 25 | 3.6 |
| Opp 10 | 4.6 |
| Opp 2 | 5.9 |
Two consequences follow that you will not get from a yardage table.
The first is that a punt can be a positive-EPA play. If a team is on fourth and 12 at its own eight, the state it is in is worth less than zero, and the state it hands the opponent, allowing for the return, is often worth less to the opponent than its own state was costing it. Punting improves the offence's expected points, in the model's terms, by getting rid of a liability. Very few statistics can express the idea that the ball is sometimes a burden.
The second is that field position swings near your own goal line are worth far more per yard than swings near midfield. The curve is steeper down there. Twenty yards of field position gained on a punt return from your own five is worth more than twenty yards gained from your own 45, which is why coaching staffs treat backed-up situations with a caution that looks excessive until you price it.
How expected points added in football is just one subtraction
Given the model, EPA is trivial. That is its best feature.
Take the state before the snap and look up its expected points. Take the state after the play has ended and look up its expected points. Subtract the first from the second. The difference is the expected points added by that play.
EPA = expected points of the state after the play - expected points of the state before the play
Three details make that subtraction work in practice.
Possession changes flip the sign. If the play ends with the other team holding the ball, the resulting state is looked up from their perspective and then negated, because their expected points are the offence's expected losses. This is why a turnover produces such a large negative number. The offence loses the value of its own state and simultaneously hands a positive state to the opponent, so the swing is the sum of two quantities rather than one.
Scores are terminal. When a play ends in a touchdown, the "after" value is the points scored rather than a lookup, because there is no subsequent state on that possession to evaluate. The EPA of a touchdown from the opponent's three-yard line is therefore smaller than the EPA of a touchdown from your own 20, and correctly so: from the three, most of the value had already been banked by whoever got the offence there.
The defence gets the mirror image. Every play produces exactly one EPA figure, and the defence's EPA on that play is its negative. There is no separate defensive model. A defence that allows minus 0.4 EPA per play is, by construction, being credited with everything the offence failed to add.
The consequence of that third point is that EPA is a zero-sum accounting of a play rather than an evaluation of it. If a defensive back falls over and a receiver walks in from 40 yards, the play generates a large positive number for the offence and an identically large negative number for the defence, and the model has no view about which of the two facts caused it.
Watching one play become an EPA value
Prose can describe the subtraction. It cannot show it. The play below is invented, chosen because the sign of the answer surprises people.
- The situation before the snapThird and eight from the offence's own 40, early in the second quarter, scores level. The model looks up this state and values it at 0.70 expected points. That figure already accounts for the likelihood the offence fails to convert and punts, which is what usually happens on third and eight.
- The playA short pass to the running back in the flat, caught four yards downfield, tackled immediately. The gain is four yards. On a stat sheet this is a four-yard reception and a completion, and it will raise the quarterback's completion percentage.
- The situation after the playFourth and four from the offence's own 44. The model looks up this state and values it at 0.15 expected points, because in almost every version of this situation the offence punts, and a punt from the 44 hands the opponent a state worth slightly more than nothing to them.
- The subtraction0.15 minus 0.70 is minus 0.55. The play cost the offence roughly half a point of expectation.
- What the defence getsPlus 0.55, automatically. The same number with the sign reversed, credited to eleven defenders and their coordinator without any further calculation.
- The same play, one down earlierRun the identical throw on second and eight and the after-state is third and four rather than fourth and four, which the model values well above the before-state. The identical four-yard completion now produces positive EPA. Nothing about the throw, the catch or the tackle has changed.
- What the box score recordedOne completion, four yards, and no indication whatsoever that the drive is now effectively over. This is the gap EPA exists to close.
The play, the situation and every expected points value in this figure are invented for this article to make the arithmetic legible. No real play is being described and no model's output is being reported. The direction of the result, negative, follows from the down and distance regardless of which model you use.
That final contrast is the whole argument in one figure. The same physical event, thrown by the same arm, caught by the same hands, is a modest success on one down and a drive-ending failure on the next, and the only statistic that notices is the one that stopped counting yards and started valuing situations.
It also explains why EPA and the traditional passing measures disagree so often about short throws. The passer rating formula prices a completion at a fixed twenty yards regardless of what the completion achieved, which means the four-yard checkdown on third and eight is rewarded twice: once for the completion, once for the yards. An expected points model charges it half a point. Those are not two views of the same event. They are two different definitions of what an event is.
Success rate is the blunt companion that keeps EPA honest
EPA has one structural weakness that shows up the moment you average it: it is dominated by its tails. A single 70-yard touchdown can carry several points of EPA, which is more than twenty ordinary plays produce between them. Average EPA per play across a game and a handful of explosive plays will set the number almost by themselves.
The standard corrective is success rate, and it is deliberately crude. A play is "successful" if its EPA is greater than zero. Count the proportion of plays that clear that bar, and that is the success rate. Every play counts once. A 70-yard touchdown and a five-yard gain on second and four are both worth exactly one success.
Older versions of the same idea used a down-based rule rather than an EPA-based one: gain 40 per cent of the yards needed on first down, 60 per cent on second, and all of them on third or fourth. The EPA-based definition has largely replaced it because it handles field position and the clock automatically, and the two agree far more often than they disagree, which is a reasonable sign that both are measuring something real.
The two metrics are used together because each covers the other's blind spot.
EPA per play answers how much value an offence produced. Success rate answers how reliably it produced any. An offence with a high EPA per play and a mediocre success rate is a boom-and-bust operation living on a few enormous plays, and it is far less stable from week to week than its average suggests. An offence with a strong success rate and a modest EPA per play grinds, converts, keeps the chains moving and rarely breaks anything long, and it will look better in the standings than its EPA implies because it punts less.
When the two agree, the conclusion is much safer. When they disagree, the disagreement is the finding, and it usually says something specific about how a team is built.
EPA per play, total EPA, and the volume problem
Total EPA sums every play a team or a player was involved in. EPA per play divides that total by the number of plays. They answer different questions and are constantly confused for one another.
Total EPA is a measure of production. It rewards volume, because more plays means more chances to add value. A quarterback who throws forty times a game will accumulate more total EPA than an equally efficient quarterback who throws twenty-five, and that is not an error: he did more.
EPA per play is a measure of efficiency, and it is the one almost every published leaderboard uses, because comparing production across teams that run different numbers of plays is meaningless. It also carries the standard hazard of any rate statistic, which is that small denominators produce unstable and occasionally absurd values. A backup who throws six passes, one of them a long touchdown, will post an EPA per play that no starter can approach.
The trap is that neither is the right answer on its own, and which one misleads depends on the question.
Ranking offences by EPA per play penalises teams that play at a fast tempo without regard to whether the tempo is generating points. Ranking them by total EPA rewards teams that simply ran more plays, which is partly a function of how often their defence got them the ball back. A leading offence in one table can be mid-table in the other and nothing has gone wrong.
The usual practice is to report EPA per play with a minimum volume qualifier, and to state the qualifier. Any figure quoted without a denominator should be treated as decorative.
- 15-yard completion on 1st and 101.1EPA
- 4-yard run on 1st and 100.15EPA
- 7-yard completion on 2nd and 60.75EPA
- 9-yard run on 1st and 100.55EPA
- 22-yard completion on 2nd and 11.65EPA
- Touchdown run from the 31.4EPA
A constructed six-play drive, invented for this article. The drive is deliberately built without a negative play so that the parts sum cleanly; a real drive contains plays with negative EPA that subtract from the total. The values total 5.60, the difference between a first and 10 at the offence's own 25 and seven points on the board.
Show the numbers
| Item | Value |
|---|---|
| 15-yard completion on 1st and 10 | 1.1EPA |
| 4-yard run on 1st and 10 | 0.15EPA |
| 7-yard completion on 2nd and 6 | 0.75EPA |
| 9-yard run on 1st and 10 | 0.55EPA |
| 22-yard completion on 2nd and 1 | 1.65EPA |
| Touchdown run from the 3 | 1.4EPA |
Two things fall out of that decomposition, and both are properties of the arithmetic rather than of the invented numbers.
The first is that EPA is additive along a drive. The sum of the per-play values equals the difference between the value of the starting state and the points actually scored. Nothing leaks. That is a genuinely useful property, and it is why EPA can be aggregated to a drive, a quarter, a game or a season without any adjustment.
The second is that credit is spread unevenly and not in proportion to yards. The 22-yard completion on second and one is worth more than the 15-yard completion that opened the drive, even though it gained less than half as much again, because it moved the offence into scoring range where the curve is steeper. The four-yard run barely registers. The touchdown itself, from the three, is worth less than the throw that got the offence there, because by that point the model already expected most of those points.
Who gets the credit is the part nobody has solved
Here is the honest limitation, stated plainly, because most explanations of EPA either skip it or bury it.
EPA is a property of a play, not of a player. When a quarterback throws a slant that produces 1.8 EPA, that figure belongs to the play. Assigning it to the quarterback is a convention, not a measurement.
Consider everything that had to happen. The coordinator called a concept that beat the coverage. The line held for the two seconds required. The receiver ran the route at the depth and pace that made the timing work, won inside leverage, caught the ball and broke a tackle for eleven yards after the catch. The quarterback identified the coverage, chose the read and delivered the ball accurately. Every one of those contributions is real and the play's EPA is a single number.
Standard practice charges the whole figure to the quarterback on a pass and to the ball carrier on a run, and the field has spent twenty years trying to do better.
Splitting on air yards and yards after the catch is the crudest approach: value the throw by where it was caught, and treat everything after that as the receiver's. It fails because yards after the catch are partly created by the throw. A ball delivered in stride to a receiver running away from coverage generates yards after the catch that a ball thrown behind him does not.
Completion probability models attack the passing half separately. Given the depth of the throw, the separation, the pressure and the receiver's position, how likely was that pass to be completed, and did it get completed? Measuring a passer against the difficulty of the throws he actually attempted isolates something closer to the quarterback's own contribution, at the cost of needing tracking data that nobody outside the league can independently verify.
Regression-based allocation treats the season as a very large system of equations, with each play's EPA on one side and indicators for every player involved on the other, and solves for the coefficients. It is the most principled approach and it suffers badly from collinearity: the same quarterback throws to the same receivers behind the same line for months, and there is no clean way to separate effects that never vary independently.
None of these is finished work. Anybody presenting a single player's EPA figure as their contribution is making an attribution assumption, and the assumption is doing more work than the model is. This is the same wall that every comprehensive value statistic runs into eventually, in every sport. Baseball's attempts to express a player's total contribution in wins get further than football's do, mostly because baseball is a sequence of individual confrontations and football is eleven people doing eleven jobs at once.
EPA describes what happened, it does not forecast what will
This distinction gets collapsed constantly, and collapsing it is the single most common misuse of the statistic.
EPA is descriptive. It measures the value of what occurred, in a currency that respects the situation. It is very good at that, and it is the right tool for the question "how well did this offence actually play".
Prediction is a different job. Predicting requires knowing which parts of past performance persist and which are noise, and a substantial share of any team's EPA total comes from components that persist poorly. Turnovers are the obvious one: interceptions and fumbles carry huge EPA swings, and both the rate at which passes are intercepted and the rate at which loose balls are recovered bounce around far more than the underlying quality that produced them. Fumble recoveries in particular are close to a coin toss, and a team that recovered most of the loose balls in a ten-game stretch carries a chunk of EPA that has no reason to repeat.
Long touchdowns are the second. A pattern of explosive plays is partly a real property of an offence and partly a run of favourable outcomes on plays that could have been tackled twenty yards earlier. Success rate helps here, which is one reason it is usually reported alongside.
Analysts who want a forecast do one of three things: strip the volatile components out and model them separately, regress EPA towards a league mean by an amount fitted from how much it historically persists, or blend it with more stable inputs. The efficiency ratings that adjust play-by-play value for opponent and situation are an attempt at a version that travels better across weeks, and they are explicit that a descriptive rating and a projection are different products.
None of that makes EPA wrong. It makes it a measurement rather than a prophecy, and measurements are still worth having. The same argument runs through every possession-value statistic in every sport, including the shot-quality models that changed how football is watched in Europe, and it always resolves the same way: the descriptive version is trustworthy, the predictive claim needs separate evidence.
Garbage time, leverage, and plays that should not count the same
Two adjustments come up whenever EPA is used seriously, and both are contested.
Garbage time. With four minutes left and a 24-point deficit, the game is decided. The defence is playing soft coverage to prevent long gains and conceding whatever is available underneath, and the offence is running plays it would never call in a competitive situation. Yardage and EPA accumulate in that environment at rates that mean almost nothing about how the team plays when the outcome is in doubt.
The standard fix is to filter, usually on win probability: discard plays where one team's chance of winning has passed some threshold, or where the score margin and time remaining make the situation non-competitive. It works, and it introduces two problems. The threshold is arbitrary, so two analysts filtering at different points get different answers about the same team. And filtering throws away real plays: the losing team's offence did produce those yards against a defence that was, whatever its coverage shell, made of professional footballers.
Leverage. Not all plays matter equally to the result, and EPA treats them as though they do. A play in a tied fourth quarter and a play in the first quarter can produce identical EPA figures while carrying wildly different consequences for who wins. Expected points is indifferent to the score by design, because it is measuring points rather than outcomes.
Weighting plays by leverage is possible, and it is what win probability added does. It also destroys one of EPA's best properties. Once you weight by game state, you can no longer aggregate cleanly, compare across games, or say that a figure describes how well a team played rather than when it happened to play well. Most practitioners keep EPA unweighted and use a separate leverage-aware measure when the question is about winning.
These are choices, made per implementation, and they are not always disclosed. When two sources report different EPA figures for the same team, the disagreement is far more often about filtering than about the underlying model.
Where expected points added and win probability part company
Win probability models look structurally identical to expected points models and answer a different question. Instead of estimating the next score, they estimate the chance of winning from a given state, which requires the score margin and the clock as first-class inputs rather than optional ones. Win probability added, WPA, is then the same subtraction applied to that model.
The two diverge exactly where football gets interesting.
A team leading by four with two minutes to play, facing third and one at midfield, has a modest expected points figure and an enormous win probability stake. Converting ends the game. Failing to convert hands the ball back with time to score. EPA prices the conversion at a fraction of a point, because points are not what is scarce. WPA prices it at a large slice of the outcome.
Run the same play in the first quarter of a level game and the numbers invert in importance: the EPA is identical and the WPA is negligible, because there is a whole game left in which to recover.
Which one to use follows from what you are asking. For evaluating how well a team or a player performed, EPA is the better instrument, because it does not punish a quarterback for playing well in a game his defence lost, or flatter one whose good plays happened to land in high-leverage moments. For evaluating an in-game decision, win probability is the only correct frame, because the coach is choosing to maximise the chance of winning rather than the expected number of points. Those two objectives genuinely conflict, most obviously at the end of games, where an offence should sometimes take a knee, decline yardage, or kick a field goal that lowers its expected points and raises its chance of winning.
Both frameworks descend from the same move, which was to stop counting events and start valuing states. The disagreements between them are about what a state is worth, not about whether states are the right thing to value.
Reading an EPA figure honestly
The number is more useful than anything it replaced and less authoritative than it is usually presented. Six checks, in order.
Ask which model. There is no official expected points model in American football, and there is no reason to expect two of them to agree. Implementations differ on which state variables enter, how many seasons of plays they are fitted on, whether they price the extra point separately, and how they handle the last two minutes of a half. A figure quoted without a source cannot be reproduced or checked.
Ask what was filtered. Garbage time, kneel-downs, spikes, plays wiped out by penalty and special teams are all included by some producers and excluded by others, and each of those choices moves the figure. This is the single largest cause of two sites disagreeing about the same team.
Ask about the denominator. EPA per play over a full season is a stable, meaningful figure. EPA per play over a small number of plays is noise with a decimal point on it. Ask how many plays, always, and be sceptical of any per-play figure computed on fewer than a few hundred.
Look at success rate next to it. If EPA per play is strong and success rate is not, a small number of explosive plays are carrying the figure and it will not persist. If success rate is strong and EPA per play is not, the offence is reliable and lacks a knockout blow. The pair tells you the shape of the distribution; the average alone does not.
Do not read a player's EPA as their contribution. It is the value of the plays they were on the field for, allocated by a convention that gives the quarterback everything on a pass. It includes the line, the scheme, the route running and the yards after the catch. It is a starting point for an argument about a player and it is not the end of one.
Remember that expected means average. The model is telling you what happened, historically, in situations resembling this one. It is not telling you what will happen next, it does not know who is playing, and on any single play it will be wrong by a large margin most of the time. That is not a flaw in the model. It is what an expectation is.
The four-yard gain that started this is still four yards. It always was. What EPA gave the sport was a way of saying, in a single signed number that can be added up across a drive, a game and a season, that on one down those four yards were worth having and on the next down they were the reason the drive ended. Every criticism above is real, and none of them sends anybody back to the yardage column.
Common questions
What is EPA in football?
EPA stands for expected points added, and it is the change in a team's expected points value across a single play. Expected points is a model estimate of the net points a team will eventually score, on average, from a given down, distance and field position. Subtract the value of the state before the snap from the value of the state after it and the difference is the EPA that play produced.
How is EPA calculated in football?
A model is fitted to a large archive of historical plays, and for each play it records what the next score in that half turned out to be. From that it learns the average net points that follow from every combination of down, distance to the sticks, yard line and clock. EPA on a play is simply the expected points of the resulting state minus the expected points of the starting state, with the sign flipped for the defence.
What does a negative EPA mean?
It means the play left the offence in a state worth fewer expected points than the state it started in. A four-yard gain on third and eight has negative EPA because it converts a chance at a first down into a punt, and an incompletion has negative EPA because it burns a down and gains nothing. Negative EPA is not the same as a bad decision; a kneel-down at the end of a winning game has negative EPA and is exactly the right call.
Is EPA predictive of future performance?
Not on its own. EPA is a descriptive measure of what already happened, and a chunk of any team's total comes from turnovers, long touchdowns and other events that repeat poorly from week to week. Analysts who want a forecast strip those volatile components out or blend EPA with more stable inputs, which is why an EPA table and a projection table rarely rank teams in the same order.
Why do different sites report different EPA figures?
Because expected points is a model rather than a rule, and every implementation makes its own choices about which state variables to include, how many seasons of plays to fit on, whether to adjust for opponent quality and how to handle the end of a half. Two credible models can disagree by a meaningful margin on the same play. Compare EPA figures only within one source, and treat a number quoted without its source as unverifiable.
Filed under American Football·nfl · american football · analytics · football statistics · quarterbacks