Analysis
DVOA explained: the football metric hidden in its name
DVOA in football, one word at a time: the success baseline, the circular opponent adjustment, the sign convention and what the model will not show you.
By CricketTaken EditorialPublished Analysis21 min read
Nobody was ever helped by being told that DVOA stands for Defense-adjusted Value Over Average. Read it as a sentence rather than as an acronym and the football statistic is explained almost completely before a single play has been watched. Four words. Four design decisions. Each with a cost attached.
Defence-adjusted. Value. Over average.
The first word says every play is discounted or credited for the quality of the opponent it was run against. The second says a play is scored against what should have happened rather than by how far the ball travelled. The third and fourth say the comparison is to the rest of the league, so zero means average and the answer arrives as a percentage rather than as a quantity of anything you could hold.
That is the whole metric. Everything else is detail, and the detail is where every argument about it actually lives.
The name is the method, in the wrong order
There is one trick to reading the name, and it is that the words are listed in the reverse of the order the model works in.
The calculation starts with value, which is a per-play grade. It then centres that grade on the league average, which supplies the zero. Only at the end does it adjust for the defence faced, because the adjustment needs the ratings of every other team in the league before it can be applied, and those ratings do not exist until the first two steps have been run for everyone.
The name puts the adjustment first because the adjustment is the selling point. Raw efficiency was never a difficult thing to compute. Correcting it for schedule was, and still is, and that is the claim the metric leads with.
The rating was created by Aaron Schatz and published for years under the Football Outsiders name before the ratings moved to FTN, and its publishers have always offered an unadjusted twin called VOA. The existence of that twin is the most useful thing about it. Compare a team's VOA with its DVOA and the difference is the entire opponent adjustment, laid out as a single number, which is the closest anybody outside the model gets to auditing the part of it that matters most.
A few things DVOA is not, since the confusions are predictable. It is not built from points scored, so a team that turns modest efficiency into a lot of touchdowns will not be flattered by it. It is not a projection, though projections are built on top of it. It is not a measure of talent, ability or effort. It is a per-play efficiency figure with a schedule correction bolted on, and it is only ever as good as those two components.
Defence-adjusted: the part that is hardest to do honestly
Start with the problem. Two offences produce identical yardage on identical downs across a season. One of them did it against the six best defences in the league, twice each. The other did it against a schedule of teams who could not stop anybody. Raw efficiency calls them the same. Everybody watching knows they are not.
The adjustment is the answer to that, and in principle it is simple. Every play is compared not just to the league baseline for that situation, but to how the specific defence on the field that day has performed in comparable situations across the whole season. Beat a defence that shuts down second-and-medium passes and the play is worth more than the same gain against a defence that surrenders them weekly. Fail against a defence nobody else fails against and the failure is forgiven, partly.
The published description of the method says the adjustment is applied with some granularity rather than as a single blunt multiplier per team. Rushing and passing are adjusted separately, down and field position enter the comparison, and passing plays are further split by the type of receiver targeted, so a defence that struggles against throws to tight ends is not credited for its coverage of wide receivers. That granularity is what makes the adjustment a real correction rather than a strength-of-schedule fudge factor.
And then you hit the problem that makes this genuinely hard.
The opponent's quality is measured by the same system. How good is that defence? It is good in proportion to how badly offences have performed against it. But how badly those offences performed is exactly what you are trying to adjust. Every rating in the system depends on every other rating in the system, and there is no fixed point outside it to anchor to. The calculation is circular by construction.
The way out is iteration, and it is worth understanding because it is the single most misunderstood part of the metric.
The first pass computes raw, unadjusted ratings for all thirty-two teams. Every play is then re-valued using those provisional ratings, which produces a second set of ratings, which are different from the first. Re-value everything again against the second set and you get a third. Repeat. Each pass moves the numbers less than the pass before, and after enough passes they stop moving to any meaningful precision. That settled state is the published rating.
- Invented team A
- Invented team B
Both teams are invented for this article and so are all six values, which exist to show the shape of the convergence rather than to describe anything that happened. Pass zero is the unadjusted rating. Team A played a hard schedule and is revised upwards; team B played an easy one and is revised down. Real ratings are recomputed to far tighter tolerances than the single decimal place shown here.
Show the numbers
| Item | Invented team A | Invented team B |
|---|---|---|
| Pass 0 | 4% | 12% |
| Pass 1 | 10.5% | 6.5% |
| Pass 2 | 8.4% | 8.9% |
| Pass 3 | 9.2% | 8% |
| Pass 4 | 8.9% | 8.3% |
| Pass 5 | 9% | 8.2% |
Two features of that picture matter more than the numbers, which are invented.
The first is the overshoot. Pass one swings much further than the final answer, then the correction oscillates inward. That is normal behaviour for a system solving itself, and it is why nobody quotes an intermediate pass. The second is that the two invented teams end up close together despite starting eight percentage points apart. That is the adjustment doing its entire job: removing a schedule difference that raw efficiency was reporting as a quality difference.
Now the honest limits of the thing.
An opponent adjustment can only redistribute credit inside a closed league. It has no access to absolute quality. If every team in a given season were worse than every team in another, the ratings would look identical, because average is defined by whoever is playing. Cross-season comparisons of DVOA are comparisons of two teams to their own contemporaries and not to each other, and anybody who forgets that ends up asserting things about eras that the metric never claimed.
The adjustment also depends on the schedule graph being connected enough to solve. Each team plays a small fraction of the league, and the ratings propagate through shared opponents. Late in a season the graph is dense enough for the iteration to mean something. In September, when teams have two or three common opponents at most, the adjustment is passing information around a network that barely exists, and the confidence it projects is not earned. The publishers know this, which is why the ratings carry warnings about the early weeks and why a separate variant exists that folds a preseason projection into the in-season number specifically to steady the first month. Blending a forecast into a measurement is an odd thing to do and an honest one, and it is an explicit admission that the raw early rating is not trustworthy on its own.
There is one more wrinkle that the casual account skips. Because a defence's rating is depressed by facing good offences, an excellent offence makes every defence it plays look slightly worse, which in turn slightly reduces the credit that offence receives for beating them. The system contains a mild self-limiting feedback loop. It is small, it is a property of any iterative rating scheme, and it is the sort of thing that gets lost in a table of ranked percentages.
Value: what a play is worth once yards stop being the unit
The second word is the one that separates this from every efficiency figure a broadcast graphic has ever shown you.
Yards are a terrible unit of football value and everybody has always known it. Four yards on third-and-three is a first down and a continued drive. Four yards on third-and-ten is a punt. The same four yards, the same box score line, opposite football outcomes. Any statistic that adds the two together is adding a success to a failure and reporting the mean.
So the model does not start with yards. It starts with a question: did this play succeed, given the situation it was run in?
The published thresholds are the clearest part of the whole method. On first down, a play succeeds if it gains 45 per cent of the yards needed. On second down, it needs 60 per cent. On third or fourth down, nothing counts except a new first down, or a score.
Those three numbers are doing something specific. They describe a drive as a series of instalments, and they ask on each play whether the offence has stayed on schedule to pay the debt. Forty-five per cent of ten yards on first down leaves five and a half from two downs, which is a routine position. Sixty per cent on second down leaves the same offence needing a manageable amount on third. Third down has no instalments left, so only settlement counts.
Look at second-and-five against third-and-two. Both need a modest gain. One is scored a success at three yards and the other is not scored a success until two, which sounds backwards until you notice that the third-down bar is absolute and the second-down bar is not. A three-yard gain on second-and-five keeps an offence on schedule. A one-yard gain on third-and-two ends the drive. The rule is producing exactly the discrimination that a yards-per-play figure destroys.
Success is not the end of it, though, and this is where the method stops being publicly checkable. A twenty-yard gain on first-and-ten is not merely a success in the same way that a five-yard gain is. Extra yardage earns extra credit, at a rate that diminishes as the gain grows, so the difference between four yards and fourteen is worth much more than the difference between forty and fifty. Turnovers carry heavy penalties, weighted by where on the field they happened. Field position enters, because gaining nine yards from your own two-yard line is not the same act as gaining nine from midfield.
The exact shape of every one of those curves is proprietary. The direction of each is published, the magnitude of none.
What the value step buys, in exchange for that opacity, is a per-play grade that already contains the situation. That is the same insight underneath the expected-points framework that values the football state before and after each play, and the two are cousins rather than rivals. The difference is the unit. Expected points added is denominated in points, an actual football quantity you can add up and reason about. The value step here is denominated in a constructed success-plus-yardage quantity that has no natural unit at all, which is precisely why the third and fourth words of the name have to exist.
Over average: zero is a team, and the units are nothing physical
Having produced a grade for every play, the model needs to tell you what a grade means. It does that the only way a constructed quantity can, by comparison.
Every play's value is measured against the average value of all plays run in comparable situations across the league. Add up the differences for a team, divide by its number of plays, and the result is reported as a percentage above or below average. Zero is exactly average. Positive is better than average for an offence. There is no maximum, no minimum, and no unit.
That last point causes more misreadings than everything else in the metric combined.
A 20 per cent offence is not 20 per cent more likely to score. It does not gain 20 per cent more yards, and it will not produce 20 per cent more points. The number means that across the season, the team's plays have averaged 20 per cent more value than the league average play in the same circumstances, where value is the constructed quantity described in the previous section. A percentage of a thing that is not itself measured in anything.
This is not a flaw. It is the honest consequence of grading plays on a scale of your own construction, since the only meaningful way to report the result is relative to the same scale applied to everyone else. But it means the magnitudes cannot be translated into football outcomes without a further model, and it means a difference of five percentage points between two teams is not five of anything. It is five per cent of an average play's value, which is a real statement and a much narrower one than the number's confident presentation suggests.
The relative baseline has a second consequence that nobody enjoys. The league average is recomputed every season from that season's plays. If passing efficiency across the whole sport rises, as it has done over decades of rule changes, the baseline rises with it and nobody's rating moves. A team that plays exactly as well as it did last year, in a league that has improved around it, will see its rating fall without having done anything differently. The metric is era-neutral by design, and era-neutrality means it cannot see eras.
- 3Units a team rating splits into
- 0The rating that means league average
- 45Share of needed yards for success on first down
- 100Share of needed yards for success on third down
Every figure here is a property of the published method rather than an observation from any season. The success proportions are the published baselines; the three units are offence, defence and special teams; zero is the definition of average and not an empirical result.
Watching one play become a DVOA contribution
The theory is easier to hold once you have pushed a single snap all the way through. Everything below is invented, including the situation, the gain and every number attached to it. The steps follow the published order of operations; the values are illustration.
- Record the situation, not just the resultSecond and seven, own 34-yard line, second quarter, scores level. A completed pass gains nine yards and a first down. Down, distance, field position, score and time are all recorded before anything is valued, because every one of them changes what the play was worth.
- Find the success barSecond down requires 60 per cent of the distance needed. Sixty per cent of seven is 4.2 yards. The play gained nine, so it clears the bar and is scored a success rather than a partial credit or a failure.
- Add credit for the yards beyond the barThe 4.8 yards past the threshold earn additional value at a diminishing rate, so each is worth less than the yards that got the play to the bar in the first place. The exact curve is proprietary; only its direction is published.
- Compare with the league in the same situationThe play is measured against the average value of all second-and-medium plays from similar field position across the league. Suppose the average such play is worth 0.10 on the model's internal scale and this one is worth 0.34. The raw margin is 0.24.
- Weight the situationScores level in the second quarter, so the play is fully weighted. Had this been the fourth quarter with the game already decided, the surplus would have been discounted heavily and the play would have contributed far less.
- Adjust for the defence facedThe opposing defence has been better than average against second-and-medium passes all season, so the surplus is credited upwards, say from 0.24 to 0.29. Against a defence that surrenders those throws weekly, the same play would have been discounted instead.
- Add to the season pile and divideThe adjusted surplus joins every other offensive play the team has run. The sum divided by the number of plays gives the mean surplus per play, which expressed against the league average play produces the published percentage.
- Then throw it away and do it againThe opposing defence's own rating has just changed, because this play was one of the plays used to compute it. Every play in the league is re-valued against the new ratings, repeatedly, until the numbers stop moving.
The play, the teams, the gain and all figures are invented for this article and no real play is being described. The sequence of operations follows the published description of the method. The intermediate values in steps four to seven are illustrative only, because the model's actual weights are proprietary and have never been published.
The final step is the one people never picture. No individual play has a fixed contribution to a rating. Its value depends on ratings that depend on it, so the number attached to that pass is provisional until the whole league's season has been solved simultaneously. A play in week two is worth something different in December than it was worth in September, and nothing about the play has changed.
Garbage time, and the case for keeping the data you are about to bin
The situational weighting deserves its own fight, because it is the part of the model that deletes information on purpose.
The reasoning is sound and easy to state. A team trailing by 24 in the fourth quarter faces a defence playing deep zones, conceding underneath throws in exchange for the clock, and interested only in preventing the one thing that could still lose it the game. Yards gained in that arrangement are cheap and say little about how the offence would fare in a competitive situation. Counting them at full value would systematically flatter the teams that spend the most time losing badly, which is to say the worst teams in the league.
So plays late in decided games are discounted. The published account also notes that the standard for success itself shifts in the fourth quarter, because a leading team running the ball to drain the clock is satisfied with a shorter gain that stays in bounds, while a trailing team needs far more than the nominal threshold for a play to help it.
Now the objection, which is stronger than its usual airing.
Discarding data to reduce noise also discards information, and the two cannot be separated by assertion. Some of what happens in a blowout is real. An offence that moves the ball efficiently against soft coverage is demonstrating something, even if it is demonstrating less than it would against a defence trying. A quarterback finding sideline throws under a two-score deficit is performing a football skill. The discount treats all of it as approximately noise, and there is no published test showing that the specific discount applied recovers more signal than it destroys.
The distributional effect is the sharper complaint. Garbage time is not evenly shared. Bad teams and teams with volatile scoring patterns accumulate far more of it than good teams do, so a correction applied to garbage time is a correction applied disproportionately to a particular subset of the league. If the discount is a little too aggressive, the teams it hurts are systematically the same teams every year. That is not random error. That is bias with a pattern, and because the thresholds are unpublished, nobody outside the model can measure how large it is.
There is a version of this problem in every metric that trims its own inputs. It appears in the passer rating's habit of capping each component and discarding the excess, where the caps solve small samples and destroy resolution at the top of the distribution. It appears in every model that excludes kneel-downs and spikes, which almost all of them do and which is uncontroversial only because the excluded plays are genuinely not attempts at football. The garbage-time discount is a much larger intervention on a much less obvious set of plays, and it is applied with weights that are a trade secret.
The minus sign on defence, and why it defeats so many people
Three ratings come out of the model rather than one: offence, defence and special teams. The split is genuinely useful. A team with an average record can be carrying an excellent offence and a poor defence, and the record shows you neither.
Then the sign convention arrives and ruins everybody's afternoon.
Defensive DVOA is negative when the defence is good. Minus 15 per cent is a strong defence. Plus 15 per cent is a bad one. This is not perversity, it is arithmetic honesty, and once you see why it has to be that way it stops being confusing.
Defensive DVOA does not measure what the defence did. It measures what the offences facing that defence achieved, on the same scale as every other offence in the league. An offence performing 15 per cent below average against a given defence produces a rating of minus 15 for that defence. The number is an offensive efficiency figure, computed from the other side of the ball, and the sign follows from that and from nothing else.
There is a real conceptual point underneath the presentation annoyance. Football defensive statistics are almost always the complement of offensive ones, because everything measurable happens to the ball and the ball belongs to the offence. A defence's rating is a statement about the offences it faced. This is why the opponent adjustment matters more on defence than anywhere else: a defence's number is built entirely from other teams' performances, so failing to correct for which teams those were is not a refinement, it is the difference between a measurement and a coincidence.
Total team DVOA is then offence minus defence plus special teams, and the subtraction is why the sign convention cannot simply be flipped for convenience. A graphic that flips the defensive sign to make it friendlier, without flipping it back before summing, produces a total that is wrong by twice the defensive rating. That has happened in public more than once.
Special teams is the third unit and the noisiest. It covers a small number of plays with enormous individual variance, and it has always been the component most likely to swing a total rating on the basis of a handful of kicks. Anyone comparing two teams' totals should check whether the gap between them is smaller than the gap in their special teams ratings, because if it is, the comparison is largely a statement about placekicking.
Per play or per season: two metrics asking different questions
DVOA is a rate. It is value per play, so a team or player who is efficient across few snaps rates identically to one who is efficient across many.
That is the right choice for teams, which all run roughly comparable numbers of plays, and the wrong one for players, who do not. A backup quarterback with a brilliant hundred-attempt sample and a starter with a very good six-hundred-attempt season are not equally valuable, and a rate statistic has no way to say so.
The published answer is a second metric, DYAR, which is cumulative. Rather than reporting efficiency per play, it totals the value produced across the whole season and expresses it against a replacement-level baseline rather than an average one, in units of yards. The two answer different questions on purpose. The rate says how well somebody played when he played. The total says how much he contributed, which folds in how much he played and how much better he was than the man who would otherwise have been out there.
That distinction is not football-specific. It is the same rate-versus-counting split that governs baseball's attempt to express a whole career in one number, and it carries the same trap: quoting the rate when the question was about total contribution, or the total when the question was about quality.
The replacement baseline in the cumulative version deserves a flag of its own. Replacement level is a modelling assumption, not an observation. It is a judgement about what a freely available substitute would have produced, and moving it moves every cumulative figure in the system. Average, by contrast, is defined by the data itself. That is one respect in which the rate metric rests on fewer assumptions than the total does, which is the reverse of what most people assume about the pair.
Why efficiency has beaten the record at predicting the next month
The strongest practical argument for DVOA has never been that it describes the past better. It is that it has historically been the better guide to the near future than a team's record.
The reason is a sample size argument and it is close to unarguable in principle. A record is a handful of binary outcomes across a season, each of which compresses sixty minutes of football into one bit. An efficiency rating is built from every play a team has run, which is a number in the thousands. A team that has won a run of one-score games has a record built on the outcomes most sensitive to a single kick, a single bounce, or a single officiating decision, and those outcomes have very little tendency to repeat. The efficiency underneath them does tend to repeat, because it is an average of a great many loosely related events rather than a tally of a few decisive ones.
This is the whole basis of the perennial complaint that some team is a fraud. The claim is not that it is not really winning. It is that its wins were produced by the components of football that persist least, while its play-by-play efficiency, which persists more, looks ordinary. Sometimes those teams keep winning anyway, and the honest version of the argument has always allowed for it, because a rating that predicts better than a record still does not predict well in an absolute sense.
Two published variants exist to sharpen the forward-looking use. A weighted version discounts earlier games so that recent form counts for more, on the reasonable ground that a team in November is not the same team that played in September, having changed personnel, schemes and health. And the early-season blend with a preseason projection, mentioned earlier, exists to stop the first few weeks producing ratings that nobody should act on.
Both variants make the metric better at forecasting and worse at describing, which is a trade rather than an improvement. A weighted rating deliberately no longer answers the question of how a team has played this season. It answers how it is playing now. Those are different questions, and the same three letters get used for both, usually without anybody saying which is on screen.
For anything that turns on the actual standings rather than on quality, the record is the only thing that counts. The rules that decide who reaches January and in what order contain no efficiency term at all, and no rating has ever won a tiebreaker.
What DVOA cannot tell you, stated plainly
It cannot be reproduced. This is the big one. The success thresholds are published. The direction of the adjustments is published. The existence of the garbage-time discount is published. The weights are not. Nobody outside the model can take a season of play-by-play data and produce the published figure, which means nobody outside the model can check it, find its errors, or fork it. An open metric such as expected points added can be computed from public data by anybody with a laptop, so disagreements about it are arguments about method rather than arguments about a number's provenance.
It gets revised. Any model that is maintained is a model that changes, and when the method changes the historical figures are recomputed. A rating quoted in an article three years ago may not match the current published rating for the same season. That is good practice and it is corrosive to citation, and it attends every proprietary sports model equally.
It needs a month before it means much. The opponent adjustment cannot work until the schedule graph is connected, and the per-play sample cannot stabilise until there are enough plays in it. Ratings published in September are, by the publishers' own admission, not the same class of object as ratings published in December.
It does not know why. A poor rushing rating does not distinguish between the blocking and the back. A strong passing rating does not separate the quarterback from the scheme, the receivers or the protection. Answering those questions requires charting or tracking data, which is what the model that estimates how likely each individual throw was to be completed and the measures built on how often a passer is under pressure exist to do. This is a team efficiency measure that also produces player figures, and the player figures inherit every attribution problem the team ones have.
It cannot see the roster. A team that lost its starting quarterback in week three carries the first two weeks in its season rating at full weight. Injuries, suspensions, benchings and midseason trades are invisible to it, which is why anybody using a season rating to forecast a specific upcoming game has to apply the roster adjustment by hand and should say so.
It cannot see intent. A team resting starters, a team playing conservatively with a lead outside the garbage-time window, a coach calling a game to protect a tiring defence: all of it registers as efficiency. The model has no concept of a team choosing to be less efficient on purpose, which is a decision teams make constantly. That is the same blind spot showing up in the arithmetic of fourth-down decision-making, where the number and the coach are frequently answering different questions.
Its units are its own. Once more, because it is the error most often made in public: a percentage of a constructed value scale is not a percentage of points, yards, wins, or anything that has ever appeared in a box score.
Using DVOA without embarrassing yourself
Six habits, in order of how much trouble they save.
Check the date before you check the number. A rating in the first month of a season is a small sample corrected by a schedule adjustment that has not converged. Treat it as a rumour. From roughly midseason onwards it becomes a serious measurement.
Read the three components, never the total alone. A total is offence minus defence plus special teams, and three very different teams produce the same total. The split is the most useful output of the whole model and it is the part that gets left out of the graphic.
Say the sign out loud. Negative defence is good. Positive offence is good. If a comparison involves both and you have not stated which is which, you are about to get it backwards, and so is whoever you are arguing with.
Look up the unadjusted twin. VOA against DVOA gives you the size of the schedule correction at a glance. If a team's ranking moves a long way between the two, its rating is mostly a statement about who it has played, and the honest way to describe that team says so out loud.
Do not compare across seasons as though the scale were fixed. The baseline is rebuilt every year from that year's plays. A rating from one season and a rating from another are each statements about their own league and not about one another.
Ask what a difference is worth before treating it as a difference. Two teams a few percentage points apart, on a rating with real uncertainty, computed from a proprietary model that gets revised, are not reliably distinguishable. Rankings imply a precision that percentages this noisy do not possess, and a good deal of the weekly argument the sport generates is conducted over gaps well inside the model's own margin.
The metric is worth using and it is not worth trusting the way people trust it. Its virtues are real: it asks the right question about a play, it corrects the thing that most needs correcting, and it has beaten the standings at describing how teams have actually played. Its central weakness is equally real and rarely admitted, which is that a number you cannot compute, cannot audit and cannot reproduce is a number you are taking on faith from somebody else's spreadsheet.
Read the name. Then ask which of the four words is doing the work in whatever claim somebody has just made with it.
Common questions
What is DVOA in football?
DVOA stands for Defense-adjusted Value Over Average. It grades every play against what an average team would have achieved in the same situation, then discounts or credits that grade according to the quality of the opponent faced, and reports the season total as a percentage above or below league average. Zero is exactly average, so the number is a comparison rather than a quantity of yards or points.
How is DVOA calculated?
Each play is scored against a success baseline that depends on down and distance, with a play needing 45 per cent of the yards required on first down, 60 per cent on second down, and a new first down on third or fourth. Extra yardage above the bar earns extra credit at a diminishing rate, turnovers carry heavy penalties, and the situation is weighted so that garbage time counts for less. The play values are then adjusted for opponent quality, averaged across the season and expressed relative to the league, and the exact weights are proprietary and have never been published in full.
Why is a negative DVOA good for a defence?
Because defensive DVOA measures the efficiency of the offences playing against that defence, not the efficiency of the defence itself. A defence that holds opponents below the league baseline produces a negative number, and the further below zero it sits, the better it has played. Offensive DVOA runs the other way, where positive is better, which is why the two are almost never plotted on the same axis without confusing somebody.
Is DVOA better than a team's win-loss record?
For working out how a team has actually played, yes, and it has historically been the better guide to how a team will play over the following weeks. A record is built from a handful of binary outcomes, several of which turn on a single kick or a bouncing ball, while an efficiency rating draws on every play of every game. The record is still the thing that decides the standings, so the two answer different questions.
What are the main limitations of DVOA?
It is a proprietary model whose exact weights are not published, so no outsider can reproduce a figure or audit why it came out the way it did. It needs several games before the opponent adjustment settles, it cannot see injuries, personnel or coaching decisions, and it cannot tell you whether a running play succeeded because of the blocking or the back. It also gets recomputed when the method is revised, so a number quoted one year may not match the same number later.
Filed under American Football·nfl · american football · football statistics · analytics · efficiency