Analysis
WAR explained: what the baseball stat actually measures
What WAR actually measures in baseball: the replacement-level question underneath it, what each component contributes, and why the public versions disagree.
By CricketTaken EditorialPublished Analysis20 min read
Every autumn somebody wins an argument by three tenths of a win. One player finishes at 6.1, another at 5.8, and a great many of the people quoting those figures treat the gap between them as a measurement, in the way a home run total is a measurement. It is nothing of the kind. Three tenths of a win is the residue of a dozen modelling decisions, several of which are live arguments rather than settled facts.
The trouble with the WAR baseball stat explained on a broadcast graphic is that it looks like a count. It is not a count. It is the answer to a hypothetical question, and the question is odd enough that every strange feature of the statistic falls straight out of it.
So start with the question. Nothing else here makes sense without it.
The question WAR is actually asking
How many more games did this team win with this player than it would have won with a freely available substitute, given the same playing time?
Every clause in that sentence is load-bearing.
"Freely available substitute" rules out the obvious comparison. The natural instinct is to measure a player against the league average, and average turns out to be the wrong yardstick for this particular question. More on that in a moment, because it is where most misreadings begin.
"Given the same playing time" fixes the workload rather than dividing it out. WAR is a counting statistic. A player who was excellent for sixty games and then tore a hamstring did not provide as much as the same player would have across a full season, and the number says so. This is deliberate, and it is the source of a permanent argument between people who want to know who is better and people who want to know who was worth more. WAR answers the second question. It is frequently used to settle the first, which is the most common misuse of it.
"Games this team won" points at the unit. Not runs, not on-base percentage, not any of the currencies of the box score. Wins, because wins are the thing a club is actually purchasing when it signs anybody. Everything upstream in the calculation is denominated in runs, and runs are converted into wins at the very end using a rate that the season's scoring environment determines.
Notice what the question does not ask. It does not ask how talented the player is. It does not ask what he will do next year. It does not ask whether he was better than his teammate at the same position, or whether he came through in September. It asks one narrow thing about one season of playing time, and the answer to that narrow thing is what gets printed.
The counterfactual at the centre of it is a fiction. There is no specific player who would have taken those plate appearances, and if there were, he would be a different fiction for every club. Replacement level is a constructed standard that stands in for all of them, which is why the next two sections are the most important in this article.
Replacement level is not average, and everything turns on that
Average is a high bar. An average major league regular is one of the best few hundred baseball players alive, and a club that fielded a genuinely average player at every position would win roughly half its games, which is what average means. No club can assemble that for free. Average players cost money, and the whole point of the exercise is to work out what a player provided beyond what the club could have got for nothing.
Replacement level is meant to represent that nothing. It is the standard of play available at essentially no acquisition cost: the waiver claim, the minor-league free agent signed in January, the utility man already on the roster, the outfielder called up when the starter's hamstring goes. He exists at every club. He costs the league minimum. He is not terrible, because nobody terrible is in the major leagues at all, and he is clearly worse than average, because if he were average somebody would have paid for him.
Three consequences follow immediately, and all three are routinely missed.
Playing time has value on its own. A durable player who is slightly below league average still accumulates positive WAR, because every game he plays is a game the club did not have to hand to somebody materially worse. Durability is not a tiebreaker in this statistic. It is one of the inputs.
Zero is not nothing. A player at zero WAR did not fail to contribute. He contributed exactly as much as a free substitute would have. He may have hit competently, fielded his position and run the bases without incident, and still come out at zero, because the club would have been no worse off had he never been signed.
A league-average regular is worth a positive and substantial number of wins. The gap between average and replacement, multiplied across a full season of playing time, is a real quantity, and it is baked into every figure you read. Anybody who reads WAR as though average were the baseline will underrate every full-time player by the same margin.
Negative WAR is possible, and it means precisely what it says: the club would have been better off with the free substitute. It happens most often to players kept in the lineup for reasons that have nothing to do with the current season.
Nobody measured replacement level, they agreed on it
Here is the thing readers get wrong more than anything else in this subject. Replacement level is not a measurement. It is a choice, and it was made by people rather than found in data.
Why can it not simply be measured? Because the population you would measure is selected by the very judgement you are trying to calibrate against. Take every minimum-salary call-up and average their production, and you are averaging a group that clubs have already filtered. The ones who play well keep playing and accumulate a large sample. The ones who play badly lose their place within a fortnight and their sample stays tiny. Playing time is assigned by scouts and managers who are estimating exactly the quantity you want to derive, so the measurement contaminates itself before you start.
You can approach it from the other direction, through roster economics: how many players are available at the league minimum in a given winter, what the twenty-sixth man on a roster typically produces, how far a club has to reach before the talent pool runs dry. Every route to the number requires assumptions that cannot be checked.
What actually happened is that the two main public producers agreed on a shared definition rather than a shared discovery. They fixed a total quantity of wins above replacement available across the major leagues in a season, split it between the leagues in proportion to games played, and worked backwards from there. That agreement is the reason the two versions sit on the same scale at all and can be quoted in the same sentence. It is a convention adopted for comparability, not a fact about baseball.
The effect of the choice on the output is specific, and it is worth understanding because it decides which comparisons survive it.
Lower the replacement level and every WAR figure in the league goes up. Raise it and they all come down. Straightforward enough, except that the shift is not uniform. The replacement adjustment is proportional to playing time, so moving the baseline moves a 650-plate-appearance regular much further than it moves a bench player who got 200.
Which means the choice of replacement level is, in practice, a weighting of durability against rate of production. Two players with identical playing time keep the same gap between them wherever the baseline sits, because the adjustment adds the same amount to both. Two players with different playing time do not. Their gap widens or narrows depending on a number that nobody measured.
That is the honest statement of the problem. Replacement level cancels out of exactly the comparisons nobody argues about, and refuses to cancel out of the comparison people argue about most: the durable good player against the brilliant injured one.
The baseball stat explained component by component
For a position player there are five run components and then a divisor. Each is computed against a different baseline, which is part of why the whole thing is hard to hold in your head.
Batting runs. Every plate appearance outcome carries a run value taken from how much that outcome changes the runs expected in the rest of the inning. A walk is worth less than a single, a single less than a double, a strikeout is negative, a double play is very negative. Add up the run values of everything a hitter did, compare with what a league-average hitter would have produced in the same number of plate appearances, and you have batting runs above average. This is the sturdiest component by a wide margin. There are six hundred or so plate appearances in a full season, every one has a directly observed outcome with no attribution problem, and the run values themselves come from an enormous sample of real innings.
Baserunning runs. Two parts. Stolen bases and times caught stealing, priced against each other on the correct terms, which is to say that a caught stealing costs more than a stolen base gains, because outs are expensive and the break-even success rate is high. Then everything else: going first to third on a single, scoring from second, being thrown out trying, avoiding the double play. Modest in magnitude, reasonably well measured, and it decides an argument perhaps once a decade.
Fielding runs. The estimated runs a player saved or gave away compared with an average fielder at the same position. This is the weak joint in the entire structure and it gets its own section below.
The positional adjustment. A correction that accounts for the fact that fielding runs are measured against a different standard at every position. Also its own section, because what it asserts is more interesting than it looks.
The replacement adjustment. The runs a replacement-level player would have been below average across the same playing time. Adding it converts a figure that is "above average" into one that is "above replacement". Note its shape carefully: it depends only on how much the player played and not at all on how well he played. It is identical for a magnificent season and a wretched one of the same length.
On top of those sits a small league adjustment, which exists so that a league's totals reconcile. Interleague play means the two leagues are not automatically on the same footing, and without a correction the batting runs in a league would not sum to zero as they are supposed to. Tiny per player. Not zero.
Then the whole pile of runs is divided by the season's runs-per-win converter, and out comes a number with one decimal place that people will argue about for six months.
- Price every plate appearance in runsEach outcome carries a run value taken from how much it changes the runs expected in the rest of the inning. A double is worth more than a single, an out is negative, a double play is worse still.
- Compare with a league-average hitterSum the player's run values and subtract what an average hitter would have produced in the same number of plate appearances. The result is batting runs above average.
- Correct for the parks he played inNot just his home park. Half his games were on the road and no two players face the same away schedule, so the correction follows the actual venues.
- Add baserunningSteals and times caught, priced against each other, plus taking the extra base, getting thrown out doing it, and staying out of double plays.
- Add fieldingThe estimated runs saved or given away against an average fielder at the same position. The least reliable line on the sheet, and the one that moves most between versions.
- Add the positional adjustmentProrated across every position he actually played, in innings. Positive at the demanding end of the defensive spectrum, negative at the easy end, largest negative at designated hitter.
- Apply the league adjustmentA small correction so that the league's totals reconcile, since interleague play means the two leagues do not sit on the same baseline by default.
- Add the replacement adjustmentThe runs a freely available player would have been below average over the same playing time. This is the step that turns runs above average into runs above replacement, and it depends only on how much he played.
- Divide by runs per winThe converter comes from the relationship between run differential and wins, evaluated at this season's scoring level. Read the answer to one decimal place, and do not believe the decimal.
The order of the components, for a position player. Every figure quoted in the worked examples elsewhere in this article is invented and round; no league constant, replacement level or converter is reported here.
Work one through with invented figures, chosen round so the arithmetic stays legible. A shortstop who plays a full season. Say his bat was worth 25 runs above the league average once his parks were accounted for. Say his baserunning added 3. Say the fielding metric had him 8 runs better than an average shortstop. Say the positional adjustment for a full season at shortstop came to 7. Say the replacement adjustment across that much playing time came to 20.
Total: 63 runs above replacement.
- Batting runs above average25
- Replacement adjustment20
- Fielding runs at shortstop8
- Positional adjustment7
- Baserunning runs3
Every figure is invented and deliberately round, chosen so the arithmetic is easy to follow. These are not any player's real components and no league constant is reported here.
Show the numbers
| Item | Value |
|---|---|
| Batting runs above average | 25 |
| Replacement adjustment | 20 |
| Fielding runs at shortstop | 8 |
| Positional adjustment | 7 |
| Baserunning runs | 3 |
Look at the shape of that. The replacement adjustment is the second-largest block on the chart and it says nothing whatsoever about the player. It is a function of playing time and a convention. Roughly a third of the number in this invented case is an accounting entry, and if the convention moved, so would it.
Why a shortstop and a first baseman with the same bat are not worth the same
The positional adjustment strikes most people as the arbitrary part, and it is in fact among the better-founded pieces of the calculation.
The problem it solves is created by the component above it. Fielding runs are measured against an average fielder at the same position. An average shortstop scores zero. An average first baseman scores zero. They have both been graded against their own peer group, and their peer groups are doing jobs of wildly different difficulty. Without a correction, the two would enter the sum identically, which would be absurd.
What the adjustment asserts is precise, and narrower than the version people argue with. It says that the amount of run prevention a player of a given fielding ability provides depends on where you stand him, and that the size of that difference can be estimated. The estimate comes from players who played two positions in the same season: take a man who spent half a year at shortstop and half at second base, look at how his fielding runs changed when he moved, and repeat across every player who has ever done it. The gaps between positions fall out of that comparison.
The resulting order runs roughly like this. Catcher at the top, then shortstop, then centre field, then second and third base, then the corner outfield spots, then first base, and designated hitter at the bottom by a distance. That ordering is confirmed by something you can watch happening every season without any statistics at all. Players move down that list as they age and almost never move up it. Shortstops become third basemen. Third basemen become first basemen. First basemen become designated hitters. The traffic is one-way, and a hierarchy of difficulty is the obvious explanation.
Designated hitter takes the largest penalty for two separate reasons. There is no defensive contribution to credit, and there is a well-documented tendency for the same hitter to produce slightly less when he is not taking the field.
The adjustment is prorated. A player who spent two thirds of his innings at shortstop and a third at second base gets a weighted blend, and a player who moved positions in July gets a figure that reflects both stints.
Now, what is it asserting about scarcity? Not that shortstops are rare, exactly. It is asserting something more useful about roster construction. If you have a shortstop who hits like a first baseman, you have bought yourself freedom in the other eight lineup spots that your rivals do not have, because you can now afford a slugger who cannot field anywhere except first base. The value is not entirely contained in the shortstop's own run prevention. Some of it lives in what he permits you to do elsewhere, and the positional adjustment is the mechanism by which that gets priced.
Nearly two wins of difference between two players who hit identically, ran identically and fielded their own position identically. That is the adjustment doing its work, and whether you find it convincing depends entirely on whether you accept the calibration behind it.
Where it is genuinely contestable: the adjustment is set at league level and applied to everyone at that position. A first baseman who is unusually good at digging low throws out of the dirt saves runs that the fielding metric largely cannot see and the positional adjustment certainly does not. He is credited as though first base were the same job for him as for everyone else, and it is not.
Fielding is the joint where the whole structure creaks
Batting is easy to measure and fielding is close to the hardest thing in the sport to measure. Both go into the sum with equal authority, and nothing on the leaderboard tells you which is which.
Five reasons fielding is hard, and they compound.
Attribution. A batted ball into the hole is a chance for the shortstop or the third baseman or neither, and before you can judge whether a play was made you have to decide who could plausibly have made it. Every fielding system has to solve that problem and they solve it differently.
Opportunity is not chosen by the fielder. A shortstop can only field balls hit near him. A season in which the pitching staff generated more ground balls to the left side gives him more chances to accumulate credit, and none of that was his doing.
Positioning is not his either. Where a fielder stands is ordered from the bench, informed by a spray chart prepared by an analyst. When a second baseman is standing in short right field and catches a ball there, the credit belongs partly to whoever positioned him. Fielding metrics have had to decide how much of that to hand to the player, and different systems decide differently, which alone is enough to make them disagree.
Sample size. A hitter gets six hundred fully observed plate appearances. A fielder gets a few hundred plays in a season where the outcome was ever in doubt, and each of them is worth a fraction of a run. The signal accumulates far more slowly than the noise, which is why a single season of fielding data tells you much less than the tidy figure implies.
The input data has changed underneath the metrics. Some systems were built on batted-ball locations recorded by human observers. Others are built on optical and radar tracking that measures where every fielder started, how fast he moved and how much time he had. Those systems disagree partly because they are looking at different things, not merely computing differently.
The practical consequence is the single most useful thing to know about reading WAR. A defensive figure needs several seasons before it carries the information that a batting figure carries in one. A one-season fielding line is an estimate with a wide band around it, and it is added to a batting line that is an estimate with a narrow band around it, and the sum is printed without any indication that the two were not equally solid.
Which is exactly why, when two versions of WAR disagree sharply about a position player, the disagreement is almost always in the fielding line. Two methods can look at the same outfielder's season and differ by a win, and neither of them is being careless.
Catchers are the extreme case. A catcher's largest defensive contribution may well be pitch framing, the business of receiving borderline pitches in a way that gets them called strikes, and the two public versions have taken different views on whether to price it at all. Where framing is included it is worth a great deal. Catcher WAR is therefore the least comparable figure across versions of anything in the statistic, and a catcher argument conducted with one version against the other is not an argument about the player.
Then the ground moves. Restrictions on defensive positioning changed what fielders are permitted to do, which changed what the metrics are measuring from one season to the next. Comparisons across that boundary are comparisons between two different measurement problems wearing the same name.
Pitcher WAR and the fork in the road
For pitchers the two public versions do not merely differ in detail. They answer different questions, on purpose, and this is where the largest gaps between them appear.
The frame is the same: how many runs did this pitcher prevent, relative to a replacement pitcher, across the innings he actually threw? The fork is over what counts as a run he prevented.
One route uses what actually happened. Start with the runs the pitcher allowed per nine innings. Adjust for the park. Adjust for the quality of the hitters he faced, since a reliever used mostly against the bottom of lineups has an easier job. Adjust for the quality of the fielders behind him, because a pitcher working in front of a poor defence gives up runs that were not his doing and one working in front of a great defence is flattered. Convert to runs above replacement, weight by innings, done.
The other route uses only what he controlled. Strikeouts, walks, hit batters and home runs allowed are very largely the pitcher's own work, with the fielders barely involved. What happens to a ball put in play is much more the fielders' business and much noisier from year to year. Fielding-independent pitching prices only the first group and puts the result on the same scale as earned run average. Build WAR on that and you get a figure that ignores the outcomes of balls in play completely, on the grounds that they were never reliably his.
Neither is wrong. They are answers to different questions. The runs-allowed version describes what the season contained. The fielding-independent version estimates what the pitcher is likely to do next, and it predicts noticeably better. A club deciding whether to extend a contract wants the second. A voter deciding what happened last year arguably wants the first.
The gap between the two opens widest for a pitcher whose runs allowed and whose fielding-independent estimate diverge, and the gap has real causes in it as well as luck. Some pitchers genuinely suppress hard contact. Some work in front of exceptional defences. Some pitch differently with runners on base, which the season-long fielding-independent figure cannot see. Neither construction separates the real effect from the fortunate one, and that is not a solvable problem with one season of data.
Home runs are the sore point in the fielding-independent route. A home run is charged fully to the pitcher, and the rate at which a pitcher's fly balls leave the park is one of the least stable things he does from year to year. Some implementations substitute an expected home run figure derived from fly-ball rate, which is a third construction again, and it moves individual pitchers a long way. All of this sits downstream of what a pitcher is actually trying to do with each offering, which is the subject of the deliberate engineering of a pitcher's arsenal and a separate discipline entirely.
Relievers add one further complication. An inning in a tie game in the ninth changes a team's win probability far more than an inning six runs down, and both versions apply some adjustment for the leverage a reliever pitched in. Both apply it only partially, on the reasonable grounds that the manager chooses the leverage and the pitcher merely turns up. Replacement level is also set differently for relief innings than for starting innings, because relief innings are cheaper to fill from the waiver wire.
Park and league adjustments do more work than anyone credits
Casual readers treat the park adjustment as a formality applied to a couple of extreme stadiums. It is nothing like that small. The distance between the most run-friendly and most run-suppressing venues in a league is wide enough to move a hitter's batting runs by an amount comparable to the entire positional adjustment. Skip it and you systematically overrate half the league and underrate the other half, every season, in the same direction.
The mechanism is to estimate each park's run environment relative to a neutral one and correct a player's line for the parks he actually played in. Not simply his home park. Half his games were away, and no two players in a league face identical away schedules.
The complications are where the interesting part sits. Park factors are themselves estimates drawn from small samples, so they are computed over several seasons and smoothed, which means a factor in use today is partly describing a stadium as it was some years ago. They differ by handedness, since a short porch in right field is not neutral between a left-handed hitter and a right-handed one. They differ by outcome type, because a park can suppress home runs while inflating doubles. And parks change: fences get moved, walls get raised, humidors get installed, the ball itself is not constant. How a stadium's run environment is estimated is a genuinely difficult piece of work, and every WAR figure inherits its errors silently.
The league adjustment is smaller and duller and still necessary. It ensures that batting runs above average within a league actually sum to zero, and it absorbs the differences between the two leagues that interleague play would otherwise smear across the totals.
There is a broader point here about baselines. WAR is defined relative to the league in which it was earned, which is what makes it usable across eras. A six-win season in a high-scoring era and a six-win season in a low-scoring one are on the same scale by construction. That is a feature, and it has a cost: WAR cannot tell you which of the two players was better in absolute terms, because it never asked. It only ever measured distance from a contemporary baseline. Any rule change that alters the run environment, and the shortening of the game by the pitch clock is a recent example, moves the baseline for everybody at once and leaves the WAR scale intact while changing what a run is worth underneath it.
Turning runs into wins, and where that conversion comes from
Everything above is in runs. The last step converts to wins, and the converter is not a constant.
It comes from the relationship between a team's run differential and its win total. Score more than you allow and you win more than you lose, in a relationship well described by a curve rather than a straight line. Take the slope of that curve at the league's actual scoring level and you have the number of marginal runs required to buy one extra win.
The important property is that the answer depends on the scoring environment. In a high-scoring league a single run is worth less, because scores are larger and noisier and a one-run edge changes the outcome of fewer games. In a low-scoring league runs are dearer and the same run buys more win. The converter therefore moves from season to season, and it is published each year by the people who compute it rather than being fixed by convention.
Take the invented shortstop from earlier. He finished at 63 runs above replacement. Say the converter for that invented season works out at 10 runs to a win, a round figure chosen here purely to keep the arithmetic readable. His WAR is 6.3.
Strictly, the converter should reflect the run environment of the specific team rather than the whole league, since a pitcher on a low-scoring team is operating on a different part of the curve. Some implementations do exactly that for pitchers. The effect is second-order and it is one more small reason two versions can differ.
- 5Run components in a position player's total
- 8Fielding positions with their own adjustment
- 2Constructions of pitcher run prevention in use
- 1Baselines the entire number rests on
Five run components for a position player: batting, baserunning, fielding, the positional adjustment and the replacement adjustment, before the small league correction. Eight fielding positions carry their own adjustment, excluding the pitcher; the designated hitter carries one without fielding at all. Two constructions of pitcher run prevention are in common public use.
That last figure is the one to sit with. Every component, every adjustment, every correction for park and league and position, all of it hangs off a single agreed baseline that nobody measured.
Why the public versions disagree, and why the gap is not a verdict
The two main public versions agree on the frame and disagree on the inputs. Roughly in order of how much difference each source makes:
Pitcher construction. Runs actually allowed against fielding-independent estimates. Comfortably the largest, and for some pitchers it is worth multiple wins in a single season.
The fielding metric for position players. Second largest, and the reason two versions can rate the same outfielder a full win apart while both being computed carefully.
Catcher framing, priced or not priced. Enormous for the small population it affects.
Park factor construction. How many seasons go into it, how it is smoothed, whether it splits by handedness.
Everything else. Baserunning components, league adjustments, the exact treatment of double plays, rounding.
They share the replacement-level definition, which is the only reason the two scales are commensurable at all.
Now the part that matters for anybody quoting them. A gap of a couple of tenths of a win between the two versions for the same player is not information. It is the noise floor of a calculation with this many estimated inputs. Neither figure is precise to a tenth. Treating a small difference as meaningful, or picking whichever version favours your argument and presenting it as the number, is the most common bad-faith move in the whole discourse and frequently is not even conscious.
The right way to read a disagreement is as a map of where the uncertainty lives. Two versions that both land near six wins for a player is a much stronger statement than either figure alone. Two versions at 4.2 and 6.1 is telling you that this player's value sits mostly in a component nobody can measure well, which in practice means fielding or balls in play. That is a useful finding. It is just not the finding people take from it.
None of this is peculiar to baseball. It is the identical argument that trails expected goals around a football weekend, and the same one that attends golf's shot-by-shot accounting against a baseline. A baseline-relative account of what already happened keeps being read as a verdict on quality, or worse, as a forecast.
What WAR is genuinely good for
Plenty, provided the uses match the precision.
Roster accounting in one currency. Add the WAR of every player on a roster and you land close to the club's actual win total, with a residual that is itself worth examining. No other common statistic is denominated in the thing clubs are trying to buy, which is why the number now appears next to every name in baseball's statistical coverage.
Comparing players who do different jobs. A shortstop who fields against a first baseman who hits, a starting pitcher against a corner outfielder, a catcher against a designated hitter. Nothing else in the sport attempts this, and the attempt is valuable even when the answer is imprecise.
Comparing across eras. Because everything is relative to a contemporary league, a season from decades ago and one from last year can be placed on a common scale without pretending the two leagues were alike.
Sorting. This is the underrated one. WAR is excellent at separating a six-win player from a two-win player and useless at separating 6.1 from 5.8. Used as a sieve to find the group of players worth arguing about in detail, it is the best tool available. Used to order that group, it is being asked for a precision it does not have.
Careers. Over ten seasons the component errors partly cancel, the fielding noise averages out towards something real, and the total becomes a far better instrument than any single season within it. Career WAR is a much more trustworthy number than seasonal WAR, which is close to the opposite of how the two are typically deployed.
Where WAR gets misused: awards ballots and contract valuation
Two misuses dominate, and they fail for different reasons.
Single-season awards arguments. A single season is one sample of everything WAR is least confident about. If two candidates are separated by less than the plausible error in the defensive component alone, the order between them on the leaderboard is arbitrary, and quoting it to one decimal place lends a false air of settlement to what is really a judgement call. WAR is also context-neutral by design: it does not record that a double came with the bases loaded in a September pennant race rather than in the fourth inning of a blowout in May. Whether a voter should care about that is a question about values, not statistics, and WAR has quietly taken a side by not measuring it.
Contract valuation. The dollars-per-win figure that circulates every winter comes from dividing what clubs actually spent in free agency by the wins those signings were expected to produce. That is a market price, backward-looking, and it is not a valuation of a player. Applying it naively fails in at least four ways.
Wins do not come in interchangeable units. Roster spots are finite, so one six-win player plus a replacement-level filler is worth more to a club than two three-win players, who consume two spots for the same total. Clubs pay a premium for concentration and a flat rate per win cannot price it.
The marginal cost of a win is not the same for every buyer. A club already above the competitive balance tax threshold pays a surcharge on every additional dollar, so the same win costs it materially more than it costs a club with room, which is one of the reasons the tax reshapes the winter market rather than merely raising money.
Wins are not equally valuable to every club either. Moving from 84 wins to 85 is worth vastly more than moving from 68 to 69, because the probability of reaching the playoffs is not linear in wins. It barely moves at the bottom and moves steeply in the middle.
Applying a free agent price to a player who is not a free agent is a category error. Pre-arbitration players are not on the market at all, and the surplus value the calculation produces is real for the club's accounts while saying nothing about what the player is worth in an open market he is contractually barred from entering.
And then there is age. A contract pays for future wins, and WAR is a record of past ones. Paying market rate for last season's figure is buying a decline curve at the price of a peak.
How to use the number honestly
Six checks, in the order that matters.
Ask which version, then look up the other one. If they agree, you have a much stronger statement than either alone. If they disagree by more than half a win, the disagreement is the story and it will nearly always be in the fielding line or the pitcher construction.
Read the components, not the total. The total is the least informative number on the page. A five-win season built from batting is a different asset from a five-win season built from fielding, because one of them is much more likely to still be there next year.
Treat the fielding line as an estimate with a wide band. One season of defensive value is barely evidence. Three or more begins to be.
Never argue from the tenths. If the case collapses when the figure moves by half a win, there was no case, only a leaderboard.
Remember that playing time is inside it. Comparing a 140-game season with a 100-game season on WAR is comparing value delivered, not ability. If ability is the question, the question needs a rate, and WAR is not one.
Know where zero is. Zero is not a failure and it is not nothing. Zero is the level of production a club can obtain for free, and its exact position is the one assumption in the entire calculation that nobody ever measured.
The usual criticism of WAR is that it is an estimate. That is not the flaw. Every number worth having in sport is an estimate. The flaw is that it is published without its error bars, to one decimal place, on a sortable table, which is an open invitation to exactly the argument it was never built to settle.
The question underneath it stays worth asking, and it is a better question than most of the ones the box score answers. Against a player you could have had for nothing, how much better was this one?
Everything else in this article is bookkeeping in service of that.
Common questions
What does WAR actually measure in baseball?
WAR estimates how many more games a team won with a given player than it would have won with a freely available replacement, over the same playing time. It is expressed in wins because wins are what clubs are buying, and every component underneath it is calculated in runs first and converted at the end. It measures value delivered rather than talent, which is why playing time is part of the answer rather than something the number controls for.
What is replacement level in baseball?
Replacement level is the standard of play a club can obtain at essentially no acquisition cost: a waiver claim, a minor-league free agent, or the next man on the depth chart. It sits well below league average, because an average major leaguer is a genuinely good player who costs real money to acquire. Its exact position is agreed by convention among the people who publish WAR rather than measured from data, and moving it changes every player's figure, with full-time players affected more than part-time ones.
Why do FanGraphs and Baseball Reference WAR differ?
The two share a replacement-level definition but differ on inputs. The largest difference is pitchers: one version builds from runs actually allowed, adjusted for park, opposition and the fielders behind the pitcher, while the other builds from strikeouts, walks, hit batters and home runs and ignores what happened to balls in play. For position players the choice of fielding metric and the treatment of catcher framing account for most of the rest, which is why the disagreements cluster on defenders and catchers.
Is a difference of half a win in WAR meaningful?
Not on its own. The defensive component of a single season carries enough uncertainty that two careful methods can differ by that much on the same player, and neither figure is precise to a tenth of a win. WAR separates a six-win player from a two-win player reliably and cannot separate 6.1 from 5.8, so an argument that depends on the tenths is unsupported by the number it cites.
Why is a shortstop worth more than a first baseman with the same numbers?
Fielding runs are measured against an average fielder at the same position, so an average shortstop and an average first baseman both score zero despite doing jobs of very different difficulty. The positional adjustment corrects for that, adding runs at the harder positions and subtracting them at the easier ones, calibrated from how players' fielding figures change when they move between positions in the same season. Identical bats at shortstop and first base therefore produce different WAR totals, which is the statistic asserting that a shortstop who hits is harder to find than a first baseman who hits.
Filed under Baseball·baseball · mlb · baseball statistics · analytics · wins above replacement