Skip to content
CricketTaken

Analysis

Box plus minus explained, and the defence it cannot see

How BPM turns a box score into an estimate of on-court impact: the RAPM regression underneath it, the coefficients, the team adjustment, and where it fails.

By CricketTaken EditorialPublished Analysis22 min read

How this is written and checkedReport an error

A number gets quoted in an argument. Somebody's box plus minus was 6.4, therefore he was better than the player at 5.1, and the conversation moves on as though a measurement has been produced.

It has not. Box plus minus explained honestly is a chain of estimates, each one standing in for something that could not be measured directly, and the figure at the end inherits every weakness in the chain. It is a model of a model: a regression on box score rates, fitted to reproduce the output of a second regression that was itself estimating on-court impact from lineup data. Understanding what the number is good for means understanding what each of those layers was trying to solve, and what it gave up in solving it.

The problem raw plus/minus could not solve

Start with the crude version, because everything else is a repair to it.

Raw plus/minus is the scoreboard margin accumulated while a player is on the floor. It has the enormous virtue of measuring the thing everyone actually cares about, which is whether the team outscores the opposition, and it needs no assumptions about which actions matter. It also has a defect so large that it makes the raw figure close to useless for evaluating an individual: a player's margin is mostly a description of the four teammates he plays with and the five opponents he plays against.

A reserve who spends his minutes alongside the two best players on a contender will post a handsome margin. A starter on a poor team who plays the toughest available opponents in every game will post an ugly one. Neither figure is telling you much about the player.

Adjusted plus/minus is the fix, and it is the obvious one once stated. Break the season into stints, which are stretches of game time with an unchanged ten players on the floor. Each stint has a point margin and a length. Treat every player in the league as a variable and solve, across tens of thousands of stints, for the set of individual coefficients that best explains the observed margins. The teammate and opponent effects are no longer contamination; they are other terms in the same equation, and the regression separates them.

That works in principle and behaves badly in practice, for one reason. Basketball rotations are not random. Two players who share almost all of their minutes are nearly indistinguishable to the regression, because there is barely any data in which one appears without the other. The mathematics responds to that ambiguity by producing wild, unstable coefficients that swing hugely from season to season and sometimes assign one player of a pair a heroic rating and his shadow a catastrophic one.

Regularised adjusted plus/minus is the standard answer. It adds a penalty term that pulls every coefficient towards a prior, so that the regression has to earn any extreme rating with a genuine weight of evidence rather than with an artefact of collinear minutes. The results stabilise. The cost is that the ratings are shrunk towards the prior, and a player with few minutes will end up close to whatever the prior said about him rather than close to what his own stints implied.

Even a well-behaved RAPM has two hard limits. It needs play-by-play data, which does not exist for most of the history of the sport. And it needs a lot of it, which is why the serious versions are fitted over multiple seasons rather than one.

Box plus minus exists because of those two limits. It asks whether the output of a RAPM can be predicted from a box score alone, and the answer turns out to be yes, imperfectly, and in a way that is available for every season a box score survives.

What the number means before you read another word

BPM reports points per 100 possessions above league average, for the team, while that player is on the floor. League average is fixed at exactly 0.0.

It is purely a rate. The published documentation is emphatic about this: playing time does not enter the calculation at all. A player who was excellent for six hundred minutes and a player who was equally excellent for two thousand receive the same rating, and the difference between them is expressed by a separate statistic, VORP, dealt with further down.

The developer's own published scale gives the number its meaning, and it is worth committing to memory because the figures alone carry no intuition:

  • +10.0 is described as an all-time season
  • +8.0 is an MVP season
  • +6.0 is an all-NBA season
  • +4.0 is in all-star consideration
  • +2.0 is a good starter
  • 0.0 is a decent starter or solid sixth man
  • -2.0 is a bench player, and is also the level defined as replacement

Two things about that ladder deserve attention. The first is that it is compressed at the top. The distance from a good starter to an MVP season is six points per 100 possessions, which sounds small until you convert it: an elite team's regular-season efficiency margin sits somewhere around the same magnitude as a single MVP-level player's rating, and a team's best five-man lineup might reach the mid-teens. Individual ratings and team ratings are on the same scale, which is the design decision that makes the number legible.

The second is that most of the league is below zero. Because better players play more minutes, there are far more below-average player seasons than above-average ones by simple count, even though the two sides balance when weighted by minutes. A rating of 0.0 is not the middle of the population. It is the middle of the minutes.

That is a different framing from the rating that fixes league average at 15, where the scaling is chosen for familiarity rather than for interpretability, and it is closer in spirit to the per-100-possession team ratings that the metric ultimately has to reconcile itself against.

Box plus minus explained, one calculation step at a time

The published method is a sequence, and the order is not negotiable because each stage consumes the output of the previous one.

How one BPM figure is produced, in the published order
  1. Estimate the player's position and offensive roleNeither is looked up from a roster. Both are regressed from the player's share of his team's rebounds, steals, fouls, assists and blocks, and from his scoring efficiency, on a scale running from 1.0 to 5.0. This happens over the whole season before anything else can be computed.
  2. Select the coefficients those estimates implySome weights are constant across the league. Others slide linearly between a point-guard value and a centre value, or between a creator value and a receiver value, according to the two position numbers just estimated.
  3. Correct the player's points for the team's shooting contextThe team's average points per adjusted shot attempt is compared with the baseline the regression was fitted on, and a constant is added to every player on the roster to reconcile the two before any individual rating is computed.
  4. Compute the raw ratingMultiply each per-100-possession box score rate by its coefficient, sum the terms, then apply the position constant and the offensive role constant. This produces the raw BPM, and it is the only stage that looks at the player in isolation.
  5. Reconcile the roster against the teamSum the raw ratings across the roster, weighted by share of minutes, and compare that total with the team's adjusted efficiency, itself corrected for the effect of playing with a lead.
  6. Add the team adjustment to everyoneA single constant is added to every player on the team so the weighted total lands exactly on the team's efficiency. This constant also serves as the regression's intercept, which is why it is typically a large negative number rather than a small correction.
  7. Read the resultRaw rating plus team adjustment is the published BPM. A rating is therefore never a statement about a player alone; it is a statement about a player given how good his team turned out to be.

The calculation sequence set out in the BPM 2.0 methodology. Each stage refers to a component described in the article. The coefficients and constants referenced are the published ones, not estimates made here.

The starting assumption underneath all of that is unusual enough to state plainly, because it explains most of the metric's behaviour. BPM begins by assuming every player on a team contributed equally. If the team was good, everyone on it is provisionally good. The box score is then used to revise that assumption, and the revision is made relative to teammates rather than relative to the league. Does this player collect more steals than the others on his roster? Does he score more efficiently than they do? Does he turn it over more?

That relative framing is why the team adjustment is not an afterthought. It is the level from which every individual departure is measured.

The coefficients are the interesting part

The weights are published in full, and reading them is the fastest route to understanding what the regression concluded about basketball. Points carry a coefficient of 0.860 and it does not change with position. A made three carries an extra 0.389 on top of the points it already contributed, uniform across the league, and the documentation is explicit that this bonus stands in for spacing and the other benefits of three-point volume that never appear as a discrete box score event. Turnovers are charged at -0.964 and personal fouls at -0.367.

Field goal attempts are a cost rather than a credit, at -0.560 for a pure creator rising to -0.780 for a pure receiver. Free throw attempts are charged at 0.44 of the field goal attempt weight, which is the same possession-conversion convention that runs through the true shooting formula. The reasoning behind the creator-to-receiver slide is worth sitting with: a shot taken by a low-usage player is more likely to have been created by somebody else, so the regression charges him more for taking it.

Then come the weights that swing hardest with position, and they are the ones that reveal the model's view of the sport.

Published BPM coefficients that change with estimated position
  • Point guard (position 1)
  • Centre (position 5)
Steals1.371.01
Blocks1.330.7
Assists0.581.03
Offensive rebounds0.610.18
Defensive rebounds0.120.18

The BPM 2.0 regression coefficients as published in the metric's methodology, applied to per-100-possession box score rates. Values slide linearly between the position 1 and position 5 figures, so a small forward at position 3 receives the midpoint. These are model parameters, not measurements of play.

Show the numbers
Published BPM coefficients that change with estimated position
ItemPoint guard (position 1)Centre (position 5)
Steals1.371.01
Blocks1.330.7
Assists0.581.03
Offensive rebounds0.610.18
Defensive rebounds0.120.18

Read the assist row first. A centre's assist is worth nearly twice a point guard's. That is not a slight against playmakers; it is the regression noticing that point guards accumulate assists as a routine function of holding the ball, while a big man who racks them up is doing something unusual, and that the same big men tend to be better defenders. The coefficient is absorbing a correlation as much as a causal contribution, which is true of every weight in the table and is the honest way to read all of them.

Then the defensive rebound row, which is the most quietly radical number in the whole model. A defensive rebound by a guard carries a coefficient of 0.116. Almost nothing. The documentation explains why: defensive rebounds matter enormously to the team, but it rarely matters which player collects one, so the credit is split among everybody on the floor rather than assigned to the man who jumped. The value has not been denied. It has been moved into the team adjustment, where every player shares it.

Offensive rebounds run the other way. They are worth 0.613 to a guard and 0.181 to a centre, because a guard who gets one has done something his position does not normally produce, while a centre who gets one has done his job.

Steals and blocks are both worth more to small players than to large ones. A block by a centre is expected and partially priced into his position; a block by a guard is information the model did not otherwise have.

Position and role are guessed from the box score, not looked up

This is the part of the method most people are unaware of, and it has consequences.

BPM does not use a listed position. It regresses one, from the player's percentage share of his team's statistics while on the floor, on a scale from 1.0 for a point guard to 5.0 for a centre. The published coefficients are an intercept of 2.130, plus 8.668 times the share of team total rebounds, plus 0.992 times the share of fouls, plus 1.667 times the share of blocks, minus 2.486 times the share of steals, minus 3.536 times the share of assists.

Rebounds and assists do the heavy lifting, in opposite directions, which is a reasonable description of how positions actually differ. The output is then clipped to the 1.0 to 5.0 range, and a constant is applied across the roster so that the minutes-weighted team average lands on exactly 3.0. A team cannot be all centres.

A second, independent scale estimates offensive role, running from 1.0 for a pure creator to 5.0 for a pure receiver. Its published coefficients are an intercept of 6.00, minus 6.642 times the share of team assists, minus 8.544 times the share of what the method calls threshold points, meaning points scored above a threshold efficiency set 0.33 points per true shooting attempt below the team's own average.

Two independent estimates, drawn from statistics that barely overlap, which is why the documentation notes it is theoretically possible for a player to come out as a position 5.0 and a creation role 1.0 at the same time. A centre who initiates the offence is not a contradiction the model has to resolve; it is two coordinates.

The consequence for a reader is that a player whose statistical profile does not match his job will be given the wrong coefficients. A guard who rebounds heavily will be pulled towards the big-man weights and will have his defensive rebounds valued more and his steals valued less than a conventional guard's. Whether that is a bug depends on whether you think the statistical profile or the job title is the better description of what he actually did, and reasonable people differ.

The team adjustment is the largest number in the calculation

Here is the step that gets skipped in every summary and that determines more of the final figure than any coefficient.

After the raw ratings are computed, they are summed across the roster weighted by share of minutes, and a single constant is added to every player so the total matches the team's adjusted efficiency. The published documentation notes that these constants are generally around -8, and explains why: the constant is also acting as the intercept for the regression. It is not a small correction bolted on at the end. It is a structural term without which the raw numbers do not mean anything.

The target the roster is reconciled against is itself adjusted, for a reason that is one of the more elegant details in the whole method. Teams play worse when leading and better when trailing, an effect the documentation attributes to work by Jeremias Englemann and describes as linear and consistently replicated: roughly 0.35 points per 100 possessions of degradation for every point of lead. Half of that is assigned to the leading team and half to the opponent. So a team's efficiency is nudged before the players are asked to sum to it, which stops good teams from having their players systematically underrated for the crime of spending the fourth quarter comfortable.

The practical implication for anyone reading a BPM figure is direct. A player's rating depends on how good his team was. Two players with identical box scores on teams of different quality will have different ratings, by construction and on purpose, because the metric's whole logic is that the box score revises a team-level starting point rather than replacing it.

That is a defensible design. It is also the mechanism by which a strong player on a bad team is quietly penalised, and it is the reason a BPM figure should never be quoted without knowing the record of the team that produced it.

A worked example that is not mine

The methodology publishes its own calculation, which is more useful here than anything I could invent, because every figure in it can be checked against the source.

The subject is LeBron James's 2017 regular season. The position regression estimates him at 2.3 rather than his listed small forward, and his offensive role at 1.0, a pure creator. Applying the coefficients to his per-100-possession rates and summing the terms gives a total of 18.7.

Then the two constants. His position constant is calculated as (3 minus 2.3) multiplied by (-0.818 divided by 2), which is -0.3. His offensive role constant is (3.0 minus 1.0) multiplied by (-2.774 divided by 2), which is -2.8. Together they take 3.1 points off, leaving a raw BPM of 15.6.

Then the team adjustment for that Cleveland roster, published as -8.0. The final figure is 15.6 minus 8.0, which is +7.6 points per 100 possessions.

Look at the size of that last step. The team adjustment removed slightly more than half of the raw number. Anyone who describes BPM as a box score statistic and stops there has described the first two thirds of the calculation and ignored the third that moved the answer most.

The same example continues into VORP. LeBron played 70 per cent of Cleveland's minutes that season, so his VORP is (7.6 minus -2.0) multiplied by 0.70 multiplied by 82 over 82, which is 6.7.

Box plus minus explained by what it never sees

Now the limitation, and the documentation states it more bluntly than most critics do.

On the offensive side the box score is nearly sufficient. Points, attempts, assists, turnovers and offensive rebounds between them capture most of what an offensive player does, with the documentation naming screen-setting as the main omission. That is a real gap, and a substantial one for certain players, but it is one gap rather than a category failure.

Defence is a category failure. The published text is worth quoting exactly: "Such critical components of defense as positioning, communication, and the other factors that make Kevin Garnett and Tim Duncan elite on defense can't be captured, unfortunately." And the instruction that follows: "Look at the defensive values as a guide, but don't hesitate to discount them when a player is well known as a good or bad defender."

The reason is arithmetic rather than philosophical. Consider what a box score contains that relates to defending. Blocks, steals, defensive rebounds, personal fouls. Now consider what a possession of good defence usually looks like: a player is in the right place early enough that no pass is attempted, the ball handler rejects a drive that was never going to be there, the shot clock burns down, and a worse shot goes up than the offence wanted. Nothing in that sequence produces a box score event. The defender who caused it appears in the record only through the absence of things.

The metric handles this the only way it can, by assigning the unexplained defensive value to everybody on the floor through the team adjustment. That is why a good defender on a good defensive team gets some credit: not because the model saw what he did, but because his team allowed fewer points and the credit was shared. It is also why an elite defender is systematically indistinguishable from his teammates.

The published all-time leaderboard makes the point better than any argument. Split the top seasons into their offensive and defensive halves and the shape is unmistakable.

The five highest BPM seasons on the published leaderboard, split into offence and defence
  • Offensive BPM
  • Defensive BPM
LeBron James 20099.53.7
Michael Jordan 19888.84.2
Michael Jordan 19918.93.2
Stephen Curry 201610.41.6
Michael Jordan 19898.43.4

Figures as published on Basketball Reference's BPM methodology page, for its all-time top-15 table with a minimum of 1,000 minutes played, accessed 2 September 2026. Ratings are recalculated over time, so the table is a snapshot rather than a permanent record. Defensive BPM is a residual, calculated as total BPM minus offensive BPM, and is not separately regressed.

Show the numbers
The five highest BPM seasons on the published leaderboard, split into offence and defence
ItemOffensive BPMDefensive BPM
LeBron James 20099.53.7
Michael Jordan 19888.84.2
Michael Jordan 19918.93.2
Stephen Curry 201610.41.6
Michael Jordan 19898.43.4

The offensive column carries the seasons. The defensive column is narrow, and it is narrow for every season on the list, not only these five. That compression is not a finding about the players. It is a finding about the box score: the instrument has very little resolution on that axis, so everybody clusters.

DBPM is a residual, not a measurement

There is a structural detail here that changes how the defensive figure should be read, and it is easy to miss in the documentation.

Offensive BPM is produced by its own regression, using the same variables with different coefficients. Defensive BPM is not regressed at all. It is calculated as total BPM minus offensive BPM.

That means every error in the offensive estimate lands, with its sign flipped, in the defensive one. A player whose offensive contribution is overstated by the model will have his defensive contribution understated by exactly the same amount, and there is no independent check anywhere in the calculation that would catch it. The defensive number is the leftovers.

The offensive coefficients tell their own story when set beside the totals. Points are weighted at 0.605 for offence against 0.860 overall. The three-point bonus is larger for offence, at 0.477 against 0.389, which the documentation reads as three-point shooters helping the offence more than they help the team overall, implying a small defensive cost. Personal fouls are charged at -0.439 on offence against -0.367 overall, meaning fouls are a net negative that carries a slight defensive positive inside it, which is the model noticing that players who foul are often players who contest.

The offensive role constant is also far smaller for the offensive rating, at 0.860 against 2.774. The documentation draws the obvious inference: most of the value the box score fails to capture is defensive value. The model knows where its blind spot is.

VORP, replacement level, and the conversion to wins

BPM ignores availability entirely, which is correct for a rate statistic and useless for answering the question people usually want answered. Value Over Replacement Player is the counting version.

The formula is published and short: BPM minus -2.0, multiplied by the share of team possessions the player was on the floor for, multiplied by team games divided by 82.

Three things are happening in that line. Replacement level is set at -2.0, a figure the documentation attributes to a discussion framed around Tom Tango's replacement-level work in baseball, and it represents the standard of player a team can acquire for the minimum without giving anything up. Measuring above that rather than above average is what makes the statistic additive: a full roster of average players is a fiction, whereas a full roster of replacement-level players is roughly what a team would field if it stopped trying.

The second term converts a rate into a total. A player who is +4.0 for 20 per cent of his team's possessions and a player who is +2.0 for 40 per cent produce the same VORP over the same number of games, which is the trade the statistic exists to price.

The third term prorates for a shortened or partially completed season, so mid-season figures are not compared against full ones by accident.

To turn VORP into wins, the documentation gives a multiplier of 2.7. The conversion is deliberately linear, taken near league average rather than in the diminishing-returns region of the Pythagorean relationship between point differential and record, and the documentation is candid that a genuinely great player would push a team into that diminishing region in reality. This is the same structural compromise made by wins above replacement in baseball, where a runs-to-wins conversion has to be linear to be additive and is not linear in the world.

Four published constants that decide what a BPM figure becomes
  • -2Replacement level, points per 100 possessions
  • 2.7Multiplier converting VORP into wins
  • 0.35Points per 100 lost per point of lead held
  • -8Team adjustment in the published worked example

Values taken from the published BPM 2.0 methodology. The team adjustment shown is the one published for the 2017 Cleveland roster in the developer's worked example; the documentation states these constants are generally around -8 because the term also serves as the regression intercept.

The documentation also notes what replacement level would look like if it were split, at roughly -1.7 on offence and -0.3 on defence, and then explains why those splits are not published: almost every point guard would sit well below the defensive threshold and almost every big would sit below the offensive one, because the concept of replacing a component of a player does not correspond to anything a front office can do.

Where the number goes wrong, and how to tell

The documentation itself lists the reasons a player can post a rating below replacement level, and the list is a good diagnostic checklist for any surprising figure. The player may genuinely be that bad, often because he is young. He may be shooting badly in a way that will not persist. He may be in a role that minimises what he is good at. He may be being stretched deliberately by a coaching staff investing in his development. Or the metric may simply not be capturing what he contributes, which the text singles out as a particular problem for elite defenders.

Add to those a few failure modes that follow from the mechanics described above.

Small samples run into the smoothing. For game-level calculations the method regresses low-minute players towards an estimate derived from their minutes per game, using a weight that starts at 150 for a player with no minutes and falls linearly to zero at 450 minutes. A player with a handful of appearances is being described substantially by that prior rather than by his own play.

Team quality leaks in. The team adjustment is the same for everyone on the roster, so a player who changes teams mid-season has his season figure assembled from two different reconciliations, and a player on a team whose efficiency is distorted by injuries carries that distortion.

Role changes break the coefficient selection. Position and offensive role are estimated over a full season, so a player who spent thirty games as a bench scorer and fifty as a primary creator gets one set of weights for both.

Shooting variance passes straight through. Points are the largest positive term in the model at 0.860, and points are a function of shots going in. A season of unusually good or bad three-point luck moves BPM directly, with none of the regression to the mean that a shot-quality model built on tracking data would apply.

The honest summary of all of this is the one the developer gives: BPM is good at measuring offence and solid overall, and its defensive numbers are a guide rather than a verdict.

How to use it without embarrassing yourself

Three habits, and they are enough.

Treat a BPM figure as a first screen rather than a conclusion. It is superb at the job of scanning a hundred and fifty player seasons and finding the twenty worth looking at properly, which is precisely the job it was built for and precisely the job no amount of film can do at that scale.

Never compare two players on a gap smaller than the metric's resolution. A difference of a few tenths between two players on different teams is not a finding. It is the two team adjustments, the two position estimates and one hot shooting month, in some combination nobody can decompose from the outside.

And read the offensive and defensive columns as different instruments rather than as two halves of one. The offensive figure is a reasonably direct measurement of things the box score records well. The defensive figure is a residual, computed by subtraction, from a record that never saw the possession where the defender's positioning meant no shot was taken at all. Set it beside on-off numbers, beside what a matchup-aware metric says, and beside the evidence of watching him. Where those disagree with the box score, the box score is usually the one that was not looking.

For everything else the metric touches, the same discipline applies across the rest of the sport's numbers: find out what the statistic could not see before you decide what it proved.

Common questions

What is a good BPM in the NBA?

The published scale sets league average at exactly 0.0, so any positive figure is above average. On the developer's own scale, +2.0 is a good starter, +4.0 is in all-star consideration, +6.0 is an all-NBA season and +8.0 is an MVP-level one, with -2.0 defined as replacement level. Because better players play more minutes, far more player seasons fall below zero than above it.

How is box plus minus calculated?

It begins by assuming every player on a team contributed equally, then revises that assumption using box score rates measured per 100 possessions and weighted by coefficients that change with the player's estimated position and offensive role. A team adjustment constant is then added to every player on the roster so the minutes-weighted total matches the team's actual efficiency. The coefficients themselves come from a regression fitted against a Bayesian-prior RAPM dataset.

What is the difference between BPM and RAPM?

RAPM is estimated from what actually happened while a player was on the floor, using play-by-play stint data and a regularised regression. BPM is a box-score model fitted to reproduce RAPM without needing play-by-play data at all, which is why it can be calculated for seasons long before tracking existed. BPM is an approximation of RAPM, not an alternative measurement of the same thing.

Why is defensive box plus minus unreliable?

Because the box score records almost nothing about defence. Blocks, steals, rebounds and fouls are the entire defensive vocabulary available, and they miss positioning, communication, rotations and deterrence. The metric's own documentation says the defensive numbers should not be treated as definitive and should be discounted where a player is known to be a good or bad defender.

What does VORP mean in basketball?

Value Over Replacement Player converts the BPM rate into a season total by multiplying the amount by which a player exceeds replacement level, defined as -2.0, by the share of team possessions he played and by the fraction of the season his team completed. Unlike BPM it rewards availability, so a good player who plays heavy minutes outranks a slightly better one who plays fewer.

Filed under Basketball·basketball · nba · analytics · basketball statistics · plus minus