Analysis
BABIP explained: what baseball's luck number really measures
BABIP explained for baseball: the formula, the quirks in its denominator, why pitchers regress and hitters do not, and how to tell noise from a real change.
By CricketTaken EditorialPublished Analysis18 min read
A hitter is batting .238 in the middle of June. He is striking out at the same rate as last season, walking at the same rate, hitting the ball as hard as he ever has, and pulling it at the same angle. Nothing has changed except the results, and the beat writers have started using the word slump.
BABIP explained properly is the tool that tells you whether there is anything there at all. It is the rate at which a baseball hit into the field of play turns into a hit, and it is the most useful single number in the sport for separating what a player did from what happened to him. It is also the most abused, because a decade of people calling it a luck statistic has left most readers with a badly wrong idea of what is inside it.
Luck is in there. So are four other things, and three of them are skills.
What BABIP actually counts
The formula is short.
BABIP = (H − HR) ÷ (AB − K − HR + SF)
Hits minus home runs on the top. At-bats, minus strikeouts, minus home runs, plus sacrifice flies on the bottom.
Read it as a question rather than an equation and it says: of all the times this player hit a ball that a fielder had a chance to do something about, how often did it end up as a hit?
Home runs come out of both halves because no fielder gets a vote on a ball in the seats. Strikeouts come out of the denominator because there is no ball in play at all. Walks and hit batsmen never appear, because they are not at-bats. What is left is contact that entered the field of play, which is precisely the set of events that fielders, positioning, the ballpark and the bounce of the ball can influence.
That is the clean version. The denominator has three quirks that almost nobody mentions, and each one is a small lie the statistic tells.
Reached on error counts as an out. A batter who scorches a ball at the shortstop and watches it go through his legs is charged an at-bat and credited with nothing. In the arithmetic he has failed. In reality he did the hard part correctly and the fielder did not. Every hitter's BABIP is therefore slightly lower than his contact deserved, and hitters who face bad defences are penalised twice: once by the official scorer, once by the formula.
A caught foul ball counts as a ball in play. A foul pop-up to the catcher is an at-bat, it is not a strikeout, and it is not a home run, so it sits in the denominator. It can never appear in the numerator, because a foul ball cannot be a hit. It is a guaranteed zero. Hitters who foul a lot of pitches into the air are carrying dead weight in their denominator that has nothing to do with anything the statistic claims to measure.
Sacrifice flies are added back but sacrifice bunts are not. A sacrifice fly is not an at-bat under the scoring rules, so it has to be put back in by hand or it would vanish from a calculation about balls in play, which is obviously what it is. A sacrifice bunt is also a ball in play and is left out entirely, on the reasoning that it was never an attempt to get a hit. That is a defensible convention rather than a truth, and it means the statistic quietly holds a view about intent.
None of these move the number much. All of them are worth knowing, because the people who use BABIP most confidently are usually the ones who have never asked what is in the denominator.
- 300League average BABIP, in points
- 800Balls in play before a hitter's number means much
- 2000Balls in play before a pitcher's number does
- 0Home runs counted in the numerator
The league figure is the long-run central value rather than any particular season's. The stabilisation thresholds are the published rules of thumb for the point at which a sample stops being mostly noise.
How small a slice of the season it governs
Take an invented season, built from round numbers so the arithmetic is checkable. Six hundred plate appearances. Sixty walks and hit by pitches. Five sacrifice flies. That leaves 535 at-bats, of which 130 end in a strikeout and 25 in a home run.
- Balls in play from at-bats380
- Strikeouts130
- Walks and hit by pitch60
- Home runs25
- Sacrifice flies5
Constructed example with deliberately round components. The BABIP denominator is the first slice plus the sacrifice flies, 385 events in total. The other 215 plate appearances are settled without a fielder touching the ball.
Show the numbers
| Item | Value |
|---|---|
| Balls in play from at-bats | 380 |
| Strikeouts | 130 |
| Walks and hit by pitch | 60 |
| Home runs | 25 |
| Sacrifice flies | 5 |
Three hundred and eighty-five of six hundred. Just under two thirds of the season passes through the part of the game BABIP describes, and the remaining third is decided entirely between the pitcher and the batter with eight fielders standing around as spectators.
That ratio is the single most important fact about the modern game and the reason BABIP has become less powerful than it was. Every strikeout removed from that block is an event with no variance in it. A hitter with a very high strikeout rate has shrunk the part of his season that BABIP governs, which makes his BABIP swing more wildly from year to year while mattering less to his overall line.
What 40 points of BABIP is worth
Keep the same invented hitter and change nothing except the rate at which those 385 balls in play become hits.
Eighty points of BABIP moves the batting average by fifty-eight. The hitter at the bottom of that chart is dropped from the lineup. The hitter at the top is in the conversation for an extension. They have identical strikeout rates, identical walk rates and identical power, and they hit the ball exactly the same number of times.
This is why the statistic exists and why it commands attention. Nothing else in a batting line moves that far on inputs the player does not obviously control.
Everything that decides whether a ball in play becomes a hit
The word luck is doing an enormous amount of concealed work in most discussions of BABIP. Pull the concept apart and there are at least seven distinguishable things in it, most of which are somebody's skill.
- Contact qualityHow hard and at what angle the ball leaves the bat. This is the hitter's skill and the pitcher's, competing, and it is the largest single input. A ball struck at high speed in the range of angles that produce line drives is a hit most of the time regardless of who is fielding.
- Batted ball typeLine drives fall in far more often than ground balls, and ground balls far more often than fly balls. A hitter's mix of the three is stable enough across seasons to be treated as a characteristic rather than an accident.
- Where the fielders were standingPositioning is a decision made before the pitch by a coaching staff working from data. It is not the hitter's doing and it is not chance. It moves BABIP more than almost anything else on this list.
- Who the fielders wereRange, first step, arm, hands. A ball hit into the gap is a double against one centre fielder and an out against another, and this is the input the pitcher is most often being credited or blamed for.
- The ballparkFoul territory, outfield dimensions, wall heights, the surface, the altitude, the wind. A ball that is caught in one park is off the wall in another and gone in a third.
- The runnerA hitter with genuine speed converts a proportion of infield ground balls into hits that a slow hitter never will. It is a small effect on the total, but it is a real and repeatable one.
- The bounceThe ball off the lip of the infield grass, the wind that holds up a fly, the fielder who slips. This part is genuinely random and it is much smaller than its reputation.
- The official scorerHit or error is a human judgement made in a press box, and it decides which half of the fraction the event lands in. Over a season it is noise. On any given ball it is a person's opinion.
Each link is a real influence on whether a ball in play becomes a hit. Only the last two are random in any useful sense, which is why calling the whole statistic a luck number is a description of the residual rather than of the mechanism.
The residual after all of those have been accounted for is the luck. It is real and it is meaningful over a few hundred balls in play. It is nowhere near the whole number.
Where the idea came from, and why it caused a fight
The insight belongs to Voros McCracken, who worked it out independently of professional baseball and published it in the late 1990s on a Usenet baseball group before setting it out at length for Baseball Prospectus at the start of the following decade.
His claim was narrow and it sounded absurd. Pitchers differ enormously in how often they strike batters out, walk them and give up home runs. They differ far less than anyone expected in how often the balls hit off them become hits. Run the correlations from one season to the next and the strikeout rate holds up strongly, the walk rate holds up, the home run rate holds up reasonably, and the hits-per-ball-in-play rate barely holds up at all.
The reaction was hostile, and understandably so. It reads as a claim that pitching does not matter, which it is not. What it says is that the pitcher's control over the outcome is concentrated in whether contact happens rather than in what the contact does afterwards, and that the second part is dominated by fielders, parks and variance in a way that a single season cannot separate.
Out of it came the whole family of defence-independent pitching statistics, of which fielding independent pitching is the one most people now see on a broadcast graphic. Every one of them is built on the same move: throw away the balls in play, evaluate the pitcher on the three outcomes he demonstrably repeats, and treat the rest as somebody else's problem.
BABIP is the leftover. It is what you are choosing to ignore when you use FIP, which is why the two are always read together. A pitcher whose earned run average is much worse than his FIP normally has an inflated BABIP behind it, and the useful question is not whether that is unlucky but which of the eight links in the chain above is producing it.
Why a pitcher's BABIP regresses and a hitter's does not
Both regress. They regress toward different places, and that difference is the most common mistake in amateur analysis.
A pitcher's BABIP regresses toward the league central value, because the evidence that pitchers have large persistent differences in it is weak. A pitcher sitting well below .300 in July will, in the absence of an identifiable cause, drift back toward .300, and the earned run average will follow him up.
A hitter's BABIP regresses toward his own career figure, which may be forty points above or below the league's. Contact quality is a repeatable skill, speed is a repeatable skill, and the batted-ball mix is stable. A hitter with a long record of high BABIP who is running a low one is not reverting to .300, he is reverting to his own number, and treating the league average as his destination will make you wrong in a predictable direction every time.
The practical rule follows directly. For a pitcher, ask what the league does. For a hitter, ask what he has always done. Confuse the two and you will spend every season declaring good contact hitters overrated and bad ones due for a rebound.
The exceptions that survived
McCracken's claim was strong and it has been sanded down at the edges ever since, mostly in ways that make it more useful rather than less.
Pop-ups. An infield fly is very close to an automatic out, and pitchers differ measurably and repeatably in how many they generate. A high pop-up rate suppresses BABIP for a structural reason, not a lucky one, because those balls are in the denominator and can almost never reach the numerator. The rule that governs what happens on those balls exists precisely because everyone accepts they will be caught.
Extreme batted-ball profiles. A pitcher who generates ground balls at a very high rate has a different BABIP baseline from one who lives in the air, because ground balls become hits more often than fly balls do. Neither is better or worse in run prevention terms, since the fly-ball pitcher is trading a lower BABIP for more home runs, but the baselines are genuinely different and comparing the two against a single league average is a category error.
Knuckleballers. The pitch produces contact that behaves unlike the rest of the sport, and the handful of pitchers who threw it seriously tended to sit persistently away from the league figure rather than oscillating around it. A tiny population, but a real counterexample.
Contact suppression. The one McCracken could not have seen, because the data did not exist. Some pitchers demonstrably induce weaker contact than others, at lower speeds and less useful angles, and weaker contact becomes hits less often. The effect is smaller than the season-to-season swing in BABIP, which is why it stayed hidden for so long, but it is not zero and it is now measurable directly rather than inferred from outcomes.
What batted-ball tracking did to the argument
The whole BABIP framework was a workaround. It existed because nobody could see how hard a ball was hit, so the fate of the ball had to stand in for the quality of the contact, and the difference between the two had to be called luck.
Now every batted ball is tracked. Its speed off the bat and its vertical angle are recorded, and from those the league publishes an expected batting average: the rate at which balls hit at that speed and that angle have historically become hits, with the runner's own sprint speed folded in on the ones where it matters. It is a direct estimate of what the contact deserved.
That does not make BABIP obsolete. It makes it decomposable. A hitter whose BABIP has collapsed can now be sorted into one of two entirely different situations within about ten minutes.
If the exit velocities and launch angles are unchanged and the expected figures still look normal, the contact is fine and the outcomes are not. That is the case where the word luck is legitimate, and it is the case where nobody should be changing anything.
If the contact quality has fallen, the low BABIP is a symptom rather than a mystery. Something is wrong with the swing, or with the pitches he is choosing to swing at, and the BABIP is reporting it accurately.
Before tracking, those two players looked identical on a stat sheet. That is the real advance, and it is why BABIP has moved from being an answer to being a first question.
Positioning, defence and the shift
The largest non-random input on the list is where the fielders stand, and that input changed twice in a decade.
The first change was analytical. Clubs began positioning defenders according to where each hitter had actually hit the ball, rather than where the position names implied they should stand, and pull-heavy hitters found three infielders on one side of second base. Their ground balls stopped becoming hits. Their BABIPs fell, in some cases by forty or fifty points, and not one of them had got worse at hitting.
The second change was regulatory. The limits placed on infield positioning took away the extreme alignments and handed some of those hits back, which moved a large number of BABIPs upward at once for reasons that had nothing to do with the players.
Both of those episodes make the same point. A statistic that people habitually describe as luck was moved league-wide, twice, by decisions taken in offices. Anyone who had been treating BABIP as noise around a fixed mean during either period was reading a signal as static.
The same applies at the club level every season. A pitcher who moves from a strong defensive team to a poor one will see his BABIP rise, and he did nothing. This is why BABIP is one of the inputs that makes valuing pitchers hard, and why every attempt at a single all-in number, including the various versions of wins above replacement, has to make an explicit choice about how much of a ball in play to charge to the man who threw it.
The ballpark, which is not luck either
Foul territory alone can move the statistic. A park with a large amount of ground behind home plate and along the lines converts foul pops into outs that would land in the crowd elsewhere. Those outs go into the denominator and never come out.
Outfield dimensions do the rest. A deep gap turns catchable fly balls into doubles. A high wall turns home runs into singles, which sounds like a BABIP gain and is, because a ball off the wall is a ball in play and a ball over it is not. Altitude changes both the flight of the ball and the amount of ground the outfielders have to cover.
None of this is exotic and all of it is measured. Park factors exist to strip it out, and a hitter's BABIP that has not been park-adjusted is being compared against a league average that includes thirty different sets of physical conditions. The correction is not large for most players. For someone who has changed clubs between two extreme parks, it can account for the entire apparent change in his ability.
How long before a BABIP means anything
Roughly 800 balls in play for a hitter. Roughly 2,000 for a pitcher.
Those are the published thresholds for the point at which the signal in the sample overtakes the noise, and they are brutal. Eight hundred balls in play is about two full seasons of everyday work. Two thousand is closer to three seasons for a starting pitcher, which for many pitchers is most of a career.
The implication is uncomfortable and it should be. A BABIP quoted in May is very nearly meaningless as a statement about the player. A BABIP quoted in September is a statement about the season, which is a different and much weaker claim. The number people cite most often, one player's figure over one season, sits well below the threshold at which it stops being mostly sample.
Relievers are the extreme case and deserve their own warning. A pitcher throwing sixty or seventy innings a season will put a few hundred balls in play in a good year, so he will not reach the pitcher threshold until he has been healthy and effective for the better part of a decade. Practically every reliever whose earned run average has swung violently from one season to the next has a BABIP swing sitting underneath it, and practically none of those swings survive being looked at properly. A bullpen rebuilt on the strength of last year's run prevention is a bullpen bought at the top of the market.
The same trap catches platoon splits. Cut a season into what a hitter did against left-handers and right-handers and each half of the split is a sample far too small to support the conclusion people reach from it, yet those numbers get quoted in lineup arguments every week of the season.
This does not make in-season BABIP useless. It makes it useful in one specific direction: as a flag that a batting line is out of step with the process behind it, prompting you to go and look at the process. It is a question generator. Used as an answer it will mislead you roughly as often as it helps.
Regressing a half-season figure, with the arithmetic shown
The standard correction is one line of arithmetic and almost nobody does it, which is why so many confident midseason predictions are wrong by a predictable amount.
Weight the observed figure against the baseline in proportion to how much of the stabilisation threshold the sample has covered. Take an invented hitter with 200 balls in play so far, running a BABIP of .380, against a career baseline of .310 and a hitter threshold of 800.
The weight on what he has actually done is 200 divided by 200 plus 800, which is one fifth. The weight on the baseline is the other four fifths. So the estimate is one fifth of .380 plus four fifths of .310, which comes to .324.
Look at what that does. The raw number is seventy points above his career figure and looks like a breakout. The regressed estimate is fourteen points above it, which is a good but unremarkable stretch. The correction has thrown away four fifths of the apparent improvement, and it has done so purely on the basis of how little evidence there is, before anyone has watched a single swing.
Run the same sum in the other direction and it is even more useful, because a hitter running .240 over 200 balls in play against a .310 career figure comes out at .296. He is barely below average as an estimate, and he is being written about as though he has forgotten how to hit.
The threshold is not a magic boundary and the weighting is a rule of thumb rather than a theorem. What matters is the shape of it: early in a season, the honest estimate of a player's true rate is mostly his history and only slightly what you have just watched, and human beings find that almost impossible to do by instinct.
Why old BABIPs are not comparable to new ones
The record books credit Ty Cobb with the highest career BABIP of any player with a substantial number of plate appearances, at .383, and with the highest single-season figure, .444 in 1911. Both are reconstructions, computed backwards from box scores long after he retired, since nobody was calculating this in his lifetime.
They are also not achievable now, and the reasons have very little to do with Cobb.
Gloves were smaller and stiffer, so the fielding range that turns a hard ground ball into an out simply did not exist. Infields were maintained to a standard that would be unacceptable in a modern minor league. Outfields in several parks ran to distances that made a ball in the gap a certain triple rather than a probable catch. Fielders positioned themselves by convention and instinct rather than from a spray chart. And hitters were not trying to elevate the ball, because a ball in the air in that era was mostly a wasted swing.
Every one of those has moved in the direction of suppressing hits on contact. The counterweight is that pitchers now throw much harder, which produces weaker contact but also produces more strikeouts, and strikeouts leave the BABIP denominator entirely rather than joining it as outs.
The net effect is that a BABIP from a hundred years ago and a BABIP from last season are two different measurements wearing the same name. Comparing them tells you about groundskeeping, glove leather and coaching philosophy, and almost nothing about the two hitters. This is a general hazard with any rate statistic that depends on fielders, and it is the reason park and era adjustments exist at all.
Team BABIP, arbitration, and the money side
At the club level the statistic changes meaning. A team's BABIP allowed is not a pitching number in any useful sense, because it is aggregated over an entire staff and a single defence in a single park. What it mostly measures is the defence, which is exactly why front offices watch it and why it turns up in the case for signing a defensive specialist nobody else wants.
The individual player side is where it gets expensive. A hitter who has a career year on the back of a BABIP fifty points above his own norm has produced real runs, won real games, and will be paid for it. His agent will present the batting average. The club across the table will present the contact quality and the regressed estimate, and will argue that it is buying a player who is about to hand most of that back.
Both sides are being honest. The runs happened. The rate that produced them is unlikely to repeat. That gap between what a player did and what he is likely to do next is the entire negotiation, and BABIP is one of the sharpest instruments for measuring it, which is why it appears in salary arbitration hearings where the panel is being asked to price one season of performance.
The asymmetry is worth noticing. A player whose BABIP collapsed can point to the contact data and argue that he was unlucky, and if the tracking backs him, clubs will believe him. A player whose BABIP soared has the harder job, because the same data will usually show that his contact did not improve nearly as much as his results did. The statistic is more forgiving on the way down than on the way up, and free agency is full of people who found that out at the worst possible moment.
Reading a BABIP that looks wrong
- Check the size of the sample firstBelow a few hundred balls in play, stop. Nothing further in this list will tell you anything the sample size has not already ruled out, and most published BABIP takes fail here without ever getting to a second step.
- Compare against the right baselineThe hitter's own career figure, not the league's. For a pitcher, the league's, adjusted for how many ground balls and pop-ups he generates. Using the wrong reference point produces confident conclusions in the wrong direction.
- Look at the contact quality directlyIf exit velocities and launch angles are where they always were, the contact is intact. If they have moved, the BABIP is a consequence rather than a cause and the investigation is now about the swing.
- Ask what the defence behind him isFielding quality and positioning are the largest controllable input. A pitcher who changed clubs, or whose club changed its infield, has an explanation available before anyone reaches for chance.
- Ask what the park isEspecially for a player who has moved, and especially in parks with unusual foul ground or outfield dimensions. Park effects on batted balls are measured and published; there is no need to guess.
- Check the batted-ball mixA hitter who has started putting more balls in the air has taken on a lower BABIP by choice, usually in exchange for power. That is a trade, not a decline, and the batting average will look worse while the overall production does not.
- Only now call it varianceWhat is left after those six checks is the genuinely random part. It is real, and it does correct, but it is a much smaller share of the deviation than the popular framing of the statistic implies.
A sequence of checks rather than a rule. Each step either identifies a cause or hands the question to the next one, and only a figure that survives every step deserves to be called variance.
What to do with the number
BABIP is not a rating. It is a residual, and a residual is only as useful as your list of things to subtract before you look at it.
Three habits get most of the value out of it. Compare every hitter to himself rather than to the league, because contact quality is a skill and career figures encode it. Compare every pitcher to the league, but adjust the expectation for how many balls he puts on the ground and how many he pops up, because those two rates set his baseline and both are repeatable. And treat any BABIP quoted over less than a season as evidence that something is worth looking into, never as evidence that anything has been established.
The hitter batting .238 in June, with an unchanged strikeout rate and unchanged contact quality, is almost certainly fine. Watch the rest of the sport long enough and you will see that player have a very ordinary second half and a very ordinary career, and you will see the people who wrote him off in June explain that he turned it around. He did not turn anything around. The balls started landing where they had always been going, and nobody noticed the difference because a batting average does not show its working.
BABIP does. That is the entire reason to use it.
Common questions
What does BABIP stand for in baseball?
Batting average on balls in play. It is the rate at which balls hit into the field of play become hits, calculated as hits minus home runs, divided by at-bats minus strikeouts minus home runs plus sacrifice flies. Home runs and strikeouts are removed because no fielder has any say in either one.
What is a good BABIP?
For a hitter it depends entirely on the hitter, because BABIP is partly a skill. The league central value sits close to .300, elite contact hitters live well above it and slow hitters who put the ball in the air live below it, so a player's own career figure is a far better reference point than the league's.
Is BABIP just luck?
No, though it contains luck. It is a residual: everything that decided the fate of a ball in play once you have removed the pitcher's strikeouts and home runs allowed. That residual holds contact quality, defensive skill and positioning, the ballpark, the runner's speed and the official scorer's judgement, and only the remainder after all of those is random.
Why do pitchers with a low BABIP get worse?
Because most pitchers have far less influence over the fate of a ball in play than they do over whether it is put in play at all, so a low figure is usually being produced by fielders, a park, or a run of favourable bounces rather than by the pitcher. Take away the cause and the effect goes with it. Some pitchers do suppress hits on contact, but the effect is small compared with the year-to-year swing.
How many balls in play does it take before BABIP means anything?
Roughly 800 for a hitter, which is about two full seasons, and roughly 2,000 for a pitcher, which is nearer three. Below those thresholds the number is telling you more about the sample than about the player, which is why a two-month BABIP is one of the least useful figures in baseball.
Filed under Baseball·mlb · statistics · sabermetrics · pitching · hitting