Skip to content
CricketTaken

Analysis

Football set piece xG: why restarts need their own model

Why set piece xG models are trained apart from open play, what a corner is worth in expected value, and how second-phase shots are attributed.

By CricketTaken EditorialPublished Analysis18 min read

How this is written and checkedReport an error

An open-play expected goals model looks at a header struck from eight metres, dead central, and returns a healthy number. It has seen thousands of headers from eight metres and it knows roughly how often they go in. What it has not seen is that this particular header was taken with a defender's shoulder across the attacker's chest, a second defender on the line, a goalkeeper set on his six-yard line rather than scrambling, and eleven other bodies inside the width of the penalty area. Football set piece xG exists because that gap between what the model sees and what actually happened is large, systematic, and always points the same way.

The direct answer to the question people ask first is that set piece models are trained separately because the relationship between a shot's features and its chance of going in is genuinely different at a restart. It is not a small correction applied to an open-play number. The same distance, the same angle and the same body part produce a materially different conversion rate depending on whether the ball arrived from a corner or from a through ball, and a single model asked to cover both will price each one wrong in opposite directions.

What an open-play model is actually looking at

To see why the transfer fails, it helps to be precise about what a general shot-quality model takes as input. Almost all of them use some combination of the same things.

Distance from goal, and the angle subtended by the two posts from the shooting position, which together describe how much goal there is to aim at. The body part used, since a header converts at a lower rate than a foot from the same spot. The type of pass that created the chance, because a cut-back and a through ball produce different finishing conditions. Whether the shot came from a fast break. And, in models built on positional feeds, the number of defenders in the shooting lane and the goalkeeper's distance from his line.

Every one of those features is measured in a world where the defence has been pulled out of shape by the ball moving. That world is the assumption baked into the training data, and a restart violates it in five ways at once.

The defence is not out of shape. It has had thirty seconds to arrange itself exactly as it wanted. The goalkeeper has chosen his position rather than recovered into one. The number of bodies in the box is roughly double what an open-play shot faces. The ball is arriving with deliberate pace and curl from a stationary strike rather than from a pass weighted for a team-mate to control. And the attacker is very often making contact while being physically held, blocked or leaned on.

The fixed dimensions the laws impose on a restart
  • 1Radius of the corner arc, metres
  • 9.15Distance opponents must retreat at a free kick, metres
  • 1Minimum distance from a defensive wall for attackers, metres
  • 11Distance from the penalty mark to the goal line, metres

Dimensions and distances from the current Laws of the Game. Structural facts, not conversion figures.

The five ways a restart breaks the assumptions

Take those violations one at a time, because each one moves the true conversion rate in a specific direction and the net effect is not obvious.

Crowding cuts both ways. More bodies in the box means more blocks, more deflections and less clear sight of goal, all of which lower conversion. It also means more chaos, more unpredictable deflections and a goalkeeper whose view is obstructed, which raise it. An open-play model that has learned "defenders in the lane means a worse chance" applies only the first half of that.

The goalkeeper is optimally placed and badly sighted. At a corner he stands where he wants, which is a large advantage. He also has to make a decision about whether to come for the ball through a wall of players, and if he commits and misses, the shot that follows is taken into an empty net. This bimodal outcome has no analogue in open play.

Contact at the moment of the shot is normal rather than exceptional. An attacker heading a corner is very often being held, and a foul that is not given still degrades the contact. Open-play data contains far fewer shots taken while physically impeded, so the model has no way to price it.

The delivery is a different object from a pass. A corner arrives at a speed and with a spin chosen to make it difficult to defend, not to make it easy to finish. Contact quality is worse for the same distance, which pushes conversion down.

The defence is set, which removes the fast-break advantage entirely. Any model feature that rewards a chance for being unstructured is inverted at a restart.

Fit one model across both worlds and it will learn the average of two different physical processes, which is a rate that describes neither.

The problem that makes a separate model hard to train

If separating the models is obvious, the reason it took a while is equally simple. There are not many set piece shots.

A general shot model has an enormous training set, because ordinary attacking play generates shots continuously across every match in every competition a supplier covers. Restart shots are a fraction of that, and once you subdivide by restart type, by delivery type and by phase, each cell of the table gets thin very quickly. Direct free kicks from a particular band of distance and angle are rarer still.

Thin cells produce unstable models. A model fitted to a small sample will latch onto accidents in the data and report them as structure, which is how you end up with a set piece model that believes back-post headers from one specific zone are unusually valuable because eleven of the fourteen it saw went in.

The practical responses are all compromises. Pool across seasons, which risks blending eras with different laws and different defensive conventions. Pool across competitions, which risks blending standards. Reduce the number of features so that each surviving one has enough data behind it, which loses detail. Or borrow structure from the open-play model and fit only the difference, which is more stable and reintroduces some of the contamination the separation was meant to remove.

Every provider makes those choices differently, and that is a large part of why two set piece numbers for the same match disagree. It is the same phenomenon that produces differing open-play values, described in why two suppliers give different expected goals for one shot, amplified by a smaller sample.

Three different numbers all called set piece xG

The term is used for at least three distinct quantities, and confusing them is the most common error in public discussion.

Shot-level set piece xG. The value of one shot, produced by a model trained on restart shots. This is the direct analogue of the ordinary figure and is what a shot map is showing when it flags a chance as a set piece.

Sequence-level, or expected value per restart. The total expected goals generated by everything that followed one corner, including any second and third attempt from the same loose ball, divided across the restarts taken. This is a different question. It asks what a corner is worth rather than what a shot is worth, and it is the number a set piece coach is actually trying to move.

Routine-level value. The expected goals attributable to a specific rehearsed pattern, aggregated across every time a team ran it. This is almost entirely an internal club figure, because it requires knowing which routine was called, which is information no data supplier has.

A claim that a side is generating a certain amount from set pieces is unreadable without knowing which of the three is meant, and whether penalties are inside or outside the total.

What a corner is actually worth

The honest headline is that a single corner is worth very little, and the reason is that most corners produce nothing at all.

Think about the funnel rather than the goal. A corner is taken. It may be cleared before it reaches an attacker, claimed by the goalkeeper, or headed away by the first defender. Only some deliveries produce an attacking contact, only some contacts produce an attempt on goal, and only some attempts go in. Each stage removes the large majority of what entered it, and the compounding is brutal.

That arithmetic has three consequences that supporters consistently get wrong.

Corner count is not a measure of dominance. A side can take a dozen corners and generate almost nothing, because the deliveries were poor or the defence was well organised, and both of those are common. The corner count measures how often the ball went out off a defender, which is only loosely related to attacking quality.

Set piece value accumulates on a timescale of a season rather than a match. A programme that improves a team's expected value per corner by a modest amount will produce a difference that is invisible in any single game and clearly visible across a campaign. Judging the work by last weekend is judging a slow signal with a fast instrument.

And the variance is enormous. Because goals from restarts are rare events attached to a small number of high-leverage moments, a team can under-perform or over-perform its set piece expected goals by a wide margin for months without anything being wrong with the process.

Constructed example: expected value per corner accumulating across a season
  • Lower value routine
  • Higher value routine
After 10 matches2xG3xG
After 20 matches4xG6xG
After 30 matches6xG9xG
After 38 matches8xG11xG

Invented arithmetic on stated assumptions, not a measurement of any team. A hypothetical side takes six corners per match across a thirty-eight match season, and the two lines show the same volume at a lower and a higher expected value per corner. The gap, not the values, is the point.

Show the numbers
Constructed example: expected value per corner accumulating across a season
ItemLower value routineHigher value routine
After 10 matches2xG3xG
After 20 matches4xG6xG
After 30 matches6xG9xG
After 38 matches8xG11xG

First contact, second phase and where the boundary gets drawn

The most consequential modelling decision in this whole subject is one nobody outside the industry discusses: when does a set piece stop being a set piece.

A corner is delivered. A defender heads it away, but only as far as the penalty spot, where an attacker volleys it goalwards. Is that a set piece shot or an open-play shot? Both answers are defensible. It came directly from the restart and would not exist without it, which argues for set piece. The defence has made a clearance and the ball is loose in a scramble, which is a different situation from a header off a delivery, and that argues for open play.

Different suppliers draw the line in different places, and some use a time limit, some a touch count, and some a rule about whether possession was genuinely regained. The choice materially changes every set piece total that gets published.

How a corner becomes a set piece expected goals figure
  1. Classify the restartThe event feed records that a corner was taken, along with the side of the pitch, the taker, and whether the delivery swung towards or away from goal. Everything downstream depends on this label being right.
  2. Attach the positional snapshotIf a tracking feed is available, the positions of all twenty-two players at the moment of the strike are joined to the event. Without it, the model is blind to how many bodies were in the six-yard box, which is the single most informative feature available.
  3. Identify the first contactThe first player to touch the ball after the delivery is recorded, along with whether it was an attacker, a defender or the goalkeeper. A goalkeeper claim ends the sequence. The other two outcomes leave the ball live.
  4. Decide where the phase endsThe supplier applies its rule for how long a sequence stays attributable to the restart. A fixed number of seconds, a fixed number of touches, or a judgement about whether the defence properly regained control. This boundary is not standardised across the industry.
  5. Price every shot inside the phaseEach attempt within the boundary is passed to the restart-trained model rather than the general one, using its own distance, angle, body part and crowding features. Attempts outside the boundary go to the open-play model instead.
  6. Aggregate to the level being reportedShot values are summed to the sequence to give a value for that corner, or across all corners to give a value per restart, or across a season to give a team total. The three answers are different numbers and are frequently quoted interchangeably.

The generic modelling pipeline. Where the phase boundary sits in step four is a supplier decision, and it is the step that most changes the published total.

The reason second-phase shots deserve their own attention is that the situation they occur in is not like anything else in football. The defence has just jumped and landed, which means nobody is set and nobody is facing the right way. The ball is loose in the most dangerous area on the pitch. And the attacking team, if it has prepared properly, has players positioned specifically for this moment while the defenders are reacting to it.

Model that situation with an open-play shot model and it will price it as a chance from the edge of the box against an organised defence, which is close to the opposite of the truth. Exclude it from the set piece total entirely and you have removed a substantial share of the value a restart programme creates, which makes the whole exercise look less worthwhile than it is. Neither error is harmless, and both are common.

Direct free kicks need a third model again

A free kick struck at goal from twenty-five metres has nothing in common with either open play or a corner, and any provider serious about restarts fits it separately.

The situation is unique in three ways. The ball is stationary and struck with a technique that exists nowhere else in the game. There is a wall of bodies deliberately positioned in the flight path, and the laws govern its formation precisely, requiring opponents to retreat a set distance from the ball and keeping attacking players at least a metre from a wall of three or more. And the goalkeeper positions himself in advance for a shot he knows is coming, which is an advantage he never has in open play.

The features that matter are also different. Distance still matters, but the angle matters in a peculiar way, because a central position gives the best sight of goal and also permits the largest wall. Very wide free kicks are usually crossed instead, which means the shots that exist in the data are a filtered sample rather than a representative one, and any model fitted to them has to account for the selection.

The sample problem is at its worst here. Direct attempts from any specific band of distance and angle are uncommon even across several seasons of a large competition, and a taker's individual quality varies far more than it does for a header in a crowd. This is one of the few places in football modelling where individual identity genuinely belongs in the model, and where there is rarely enough data on any one player to include it responsibly.

Why penalties sit outside every set piece number

Penalties are technically a restart and are excluded from set piece totals by almost every serious framework, for a reason worth stating.

A penalty is a fixed situation. The ball is placed at the same distance every time, the goalkeeper must stay on his line until it is struck, and no other player may enter the area before contact. It has close to a single value, and that value is high relative to any other shot in football. Include penalties in a set piece total and the total becomes dominated by how many penalties a team was awarded, which is a measure of something else entirely.

Excluding them is also what makes the remaining number useful. Set piece value that excludes penalties is a measure of how well a side attacks corners, free kicks and throws, which is a coachable property. Penalty count is a measure of how often opponents foul inside their own area, which is largely a consequence of how often you get the ball into it.

Throws, indirect free kicks and the taxonomy that decides the sample

Underneath every published set piece figure sits a taxonomy, and the taxonomy determines what is being counted.

The categories that usually exist are corners, direct free kicks, indirect and crossed free kicks, throw-ins, and kick-offs. Providers disagree about whether all of them count. Long throws into the penalty area are a delivery into a crowded box with the same physical character as a corner, and the fact that no offside offence can be committed directly from a throw-in makes them tactically distinct in a way that matters, as set out in the offside law and the restarts it does not cover. Some frameworks include them, some do not.

Crossed free kicks from wide areas are the largest grey zone. A free kick played into the box from the flank is functionally a corner taken from a different spot, and treating it as open play because it was not struck at goal loses a genuine category of chances. Kick-offs are almost always excluded and almost never matter.

The rule for reading any published figure follows from this. Before comparing two teams' set piece numbers, establish which restart types are inside each total and where the phase boundary sits. If those differ, the comparison is not measuring the teams.

Constructed example: where one season's restart shots come from
39%27%19%
  • Corners, first contact44shots
  • Corners, after the first clearance31shots
  • Free kicks delivered into the box22shots
  • Direct free kicks struck at goal9shots
  • Throw-ins into the penalty area7shots

Invented composition on stated assumptions, chosen to show why the taxonomy decides the total. Not a measurement of any competition. Penalties are excluded, as they are from most serious frameworks.

Show the numbers
Constructed example: where one season's restart shots come from
ItemValue
Corners, first contact44shots
Corners, after the first clearance31shots
Free kicks delivered into the box22shots
Direct free kicks struck at goal9shots
Throw-ins into the penalty area7shots

The defensive side, which is the half everybody forgets

Every attacking set piece number has a defensive twin, and the defensive twin is easier to improve.

Set piece expected goals conceded measures the quality of chances a side gives up from opposition restarts. It behaves better than the attacking version in one important respect: a defence faces restarts of every type from every opponent across a season, which produces a larger and more varied sample than any single team's attacking routines generate.

It also isolates a specific failure that goals conceded cannot. A side can concede very few set piece goals while giving up excellent chances from every corner, and the difference will show up as a gap between goals conceded and expected goals conceded from restarts. That gap is a warning, because the goalkeeper cannot keep saving them.

The defensive number is also where the marking system argument becomes empirical rather than rhetorical. The long dispute between man-oriented and zonal approaches, laid out in the comparison of marking systems at set pieces, is not settled by counting goals conceded, because the sample is far too small and every goal conceded zonally is remembered while every goal conceded to a lost man-marker is blamed on the individual. Chance quality conceded across a season is a fairer instrument, and it is the one clubs actually use.

The measurement came first, then the specialist

The rise of the dedicated set piece coach is often described as a tactical trend. It is better understood as a consequence of the numbers becoming good enough to justify a salary.

The chain runs like this. Once a club can put an expected value on a restart, it can measure whether a change to its routines moved that value. Once it can measure the movement, it can compare the cost of achieving it through coaching against the cost of achieving the same improvement in goal difference through the transfer market. That comparison is not close, because coaching a restart requires training time and attention while the equivalent improvement in open play requires signing players, and the second is expensive and contested by every rival.

The measurement also changed what the specialist is evaluated on. Goals from restarts are far too rare to grade a coach on within a season, so the process measures took over: delivery accuracy into the intended zone, whether the intended attacker was actually free at the moment of the strike regardless of the outcome, share of first contacts won at both ends, share of loose balls recovered, and chance quality created and conceded. Those signals accumulate quickly enough to be acted on, and they are the ones the coach can genuinely influence. The broader tactical craft that sits on top of them is covered in how set piece coaching actually works in practice.

There is a limit that the enthusiasm tends to skip past. Restart value is a marginal gain, and marginal gains decide matches only when everything else is roughly level. No corner routine rescues a side that is being outplayed for ninety minutes, and treating set piece work as a substitute for a functioning team is a misreading of what the numbers say.

Delivery type is a feature, and it is the one with the worst data

If you were building a restart model from scratch, the delivery would be near the top of your list of inputs. It is also the input the available data describes worst, and the gap between the two is a good illustration of why these models remain cruder than their users assume.

The direction of swing changes the physics of the contest. A ball curling towards the goal arrives on a trajectory that pulls the goalkeeper and the defenders backwards, and any glancing touch is heading in the right direction. A ball curling away from goal arrives moving away from the keeper, which keeps him out of the contest, and requires the attacker to generate all of the pace himself. Those are two different chances even when the contact point is identical, and a model that treats them as the same delivery from the same spot is throwing away real information.

Height and pace change it again. A driven, flat delivery gives the defence less time to organise and gives the attacker less margin for error. A floated ball does the opposite. There is a genuine trade in there and no obviously correct answer, which is exactly the kind of question a well-specified model should be able to settle.

The problem is that public event feeds describe the delivery in coarse categories at best, and often only record that a cross was attempted from the corner arc. Swing direction can sometimes be inferred from the taker's foot and the side of the pitch, which is a reasonable heuristic that fails whenever a player takes a corner on his unnatural side deliberately. Ball flight height and speed require a positional feed, and the feeds that carry ball trajectory reliably through a crowded penalty area are the scarcest data in the sport, for the reasons set out in how ball and player positions are actually captured.

So the feature that a coach would rank first is the feature the model knows least about. That has two effects. Published set piece numbers are systematically less sensitive to delivery quality than they should be, which flattens the difference between a side with an excellent taker and one without. And clubs with access to trajectory data hold a genuine edge here, because they can price a delivery while everybody outside is pricing a location.

The short corner deserves a brief note in the same vein. Because it produces no immediate delivery into the box, a naive model sees a corner that generated nothing and records a zero. What actually happened is that the attacking side converted two players into a numerical advantage in a wide area and moved the point of delivery, which changes every subsequent geometry. Sequence-level valuation handles this correctly, because it follows the possession until it ends. Shot-level counting does not, and it is the reason short corners look worse in public data than they are in club data.

What set piece xG still cannot see

Four blind spots survive even in the best implementations.

The routine is invisible. A supplier's model knows a corner was in-swinging to the near post. It does not know that the near-post run was a decoy and the intended target was the player arriving late at the far post, because nobody outside the club knows what the plan was. Every public set piece number grades execution against an unknown intention.

Blocking and holding are barely captured. Legal screening is central to attacking corners and consists mostly of players standing still in useful places, which is exactly the sort of non-event no feed records. The most skilful thing an attacker does at a corner may leave no trace at all.

Individual aerial ability is under-modelled. A delivery into a zone occupied by a dominant header of the ball is worth more than the same delivery into a zone occupied by an average one, and most public models treat the zone rather than the occupant.

Fatigue and match state are absent. A corner in the ninetieth minute with a side chasing an equaliser is a different situation from the same corner at nil-nil in the twentieth, because the defending team is deeper and the attacking team has sent its goalkeeper forward. The models rarely price that, and the coaches certainly do.

Reading a set piece xG claim without being had

Five checks handle nearly everything.

Which of the three numbers is it? A shot value, a value per restart, or a season total. They are not interchangeable and the same phrase covers all three.

Are penalties in it? If nobody says, assume the total is unreadable and ask.

Where is the phase boundary? A figure that stops at first contact and one that runs until the defence regains control are measuring different things, and the second will always be larger.

Is the sample big enough to say anything? A team's set piece figures over ten matches are almost pure noise. Over a season they begin to mean something, and they are still the noisiest section of any team profile.

Attacking or defending? The defensive number is more stable, more actionable and much less quoted, which makes it the more useful one to look up. It is also the one a supporter can sanity-check by watching, since a defence giving up free headers is visible long before the goals arrive.

Get those five right and the restart numbers become one of the more honest parts of football analytics, because the modelling problems are visible rather than hidden. More on tactics, the laws and the economics behind them sits in the football archive, and the full set of sports explainers is in the blog index.

Common questions

What is set piece xG?

It is an expected goals value produced by a model trained only on shots that follow a dead-ball restart, rather than by the general model used for open play. Because the situations differ so completely in crowding, defensive positioning and how the ball arrives, an open-play model systematically mis-prices a set piece shot. Some providers go further and report an expected value for the restart itself, covering every shot the routine generates rather than only the first one.

Why can't the normal expected goals model be used for corners?

Because the features that drive an open-play model behave differently at a restart. Distance and angle still matter, but the number of bodies between the shot and the goal is far higher, the goalkeeper starts from a position he chose rather than one he was forced into, and the ball arrives with pace and spin from a delivery rather than from a pass. A model that has never seen those conditions will read a header in a crowded six-yard box as a straightforward chance.

How much is one corner worth?

Far less than most supporters assume, because the large majority of corners produce no shot at all. The value of a corner is only meaningful as an average across many of them, and it accumulates slowly, which is why a side can win a match having taken twelve corners and created nothing from any of them. Judging a set piece programme on a single match's corner count is a category error.

What is a second phase set piece shot?

It is a shot taken after the first contact from the delivery has been won, cleared or deflected, while the ball is still loose from the same restart. A defensive header cleared to the edge of the penalty area and struck first time is the standard case. Whether that shot is filed as a set piece or as open play is a decision the data supplier makes, and different suppliers draw the boundary in different places.

Are penalties included in set piece xG?

They are almost always excluded and reported separately. A penalty is a fixed, repeated situation with a single value that hardly varies, so including it would swamp the rest of the set piece total and tell you only how many penalties a team won. Any set piece figure quoted without saying whether penalties are in it should be treated as unreadable.

Did the analytics cause the rise of set piece coaches?

The measurement came first and the specialism followed. Once clubs could price a restart in expected goals and compare it against the cost of buying the same improvement in the transfer market, the case for a dedicated coach became an ordinary budgeting decision rather than a tactical fashion. The staffing change is downstream of somebody being able to put a number on it.

Filed under Football·football · expected goals · set pieces · analytics · corners · modelling