Analysis
The sports analytics department explained, data to decision
What a sports analytics department does: the five workstreams, the cleaning nobody sees, why the reporting line decides everything, and how to judge one.
By CricketTaken EditorialPublished Analysis21 min read
A model that is correct and ignored is worth precisely nothing. That single sentence explains more about how a sports analytics department succeeds or fails than any amount of discussion about which metric is best, and it is the sentence most job descriptions in the field manage to avoid.
The work of a sports analytics department, explained honestly, has two halves. The first is producing an answer that is true. The second is getting somebody with authority to do something differently because of it. The first half is a technical problem with known methods and a growing supply of people who can solve it. The second half is a political, organisational and rhetorical problem, and it is where most departments fail.
This matters more each year, because the technical half has become cheap. Data that was proprietary is now bought off a shelf. Methods that were novel are now a weekend project. What has not become cheap is the ability to change a decision inside an organisation where the person making it has spent thirty years learning to trust their eyes, will be judged on the outcome personally, and has been told before by a confident person with a spreadsheet who turned out to be wrong.
What the department is actually for, workstream by workstream
Strip away the tooling and a department serves five kinds of decision. They differ in speed, in reversibility and in how well data performs on them, and any assessment of a department has to be made workstream by workstream rather than in general.
Recruitment and squad planning. Who to sign, who to sell, what to pay, when to move, and what the squad needs to look like in three years. This is the largest single use of analytics resource at most well-run clubs and the one with the clearest evidence behind it.
Opposition analysis. What the next opponent does, under what triggers, and where they can be hurt. This is the oldest analytical function in sport and it long predates data: it was video, and before that it was a scout with a notebook.
In-game decision support. Substitution timing, the aggression thresholds that govern whether you take a risk or protect a position, and the pre-agreed triggers for changing shape. In some sports this is very well developed. American football's fourth-down decisions have been reduced to something close to a lookup table, and the arithmetic behind going for it is now argued about by supporters rather than only by coaches.
Workload, availability and return to play. Training load prescription, monitoring, and the attempt to keep players available. This is where the most data is collected per athlete and where the analytical claims should be made most carefully, because the honest position on individual injury prediction is far weaker than the volume of data suggests.
Set pieces. Corners, free kicks, throw-ins, and their equivalents in other sports. This is a small workstream with an outsized reputation, for good reasons covered below.
A sixth function exists in most organisations and is usually separate: the business side, covering ticketing, pricing, sponsorship valuation and broadcast. It uses similar people and completely different data, and conflating the two is one of the reliable signs of an organisation that has bought a job title rather than a capability.
Most of the work is cleaning, and pretending otherwise is how departments get under-resourced
Ask an analyst what they did last week and the honest answer is rarely modelling. It is reconciling.
A club consuming data at any serious level is taking in several separate feeds that do not agree with each other. Event data, which records what happened on the ball and is produced by a mixture of automated systems and human taggers. Tracking data, which records where all the players were several times a second, from either optical camera systems or wearable devices. Physical output from training, from a different vendor again. Medical records. Video. Scouting reports written in prose. Contract and transfer information. Wellness questionnaires the players fill in badly at eight in the morning.
Each of these arrives with its own identifiers, its own definitions and its own gaps, and none of them was designed to be joined to the others.
The identifier problem alone is a permanent tax. The same player appears under different names in different systems, with transliteration differences, with a middle name in one feed and not another, occasionally with two different dates of birth. A player who moves between leagues acquires a new identifier and loses his history unless somebody maps it. Building and maintaining a reliable master list of players, clubs, competitions and seasons is boring, unending, and the foundation everything else stands on. Departments that skip it produce analysis that is wrong in ways nobody notices for months.
Definitions are worse, because they fail silently. Two providers can both sell you "pressures" and mean different things. A duel, a key pass, a progressive carry, a high turnover: each of these is a human decision about where to draw a boundary, and each provider draws it somewhere slightly different. Change provider and every trend in your database develops a step change on the switchover date that has nothing to do with any player. Providers also revise their taxonomies, sometimes mid-season, and the revision is announced in release notes nobody read. Understanding exactly what you are buying is most of the value in knowing how the data providers differ.
Physical data has the same problem with a governing body attached. Football's world governing body runs a quality programme for electronic performance and tracking systems, testing certified devices against a reference motion-capture setup and publishing how accurate each one is across different speed bands. Read that carefully and the useful implication is the opposite of reassuring. Certification tells you a system's accuracy has been measured and published. It does not tell you that two certified systems will give you the same number for the same sprint. A club that changes tracking supplier over a summer will see its sprint counts move, and somebody in the building will interpret that as a change in the players.
- Ingesting, reconciling and fixing data14
- Routine reporting that somebody expects9
- Answering ad hoc questions from staff6
- Actual analysis of a live decision6
- Meetings, briefings and being in the room5
Invented hours, chosen to sum to 40, illustrating a common shape rather than describing any real department.
Show the numbers
| Item | Value |
|---|---|
| Ingesting, reconciling and fixing data | 14 |
| Routine reporting that somebody expects | 9 |
| Answering ad hoc questions from staff | 6 |
| Actual analysis of a live decision | 6 |
| Meetings, briefings and being in the room | 5 |
Those numbers are invented, but the shape is the argument. If the pipeline is unstaffed, the analysts maintain it, and the hours come out of the only column that changes anything. This is why a club that hires three analysts and no data engineer has, in practice, hired three data engineers who are unhappy about it.
- CollectionEvent data, tracking data, physical output from training, medical notes, video and scouting text arrive from several vendors on different schedules.
- Identity resolutionEvery player, club, competition and season is mapped to one internal identifier, so the same person is the same person across all feeds.
- Definition mappingEach provider's version of a metric is translated into the club's own definition, and the differences are documented rather than averaged away.
- Quality checksMissing matches, dropped tracking frames, impossible values and unit mismatches are caught automatically, because they will not be caught by eye.
- Context attachmentEvery observation is tagged with score state, opposition strength, minutes played, pitch and competition, because a number without context cannot be interpreted.
- Storage and accessOne place, one set of definitions, queryable by anybody in the department without asking the person who built it.
- AnalysisOnly now. Everything above is a precondition, and it is where the working week actually goes.
- DeliveryA recommendation in a form the recipient can act on, at the moment they are making the decision, not afterwards.
The pipeline stages every club department runs, whether or not it has named them.
Reporting is not decision support, and the difference is the whole job
A department can be busy, productive, well-liked and completely useless, and the mechanism is always the same. It produces reports.
Reporting answers the question "what happened". It is safe, it is visible, it fills a weekly slot, it makes the department look industrious, and it is almost never the thing anybody needed. A dashboard showing last match's numbers is consumed the way weather is consumed. People glance at it, feel informed, and do exactly what they were going to do.
Decision support answers a different question: "here is a choice you are about to make, here are the options, here is what we think each one is worth, and here is how sure we are". It is uncomfortable to produce because it commits the department to a position. It is uncomfortable to receive because it makes the recipient's judgement explicit. It is also the only version of the job that justifies the salaries.
The recent research literature has converged on this point with unusual clarity. Kwon's 2026 review of the analytics-practice gap identifies four causes, and the first is metrics that describe what happened without guiding what to do next. The others are missing context, so that the same number means different things in different situations; fragmentation, so that physical, tactical and medical information never get interpreted together; and practical constraints of time, expertise and communication. The proposed correction is to invert the sequence: start with the coaching problem, then choose metrics, then interpret in context, then communicate, then translate into a change in training or selection, then check whether behaviour actually changed.
That last step is the one nobody does. Evaluating whether a recommendation changed anything requires admitting that most of them did not.
The reporting trap has an organisational cause as well as an intellectual one. Reports are legible to people who cannot evaluate analysis. An executive who cannot judge whether a valuation model is any good can absolutely judge whether the weekly pack arrived on Monday. So the department is measured on the thing that can be measured, and it optimises accordingly, and within two years it is a publishing house.
The reporting line decides what the department is allowed to be
Where the department sits on the organisation chart determines its remit more completely than any strategy document, because a department ends up working on whatever its boss is accountable for.
Reporting to the head coach. The horizon becomes the next match. Opposition analysis is served well, in-game preparation is served well, and recruitment is served only where it overlaps with what the coach wants this season. The deeper problem is tenure. The department belongs to a person whose job is unusually insecure, and when he goes, the incoming staff arrive with their own analysts and their own opinions about the previous regime's work. Institutional memory resets, the models are abandoned, and the club starts again. Clubs in this configuration have a data department the way they have a fitness coach: it is part of the manager's kit.
Reporting to a sporting or technical director. The horizon becomes multi-season. Recruitment becomes the centre of gravity, squad planning becomes possible, and the department survives a coaching change because it does not belong to the coach. The cost is a permanent tension with the first team, because the analysts now work for the person who signs players the coach did not ask for. Most organisations that get sustained value from analytics are arranged this way, and most of the friction stories come from it too.
Reporting to the chief executive or owner. Independence, budget, and a fatal access problem. A department that has authority but is not physically present at the training ground gets treated as surveillance, and the football staff will route around it. Analysts in this position often find their reports being read carefully by people who cannot implement anything and ignored by everybody who can.
Reporting into performance or medical. The department becomes a sports science function, deep on load and availability, absent from recruitment and tactics. This is a coherent structure and it is honest about its scope, which is more than can be said for some of the alternatives.
There is no correct answer, but there is a diagnostic. Look at what the department is measured on and who can overrule it, and you will know what it will produce in eighteen months regardless of what it was hired to do.
The communication problem: one sentence against a distribution
Here is the shape of the problem, stated as plainly as possible.
A coach standing on the touchline with eleven minutes left needs one sentence. Make the change or do not. The model, if it is any good, does not produce a sentence. It produces a distribution over outcomes, conditional on assumptions, with uncertainty that is often wide enough to include both answers.
Somebody has to compress the second into the first. That compression destroys information by design, and the analyst has to choose what to destroy. Do it badly and you have either an unusable hedge or a false certainty. Do it well and you have a recommendation that carries its confidence in its phrasing rather than in a number.
Several practical things follow, and they are the difference between departments that get listened to and departments that get thanked.
Pre-agree the rule, do not argue in the moment. The single most effective format is a threshold decided calmly during the week: if this situation arises, we do this. In-game persuasion almost never works, because the decision window is shorter than the explanation, and because a coach in the eightieth minute is managing eleven people and a crowd. Sports where the decision points are discrete and repeated, from fourth downs to declarations, are exactly where pre-agreed thresholds have travelled furthest.
Give a recommendation, not a range, and put the uncertainty in the language. "I would do it, and it is close" is honest and usable. "The expected value is marginally positive under most parameterisations" is neither.
Never present a probability to somebody who will hear it as a promise. A seventy per cent chance is heard as yes by most listeners, and then remembered as a wrong prediction three times out of ten. Where a number has to be given, give the frequency in a form people understand and state the failure case out loud before it happens.
Use their units. Distances in the units the coaching staff already uses, positions in their names for positions, phases in their names for phases. A finding delivered in the vocabulary of statistics is asking the listener to do the translation, and they will not.
Timing beats content. An excellent insight delivered after the team meeting is worth nothing. Departments that matter know exactly when each decision is made and work backwards from it, which usually means the deadline is earlier and the format is shorter than the analyst would like.
The credibility economy, and why being right is not enough
Inside a club there is a currency, and it is not accuracy. It is standing.
Standing is earned slowly, in the coin the football staff recognise. Being useful about things they already care about. Being present at the training ground rather than in an office. Turning around a boring request quickly and without complaint. Being right about something checkable within a week, repeatedly, in front of people.
It is spent quickly. A confident recommendation that fails publicly costs several times what a correct one earns, and the asymmetry is not irrational. A conventional decision that fails is absorbed as bad luck, because everybody would have made it. A data-led decision that fails is evidence that the data was wrong, because there was an identifiable alternative and an identifiable person who argued against it. That asymmetry is a real feature of how blame is distributed in sport, and any analyst who ignores it will be surprised by the reception their best work gets.
The practical consequence is that a department has to build a balance before it spends one. Volunteering for the unglamorous work, the opposition set-piece clips and the availability reports, is not careerism. It is how you acquire the right to be heard on the recruitment question that actually matters.
The failure mode here is the single big call. An analyst who saves their credibility for one high-stakes recommendation, delivered as a confrontation, will lose. Not because they are wrong, but because the organisation has no history with them and no reason to accept the risk. The departments that win do so by changing many small decisions that nobody argues about, until the accumulated record makes the large recommendation unremarkable.
There is a second-order effect worth naming. A department that is never allowed to say no becomes a public relations function. If the analysis always supports what the club had already decided, either the club is uncannily good at deciding or somebody has learned which answers are welcome. The ability to deliver an unwelcome finding and survive is the clearest single indicator that a department is real.
Recruitment is the clearest win because the decision is slow and reversible
Of the five workstreams, recruitment is where analytics has the strongest case, and the reasons are structural rather than technological.
The decision is slow. Weeks or months pass between identifying a target and signing him, which is enough time to build an argument, be challenged, and revise. Nothing else the department does has that luxury.
The default is inaction. Not signing a player costs nothing directly, which means the analyst's recommendation to avoid somebody is cheap to accept. Recommendations that prevent a mistake are the most undervalued output in the field and the easiest to get adopted, because the organisation is being asked to do nothing.
The decision repeats. A club considers hundreds of players a year and signs a handful. Repetition is what makes probabilistic reasoning pay: a process that is right more often than the alternative will show up in the results, given enough decisions, and recruitment is the only part of the football operation with enough decisions.
And the candidate pool is far larger than any human can watch. A scouting department can properly assess some hundreds of players a season. A data pipeline can filter tens of thousands, across leagues nobody has the budget to travel to. This is the genuine comparative advantage, and it is a filtering advantage rather than a judging one.
Read that funnel and the division of labour becomes obvious. The first three stages are the analysts' work and they are almost entirely about exclusion. The last three are human judgement, and no serious department claims otherwise. The department's product is a shortlist that is smaller, wider-ranging and less biased by who happened to be watched, and that is a genuinely large contribution that does not require the model to be right about any individual player.
Underneath sits the harder and more valuable piece: valuation. Not what a player is worth in an abstract sense, but what the market is currently overpaying for and what it is ignoring. Goals from a striker in a strong league are priced efficiently, because everybody can see them. Defensive contribution from a full back in a mid-table side in a smaller league is priced badly, because fewer people can see it and fewer still can compare it across competitions. The department's real asset is a defensible way of comparing players across leagues of different quality, and that comparison is a modelling problem with no clean answer, which is exactly why the advantage lasts.
The failure mode is specific and expensive. A player signed on analytical grounds whom the coach did not want will not play, will not settle, and will be sold at a loss inside two years, at which point the model gets the blame for a decision that was actually an organisational failure. The recommendation was never wrong on its own terms. It was wrong about the organisation it was made in.
Where sports analytics genuinely fails, stated without hedging
A department that claims competence everywhere is telling you it has not been tested anywhere. There are areas where the honest answer is that the methods are weak, and saying so is what makes the strong claims believable.
Sample size, everywhere. A domestic football season is a few dozen matches. A manager's tenure is often shorter than the time required to distinguish a good process from a lucky one. This is not a technical problem to be solved with a better model, it is a hard limit on what can be known, and it applies with even greater force to cup competitions and short tournaments.
Attributing outcomes to tactics. There is no control group. A change of shape coincides with a change of personnel, a change of opponent, a change in fitness and a change in confidence, and nothing separates them. Claims that a tactical instruction caused an improvement are almost always stories told over noise, however well the numbers behind them are constructed.
Individual injury prediction. Group-level risk factors are well established, and load management on that basis is sensible practice. Predicting which specific player will get injured in which specific week is a different problem, and the published evidence for it is much weaker than the marketing around it. The responsible position is that data supports better decisions about training design without conferring the ability to see individual injuries coming, and this remains genuinely contested territory in the science of keeping athletes available.
Psychology, character and dressing rooms. Not measured, frequently claimed. The absence of data does not make these factors unimportant, and a department that treats what it cannot measure as though it does not exist will be wrong in ways it cannot detect.
Youth projection. The gap between a fifteen-year-old's output and a professional's is filled with physical maturation, and the relative age effect distorts every junior sample before it is even collected. Projection at that age is a genuinely open problem.
Live in-game analysis. The window between an event and the decision it should influence is often measured in seconds, and no analyst in a stand is going to change what happens in that window. What works in-game is what was agreed beforehand.
Set against those, two areas work unusually well, and it is worth understanding why. Set pieces are a closed system: a fixed starting position, a small number of options, high repetition, and outcomes attributable to the routine rather than to the flow of play. Goalkeeping and shot quality are similar, which is why shot-value models were the first analytical idea to cross into the mainstream in several sports, from football to the version used in hockey. The common feature is that the situation repeats in a comparable form, which is the precondition for learning anything from data at all. It is the same reason effects like the size and persistence of home advantage can be measured with confidence while the mechanisms behind them are still argued about.
How to judge whether a sports analytics department is working
Headcount is not the test. Neither is software spend, nor the number of dashboards, nor whether the club talks about data in interviews. Six questions do most of the work, and all of them can be asked by somebody outside the department.
Name three decisions in the last year that went differently because of the department, and say what happened. This is the whole test in one question. A functioning department has answers, with dates, and at least one of them will be an outcome that went badly, because a department that only remembers its wins is not keeping records.
Are recommendations written down before the result is known? A decision log, recording what was recommended, what was decided and why, is the only defence against hindsight. Without it, every past decision was obviously correct and every failure was somebody else's.
Does it survive a change of head coach? If a new manager arriving means the department is rebuilt, it was never a club function.
Is the analyst in the room? Physically, at the moment the decision is made. Departments that email into decisions do not influence them.
Does anybody ask them questions unprompted? Inbound demand from coaching or recruitment staff is the strongest single signal that the department has earned standing. Outbound reports nobody requested is the strongest signal that it has not.
Is the pipeline staffed separately? If the analysts maintain the data, they are not analysts. This is the most fixable of the six and the most commonly ignored.
Two further tests are worth applying to the organisation rather than to the department. Is there a named person accountable for whether analytics improves outcomes, as opposed to accountable for the analytics existing? And has the department ever been permitted to deliver an answer the club did not want, and remained intact afterwards? Where the answer to both is yes, the technical quality of the work almost stops mattering, because the organisation will notice and fix bad analysis. Where the answer is no, excellent analysis will change nothing at all.
The pattern holds across sports, which is why the same argument recurs everywhere from baseball front offices to cycling teams to national governing bodies, and why so much of how modern sport is actually run now depends on an organisational question rather than a technical one. The methods travel easily. The willingness to be told something inconvenient does not, and the same tension appears wherever a decision-maker's judgement meets a measurement of it, including in the study of how officials decide.
The uncomfortable summary for anyone entering the field is that the hardest skill is not the one being taught. Learning to build the model is a solved problem with abundant materials. Learning to hand a coach one sentence he will act on, at the right moment, in his own vocabulary, having earned the standing to be listened to, is not taught anywhere, and it is the entire job.
Common questions
What does a sports analytics department actually do?
It supports five kinds of decision: who to sign and at what price, how to prepare for a specific opponent, what to do during a match, how to manage training load and availability, and how to attack and defend set pieces. Underneath all five sits a data pipeline that consumes most of the department's working hours. The output is not a model, it is a recommendation that somebody with authority acts on.
Why do coaches ignore analytics?
Usually because the finding arrives in a form they cannot act on. A coach needs one clear instruction before a decision, and a model produces a range of outcomes with uncertainty attached, so somebody has to compress the second into the first and take responsibility for what is lost. Coaches also carry the blame for decisions personally, and a data-led decision that fails is criticised far more harshly than a conventional one that fails.
Who should a sports analytics department report to?
The reporting line decides the department's real remit, because a department ends up working on whatever its boss is accountable for. Reporting to the head coach makes the horizon the next match and ties the department's survival to the coach's. Reporting to a sporting or technical director gives it a multi-season horizon and makes recruitment its centre of gravity, which is why most clubs that get value from analytics are structured that way.
Is data analysis better at recruitment than at coaching?
Recruitment is the clearest win, because the decision is slow, repeated hundreds of times a year, and the default option is to do nothing. That gives an analyst time to build a case, and it means being right on average pays off. In-game and coaching decisions are fewer, faster and harder to attribute, so the same quality of work produces far less measurable benefit.
How can you tell if a sports analytics department is working?
Ask anybody senior to name three decisions in the last year that went differently because of the department, and to say what happened. If nobody can name them, the department is producing reports rather than changing outcomes, however good its models are. The other honest tests are whether it survives a change of head coach and whether its recommendations are written down before the result is known.
Filed under Across Sport·analytics · data · recruitment · coaching · sports science