Last season our squad, a top-flight professional rugby team, had more hamstring and calf strains than I was happy with. Across the two years in this piece they cost us roughly seventy player-weeks of availability: a typical spell around two weeks, the worst well over a month. We monitored load the way most clubs do, the numbers mostly looked fine, and players were still going down. You know how it is: everyone looks great, the boys are training well, then someone grabs the back of his leg and you think, where the hell did that come from?
So when the season ended I went back through all of it, every GPS file and every medical sheet, to work out what our monitoring was seeing and what it was missing. With the help of AI I pulled every file into one database, which let me run statistical models I couldn't have run on my own before.
This is what came out of it, including the parts that killed my own earlier work.
What the workload ratio missed
Start with the standard tool. Almost everyone runs an acute:chronic workload ratio, this week's load over the average of the last month. We ran different windows, and the smoothed EWMA version as well. Above about 1.5, be careful. Around 1.0, all good. On our data, on days players actually trained, it picked the right player about as well as a coin flip.
The moment it actually lost me wasn't a summary statistic though. It was one player, and not one anyone was watching: in a squad of about sixty he'd have been well down the priority list. I wasn't even looking for him at the time — I was pulling load charts together for a different piece and his season stopped me about halfway through building it. I saw it by eye, in the plain bar chart, before any model existed: three weeks in a row above his own baseline going into that final lighter week, the load stacking with nothing draining, right there in front of me. I knew that run-up was the problem; the accumulated fatigue was what got him. What I couldn't do was put a number on it, and nothing we were monitoring agreed with me. Real data all through this post, so I'll just call him Player A.
One definition before we go on. A "hamstring injury" here means a time-loss soft-tissue injury recorded by our medical lead: the player missed at least one day of training because of it. Some cost a single day, some cost a month; the median spell was about two weeks, so "injury" here spans anything from a one-day strain to a month on the sidelines.
Player A put together six weeks between 1,300 and 2,500 metres of high-speed running (anything above 5.5 metres a second, here and throughout), most of them above a personal norm of about 1,340. Then his last week came down. A planned lighter week, the kind we'd all sign off on. He injured his hamstring at the end of it.
Here's what the ratio made of those same twenty weeks.
The week he was injured, his acute:chronic came out at 0.8, right on the floor of the sweet spot. Read that either way you like: fine, or the undertrained flag some frameworks raise below 0.8. Neither reading is "this man is carrying a month of work." A ratio of this week against last month rewards a light week no matter what it lands on, because it has no memory of what it landed on.
That's not a cherry-picked bad reading. When I later ran a much wider search, no acute:chronic cut-off on any metric ever made the top of the table. The problem isn't where anyone set the threshold; no threshold can give the ratio a memory.
How long a hard week actually lasts
The picture I couldn't get out of my head was the stacked one in that first figure. Each hard week throws something forward, and the next hard week lands on top of it rather than on zero.
If the statistics feel abstract, the concept isn't. It's rest periods in the gym. Back-to-back sets of deadlifts on thirty seconds' rest: the fatigue from set one is still sitting there when set three starts, and the sets stack. Give it five minutes and it dissipates, and set three starts closer to fresh. Same mechanism here, stretched from minutes out to weeks. Except nobody times the rest periods between hard training weeks, because nothing on the standard dashboard shows them stacking.
There's a proper statistical tool for exactly this idea. It's called a distributed lag model (DLNM), and it's mostly used in epidemiology to ask things like "how long after a heatwave do hospital admissions stay elevated?" I pointed it at our data and asked the same shape of question: how long after a hard running week does the injury risk stay elevated?
A hard week turned out to be strongest over the following fortnight, roughly halved by week three, and gone by about week seven. I've come to see it as fatigue with a half-life: the cost of a big week decays rather than expiring, halving every few weeks, still faintly there long after the work that caused it. The biggest week of pre-season? It carries six to seven weeks of shadow, and the shadows add up. Which is exactly why Player A's light week didn't save him. The bar was small, but the layers underneath were still draining on their own schedule.
A scope note before going further: everything in this piece is running load into hamstrings. Metres above 5.5 a second (about 20 km/h, the traditional high-speed cut), one signal, one tissue. A hard bike, rowing or pool week works the engine without the same impact and eccentric strain, so it carries its own lag window, almost certainly a much shorter one; that would be a model of its own and beyond the scope here, so I can't put a number on it. But it is the logic behind cross-training a man while his running stack drains: you are resting the tissue, not the athlete. And none of the mechanism is rugby-specific: hamstrings in soccer, Australian rules and the NFL live on the same high-speed running, so the shape should travel even though the thresholds here are ours and yours would differ.
I tested this the hardest way I could find: comparing each injured player only against himself at other points of the same season, so nothing about who he is or when in the year it was could sneak in. The shape held. Two other things came out of the same testing that I didn't expect: the danger zone was in absolute metres, roughly 1,000 to 1,600 a week, at the same metres whether the player was a high-volume runner or not.
That sounds odd until you think in metres rather than percentages. A jump from 200 to 300 high-speed metres and a jump from 2,000 to 3,000 are the same 50 percent rise, but one asks the tissue for an extra hundred metres of sprinting and the other for an extra thousand. Ratios have their place; the hamstring only experiences metres.
And the players who got hurt weren't the spiky ones. They were the steady ones, grinding out much the same number week after week. The flat line I'd have called well-managed was the riskier pattern.
The six markers
A fitted statistical surface is not something you can run a Monday meeting off. So the practical version became six yes/no questions about each player, each one a marker that earned its place on the data:
weekly HSR inside 1,000–1,600 m — a band, not a floor
3+ consecutive weeks above his own baseline
six-week excess over his baseline, top third
weekly totals barely moving — flatness, not spikes
any prior soft-tissue strain
a ≥90% top-speed effort in the last 7 days
Five of the six compare a player to himself; only the metres band is absolute. None of these is an alarm on its own; the most common one is on for four players in ten on any given day. What matters is how many are on at once, and that count turned out to be worth more than anything else I built, including the fancy model it came from:
Nought to three lights, risk roughly doubles at each step and stays under one percent. Under one percent I'm pushing those players to train hard, build robustness and work through soreness — classic rugby coach chat, foot on the throat. Three to four, it multiplies by five, and that step is the whole reason the review line sits at four and not three. At four or more lights you're looking at about three names out of a squad of sixty on a Monday morning, and across our two seasons, twelve of the seventeen hamstring injuries had been on that list inside the previous fortnight.
Worth being clear about what that number is and isn't. Roughly one man in twenty on that list gets injured in the next two weeks, which means nineteen don't. It isn't a prediction machine and I don't use it as one. It tells you where your next two hours of attention are worth most, about eight times better than spreading them evenly. Run the same exercise with acute:chronic and the list is no better than drawing three names from a hat; that is what a coin-flip ranking means in practice.
I should also say why I trust it, because I've been burned by my own numbers in this project. An earlier version of this work found a single rule that looked spectacular, and it passed significance testing easily. Then I made the test harder: made the shuffled comparison respect the calendar, and counted every variant I'd tried against it. The rule died completely. The light count survived the same treatment at every level of strictness I could construct, and no single light is carrying it. I've also written the whole specification down, thresholds frozen, committed before next season's data exists, so next year is a real test rather than a re-fit. If it fails, I'll publish that too.
Calf turned out to be different
Everything above is hamstring. When I began this I hoped all soft tissue would fit in one bucket, and I assumed calf would just need its own thresholds. It didn't. The same six lights did nothing on our twenty-eight calf strains, and every other load formulation I tried came back just as empty.
What calf tracked instead was two things load can't see. Recurrence: over a third of our calf strains were re-injuries, and most of those arrived within eleven weeks of the previous one. And age, which turned out to be a major factor.
Twenty-two players under 25 produced one calf strain between them in two seasons. The 30-and-overs carried around four and a half times the rate of the under-30s, and the climb starts at about 28. That's a comparison between players rather than within them, on 28 events, so hold it accordingly; what makes me trust it is that it lands exactly where the clinical calf research has pointed for years, and I compiled every birthdate from published squad lists before looking at the injury data. Meanwhile hamstring completely ignored age, flat across every band. Same squad, same GPS, two tissues. Hamstring cares about your training pattern and not your birthday; with calf it's the other way round.
That changed how I'd manage calf entirely, and it deserves its own piece rather than a corner of this one — the short version is that calf runs off the roster and the medical record, not the GPS. From here I'll keep to hamstrings.
What I'd do differently now
Put together, the morning view is one table. Six markers per player, the count, what the count is worth, his trend, and his recent peak capacity for context. The board starts conversations, it doesn't finish them: five lights is a discussion with the performance staff, not an automatic cut to a player's load. It's a starting point, not a final decision. And one rule of use: this is a staff tool. Risk numbers shown to players become their own kind of load, and nothing here is precise enough to put in a player's head.
| Player | Weekly HSR band 1,000–1,600 m | Streak wks above normal · ≥3 | 6-wk excess >1,426 m | Variation <0.41 = flat | Prior strains ≥1 | Last sprint since ≥90% · ≤7d | Lights | 14-day risk | 12-wk trend | Cap 180d peak · context |
|---|---|---|---|---|---|---|---|---|---|---|
| Player A ← injured next day | 1,356 | 4 | 2,088 | 0.29 | 0 | 13 | 4 | 4.67% 1 in 21 | 2,674 MID · FIRST CALL | |
| Player B | 1,134 | 7 | 1,700 | 0.28 | 0 | 15 | 4 | 4.67% 1 in 21 | 3,695 HIGH | |
| Player C | 1,044 | 2 | 351 | 0.21 | 0 | 1 | 3 | 0.88% 1 in 113 | 2,408 MID | |
| Player D | 1,849 | 1 | 784 | 0.27 | 0 | 6 | 2 | 0.42% 1 in 236 | 3,092 HIGH | |
| Player E | 1,833 | 1 | 2,436 | 0.56 | 0 | 30 | 1 | 0.25% 1 in 401 | 3,213 HIGH | |
| Player F | 895 | 3 | 2,334 | 1.11 | 2 | 144 | 3 | 0.88% 1 in 113 | 2,001 MID |
Planning the next block
The last step was turning the hamstring side into something you can plan with, not just review with. Set the four weeks a player has just done, set the four you intend, and it scores every week in lights.
The job is one sentence: drag until no planned week reaches four lights. Catching lit-up players is half the job; planning blocks that never light anyone up is the better half, and seeing how the last month's lag will interact with the next one is the starting point. Have a play with it.
The default block loaded above is worth a look before you touch it. A flat 1,200-metre block, for a player with a previous strain who sprints weekly, scores four lights with zero accumulated overload, because flatness is itself one of the markers. Vary the weeks and it clears. That surprised me the first time I saw it. Then it stopped surprising me, because it's something every coach already half-knows from the gym anyway: waves beat grinding.
The other side of the coin
Everything above is about catching risk. The better use is never accumulating it. The board and the planner imply a set of programming habits, most of them free, and none of them trains the squad less: the metres stay, the shape changes.
Break the run before it reaches four. The big one. If a player has been above his own baseline three weeks straight, the fourth week is a decision, not a drift. A genuine below-baseline week there resets the streak and starts the stack draining. In the gym you'd never program a fourth straight top set without backing off; same instinct, week-scale.
Plan the fluctuations, don't wait for the reload. Stacking consistent hard weeks is where the danger lives; planned swings in volume and intensity, which is just proper periodisation, pay for themselves later. The shadow of a big week is still half there three weeks on, and every coach learns early that players get injured in week three; this is that saying, visualised. A light week planned before the debt crosses roughly 1,400 metres of six-week excess does its job. The same light week after the debt is built changes almost nothing that fortnight, which is exactly what Player A's last week showed. The time to insert the easy week is when it feels slightly too early.
Wave it, don't grind it. Flat blocks, the same weekly total month after month, show up on the variation marker, and on this data the steady grinders were the ones getting injured. Hard-easy waves break both the streak and the flatness at once, and in the planner they beat every flat block at the same total metres. This was the finding that most went against my instincts, and it's the cheapest to act on.
Treat the 1,000–1,600 m band as spending, not drifting. That band is where the adaptation happens and where the risk lives, both at the same metres, for every kind of runner. Weeks in the band should be placed on purpose, with a plan for what follows them, not arrived at by accident because the schedule filled up.
Give the previously-injured a head start on the count. A player with a prior strain wakes up at one light every day. His review level is effectively three, not four. In practice he gets first call on the varied weeks and the early deloads, not the leftover schedule.
Keep the sprinting; place it. The sprint marker is exposure, not sin. Players need top-speed work, pulling it out entirely creates a different problem, and nobody is recommending that. What the board changes is the placement, which did surprise me: a ≥90% session in a week already carrying three lights is the avoidable version. Longer term it circles back to better planning, more robust athletes, and exposure as protection, all things I'm a huge fan of.
A caveat on all of the above: these are the habits the associations point at, not proven interventions. Nobody has run the trial where you act on the board and count what doesn't happen. That's the study I'd most like to run next.
Where machine learning fits
People ask why this isn't an AI model, and the question deserves a proper answer, because injury prediction is drowning in machine learning right now.
The short answer is seventeen injuries. Every flexible model, gradient boosting, random forests, neural networks, buys its cleverness with data, and the currency is events, not GPS rows. I have twenty-six thousand player-days but only seventeen injuries. Point a big model at that and it either memorises noise around seventeen accidents or, tested properly, learns nothing. Rather than assert that, I built the strong machine-learning version and made it sit the same exam as the six counted markers.
The result reads the same in both directions. The fitted model wins easily on players it trained on, which is the number every product demo quotes, and falls below the plain count the moment it meets players it never saw. At three names a day it catches five of seventeen against the count's twelve. And there was one detail I genuinely enjoyed: left to learn whatever it wanted, the model's strongest interpretable term came out as consecutive weeks above baseline, weight +0.39. The machine rediscovered the coaching hypothesis this whole piece started from. The full story of that exam, including what the fitted model can do that counting can't, is its own short piece.
What I did add is a Bayesian version of the same model, which is less about predicting and more about knowing how sure to be. Three things came out of it. The strange finding that weeks over 1,600 metres look safer than weeks at 1,300 is probably real: 94 percent probability, under a fitting method that pulls in the opposite direction. The shadow's half-life is about four weeks; it is still there at week six (87 percent), and week seven is a coin flip, so that is where the claim ends. And after everything training load explains, players still differ several-fold in how easily they get injured. That last one is the upgrade path: a per-player susceptibility the model learns slowly as seasons accrue. I've deliberately left it off the board for now, because adding it on seventeen events would be fitting, not learning.
That frailty number sent me back to the data with a question a coach asked me directly: can you see robustness from the outside? The obvious candidate is what a player has already proven he can do. So I took each man's biggest running week from his recent past and asked whether the same red state costs a high-capacity player less. It leans that way everywhere I looked.
At four or more lights, the high-capacity third of the squad sat at the lowest rate under every lookback I tried, somewhere between one in thirty and one in a hundred against roughly one in fifteen for everyone else. Hold it loosely: the intervals overlap, each cell rests on a handful of injuries, and there is a trap in the middle of it. Fragile players never get the chance to bank a 3,000-metre week, so high capacity is partly a consequence of being robust rather than a cause of it. The data cannot tell those apart. Four-plus is as fine as seventeen events can split; at exactly five lights we have four injuries in total, and a split that small cannot be trusted either way.
So what do you actually do with the number on a Monday? Use it to order the list, never to shorten it. Four-plus lights puts a man on the review list regardless; capacity decides who you talk to first. Four lights without a big recent peak behind him is the priority conversation, roughly one in sixteen on our data. Four lights on a man who has recently shown he can carry 3,000-metre weeks still gets the conversation, just with more headroom, roughly one in thirty. The intervals overlap, so treat the ordering as a lean rather than a rule. On the board above, Player A was the first kind and Player B was the second, and the season resolved them exactly that way. Capacity stays a context column rather than a seventh light, and if the effect replicates next season it graduates into the per-player starting point the Bayesian model is already set up to learn.
When would heavier machine learning earn its place? At around a hundred events, which means several clubs or several more seasons, boosted trees under grouped cross-validation become worth running. In the thousands, sequence models might start out-learning hand-built features. Nobody in rugby has that data yet. And the objective matters as much as the sample. For a Monday meeting, a model a coach can read, question and disagree with beats one that is two points more accurate and mute.
Where this actually stands
Two seasons, one squad, seventeen hamstring injuries. That's what all of this rests on, and I'm not going to pretend it's more. The thresholds were found in this data, so they'll perform worse on the next squad, the way found thresholds always do. Everything here is also retrospective, worked out after the injuries with the outcomes on the table. I'm not claiming I could have predicted these seventeen, and I'm certainly not claiming the next seventeen. The claim is smaller: knowing what I know now, I would plan differently, and the frozen specification is how I find out whether that's knowledge or hindsight.
One check worth reporting on the twelve the board did flag: how many lights a player showed had no relationship with how long he ended up out. The two longest absences were both flagged, and two of the five misses cost a single day each, so the catch count isn't padded with trivial cases. Nothing here shows that acting on a flag prevents anything; that's a different study, and nobody's run it yet, including me. And the calf numbers are description, not a model.
What I'd stand behind: a hard week keeps contributing for about six weeks, and it stacks. The ratio most of us run cannot see that, and on this data it never came close. Counting six simple markers beat everything more sophisticated I built, and survived tests that killed things which looked better. The specification for next season is frozen and time-stamped, so in a year I'll know whether this replicates, and so will anyone who asks.
If you work in this space and want to argue with any of it, I'd enjoy that. The methods are the part I'm most confident in, and they only got that way from being wrong repeatedly on the way here.
References
A few pieces of literature sit behind the choices in this piece:
- Gasparrini A, Armstrong B, Kenward MG (2010). Distributed lag non-linear models. Statistics in Medicine 29(21), 2224–34. The method behind the lag surface. doi.org/10.1002/sim.3940
- Impellizzeri FM, Tenan MS, Kempton T, Novak A, Coutts AJ (2020). Acute:chronic workload ratio: conceptual issues and fundamental pitfalls. International Journal of Sports Physiology and Performance 15(6), 907–13. The formal version of the case this piece makes against the ratio. doi.org/10.1123/ijspp.2019-0864
- Malone S, Roe M, Doran DA, Gabbett TJ, Collins K (2017). High chronic training loads and exposure to bouts of maximal velocity running reduce injury risk in elite Gaelic football. Journal of Science and Medicine in Sport 20(3), 250–54. The U-shaped finding: both under- and over-exposure to top speed raise risk, which is why the sprint marker is about placement rather than avoidance. doi.org/10.1016/j.jsams.2016.08.005
- Green B, Bourne MN, van Dyk N, Pizzari T (2020). Recalibrating the risk of hamstring strain injury. British Journal of Sports Medicine 54(18), 1081–88. The meta-analysis behind the previous-strain marker: any history of hamstring injury multiplies risk around 2.7-fold. It also lists older age as a hamstring risk factor, which our squad didn't show; two seasons of one team is exactly the sample size where a real effect can hide. doi.org/10.1136/bjsports-2019-100983
- Green B, Pizzari T (2017). Calf muscle strain injuries in sport: a systematic review of risk factors. British Journal of Sports Medicine 51(16), 1189–94. Age and previous strain as calf's leading risk factors, which is where our calf data landed. doi.org/10.1136/bjsports-2016-097177