Louis Dallimore //Strength & Conditioning
Essay//Modelling Accumulated Load, Lag, and Soft-Tissue Injury RiskAnalytics

Modelling Accumulated Load, Lag, and Soft-Tissue Injury Risk

Twelve of seventeen hamstring injuries across two seasons were flagged inside the prior fortnight by counting six simple training markers. The model, the checks that broke earlier versions of it, and the board and planner that came out the other side.

Last season our squad, a top-flight professional rugby team, had more hamstring and calf strains than I was happy with. Across the two years in this piece they cost us roughly seventy player-weeks of availability: a typical spell around two weeks, the worst well over a month. We monitored load the way most clubs do, the numbers mostly looked fine, and players were still going down. You know how it is: everyone looks great, the boys are training well, then someone grabs the back of his leg and you think, where the hell did that come from?

So when the season ended I went back through all of it, every GPS file and every medical sheet, to work out what our monitoring was seeing and what it was missing. With the help of AI I pulled every file into one database, which let me run statistical models I couldn't have run on my own before.

This is what came out of it, including the parts that killed my own earlier work.

What the workload ratio missed

Start with the standard tool. Almost everyone runs an acute:chronic workload ratio, this week's load over the average of the last month. We ran different windows, and the smoothed EWMA version as well. Above about 1.5, be careful. Around 1.0, all good. On our data, on days players actually trained, it picked the right player about as well as a coin flip.

The moment it actually lost me wasn't a summary statistic though. It was one player, and not one anyone was watching: in a squad of about sixty he'd have been well down the priority list. I wasn't even looking for him at the time — I was pulling load charts together for a different piece and his season stopped me about halfway through building it. I saw it by eye, in the plain bar chart, before any model existed: three weeks in a row above his own baseline going into that final lighter week, the load stacking with nothing draining, right there in front of me. I knew that run-up was the problem; the accumulated fatigue was what got him. What I couldn't do was put a number on it, and nothing we were monitoring agreed with me. Real data all through this post, so I'll just call him Player A.

Figure 01 — Player A · twenty weeks · load above, carried load belowweekly HSR · injury 04-02
WEEKLY HSR · DASHED = HIS OWN BASELINECARRIED LOAD, STACKED · DEEPER COPPER = MORE ACCUMULATEDINJURY11-2012-1101-0101-2202-1203-0403-25
Bars: weekly high-speed metres vs his own 16-week baseline. The fill below: every above-baseline week’s fading contribution, stacked — the deeper the colour, the more accumulated load at that moment. White seams trace the layers. He is injured the day after the last column.

One definition before we go on. A "hamstring injury" here means a time-loss soft-tissue injury recorded by our medical lead: the player missed at least one day of training because of it. Some cost a single day, some cost a month; the median spell was about two weeks, so "injury" here spans anything from a one-day strain to a month on the sidelines.

Player A put together six weeks between 1,300 and 2,500 metres of high-speed running (anything above 5.5 metres a second, here and throughout), most of them above a personal norm of about 1,340. Then his last week came down. A planned lighter week, the kind we'd all sign off on. He injured his hamstring at the end of it.

Here's what the ratio made of those same twenty weeks.

Figure 02 — The same twenty weeks, read by ACWRacute : chronic · final reading 0.8
SWEET SPOT · 0.8–1.3“UNDERTRAINED” · <0.8“DANGER” · >1.50.51.01.52.00.8 — ON THE FLOOR OF THE SWEET SPOT11-2012-1101-0101-2202-1203-0403-25
The traditional zones, in the traditional colours. He spends the whole run-up in and around the green, and the week he is injured it reads 0.8. Call that fresh or call it undertrained; neither reading sees the month he is carrying.

The week he was injured, his acute:chronic came out at 0.8, right on the floor of the sweet spot. Read that either way you like: fine, or the undertrained flag some frameworks raise below 0.8. Neither reading is "this man is carrying a month of work." A ratio of this week against last month rewards a light week no matter what it lands on, because it has no memory of what it landed on.

That's not a cherry-picked bad reading. When I later ran a much wider search, no acute:chronic cut-off on any metric ever made the top of the table. The problem isn't where anyone set the threshold; no threshold can give the ratio a memory.

How long a hard week actually lasts

The picture I couldn't get out of my head was the stacked one in that first figure. Each hard week throws something forward, and the next hard week lands on top of it rather than on zero.

If the statistics feel abstract, the concept isn't. It's rest periods in the gym. Back-to-back sets of deadlifts on thirty seconds' rest: the fatigue from set one is still sitting there when set three starts, and the sets stack. Give it five minutes and it dissipates, and set three starts closer to fresh. Same mechanism here, stretched from minutes out to weeks. Except nobody times the rest periods between hard training weeks, because nothing on the standard dashboard shows them stacking.

There's a proper statistical tool for exactly this idea. It's called a distributed lag model (DLNM), and it's mostly used in epidemiology to ask things like "how long after a heatwave do hospital admissions stay elevated?" I pointed it at our data and asked the same shape of question: how long after a hard running week does the injury risk stay elevated?

A hard week turned out to be strongest over the following fortnight, roughly halved by week three, and gone by about week seven. I've come to see it as fatigue with a half-life: the cost of a big week decays rather than expiring, halving every few weeks, still faintly there long after the work that caused it. The biggest week of pre-season? It carries six to seven weeks of shadow, and the shadows add up. Which is exactly why Player A's light week didn't save him. The bar was small, but the layers underneath were still draining on their own schedule.

A scope note before going further: everything in this piece is running load into hamstrings. Metres above 5.5 a second (about 20 km/h, the traditional high-speed cut), one signal, one tissue. A hard bike, rowing or pool week works the engine without the same impact and eccentric strain, so it carries its own lag window, almost certainly a much shorter one; that would be a model of its own and beyond the scope here, so I can't put a number on it. But it is the logic behind cross-training a man while his running stack drains: you are resting the tissue, not the athlete. And none of the mechanism is rugby-specific: hamstrings in soccer, Australian rules and the NFL live on the same high-speed running, so the shape should travel even though the thresholds here are ours and yours would differ.

I tested this the hardest way I could find: comparing each injured player only against himself at other points of the same season, so nothing about who he is or when in the year it was could sneak in. The shape held. Two other things came out of the same testing that I didn't expect: the danger zone was in absolute metres, roughly 1,000 to 1,600 a week, at the same metres whether the player was a high-volume runner or not.

That sounds odd until you think in metres rather than percentages. A jump from 200 to 300 high-speed metres and a jump from 2,000 to 3,000 are the same 50 percent rise, but one asks the tissue for an extra hundred metres of sprinting and the other for an extra thousand. Ratios have their place; the hamstring only experiences metres.

And the players who got hurt weren't the spiky ones. They were the steady ones, grinding out much the same number week after week. The flat line I'd have called well-managed was the riskier pattern.

The six markers

A fitted statistical surface is not something you can run a Monday meeting off. So the practical version became six yes/no questions about each player, each one a marker that earned its place on the data:

Figure 03 — The six lights26,013 player-days · rate red vs not
WEEKLY HSRred 20% of days

weekly HSR inside 1,000–1,600 m — a band, not a floor

2.6× the rate when red
STREAKred 17% of days

3+ consecutive weeks above his own baseline

4.5× the rate when red
6-WK EXCESSred 33% of days

six-week excess over his baseline, top third

2.2× the rate when red
VARIATIONred 30% of days

weekly totals barely moving — flatness, not spikes

1.5× the rate when red
PRIOR STRAINSred 17% of days

any prior soft-tissue strain

3.5× the rate when red
LAST SPRINTred 39% of days

a ≥90% top-speed effort in the last 7 days

4.8× the rate when red
Five of the six compare a Player to himself. None is an alarm on its own — the weakest (variation) is also the most expensive to remove, because it is the most independent of the others. What matters is how many are on at once.

Five of the six compare a player to himself; only the metres band is absolute. None of these is an alarm on its own; the most common one is on for four players in ten on any given day. What matters is how many are on at once, and that count turned out to be worth more than anything else I built, including the fancy model it came from:

Figure 04 — Chance of an injury in the next fortnight, by lights showingobserved, two seasons · 17 injuries
0%5%10%15%0.14%01 in 72719.8% of days0.25%11 in 40132.8% of days0.42%21 in 23625.7% of days0.88%31 in 11315.8% of days4.67%41 in 215.1% of days6.36%51 in 150.9% of daysTHE STEP THAT MATTERS
Nought to three roughly doubles each step and stays under one percent; three to four multiplies by five. Whiskers are 95% credible intervals, clustered by event — the five-light bin rests on four injuries and runs 2 to 14 percent, so read its height loosely. The step at four is the solid part: P(real) exceeds 0.999, and it is at least a tripling with probability 0.92.

Nought to three lights, risk roughly doubles at each step and stays under one percent. Under one percent I'm pushing those players to train hard, build robustness and work through soreness — classic rugby coach chat, foot on the throat. Three to four, it multiplies by five, and that step is the whole reason the review line sits at four and not three. At four or more lights you're looking at about three names out of a squad of sixty on a Monday morning, and across our two seasons, twelve of the seventeen hamstring injuries had been on that list inside the previous fortnight.

Worth being clear about what that number is and isn't. Roughly one man in twenty on that list gets injured in the next two weeks, which means nineteen don't. It isn't a prediction machine and I don't use it as one. It tells you where your next two hours of attention are worth most, about eight times better than spreading them evenly. Run the same exercise with acute:chronic and the list is no better than drawing three names from a hat; that is what a coin-flip ranking means in practice.

I should also say why I trust it, because I've been burned by my own numbers in this project. An earlier version of this work found a single rule that looked spectacular, and it passed significance testing easily. Then I made the test harder: made the shuffled comparison respect the calendar, and counted every variant I'd tried against it. The rule died completely. The light count survived the same treatment at every level of strictness I could construct, and no single light is carrying it. I've also written the whole specification down, thresholds frozen, committed before next season's data exists, so next year is a real test rather than a re-fit. If it fails, I'll publish that too.

Calf turned out to be different

Everything above is hamstring. When I began this I hoped all soft tissue would fit in one bucket, and I assumed calf would just need its own thresholds. It didn't. The same six lights did nothing on our twenty-eight calf strains, and every other load formulation I tried came back just as empty.

What calf tracked instead was two things load can't see. Recurrence: over a third of our calf strains were re-injuries, and most of those arrived within eleven weeks of the previous one. And age, which turned out to be a major factor.

Figure 05 — Injury rate by age band · hamstring vs calfper 1,000 player-days · 56 players with birthdates
0.01.02.00.60.1<2522 players0.90.625–2716 players0.42.228–307 players0.81.831+11 playersHAMSTRINGCALF
Hamstring is flat across every age band. Calf climbs with age: the 30-and-overs ran at about four and a half times the under-30 rate, and twenty-two players under 25 produced one calf strain between them in two seasons.

Twenty-two players under 25 produced one calf strain between them in two seasons. The 30-and-overs carried around four and a half times the rate of the under-30s, and the climb starts at about 28. That's a comparison between players rather than within them, on 28 events, so hold it accordingly; what makes me trust it is that it lands exactly where the clinical calf research has pointed for years, and I compiled every birthdate from published squad lists before looking at the injury data. Meanwhile hamstring completely ignored age, flat across every band. Same squad, same GPS, two tissues. Hamstring cares about your training pattern and not your birthday; with calf it's the other way round.

That changed how I'd manage calf entirely, and it deserves its own piece rather than a corner of this one — the short version is that calf runs off the roster and the medical record, not the GPS. From here I'll keep to hamstrings.

What I'd do differently now

Put together, the morning view is one table. Six markers per player, the count, what the count is worth, his trend, and his recent peak capacity for context. The board starts conversations, it doesn't finish them: five lights is a discussion with the performance staff, not an automatic cut to a player's load. It's a starting point, not a final decision. And one rule of use: this is a staff tool. Risk numbers shown to players become their own kind of load, and nothing here is precise enough to put in a player's head.

Figure 06 — The board I’d build now · one real morning, six of the squad2024-04-01 · actual data, anonymised
PlayerWeekly HSR
band 1,000–1,600 m
Streak
wks above normal · ≥3
6-wk excess
>1,426 m
Variation
<0.41 = flat
Prior strains
≥1
Last sprint
since ≥90% · ≤7d
Lights14-day risk12-wk trendCap
180d peak · context
Player A ← injured next day1,35642,0880.2901344.67%
1 in 21
2,674 MID · FIRST CALL
Player B1,13471,7000.2801544.67%
1 in 21
3,695 HIGH
Player C1,04423510.210130.88%
1 in 113
2,408 MID
Player D1,84917840.270620.42%
1 in 236
3,092 HIGH
Player E1,83312,4360.5603010.25%
1 in 401
3,213 HIGH
Player F89532,3341.11214430.88%
1 in 113
2,001 MID
Real rows from one morning, players anonymised — the actual number in every cell, shaded by whether it crosses its threshold. Player A: in the HSR band, a four-week streak, 2,088 m of excess, dead-flat weeks — four reds from load pattern alone. Player F: three reds from history and accumulation, nothing to do with this week. Player A was injured the next day; Player B, at four lights a month straight, did not — and B carries the biggest recent peak on the board (see Figure 07). Capacity never changes the count; at four-plus lights it orders the list, and the tinted cell marks the first call. Last-sprint cells go grey, not green, when clear: a long gap is not safety, it is an unassessed tissue.

Planning the next block

The last step was turning the hamstring side into something you can plan with, not just review with. Set the four weeks a player has just done, set the four you intend, and it scores every week in lights.

Figure 09 — Forward planner · score your block in lightsinteractive · drag the weeks
wk -4
900
2
wk -3
1,400
3
wk -2
1,600
3
wk -1
1,250
4
plan 1
1,200
4
plan 2
1,200
4
plan 3
1,200
4
plan 4
1,200
4
CARRIED LOAD, STACKED · TAIL = THE BLOCK DRAINING AFTER IT ENDSRISK · OBSERVED 14-DAY RATE AT EACH WEEK’S COUNT0%3%6%wk -4wk -2plan 1plan 3+1+3+5
4 lights
worst planned week (wk 1) — at review level, replan
4.67%
observed 14-day rate at that count · 1 in 21
4,800 m
planned HSR over the four weeks
The HSR band, streak and six-week excess use the pre-registered definitions exactly; the variation marker is approximated from the visible weeks; weeks before the window are assumed at baseline. The rate is two-season calibration, not prophecy. The job: drag until no planned week reaches four.

The job is one sentence: drag until no planned week reaches four lights. Catching lit-up players is half the job; planning blocks that never light anyone up is the better half, and seeing how the last month's lag will interact with the next one is the starting point. Have a play with it.

The default block loaded above is worth a look before you touch it. A flat 1,200-metre block, for a player with a previous strain who sprints weekly, scores four lights with zero accumulated overload, because flatness is itself one of the markers. Vary the weeks and it clears. That surprised me the first time I saw it. Then it stopped surprising me, because it's something every coach already half-knows from the gym anyway: waves beat grinding.

The other side of the coin

Everything above is about catching risk. The better use is never accumulating it. The board and the planner imply a set of programming habits, most of them free, and none of them trains the squad less: the metres stay, the shape changes.

Break the run before it reaches four. The big one. If a player has been above his own baseline three weeks straight, the fourth week is a decision, not a drift. A genuine below-baseline week there resets the streak and starts the stack draining. In the gym you'd never program a fourth straight top set without backing off; same instinct, week-scale.

Plan the fluctuations, don't wait for the reload. Stacking consistent hard weeks is where the danger lives; planned swings in volume and intensity, which is just proper periodisation, pay for themselves later. The shadow of a big week is still half there three weeks on, and every coach learns early that players get injured in week three; this is that saying, visualised. A light week planned before the debt crosses roughly 1,400 metres of six-week excess does its job. The same light week after the debt is built changes almost nothing that fortnight, which is exactly what Player A's last week showed. The time to insert the easy week is when it feels slightly too early.

Wave it, don't grind it. Flat blocks, the same weekly total month after month, show up on the variation marker, and on this data the steady grinders were the ones getting injured. Hard-easy waves break both the streak and the flatness at once, and in the planner they beat every flat block at the same total metres. This was the finding that most went against my instincts, and it's the cheapest to act on.

Treat the 1,000–1,600 m band as spending, not drifting. That band is where the adaptation happens and where the risk lives, both at the same metres, for every kind of runner. Weeks in the band should be placed on purpose, with a plan for what follows them, not arrived at by accident because the schedule filled up.

Give the previously-injured a head start on the count. A player with a prior strain wakes up at one light every day. His review level is effectively three, not four. In practice he gets first call on the varied weeks and the early deloads, not the leftover schedule.

Keep the sprinting; place it. The sprint marker is exposure, not sin. Players need top-speed work, pulling it out entirely creates a different problem, and nobody is recommending that. What the board changes is the placement, which did surprise me: a ≥90% session in a week already carrying three lights is the avoidable version. Longer term it circles back to better planning, more robust athletes, and exposure as protection, all things I'm a huge fan of.

A caveat on all of the above: these are the habits the associations point at, not proven interventions. Nobody has run the trial where you act on the board and count what doesn't happen. That's the study I'd most like to run next.

Where machine learning fits

People ask why this isn't an AI model, and the question deserves a proper answer, because injury prediction is drowning in machine learning right now.

The short answer is seventeen injuries. Every flexible model, gradient boosting, random forests, neural networks, buys its cleverness with data, and the currency is events, not GPS rows. I have twenty-six thousand player-days but only seventeen injuries. Point a big model at that and it either memorises noise around seventeen accidents or, tested properly, learns nothing. Rather than assert that, I built the strong machine-learning version and made it sit the same exam as the six counted markers.

Figure 08 — The six markers vs the fitted model, examinedsame data · engine scored out-of-sample
Ranking players it trained onin-sample AUC0.727BOARD0.849ENGINERanking players it never sawplayer-grouped cross-validation0.727BOARD0.654ENGINEInjuries caught at the working budget, three names a dayof 17, flagged inside the prior fortnight · ±1 from tie-breaking12/17BOARD5/17ENGINE, unseen players10/17engine if allowed to memorise
On the players it trained on the fitted model wins easily. On players it never saw it falls below the count, which cannot overfit because nothing in it is fitted. The gap between the model’s two scores is memorisation, measured. Within a fixed count its leftover signal separates nothing, so blending the two only dilutes.

The result reads the same in both directions. The fitted model wins easily on players it trained on, which is the number every product demo quotes, and falls below the plain count the moment it meets players it never saw. At three names a day it catches five of seventeen against the count's twelve. And there was one detail I genuinely enjoyed: left to learn whatever it wanted, the model's strongest interpretable term came out as consecutive weeks above baseline, weight +0.39. The machine rediscovered the coaching hypothesis this whole piece started from. The full story of that exam, including what the fitted model can do that counting can't, is its own short piece.

What I did add is a Bayesian version of the same model, which is less about predicting and more about knowing how sure to be. Three things came out of it. The strange finding that weeks over 1,600 metres look safer than weeks at 1,300 is probably real: 94 percent probability, under a fitting method that pulls in the opposite direction. The shadow's half-life is about four weeks; it is still there at week six (87 percent), and week seven is a coin flip, so that is where the claim ends. And after everything training load explains, players still differ several-fold in how easily they get injured. That last one is the upgrade path: a per-player susceptibility the model learns slowly as seasons accrue. I've deliberately left it off the board for now, because adding it on seventeen events would be fitting, not learning.

That frailty number sent me back to the data with a question a coach asked me directly: can you see robustness from the outside? The obvious candidate is what a player has already proven he can do. So I took each man's biggest running week from his recent past and asked whether the same red state costs a high-capacity player less. It leans that way everywhere I looked.

Figure 07 — Rate at 4+ lights, split by recent peak capacityexploratory · whiskers: 95% CI
0%5%10%15%5.2%low8.9%mid1%high90-day lookback6.4%low6.1%mid3.1%high180-day6.9%low6.9%mid1.9%high365-day
The high-capacity third (copper) sits at the lowest rate under every lookback tried — six cells of six lean the same way — but the intervals overlap and each cell holds two to six injuries. The lean is consistent but unproven, so it stays off the scored board.

At four or more lights, the high-capacity third of the squad sat at the lowest rate under every lookback I tried, somewhere between one in thirty and one in a hundred against roughly one in fifteen for everyone else. Hold it loosely: the intervals overlap, each cell rests on a handful of injuries, and there is a trap in the middle of it. Fragile players never get the chance to bank a 3,000-metre week, so high capacity is partly a consequence of being robust rather than a cause of it. The data cannot tell those apart. Four-plus is as fine as seventeen events can split; at exactly five lights we have four injuries in total, and a split that small cannot be trusted either way.

So what do you actually do with the number on a Monday? Use it to order the list, never to shorten it. Four-plus lights puts a man on the review list regardless; capacity decides who you talk to first. Four lights without a big recent peak behind him is the priority conversation, roughly one in sixteen on our data. Four lights on a man who has recently shown he can carry 3,000-metre weeks still gets the conversation, just with more headroom, roughly one in thirty. The intervals overlap, so treat the ordering as a lean rather than a rule. On the board above, Player A was the first kind and Player B was the second, and the season resolved them exactly that way. Capacity stays a context column rather than a seventh light, and if the effect replicates next season it graduates into the per-player starting point the Bayesian model is already set up to learn.

When would heavier machine learning earn its place? At around a hundred events, which means several clubs or several more seasons, boosted trees under grouped cross-validation become worth running. In the thousands, sequence models might start out-learning hand-built features. Nobody in rugby has that data yet. And the objective matters as much as the sample. For a Monday meeting, a model a coach can read, question and disagree with beats one that is two points more accurate and mute.

Where this actually stands

Two seasons, one squad, seventeen hamstring injuries. That's what all of this rests on, and I'm not going to pretend it's more. The thresholds were found in this data, so they'll perform worse on the next squad, the way found thresholds always do. Everything here is also retrospective, worked out after the injuries with the outcomes on the table. I'm not claiming I could have predicted these seventeen, and I'm certainly not claiming the next seventeen. The claim is smaller: knowing what I know now, I would plan differently, and the frozen specification is how I find out whether that's knowledge or hindsight.

One check worth reporting on the twelve the board did flag: how many lights a player showed had no relationship with how long he ended up out. The two longest absences were both flagged, and two of the five misses cost a single day each, so the catch count isn't padded with trivial cases. Nothing here shows that acting on a flag prevents anything; that's a different study, and nobody's run it yet, including me. And the calf numbers are description, not a model.

What I'd stand behind: a hard week keeps contributing for about six weeks, and it stacks. The ratio most of us run cannot see that, and on this data it never came close. Counting six simple markers beat everything more sophisticated I built, and survived tests that killed things which looked better. The specification for next season is frozen and time-stamped, so in a year I'll know whether this replicates, and so will anyone who asks.

If you work in this space and want to argue with any of it, I'd enjoy that. The methods are the part I'm most confident in, and they only got that way from being wrong repeatedly on the way here.

References

A few pieces of literature sit behind the choices in this piece:

New essays in your inbox.

Roughly one a fortnight. Programming, GPS, return-to-play, applied ML. No spam, unsubscribe anytime.