A young runner checking a training readiness score on a watch, illustrating false precision in load monitoring

Load, Readiness and False Precision: the limits of algorithmic training prescription

Paper five of six in the Athleet.AI Working Papers on AI in coaching and athlete development, on what a colour on a screen does to a 15-year-old before a championship.

What is false precision in load monitoring?

A contested measure, presented as a settled number, to a reader who has no way to see the argument behind it. A readiness score is a claim about the future dressed as a description of the past.

The paper opened with an invented case. Maya, 15, runs 1500 metres for a club with two volunteer endurance coaches. Her mother pays for a consumer app linked to her watch. On the Thursday before the county championships, after a week with an added school cross-country fixture, the app shows her load ratio in red, says her injury risk is high, and suggests cutting training by a set percentage. Maya reads it on the bus, tells her coach she is worried, and asks whether she should race at all. The coach of four years thinks she is exam-tired but well, knows her volume is modest, and has no way to see how the app reached the figure.

The coach then spends the conversation undoing a worry the evidence does not support. That is the cost, and it lands on the athlete rather than on the vendor.

What in load monitoring holds up, and what does not?

A good deal holds up. The distinction between internal and external load is sound, with internal load being the stimulus that drives the adaptation. Session rating of perceived exertion multiplied by duration is a defensible index. Subjective self-report of mood, stress and recovery tracks training load changes with greater sensitivity and consistency than blood markers or heart rate indices. The observation that successful endurance athletes spend roughly 80 per cent of sessions at low intensity is robust and useful.

The layer built on top is where the trouble starts. The acute to chronic workload ratio, which compares roughly a week of load with roughly a month, was popularised with a range of about 0.8 to 1.3 associated with lower injury risk and values around 1.5 and above with higher risk. It was then taken apart in the same journals that carried it. Because the acute week sits inside the chronic period, the two terms are mathematically coupled, which produces spurious correlation. A second critique went further, arguing the ratio lacks a sound conceptual basis, does not normalise load in the way its users assume, and should be dismissed as a framework for managing injury risk.

Screening fares no better. A risk factor strongly associated with injury across a group can still discriminate poorly between individuals, which is why screening tests are unlikely ever to predict injury. Reviews of injury prediction models have found most at high risk of bias, with small samples and poor reporting, and none externally validated.

Why does a red warning mislead a young athlete in particular?

Base rates. Where weekly injury is uncommon, even a reasonably discriminating model issues far more warnings to athletes who stay healthy than to athletes who get hurt. A young athlete looking at a red square is more likely to be fine than not, and has no way to know it.

The paper traced a six-step chain from research to belief. A contested metric becomes a single score. The score gets a threshold and a colour. The colour acquires confident wording. The wording becomes a prescription. The athlete or the parent acts on it without the coach. An honest model shown through a misleading interface still misleads.

There is a scale problem underneath as well. Software thinks in days, while the timescale that matters in long-term development is years. Much apparent individual response to training is measurement error and within-subject variation, and separating a true responder from noise requires repeated measurement and repeated interventions. A tool claiming to have learned how a particular athlete responds may therefore be describing noise with great confidence.

What should a tool be allowed to say?

The paper set out what each measure can support.

  1. External load. Yes, as a record of what was done.
  2. Session RPE and self-report. Yes, as reported, and treated as the most sensitive routine signal available.
  3. Time in intensity zones. Yes, with the date the zones were set shown alongside.
  4. Acute to chronic ratio. Trend only, and never as a risk label.
  5. Injury risk score or probability. No, absent validation in that population.
  6. Recommended next-week load. An option for the coach, never an instruction to the athlete.

Five design principles follow. Describe before predicting. Label any classification as descriptive and non-clinical, avoiding the words risk, danger and warning. Make uncertainty visible through ranges, the date baselines were set, and flags for missing data. The coach owns the decision, which means no direct prescriptions to minors. No injury-risk claim without validated evidence for that population.

What should a club or governing body ask before buying?

Four questions. What does each score contain and how are the inputs weighted? Has any risk classification been validated, on whom, and published where? How is uncertainty shown, and what happens when data is missing? Does the tool message under-18s directly, and can a coach moderate it?

For a governing body there is a second job, and the author’s view is that it would do more good, sooner, than any improvement in the underlying models: publish short guidance for volunteer coaches on what to say when an athlete arrives worried about a colour. Set conditions for use with developing athletes, meaning no injury-risk labels without validated evidence for that population, visible uncertainty, descriptive language, and the coach owning the decision.

For a coach, the version that works this week is simpler. Ask athletes to share the underlying record alongside any score. Set aside a few minutes weekly to talk through how training actually felt, which is the measure with the best evidence behind it. Tell athletes and parents plainly that an app’s injury-risk message is not medical advice and should come to the coach first, and route anything involving pain or illness to a physiotherapist or medical lead.

Cover of Load, Readiness and False Precision, an Athleet.AI whitepaper

Free download

Get the full paper

Read Load, Readiness and False Precision in full, with every claim carrying an evidence grade and the full reference list. Give your name and email and the PDF opens straight away.

We use your details to send you this paper and to tell you when the next one is published. Nothing else, and no third parties.

Frequently asked questions

Is the acute to chronic workload ratio still usable?

As a descriptive trend, with care. As a label telling an athlete their injury risk, the statistical criticisms leave very little room. The two terms are mathematically coupled, which generates correlation that is an artefact of the arithmetic.

What is a readiness score actually measuring?

There is no agreed research definition of readiness, and no way for a user to see how a given composite weights its inputs. That inability is itself the point the paper is making.

Should a teenager see an injury risk figure at all?

The paper argues no, absent validation in that population, and argues against tools messaging minors directly in any case. The coach should own the decision and the conversation.

What is the best-evidenced thing a coach can monitor?

Self-report. Mood, stress and recovery track changes in training load with greater sensitivity and consistency than the objective markers that cost far more to collect.

Does machine learning improve injury prediction?

Not on the published record the paper reviewed. Eligible studies were few, predictive performance ranged from poor to strong, and the one clear success used data from a single professional club over a single season.

Legal and regulatory notice. This article describes research findings and product design practice. It is not medical advice and nothing here constitutes legal advice, and the provision of legal advice sits outside the terms of any engagement with the author or with Athleet.AI. The material is presented to support discussion and further review by qualified advisers. Pain, illness and injury should be referred to a qualified clinician.

Share:
Whitepapers