Paper four of six in the Athleet.AI Working Papers on AI in coaching and athlete development, on the kind of fairness that general AI ethics does not cover.
What is developmental fairness?
The question of whether an output, at the moment it is used, prematurely narrows the options of someone who has not finished growing. It is distinctive because it concerns time, and it is the dimension that general AI fairness work leaves out.
The paper opened with two 13-year-olds on adjacent lines of a dashboard, each with a composite potential score of 71, produced by a purchased tool blending sprint times, a jump test, coach ratings and match statistics. The first was born in the opening weeks of the selection year and has started her growth spurt early. She is taller, heavier and among the quickest at every age group. The second was born ten months later, has not started her spurt, and reaches 71 through first touch and decision-making. With fewer places than candidates and nothing on the page to separate them, the panel picks the one who looks further ahead. Her club moves the second girl to a lower training group.
No figure on that dashboard is wrong. The second athlete’s 71 is the stronger signal, because it was earned against a physical handicap that will shrink, and the tool has no way to say so.
Why does average accuracy tell you nothing about fairness?
Because it describes whoever dominates the test set. The benchmark study on facial analysis remains the clearest demonstration: error rates reached 34.7 per cent for darker-skinned women against a maximum of 0.8 per cent for lighter-skinned men, on benchmarks that were roughly 80 to 86 per cent lighter-skinned. Average accuracy is not a fairness test.
Sport carries its own version of that problem. A review of 5,261 publications across six leading journals found 31 per cent of studies used male samples only and six per cent female only, with 66 per cent of more than 12 million participants being male. A tool developed on a boys’ programme and extended to girls on a smaller sample is the ordinary case rather than the exception.
Biased proxies do the rest of the damage. A health algorithm that predicted cost as a stand-in for illness produced results so skewed that correcting it would have raised the share of Black patients flagged for extra help from 17.7 per cent to 46.5 per cent. Current performance as a stand-in for potential is the same move, made in a population where current performance is substantially a function of birth month and maturity.
What are the four dimensions of fairness, and how is each tested?
Each has a characteristic failure and a test that can be run before deployment rather than after a complaint.
- Representational. Who is in the data? Girls and para athletes are thinly sampled. Test by requiring subgroup counts and results from the vendor.
- Allocative. Who gets the resources? Places tilt towards relatively older athletes. Test by replaying a past cohort and comparing the allocations the tool would have made.
- Developmental. Does the output close options too early? Growth-spurt scores get used for deselection. Test by checking how maturity is reported and what the rescore rules are.
- Procedural. Can families understand and challenge it? Often no record exists that a tool was used at all. Test by walking a mock appeal through the process.
There is a mathematical reason none of this resolves into a single fairness setting. Competing fairness conditions cannot all hold together except in special cases such as perfect prediction or identical base rates, and in youth sport the base rates are themselves shaped by relative age and maturity. The response is procedural rather than technical.
What does the pre-deployment test look like?
Six steps, deliberately procedural, none of which asks a pathway to build a model. Ask for the card, meaning documentation of the training population and subgroup performance. Check the target the model actually predicts. Replay a past cohort. Set the floor, meaning a minimum subgroup size below which the tool is treated as untested for that group. Rehearse a challenge. Schedule a rescore.
Bio-banding supplies three design rules alongside it. Separate current performance from maturity status and report them apart. Compare within bands, ranking against similar maturity and flagging large maturity gaps. Revisit, rescoring after the growth spurt and treating earlier scores as provisional. Age-group context is kept alongside, rather than replaced.
Maturity estimation is good enough to support this and not good enough to be treated as exact. The standard non-invasive equations were developed on 152 children and carry a standard error of roughly half a year, which their own authors advised treating categorically rather than as a precise figure.
What should a governing body require?
Seven things, each of which fits inside a procurement specification or a selection policy within a season. Subgroup evaluation with counts and uncertainty ranges by sex, relative age quartile and maturity band before any selection use. Establish what each tool predicts, preferring targets that do not simply reproduce past selection decisions. Set and record a minimum subgroup size. Require current performance and estimated maturity to be reported separately, with growth-spurt outputs marked provisional and given a fixed rescore date. Replay each tool on at least one past cohort. Tell families when a tool contributed, give them the child’s inputs, and name someone who can overturn it. Review annually against the four dimensions, and suspend selection use where subgroup performance cannot be shown.
Underneath all seven sits a position on which error is worse. Wrongly excluding a late developer costs more than wrongly including one, because a wrongful inclusion is corrected at the next review while an exclusion removes the athlete from the evidence altogether. In a pathway, the characteristic harm is a closed door.
The paper’s most uncomfortable open question is whether a single pathway can even detect the feedback loop in which a tool’s outputs shape the coaching that later confirms them. Its author suspects not. A coach’s hunch can be questioned over a cup of tea, and a score on a dashboard arrives looking finished.
Free download
Get the full paper
Read Maturation, Bias and the Training Set in full, with every claim carrying an evidence grade and the full reference list. Give your name and email and the PDF opens straight away.
Frequently asked questions
What should a coach ask before acting on a composite score?
Where the maturity figure is. A composite read as potential without maturity beside it is being read as something it cannot support. Any output produced during a growth spurt should carry a rescore date.
What is a minimum subgroup size and who sets it?
The number below which a tool’s performance for that group cannot be shown, so the tool is treated as untested for them. No published guidance exists for youth sport, so each pathway must set its own and record it.
What can a parent reasonably expect?
To be told a tool was used, to see their child’s inputs, to be told that outputs from that age are provisional, and to be given the name of a person with the authority to order a rescore or a fresh assessment.
Does the relative age effect apply to girls as well?
The evidence base is thinner for female and culturally diverse settings, which the original meta-analysis flagged and called for work on. Thin evidence is a reason for caution in extending a tool across sexes, rather than a reason to assume the effect is absent.
Has any commercial youth sport tool been audited by subgroup?
The paper knew of no published audit, and recorded that absence as a finding. Asking the vendor for one, and treating a refusal as an answer, remains the practical move available to a pathway.
Legal and regulatory notice. This article discusses fairness and selection practice in general terms. Data protection obligations where children’s personal data is involved sit outside its scope. Nothing here constitutes legal advice, and the provision of legal advice sits outside the terms of any engagement with the author or with Athleet.AI. The material is presented to support discussion and further review by qualified advisers.