Skip to content

Method

How we decide that a number is outside the norm

The method choices, each with its price written beside it. There is no absolutely right way to read a footballer’s load: there is a declared way, and a hidden one. This is ours, declared.

1The reference is the athlete, not the team

A full-back runs twice as far as a centre-back, and a goalkeeper has nothing to do with either. Comparing a player against the squad average shows the difference between POSITIONS dressed up as a difference in load: the number is true, the question it answers is another one.

That is why the default reference is the personal norm — that athlete’s own history, not the team’s — and the order is fixed: themselves first, then their position group if it is large enough, never the team average as the yardstick for one individual. When it falls back on the group, the program says so.

The price: a new athlete has no norm until they have a history. The program says “I don’t know” instead of lending them somebody else’s.

2A Tuesday is not a Tuesday

In football the week turns around the match, and every day of the microcycle carries a different expected load by design: MD-3 is the hard day, MD-1 is the pre-match session, MD+1 is recovery. The calendar day, on its own, says nothing — a Tuesday can be MD-4 or MD+2, and those are two different worlds.

So comparisons are made between MATCHING days: this MD-3 against your previous MD-3 days. On the test season the difference is measurable — within the same MD, load variability drops by 37% on distance and by 47% on sRPE: almost half of what looked like noise was the training programme.

The labels the staff write always beat the automatic calculation: inside a mark made by hand there is a judgement about an odd week — a friendly that does not count, a postponed match — that no rule can infer.

3Not every source is measured as often

GPS and wellness are measured every day; force plates every seven or ten. A window expressed in DAYS that works for load is empty for the jump: twenty-eight days hold three tests, and from three values a standard deviation cannot be estimated.

It is a mistake we made and measured: with the load window, the rule “high load plus falling jump” stayed unassessable across the whole squad — eleven athletes out of eleven, even though each of them had thirty-four tests. Every window has its own length, and where the source is too sparse the program declares it instead of computing anyway. Since every window is derived from that metric’s real cadence — and where a norm does not yet exist the last test is compared with the previous one — ninety-five per cent of the force-plate days that used to stay silent now say something.

The price is time: a window that stretches to seventy-two days asks for seventy-two days of history, so on a sparse metric a new club sees nothing for two months. And what decides is not the source but how many values actually fall inside: the cadence is measured on the squad, not on the individual, because two team-mates judged with different rulers would no longer be comparable in the very ranking that says who to look at first.

4A gap is not a rest day

A zero is a measurement: it means that on that day the athlete did not work. A gap is the absence of a measurement, and it can mean they were away with the national team, that the GPS did not start, or that nobody uploaded the file. Confusing the two distorts ACWR, monotony and z-score invisibly — and in the more dangerous direction, because a missing load read as a zero lowers the norm and switches the alerts off.

The calendar decides: a day marked as rest with no measurements counts as zero, a training or match day with no measurements stays a declared gap. A day of unknown type behaves as before, because an incomplete calendar is the norm in a club’s first month.

On the charts the gap can be seen: the line breaks, and the width of the empty space is the length of the absence. A break of twenty-eight days is not a one-day sliver.

5When a change is real, and when it matters

These are two different questions and they have to be kept apart. The first belongs to the instrument: from the repeated trials of a test — the three trials of a CMJ are separate rows, not an average already taken — you get the typical error, and from there the smallest detectable change. Below that number, two measurements are not different: they are the same measurement repeated.

The second is a decision, not a property: how much it has to change to be worth looking at. The program proposes it from your own data, in two ways that answer two different questions, and leaves the declaring to you.

That the two numbers diverge is information, not a fault: if the noise is larger than the threshold that matters, that instrument cannot see what would interest you — and it is better to know before building a decision on top of it.

6The threshold knows how many values it stands on

A z-score reads as if the mean and standard deviation were known. They are not: they are estimated from the athlete’s history, and on a metric tested every seven days that history is a dozen values. A fixed threshold applied to such an estimate is more confident than it has any right to be.

So the thresholds are corrected for how many values the window really holds — the right distribution is a Student’s t, not a normal. On daily data the attention threshold moves from 1.5 to about 1.57 deviations; on the force plates, with twelve values, to 1.69. And the program always shows the thresholds actually used, never the nominal ones.

The price: a few fewer reds. On the test season about three per cent of the judgements change, almost all reds turning amber — and it is the honest direction, because a red built on twelve values was too sure of itself.

7Few alerts, or nobody looks at them

A system that flags forty per cent of the days flags nothing: it teaches people to ignore the colour. It happened to us too — with a percentage-change threshold the alerts reached 66% of athlete-days, because inside a microcycle the hard day sits NATURALLY twenty per cent above its reference.

Moving to the z-score against the personal norm — which divides by how much that athlete varies, and therefore absorbs the structure of the week — the alerts fell to 8.6%, and the ratio between reds and oranges straightened out. The alerts that remain cross TWO sources: high load with falling wellness, high load with falling jump, a return with load already high, asymmetry getting worse.

Every alert carries the numbers that lit it up beside it, and every rule that could not be assessed declares it, with the reason and the remedy. “I could not look” is not “everything is fine”.

8Why it does not predict, and will not

There is no risk score, no injury probability, no recommendation on fitness to play. This is not commercial caution: a system that predicts is a clinical decision support tool, with everything that entails — and prediction promises, in this field, have not held.

What the program says is: this value sits this far from this athlete’s norm, and here is the raw data behind it. The vocabulary is chosen on purpose — “outside their norm”, “falling”, “worth a look” — and it does not contain “at risk”.

Interpretation stays with the qualified practitioner, and the sentence sits INSIDE the product: next to the traffic light, and on the screen where the thresholds are configured, which is the moment when a person believes most in what they are building.

9What it takes to start

A CSV file from any vendor, and the squad. The columns are mapped once and the profile is saved; the season dates and the fixtures in the calendar are what the microcycle needs, and they go in over one afternoon.

In the first weeks the program will often say “I don’t know”, and that is correct: a personal norm needs history. With four weeks of daily load, trends and deviations begin to say something; the uncoupled ACWR asks for thirty-five days, and writes it.

You do not need to change vendor, you do not need a connector, and you do not need to decide anything before seeing the first number: the defaults are sensible, declared and all of them adjustable.

If the method convinces you, the rest takes half an hour to see.