Regression to the mean is the tendency for an extreme measurement to be followed by one closer to average, for no reason other than that extremes usually involve some luck. It is a statistical certainty, it is invisible without deliberate attention, and it produces false conclusions across medicine, education, sport, and management with remarkable regularity.

Any measurement combines a stable component and a fluctuating one. An unusually high result is more likely to have had favourable fluctuation contributing to it than an average result. Fluctuation does not persist, so the next measurement will on average be less extreme. The stable part has not changed and nothing has been corrected; the arithmetic simply reasserts itself.

The effect is stronger the less reliable the measurement, and it is absent only if the measurement is perfectly reliable.

Francis Galton, who identified the effect while studying the heights of parents and children and originally called it regression toward mediocrity.
Francis Galton, who identified the effect while studying the heights of parents and children and originally called it regression toward mediocrity.Credit: not stated (Public domain).

Francis Galton found it in the 1880s studying heredity: tall parents had tall children, but on average less tall than themselves. He initially read it as a biological force pulling toward an average type, and only later recognised it as a property of imperfectly correlated measurements rather than a fact about inheritance.

Galton's quincunx, which he built to demonstrate how a distribution arises from accumulated small chance effects. The device made visible the mechanism behind the pattern he had found in his heredity data.
Galton's quincunx, which he built to demonstrate how a distribution arises from accumulated small chance effects. The device made visible the mechanism behind the pattern he had found in his heredity data.Credit: w: Francis Galton (Public domain).

The pattern is always the same: select on an extreme, intervene, measure again, and credit the improvement to the intervention.

In medicine, patients enter trials when their symptoms are at their worst, which is precisely when the next measurement will tend to be better regardless of treatment. This is a major part of what an untreated control group exists to capture, and it is distinct from the placebo effect although the two are routinely conflated.

In education, the worst-performing schools are selected for an intervention and improve. In business, the worst quarter is followed by a management change and a better quarter. In sport, the phenomenon known as the second-season slump follows an exceptional first season.

Daniel Kahneman described a striking case with Israeli flight instructors, who had observed that praising a good landing was followed by a worse one and criticising a bad landing by a better one, and had concluded that criticism works and praise does not. The pattern is what regression predicts with no effect from either.

Three defences. Use a control group selected the same way, so regression affects both arms equally and cancels in the comparison. Do not select on the same measurement used to assess the outcome; use an independent measure or an earlier one. And where selection on an extreme is unavoidable, expect regression and estimate its size rather than reading the whole change as an effect.

The most consequential lesson is that a before-and-after comparison on an extreme group is not evidence of anything, which invalidates a large amount of published evaluation.