Average scores on intelligence tests rose steadily throughout the twentieth century, by roughly three points per decade in the countries with long records. The rise is not disputed. What it measures is, and the effect has recently reversed in several places.

Gains on the Wechsler intelligence scales over time. The rise is steady, large over a century, and appears wherever long test records exist.
Gains on the Wechsler intelligence scales over time. The rise is steady, large over a century, and appears wherever long test records exist.Credit: Mohamed Nagdy (CC BY 4.0).

James Flynn documented the pattern across dozens of countries from the 1980s, building on scattered earlier observations. It is named for him and he was consistently careful about what it showed.

The evidence is the test manuals themselves. Intelligence tests are standardised so the population mean is 100, and they are restandardised periodically. Each restandardisation required the raw score for 100 to be raised, because the old norms made the current population look above average. Administering an old version to a modern sample produces inflated scores, and the gap grows with the age of the version.

Over the twentieth century the cumulative rise is roughly thirty points in some countries, which is the difference between an average score and the threshold once used for giftedness.

The rise is not spread evenly across the test, and the pattern is what makes the interpretation hard.

Gains are largest on abstract reasoning tasks: matrix problems, similarities, pattern completion. Raven's Progressive Matrices, designed to be culture-free and to measure reasoning as directly as possible, shows the largest gains of all.

Gains are smallest, in some cases negligible, on tests of vocabulary, arithmetic and general knowledge, which are the tests most directly reflecting schooling.

That distribution rules out several simple explanations. If people were getting better educated in the ordinary sense, the knowledge subtests should have risen most, and they rose least.

A public examination. Gains are smallest on the subtests most directly tied to schooling and largest on abstract reasoning, which is the pattern any explanation has to account for.
A public examination. Gains are smallest on the subtests most directly tied to schooling and largest on abstract reasoning, which is the pattern any explanation has to account for.Credit: The original uploader was Marcin Otorowski at Polish Wikipedia. (CC BY-SA 3.0).

Flynn's own account was that modern life demands abstraction. Work, media and schooling increasingly require classifying, hypothesising and reasoning about categories rather than manipulating concrete particulars. He illustrated it with interviews of rural Russians from the 1930s who, asked what dogs and rabbits have in common, answered in terms of use rather than category. The abstract answer now seems obvious; it is a habit of thought, and habits of thought spread.

Nutrition and health explain part of it. Better childhood nutrition, reduced disease burden, and the removal of lead from the environment all plausibly raise cognitive performance, and they predict the largest gains at the bottom of the distribution, which is where some studies find them.

Familiarity with testing accounts for some, though probably not much of a century-long trend.

Smaller families and more adult attention per child, longer schooling, and more cognitively demanding work have all been proposed. Most researchers hold that several causes operate together and that their shares are not currently separable.

Scores stopped rising in several developed countries around the 1990s and have since fallen. Norway, Denmark, Finland, France and the United Kingdom have all recorded declines in national test data, and a 2018 Norwegian study using conscription records, which cover nearly the whole male population, found the turning point within families as well as between them. That within-family finding matters: it rules out explanations based on changing population composition, such as immigration or differential birth rates.

Nobody knows why. Changes in schooling, in reading habits, in media use and in what the tests reward have all been proposed without resolution.

The effect creates a genuine difficulty for the concept. Whatever the tests measure rose by an amount that, taken at face value, would place the average person of 1920 near the bottom of today's distribution, which is not credible as a claim about their actual capability.

Alfred Binet, who developed the first practical intelligence test in 1905 and warned against treating its score as a fixed quantity of an individual.
Alfred Binet, who developed the first practical intelligence test in 1905 and warned against treating its score as a fixed quantity of an individual.Credit: Unidentified photographer (Public domain).

Two readings follow. Either the tests measure something narrower than intelligence, which changed while intelligence did not, or intelligence in a meaningful sense really did rise and the population of a century ago was substantially less able at abstract reasoning than people are now. Flynn came to hold something close to the second, restricted to the specific skills the tests reward.

The effect also has direct practical consequences. Norms must be updated, and old norms inflate scores. In the United States this affects capital cases, where an intellectual disability diagnosis bars execution, and courts have had to rule on whether scores should be adjusted for the age of the test used.