How much of the variation in measured intelligence is attributable to genetic variation is among the most technically misunderstood questions in science, and among the most politically loaded. Much of the public argument turns on a statistical term that does not mean what it is usually taken to mean.

This capsule sets out what the measurements show, what they cannot show, and where researchers genuinely disagree. It does not adjudicate, and several of the disputes below are unresolved within the field itself.

Heritability is the proportion of the variation within a particular population, in a particular environment, that is statistically associated with genetic variation. Four consequences follow, and each is routinely lost in popular accounts.

It is not the proportion of a trait that is genetic in an individual. It says nothing about any one person.

It is a property of a population, not of a trait. The same trait can have different heritability in different populations, and the figure changes when the environment changes. If every child received identical schooling, heritability of educational outcomes would rise, because environmental variation would have been removed.

High heritability does not mean fixed. Height is highly heritable and rose substantially across the twentieth century with nutrition. Vision is heritable and corrected with glasses.

And heritability within groups says nothing about the cause of differences between groups. This is the most consequential misunderstanding, and it is a matter of statistical logic rather than politics: seed sown in poor and rich soil will show high heritability of height within each plot while the difference between plots is entirely environmental.

Twin and adoption studies, and more recently direct genomic methods, consistently find that a substantial share of variation in IQ scores within studied populations is associated with genetic variation. Estimates rise with age, from roughly a quarter in early childhood to something over half in adulthood, a pattern known as the Wilson effect and generally attributed to people increasingly selecting environments that suit their dispositions.

Twins. Comparing identical and fraternal pairs is the classical method for estimating heritability, and it rests on assumptions about shared environment that are themselves debated.
Twins. Comparing identical and fraternal pairs is the classical method for estimating heritability, and it rests on assumptions about shared environment that are themselves debated.Credit: St2671 (CC BY-SA 4.0).

Do the methods hold? Twin studies assume identical and fraternal pairs share environments to a similar degree, which critics argue is false since identical twins are treated more alike. Defenders point to adoption studies and to twins reared apart as independent checks; critics note those samples are small and non-random.

The missing heritability problem. Genome-wide studies with very large samples identify enormous numbers of variants each with a minuscule effect, and polygenic scores for cognitive traits currently predict far less variance than twin studies imply should be there. Whether this reflects rare variants, gene-environment interaction, or an overestimate from the twin method is unsettled.

What is being measured. IQ tests predict educational and occupational outcomes reasonably well, which is not in dispute. Whether they measure a general cognitive capacity, or a culturally specific set of skills that correlates with success in societies organised like ours, is disputed, and the Flynn effect, the substantial rise in scores across generations far too fast for genetic change, is central to that argument.

Group differences. Observed average differences between populations are documented; their causes are not established, and the between-group inference does not follow from within-group heritability. Most researchers attribute them to environmental and historical factors including poverty, discrimination, education, and test context. The question has a long history of being answered before it was investigated, in service of conclusions decided in advance, which is a reason for particular care rather than for silence.

Alfred Binet, who devised the first practical intelligence test in 1905 to identify children needing additional help, and who warned explicitly against treating the resulting number as a fixed measure of a person's capacity.
Alfred Binet, who devised the first practical intelligence test in 1905 to identify children needing additional help, and who warned explicitly against treating the resulting number as a fixed measure of a person's capacity.Credit: Thomas Hunt Morgan (Public domain).

Alfred Binet, whose 1905 scale is the ancestor of modern IQ testing, designed it to identify children who needed extra teaching, and objected to the idea that intelligence is a fixed quantity a test could measure once and for all. That warning was issued at the outset and has been disregarded regularly since.