Selecting part of a population in order to draw conclusions about the whole. It is what makes measurement of large populations possible, and how the sample is chosen matters far more than how large it is.

Examining everyone is usually impossible, prohibitively expensive, or destructive, as when testing whether a component fails.

A sample allows inference about the whole, and the inference is valid only if the sample is representative in the relevant respects.

The key requirement is that selection be governed by chance rather than by anything related to what is being measured. Random selection means every member has a known probability of inclusion, which is what allows the uncertainty in the result to be quantified.

Simple random sampling. Every member of the population has an equal chance of selection, which is the baseline against which other methods are assessed.
Simple random sampling. Every member of the population has an equal chance of selection, which is the baseline against which other methods are assessed.Credit: Dan Kernler (CC BY-SA 4.0).

Simple random sampling gives every member an equal chance. It is the conceptual baseline and is often impractical, since it requires a complete list of the population.

Systematic sampling takes every nth member from an ordered list. It is easy to implement and fails if the ordering contains a periodicity matching the interval.

Stratified sampling. The population is divided into groups and sampled within each, which guarantees representation and reduces uncertainty when the groups differ.
Stratified sampling. The population is divided into groups and sampled within each, which guarantees representation and reduces uncertainty when the groups differ.Credit: Dan Kernler (CC BY-SA 4.0).

Stratified sampling divides the population into groups and samples within each, guaranteeing that each is represented and reducing uncertainty when the groups differ from one another. Sampling can be proportional to group size, or deliberately oversample small groups so that estimates for them are precise enough to use.

Cluster sampling. Whole groups are selected and sampled within, which is far cheaper for dispersed populations and yields less precision for a given sample size.
Cluster sampling. Whole groups are selected and sampled within, which is far cheaper for dispersed populations and yields less precision for a given sample size.Credit: Dan Kernler (CC BY-SA 4.0).

Cluster sampling selects whole groups, such as schools or districts, and surveys within them. It is far cheaper when the population is geographically dispersed, and it yields less precision for the same number of respondents because members of a cluster resemble one another.

Multistage designs combine these, and most large national surveys use them.

The precision of an estimate improves with the square root of sample size, which has two consequences.

Quadrupling the sample halves the margin of error, so precision becomes expensive quickly.

Population size is nearly irrelevant once the sample is a small fraction of it. A properly drawn sample of a thousand gives similar precision for a city of a million and a country of three hundred million, which is counterintuitive and correct.

The crucial point is that size does not correct bias. A biased sampling method produces a wrong answer with high confidence when the sample is large, and increasing the sample narrows the interval around the wrong value.

The 1936 Literary Digest poll predicted a landslide for Alf Landon in the American presidential election. Franklin Roosevelt won overwhelmingly.

The poll collected over two million responses, which was enormous. Its sample came from telephone directories, magazine subscribers and vehicle registrations, which in 1936 skewed toward the wealthy, and only a fraction of those contacted responded, with respondents differing systematically from non-respondents.

George Gallup predicted the result correctly with a sample of a few thousand, selected to reflect the population. The episode is the standard demonstration that method beats size.

The 1948 Truman and Dewey polls failed differently, through quota sampling that allowed interviewers to choose respondents within categories, and through stopping polling too early.

Selection bias appears constantly in less obvious forms. Survivorship bias, in which only surviving cases are examined, is the best known: the analysis of returning aircraft in the Second World War, which concluded that armour should be added where returning planes were not hit, because those hit there did not return.

Response rates have collapsed. Telephone survey response rates that were above seventy per cent decades ago are now frequently in single digits, which makes non-response bias the dominant concern.

Low response does not automatically produce bias, and it does so if the reasons for not responding are related to what is being measured. Establishing that they are not is difficult and is the main challenge facing survey research.

Weighting adjusts results so that the sample matches known population characteristics, and it corrects only for the characteristics used. Polling errors in several recent elections have been attributed to non-response correlated with political preference in ways that standard demographic weighting did not capture.

Online panels are cheap and are not probability samples, since members opt in. Some perform well and their validity rests on modelling assumptions rather than on the mathematics of random selection.

Big data sources are frequently large and unrepresentative. Social media analysis measures people who post, which is a specific and unusual group, and volume provides no protection against that.

Sampling is what allows anything to be known about a population without examining all of it, which covers election polling, official statistics, medical trials, quality control and market research.

Its central lesson is durable and constantly ignored: a small well-drawn sample beats a large badly drawn one, and the confidence attached to a result depends entirely on how the sample was selected rather than on how much data was collected.