Statistical significance is a threshold used to decide whether a result is worth taking seriously. The conventional cutoff, a p-value below 0.05, governs what gets published across much of science, and a large number of statisticians think it is doing serious damage.

A p-value is the probability of obtaining a result at least as extreme as the one observed, assuming the null hypothesis is true. That definition is precise and almost universally misreported.

It is not the probability that the hypothesis is true. It is not the probability the result occurred by chance. It is not one minus the probability of replication. Each of those is a different quantity, and surveys repeatedly find that researchers, including people teaching statistics, endorse at least one of them.

Ronald Fisher, who introduced p-values in the 1920s and popularised 0.05 as a convenient threshold, while explicitly warning that no fixed level should be applied mechanically.
Ronald Fisher, who introduced p-values in the 1920s and popularised 0.05 as a convenient threshold, while explicitly warning that no fixed level should be applied mechanically.Credit: Unknown authorUnknown author (Public domain).

Ronald Fisher suggested it in the 1920s as a convenient rule of thumb, and said so plainly: it was one option among several, to be chosen by the investigator according to circumstances. He warned against exactly the mechanical use it acquired. The threshold hardened into a rule because journals needed a decision procedure, not because anything justifies 0.05 over 0.04 or 0.06.

John Arbuthnot, whose 1710 analysis of London christening records is often cited as the first significance test, well before the machinery had a name.
John Arbuthnot, whose 1710 analysis of London christening records is often cited as the first significance test, well before the machinery had a name.Credit: Godfrey Kneller (Public domain).

It answers the wrong question. A researcher wants to know how much to believe a hypothesis given the data. A p-value reports how surprising the data would be if the hypothesis were false. Those are not the same, and converting between them requires a prior the p-value does not contain.

It creates a publication cliff. With significance as a gate, results just under the threshold are published and results just over it are not, which biases the literature toward findings that cleared an arbitrary line. Analyses of published p-values find a conspicuous pile-up just below 0.05.

It invites p-hacking. Analytic choices are numerous, and trying several until one crosses the line is easy, often unconscious, and undetectable in the published paper. This is a principal suspect in the replication crisis.

Significance is not importance. A large enough sample makes a trivial effect significant. A small sample can miss a large one. The p-value carries no information about the size of an effect, which is usually what matters.

Non-significant does not mean no effect. Treating a p of 0.06 as proof of absence is one of the most common errors in applied research.

A one-tailed critical region. Where the threshold sits is a convention rather than a property of the data, and 0.05 was proposed as a rule of thumb rather than derived.
A one-tailed critical region. Where the threshold sits is a convention rather than a property of the data, and 0.05 was proposed as a rule of thumb rather than derived.Credit: Hz587 (CC BY-SA 4.0).

In 2016 the American Statistical Association issued its first ever formal statement on a specific method, warning against the misuse of p-values. In 2019 several hundred scientists signed a call in Nature to abandon statistical significance as a category, and the American Statistician devoted a special issue to alternatives.

Proposals include lowering the threshold to 0.005 for new claims, reporting effect sizes and confidence intervals as the primary result, adopting Bayesian methods that give the probability of a hypothesis directly, and requiring preregistration so that analytic choices are fixed before the data is seen.

Defenders reply that the tool is not the problem. A p-value is a well defined quantity that answers a specific question honestly, and the failures are failures of interpretation and incentive that no replacement would fix, since a Bayesian analysis can be gamed as readily. The dispute is unresolved and the 0.05 convention remains dominant in practice.