The replication crisis is the finding that a substantial share of published research results cannot be reproduced when the studies are repeated. It surfaced most visibly in psychology and medicine from around 2010, and how deep the problem runs, and what it means, is still argued.
In 2015 the Open Science Collaboration repeated 100 psychology experiments from leading journals. Around a third to a half produced significant results in the same direction as the original, depending on the criterion used, and effect sizes were on average about half those originally reported.
Independent efforts found similar patterns elsewhere. Amgen reported being able to confirm a small minority of landmark preclinical cancer studies. Bayer reported comparable figures. The Reproducibility Project in cancer biology, completed in 2021, found effect sizes far smaller than the originals and encountered so much difficulty obtaining protocols and materials that many intended replications could not be attempted at all.
Several individually famous results have failed to hold, including social priming effects, ego depletion at its original magnitude, and the power pose. The Stanford Prison Experiment was undermined by archival work rather than replication.

John Ioannidis argued in 2005, before the empirical work, that the problem follows from the system's design. Publication favours novel positive findings, so null results go unpublished, which biases the visible literature. Small samples produce noisy estimates, and a noisy significant result overstates the effect. Analytic flexibility, later called the garden of forking paths, lets researchers reach significance without any conscious dishonesty. Replications have been hard to publish and confer little career benefit, so almost nobody attempted them.
Outright fraud exists and is not the main story. Most of the failure is explained by ordinary incentives acting on people behaving normally.

There is no agreement on how to read the results.
The pessimistic reading holds that a large part of several literatures is unreliable, that textbooks and policy built on it need revisiting, and that the failures indicate a systemic problem rather than bad luck.
The moderate reading holds that failure to replicate is not proof the original was wrong, since replications differ in population, context, and procedure, and that effect sizes shrinking toward more modest values is what an honest correction looks like rather than a collapse.
The optimistic reading holds that this is self-correction working publicly, that the fields with the worst numbers are the ones that bothered to measure, and that the reforms already adopted are substantial.
Those reforms are real: preregistration, registered reports where a journal accepts a design before results exist, data and code sharing, larger samples, multi-laboratory collaborations, and open replication. Whether they have changed the published literature enough to show up in the numbers is itself now a research question, and the early answers are mixed.
