Assign people to treatments by chance, and the groups differ only by chance. It is a simple idea, it is the most reliable method medicine has for finding out whether something works, and it arrived remarkably late.
Comparing people who took a treatment with people who did not tells you very little, because the two groups differ in why they took it.
Sicker patients get treated more aggressively, which makes treatments look harmful. Patients who adhere to a regimen are more conscientious in every other respect, which makes treatments look beneficial. Doctors choose treatment based on prognosis, which builds the expected answer into the comparison before anything is measured.
Adjusting statistically for known differences does not fix this, because the differences that matter are frequently unknown and unmeasured.
Allocating by chance means that every characteristic, known and unknown, measured and unmeasured, is distributed between the groups by the same random process.
This is the property no other method has. Adjustment handles what has been thought of; randomisation handles what has not. Any remaining imbalance is chance imbalance, whose size is calculable and shrinks as the trial grows, which is precisely what statistical tests are built to quantify.

Blinding, when possible, removes two further sources of bias: patients who know they are treated report improvement, and assessors who know the allocation measure outcomes differently. Concealing the allocation sequence from whoever enrols participants is separately important, since a clinician who can predict the next assignment can steer patients toward it.
The 1948 Medical Research Council trial of streptomycin for pulmonary tuberculosis is generally regarded as the first.

Austin Bradford Hill designed it. Streptomycin was in extremely short supply, which made withholding it ethically defensible, and Hill used sealed envelopes with a concealed random sequence so that neither patient nor clinician could influence assignment.
The result was clear: 7 percent mortality in the treated group against 27 percent in the control. More importantly, the method was clear, and it was published in enough detail to be copied.
Hill later formulated the criteria for assessing causation from observational data, and he and Richard Doll used them in establishing that smoking causes lung cancer, a case where randomisation was impossible.
Before randomised trials, medical practice rested on authority, case series and physiological reasoning. A great many treatments in wide use turned out, once tested, to be useless or harmful.
Bloodletting persisted for two thousand years. Radical mastectomy was standard for decades before trials showed less extensive surgery worked as well. Hormone replacement therapy was believed to prevent heart disease on the basis of observational data until the Women's Health Initiative trial in 2002 found it did not. Antiarrhythmic drugs suppressed the irregular heartbeats that predicted death after myocardial infarction, and the CAST trial found they increased mortality.
Each of these was supported by plausible mechanism and by observational evidence. The pattern is consistent enough to be the central argument for the method.

Randomised trials are not a universal solvent, and treating them as one causes its own problems.
They answer average questions. A trial establishes the mean effect in the population studied, and individuals vary. Subgroup analysis to find who benefits most is statistically treacherous and generates false findings routinely.
External validity is often weak. Trial populations are selected, typically younger and healthier than the patients eventually treated, with fewer other conditions and better adherence.
Some questions cannot be randomised. Whether smoking causes cancer, whether parachutes work, whether a rare exposure is harmful: in each case randomisation is impossible or unethical, and other methods are required and are perfectly capable of establishing causation when used well.
Trials are expensive, which shapes what gets tested. Interventions with a commercial sponsor are tested thoroughly; surgical techniques, diet, exercise, and public health measures are tested far less.
And they can be gamed. Selective outcome reporting, publication of favourable results only, and analysis choices made after seeing the data all corrupt the method's guarantee. Mandatory prospective registration of trials and pre-specification of the primary outcome exist because these were common, and they remain imperfectly enforced.
The trial is the strongest tool available for the questions it can answer, and its authority rests on the randomisation itself rather than on anything about medicine. The same logic now underpins evaluation in economics, education, agriculture and public policy, where field experiments have become standard, and the 2019 Nobel Prize in Economics recognised their use in development research.