The set of practices by which scientific claims are tested against evidence. It is not a single fixed procedure, and its most important components are the ones that guard against the investigator's own expectations.
The version taught in schools runs: observe, form a hypothesis, predict, experiment, analyse, conclude.
It is a reasonable first approximation and it is not how most science proceeds. Real work is iterative rather than linear, hypotheses often follow data rather than preceding it, much science is observational because experiment is impossible, and discovery frequently begins with an anomaly nobody was looking for.
Fields differ substantially. Experimental physics, field ecology, epidemiology and palaeontology use methods that have little in common procedurally.
What they share is not a sequence of steps but a commitment: claims must be checkable against evidence by people who did not make them, and must be capable of being shown wrong.

Systematic observation and mathematical reasoning appear in many traditions, including Greek, Islamic, Chinese and Indian, and Ibn al-Haytham's work on optics in the eleventh century is frequently cited as an early instance of experimental testing.
Francis Bacon argued in the early seventeenth century for building knowledge from systematic observation rather than deducing it from accepted authority, and for deliberately seeking evidence that would contradict a belief.
The seventeenth century also produced the institutional apparatus: scientific societies, published journals, and the convention of describing methods in enough detail that others could repeat the work.
That last convention is the one doing most of the work. A result nobody else can reproduce is not yet a scientific result.

Testability. A claim that no observation could contradict is not being assessed by this method, whatever else may be said for it. Karl Popper made falsifiability the centre of his account, and while few philosophers now accept it as a complete demarcation between science and non-science, the practical requirement that a claim rule something out remains.
Controls. Comparing an intervention against an otherwise identical condition without it is what distinguishes an effect from a coincidence. The randomised controlled trial is the developed form.
Blinding. Where measurement involves judgement, the person measuring must not know which condition they are assessing. The N-rays and water memory capsules record what happens without it, and the 1784 commission on mesmerism, described in its own capsule, invented the technique.
Quantification and error. Reporting an estimate with its uncertainty, rather than a bare number, is what allows results to be compared and combined.
Replication. Independent repetition is the actual test, and the replication crisis capsule records what happened when a field discovered how rarely it had been done.
Peer review filters before publication and is a weaker check than it is popularly assumed to be, as its own capsule sets out.

The tidy account understates the role of accident. Penicillin, X-rays, the cosmic microwave background and several others were noticed rather than sought.
What distinguishes these from anecdotes is what followed. Fleming's contaminated plate was an accident; establishing that the mould produced a substance that killed bacteria, isolating it, and testing it was method.
Hypotheses also come from anywhere: analogy, mathematics, dreams in at least one well-known account. The method governs how a hypothesis is tested, not where it comes from, and this distinction between the context of discovery and the context of justification is standard in philosophy of science.
The method does not deliver certainty. It delivers claims that have survived attempts to refute them, which is a different and weaker thing, and the history of science is substantially a history of well-supported claims being replaced.
It is performed by people with careers, funding and prior commitments, and the sociology of that has measurable effects on what gets studied, published and believed.
Publication bias, in which positive results are more likely to be published, distorts the literature independently of any individual's conduct.
And the method says nothing about which questions are worth asking, which is decided by funders, institutions and prevailing interest.
Recognising these does not undermine the method. It is the reason for the specific safeguards, and each safeguard exists because the failure it prevents actually occurred.
The scientific method is the most reliable procedure yet found for correcting beliefs about the physical world, and its distinguishing feature is not that its practitioners are unbiased but that it is arranged so that being wrong eventually becomes visible.
The practical test of whether a claim is being handled scientifically is straightforward: what observation would its proponents accept as showing it to be wrong, and has anyone looked.