Testing people without symptoms to detect disease early. It saves lives in specific programmes and causes harm in others, and distinguishing between the two requires arithmetic that is frequently not done.

A mammogram. Screening applies a test to people who feel well, which changes the balance of benefit and harm compared with testing someone who has symptoms.
A mammogram. Screening applies a test to people who feel well, which changes the balance of benefit and harm compared with testing someone who has symptoms.Credit: National Cancer Institute (Public domain).

Testing someone with symptoms is diagnosis. Testing someone who feels well is screening, and the difference matters because the population is different.

In a symptomatic population, the condition is common enough that a positive result is likely to be correct. In a screening population, the condition is rare, and this changes the meaning of a positive result entirely.

The criteria for a worthwhile programme were set out by Wilson and Jungner for the World Health Organization in 1968 and remain the standard reference. They require that the condition be important, that there be a recognisable early stage, that treatment at that stage be more effective than later, that a suitable test exist, and that the programme be acceptable and economically reasonable.

The condition most often overlooked is that early treatment must actually work better. Detecting a disease earlier is only useful if doing something about it earlier changes the outcome.

Consider a test that is ninety nine per cent accurate for a condition affecting one person in a thousand.

Screen a hundred thousand people. A hundred have the condition and ninety nine are correctly identified. Of the ninety nine thousand nine hundred without it, one per cent test positive, which is around nine hundred and ninety nine people.

So roughly a thousand and ninety eight test positive and only ninety nine have the condition. A positive result means about a nine per cent chance of having it.

This is not a defect of the test. It follows from the condition being rare, and it is why screening programmes generate large numbers of false positives however good the test is.

The consequences of those false positives are real: anxiety, further investigation, and sometimes invasive procedures with their own risks.

Three statistical effects make screened populations appear to do better even when screening has no effect at all, and all three must be corrected for.

Lead time bias. Detecting a disease earlier means the patient knows about it for longer. Survival measured from diagnosis increases even if the date of death does not change.

Length bias. Screening at intervals preferentially detects slow-growing disease, because fast-growing disease appears and progresses between screens. Screen-detected cases are therefore systematically less aggressive than symptomatic ones, and compare favourably for that reason alone.

Overdiagnosis is the extreme form: detecting disease that would never have caused symptoms in the person's lifetime. Every overdiagnosed case is counted as a life saved in survival statistics while receiving only the harms of treatment. The overdiagnosis capsule treats the evidence.

Because of these, survival rates are the wrong measure. The only reliable measure of whether screening works is a randomised trial with mortality as the endpoint, comparing all deaths in screened and unscreened groups.

Cervical screening. Detection and removal of pre-cancerous changes is the clearest case, since the intervention prevents cancer rather than finding it earlier.
Cervical screening. Detection and removal of pre-cancerous changes is the clearest case, since the intervention prevents cancer rather than finding it earlier.Credit: Mre80 (CC BY-SA 4.0).

Cervical screening is the strongest case. It detects pre-cancerous changes that can be removed, which prevents cancer rather than finding it earlier, and incidence and mortality have fallen substantially where programmes are established. Screening based on human papillomavirus testing has largely replaced cytology as more sensitive.

Colorectal screening has good trial evidence for mortality reduction and, like cervical screening, can remove pre-cancerous lesions.

Newborn screening for a small number of treatable metabolic conditions is highly effective, since the conditions are otherwise devastating and treatment is straightforward.

Breast screening reduces breast cancer mortality, and the size of the benefit and the extent of overdiagnosis are genuinely disputed, with credible estimates differing several-fold and reasonable people reaching different conclusions about the balance.

Prostate screening using PSA is the clearest cautionary case. It detects a great deal of cancer that would never have caused harm, and trials show at best a small mortality benefit against substantial overdiagnosis and treatment harms including incontinence and impotence. Recommendations shifted from routine screening to individual decision-making as a result.

Whole-body scans and many commercially marketed screening packages have no evidence of benefit and clear evidence of harm through incidental findings.

A mobile screening unit. Reaching people who would not otherwise attend is a large part of whether a programme achieves what its trials showed.
A mobile screening unit. Reaching people who would not otherwise attend is a large part of whether a programme achieves what its trials showed.Credit: National Institute for Occupational Safety and Health (NIOSH) from USA (Public domain).

An organised programme with defined intervals, quality assurance and follow-up outperforms opportunistic testing.

Informed choice rather than exhortation, since the balance of benefit and harm differs between individuals and communicating both is now regarded as an obligation rather than an option.

Presenting risk as natural frequencies, meaning counts out of a hundred or a thousand people, rather than as percentages or relative risks, which measurably improves comprehension as the numeracy capsule describes.

Monitoring, since a programme that works in a trial may not work as implemented.

Screening is among the few medical interventions applied to healthy people at population scale, which sets a higher standard: the default is to do nothing, and a programme must demonstrate that acting is better.

It is also the clearest practical application of conditional probability. Whether a positive result means anything depends on how common the condition is, and the widespread failure to reason about this, by patients and clinicians alike, has direct consequences for how many people are treated for disease they do not have.