Formal tests used to assess knowledge or ability, generally to certify attainment or to select. They are far older than modern schooling, they determine access to opportunity in most societies, and what they measure is persistently contested.
Certification records that a person has reached a defined standard, which allows employers and institutions to rely on it without retesting.
Selection allocates scarce places, which is a different purpose and produces different design: a certifying exam should distinguish competent from not, and a selecting exam must discriminate finely between people who are all competent.
Diagnosis identifies what a learner has not understood, which is formative rather than summative and is used to guide teaching rather than to record outcomes.
Accountability uses aggregate results to judge institutions or systems, which is a further purpose again and is where most of the distorting effects arise.
These purposes conflict. A test optimised for one is generally poor at another, and much criticism of examinations concerns a test being used for a purpose it was not designed for.

The imperial examination system, developed from the Sui dynasty around 605 and running with interruptions until 1905, is the earliest large-scale merit selection system.
Candidates were tested on the Confucian classics, on composition and on policy questions. Success brought entry to the civil service and substantial social standing.
The apparatus was elaborate. Candidates sat in individual cells for several days, papers were recopied by clerks so that handwriting could not identify the author, and elaborate precautions against cheating were maintained.
The system was genuinely open in principle, and in practice preparation required years of study that only some families could support, so mobility was real but limited.

European observers reported on the system from the sixteenth century, and it influenced the development of competitive civil service examinations in Britain and elsewhere in the nineteenth century, which replaced appointment by patronage.
Written examinations became standard in European universities during the nineteenth century, replacing oral disputation.
Standardised testing, with identical questions and mechanical scoring, developed in the twentieth century alongside statistical methods for analysing results.
Multiple choice was introduced for large-scale administration and is efficient to score and limited in what it can assess, particularly for reasoning that requires construction rather than selection.
Computer-based and adaptive testing adjusts difficulty according to performance, reaching a precise estimate with fewer questions.

Continuous and coursework assessment spreads judgement across a period rather than concentrating it, which reduces the effect of a single occasion and introduces different problems concerning authorship and comparability.
Reliability, meaning consistency of result, is generally good for well-constructed standardised tests and weaker for essay marking, where inter-marker agreement is a recognised problem addressed by moderation.
Validity, meaning whether the test measures what it claims, is the harder question. A test predicts later performance to some degree, and the correlations are moderate rather than strong, which means examination results carry real information and considerable error.
Preparation effects are substantial. Coaching improves scores measurably, which means results partly reflect access to preparation rather than underlying attainment, and this is the mechanism by which examinations reproduce advantage while appearing neutral.
Teaching to the test is the predictable response to high-stakes accountability use. Where results determine institutional consequences, effort concentrates on what is measured, and Campbell's law states the general case: the more a quantitative indicator is used for decision-making, the more it will distort the process it monitors.
Anxiety affects performance, and the size of the effect and how far it differs between groups is studied and disputed.
Portfolios, projects and practical assessment measure things examinations cannot, and are harder to standardise and to compare across institutions.
Teacher assessment correlates reasonably with examination results and shows systematic differences by pupil characteristics in several studies, which is the objection to relying on it alone.
The 2020 cancellation of examinations in several countries during the pandemic produced a natural experiment in substituting other methods, and the difficulties encountered, particularly with algorithmic moderation of teacher estimates, illustrated why the problem is genuinely hard rather than merely neglected.
Examinations allocate access to education and employment in most societies, which makes their design a question about opportunity rather than only about measurement.
The Chinese system also demonstrates the enduring appeal of the underlying idea. Selecting by examination rather than by birth or patronage was a substantial advance, and its limitations, that preparation is unequally available and that what is examined comes to define what is learned, were visible within that system and remain the same criticisms made today.