The Turing test is a proposal, made by Alan Turing in 1950, for sidestepping the question of whether a machine can think. Rather than defining thought, Turing suggested a practical substitute: if a human interrogator conversing through text cannot reliably tell a machine from a person, the question of whether the machine is thinking has lost its force. Whether that substitution is legitimate has been argued ever since.

Turing's paper opens by dismissing the question "can machines think?" as too meaningless to deserve discussion, and replaces it with what he called the imitation game. An interrogator communicates by text with two hidden participants, one human and one machine, and tries to determine which is which. Turing predicted that by around the year 2000 machines would fool an average interrogator roughly thirty per cent of the time in a five minute conversation.

A diagram of the imitation game. An interrogator communicates by text with a hidden human and a hidden machine, and attempts to tell them apart.
A diagram of the imitation game. An interrogator communicates by text with a hidden human and a hidden machine, and attempts to tell them apart.Credit: Juan Alberto Sánchez Margallo (CC BY 2.5).
Alan Turing, who proposed the test in his 1950 paper Computing Machinery and Intelligence, deliberately replacing the question of whether machines think with a question that could be settled by observation.
Alan Turing, who proposed the test in his 1950 paper Computing Machinery and Intelligence, deliberately replacing the question of whether machines think with a question that could be settled by observation.Credit: Elliott and Fry (Public domain).

The first program to trouble people was ELIZA, written by Joseph Weizenbaum in 1966, which imitated a psychotherapist by reflecting the user's statements back as questions. It contained no understanding of any kind and was, by Weizenbaum's own account, a demonstration of how superficial the illusion could be. He was disturbed to find that people confided in it and formed attachments to it even after the mechanism was explained to them, an effect now called the ELIZA effect. Weizenbaum spent much of his later career arguing against the uses to which such systems were being put.

A conversation with ELIZA, written in 1966. The program had no understanding whatever and worked by reflecting the user's own words back as questions, yet people confided in it readily.
A conversation with ELIZA, written in 1966. The program had no understanding whatever and worked by reflecting the user's own words back as questions, yet people confided in it readily.Credit: Unknown authorUnknown author (Public domain).
Joseph Weizenbaum, ELIZA's author, who became one of the sharpest critics of the field his program helped inspire.
Joseph Weizenbaum, ELIZA's author, who became one of the sharpest critics of the field his program helped inspire.Credit: Rochester Institute of Technology (Public domain).

The central objection is John Searle's Chinese Room argument of 1980: a person who does not speak Chinese, following an elaborate rulebook to produce Chinese replies, would pass such a test while understanding nothing. Searle concludes that symbol manipulation cannot amount to comprehension. Defenders reply that understanding may be a property of the whole system rather than of the person inside it, and the exchange has not been resolved.

A second objection is that the test rewards deception rather than intelligence. A machine passes by imitating human limitations, including making arithmetic errors and feigning ignorance, so it measures a capacity for convincing impersonation that may be unrelated to reasoning.

A third, more pressing since large language models became capable of fluent conversation, is that the test may simply be easier than Turing expected. Systems now produce text that many people cannot distinguish from human writing, without anything most researchers would call understanding. Opinions divide on whether this means the machines have achieved something significant, or that the test was never a good measure. Most working researchers now treat it as a historically important thought experiment rather than a benchmark.