Naming the colour a word is printed in becomes markedly harder when the word itself names a different colour. The delay is a few hundred milliseconds, it is essentially universal, and it is one of the most reliably replicable findings in psychology.

Stroop stimuli. Naming the ink colour is fast when word and colour agree and reliably slower when they conflict, by a few hundred milliseconds.
Stroop stimuli. Naming the ink colour is fast when word and colour agree and reliably slower when they conflict, by a few hundred milliseconds.Credit: Unreferierbar (CC0).

John Ridley Stroop published the effect in 1935 in his doctoral work. Participants named the ink colour of words under three conditions: colour words printed in matching ink, colour words printed in conflicting ink, and neutral stimuli.

Naming the ink of a conflicting word took substantially longer and produced more errors. Reading the word aloud, by contrast, was unaffected by the ink colour.

That asymmetry is the crux. Reading interferes with colour naming; colour naming does not interfere with reading.

Stroop's paper is among the most cited in the history of psychology, and the effect has been replicated thousands of times across languages, writing systems, ages and populations. In a field where replication has been a serious problem, the Stroop effect is one of the results that never fails.

The dominant account is automaticity. Reading, for a literate adult, is not optional. The word's meaning is extracted whether or not it is wanted, and it arrives faster than the colour name can be retrieved.

The conflict is at the level of response. Both processes converge on the same output, saying a colour name, and one of them must be suppressed. Suppression takes time, and the time it takes is the effect.

Colour filters. Naming a colour is a slower operation than reading a word for a literate adult, which is the asymmetry the effect depends on.
Colour filters. Naming a colour is a slower operation than reading a word for a literate adult, which is the asymmetry the effect depends on.Credit: User Stefan.lila on sv.wikipedia (CC BY-SA 3.0).

Several observations support this. The effect is small in children who are learning to read and grows as reading becomes fluent. It is reduced when the word is printed in an unfamiliar script the reader knows less well. And it is smaller in a second language than a first, in proportion to fluency.

Variants generalise the finding well beyond colour. The numerical Stroop task uses digits of different physical sizes, so that a large 2 beside a small 8 slows judgement of which is numerically larger. The emotional Stroop uses threatening words, which slow colour naming in anxious participants specifically. In each case an automatic process intrudes on a controlled one.

The task is a standard measure of executive function, specifically of inhibitory control, and it is used in clinical assessment and in a great deal of research on attention.

Neuroimaging consistently implicates the anterior cingulate cortex, which shows increased activity on conflicting trials and is generally interpreted as detecting the conflict, and the dorsolateral prefrontal cortex, which is interpreted as implementing the control.

Performance declines with damage to prefrontal regions, in schizophrenia, in ADHD, and with normal ageing. It is sensitive to fatigue and to alcohol. Its reliability as an individual difference measure is weaker than its reliability as a group effect, which is a general problem for cognitive tasks used as personality-like measures and is worth stating: the effect is rock solid, the score's usefulness for characterising a particular person is much less so.

Two claims frequently attached to the effect go beyond it.

That it measures a single unitary capacity called inhibition is doubtful. Different conflict tasks correlate with each other far more weakly than a single underlying ability would predict, which is one of the more awkward findings in the executive function literature.

That training on Stroop-like tasks improves general self-control is not supported. Improvement on the trained task is reliable; transfer to anything else is not, which is the same pattern found across most cognitive training research.

The effect's real significance is more basic and more interesting than either. It demonstrates that a substantial part of mental processing is not under voluntary control, that skills become automatic in a way that cannot be switched off, and that attention works by suppressing what is already happening rather than by selecting what to start.