A reaction-time task intended to measure automatic associations that a person may not report or may not be aware of. It is among the most widely taken psychological instruments in the world, and what it measures and predicts has been disputed for two decades.

The participant sorts items into categories as fast as possible using two keys. In one block, one key might be shared by a social group and by pleasant words, and the other key by a second group and unpleasant words. In another block the pairings are switched.

The measure is the difference in average response time between the two arrangements. Faster sorting when a group is paired with pleasant words than when it is paired with unpleasant words is taken to indicate a stronger positive association with that group.
The reasoning is that pairings consistent with an existing association are easier, and that this shows up as a difference of tens or hundreds of milliseconds. The task itself is not in dispute; the difference is real and reliably produced.
The test was introduced by Anthony Greenwald, Debbie McGhee and Jordan Schwartz in 1998 and developed with Mahzarin Banaji and Brian Nosek.
Project Implicit made versions freely available online and has collected many millions of completions. The instrument moved rapidly into diversity training, into public discussion of prejudice, and into court and policy argument.

Its cultural effect is hard to overstate: it supplied a concrete, quantitative meaning for implicit bias, a phrase that was rare before it and is now ordinary.
Three separate problems are often run together and should be kept apart.
Reliability. Test-retest correlations are typically around 0.4 to 0.6, which is low for an instrument used to tell individuals something about themselves. A person can take it twice and get materially different results, which means individual scores are unstable regardless of what they measure.
Predictive validity. Meta-analyses have found that scores correlate only weakly with measured discriminatory behaviour. An influential 2013 meta-analysis by Frederick Oswald and colleagues reported correlations low enough to question practical use, and a large 2019 meta-analysis by Patrick Forscher and colleagues found that procedures which do change implicit measures do not reliably change behaviour. That second finding is the more damaging one for applications, because it breaks the causal chain the interventions assume.
Construct validity. What the score reflects is not agreed. It may index a personal attitude, or knowledge of a cultural association that the person does not endorse, or familiarity, or aspects of the task itself such as the order of blocks and the salience of the categories. A person who is perfectly aware of a stereotype and rejects it may still sort faster along it.
Defenders hold that small correlations aggregated across many people and many decisions can still produce substantial disparities, that behavioural measures in this field are themselves noisy which attenuates observed correlations, and that discarding the construct because one instrument is imperfect would be an overcorrection.
Critics hold that an instrument this unreliable at the individual level cannot support the uses it has been put to, that implicit bias training built on it has not been shown to change behaviour or outcomes, and that the concept's public authority rests on a measurement claim the data do not sustain.
Both sides largely agree on the underlying numbers. The dispute is over what follows from them.
The test is a case study in the distance between a measurement and the thing it is taken to measure, and in how quickly a laboratory instrument can acquire institutional weight.
It also has a practical consequence. Considerable resources are spent on interventions premised on shifting implicit associations, and the best available evidence indicates that shifting them does not reliably shift behaviour. Whether that argues for better interventions or for a different approach to discrimination entirely is exactly what is in dispute.