04 / 07
The Stroop in Neuropsychological Testing — What It Can and Cannot Show
How clinical Stroop tests such as the Golden version, the Victoria version and the D-KEFS are structured and used, and why a single score cannot make a diagnosis. This site is not a diagnostic tool.
Beyond the laboratory, the Stroop task has long been used in clinical neuropsychological assessment. However, the test on this site is not a standardized test and cannot diagnose or rule out any condition. What follows is general information about how clinical tests are used.
Common clinical Stroop tests
- Stroop Color and Word Test (Golden version) — Developed by Golden (1978). It counts how many items you complete in 45 seconds on each of three pages: reading color words, naming colors, and naming colors of incongruent words.
- Victoria Stroop Test — A short version used at the University of Victoria in Canada, with three cards — dots, common words and color words — of 24 items each. Administration and data are included in a compendium of neuropsychological tests (Strauss, Sherman & Spreen, 2006).
- D-KEFS Color-Word Interference Test — Part of the executive function battery by Delis, Kaplan & Kramer (2001). In addition to the basic conditions, it includes an inhibition/switching condition in which the rule changes.
Most of these tests use vocal responses, with a trained examiner timing performance and recording errors. Results are then interpreted against norms by age and education level.
What are they trying to measure?
Clinically, Stroop tests are mainly used to look at inhibitory control (the ability to suppress an automatic response), selective attention and processing speed. They are often included as part of an executive function battery and have been used in research on many conditions, including brain injury, attention-deficit/hyperactivity disorder (ADHD) and dementia.
Why a single score cannot make a diagnosis
- Several abilities are mixed together — Slow color naming may reflect a problem with inhibition, or simply processing speed, eyesight or color vision. That is why these tests measure baseline conditions such as reading and color naming separately for comparison.
- Reading ability and education — People who are less practiced readers process words less automatically, so their interference can actually come out smaller.
- Not specific to any one condition — Large interference can occur in many different conditions, so on its own it does not point to any particular illness. Fatigue, anxiety or lack of sleep alone can change it.
- Reliability of difference scores — Scores built from the difference between two conditions, like interference, are clear at the group level but have been criticized as hard to measure reliably for individual differences (Hedge, Powell & Sumner, 2018).
Clinical judgment is therefore made by a professional who integrates interviews, multiple tests, medical history and everyday functioning.
How this site differs from clinical tests
| Item | Clinical test | This site |
|---|---|---|
| Administration | Trained examiner, standard procedure | Self-administered, varies by device |
| Response | Mostly vocal | Buttons / keyboard |
| Measure | Time or item count per page (card) | Reaction time per trial (ms) |
| Interpretation | Norms by age and education | No norms; compare with your own results |
If you notice a worrying change — for example, concentrating has suddenly become difficult or your memory feels different — do not judge it by an online test result; talk to a healthcare professional.
References
Golden, C. J. (1978). Stroop Color and Word Test: A manual for clinical and experimental uses. Stoelting.
Delis, D. C., Kaplan, E., & Kramer, J. H. (2001). Delis-Kaplan Executive Function System (D-KEFS). The Psychological Corporation.
Strauss, E., Sherman, E. M. S., & Spreen, O. (2006). A Compendium of Neuropsychological Tests (3rd ed.). Oxford University Press.
Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186.