Stroop Test

07 / 07

How to Read Your Stroop Test Results — Interference Effect, Accuracy and Scatter Plot

How to read the interference effect, facilitation and inhibition cost, error rate difference, mean and median, and the trial-by-trial scatter plot on the results screen — plus the limits of measurement accuracy and how many runs you need for a reliable picture.

When the test ends, you see the mean and median reaction time per condition, accuracy, the interference effect and a trial-by-trial scatter plot. Here is what each number means and how far you can trust it.

How it is calculated

  • Reaction time — The time (ms) from the frame in which the word is drawn on screen to the button or key input event. Means and medians use correct trials only.
  • Exclusion rules — Responses faster than 200 ms are treated as anticipatory presses made before judging the word, and are excluded from both reaction time and accuracy. Timeouts count as errors and are excluded from reaction time. Presses made before the stimulus appears (while the + is shown) are ignored, and only their count is reported.

Interference effect = incongruent mean − congruent mean

Facilitation = neutral mean − congruent mean · Inhibition cost = incongruent mean − neutral mean

Reading the interference effect

A positive value means you were that much slower in the incongruent condition. Most people get a positive value; if yours is close to zero or negative, first consider the possibility that it fluctuated by chance because there were few trials.

This site does not compare you with a “top X%” or with averages by age. There are no validated norms for online button-press tasks, and because latency varies from device to device, your numbers cannot be compared directly with other people's. Make comparisons only between your own results measured on the same device, in the same language, with the same settings.

Look at accuracy too

If the reaction time difference on incongruent trials is small but you made many errors, you may have traded accuracy for speed (the speed–accuracy trade-off). Check the error rate difference (incongruent − congruent) on the results screen as well. If incongruent trials are worse in both reaction time and error rate, you can consider the interference clear.

Mean and median

One or two trials where you zoned out and pressed late can swing the mean considerably. The median is less sensitive to such outliers. If the mean-based and median-based interference differ a lot, that signals an outlier — look for an isolated point high up in the scatter plot.

Reading the scatter plot

The horizontal axis is trial order and the vertical axis is reaction time. ○ is congruent, ■ is incongruent, △ is neutral, × is an error, timeout or anticipatory press, and the dashed lines are the means for each condition.

  • ■ sits higher than ○ overall — There is interference.
  • Early points are high and drop toward the end — This is a practice effect as you get used to the task. Turning on the 8 practice trials reduces it.
  • Points rise or scatter toward the end — Suspect fatigue or a drop in concentration.
  • A cluster of × — A stretch where you confused the rule or rushed.

How many runs until it is reliable?

The 24-trial setting has only 8 trials per condition, so values fluctuate a lot. If there are fewer than 5 correct trials per condition, the results screen shows a warning. A difference between two means, like interference, is noisier than each mean on its own, and Hedge, Powell & Sumner (2018) reported that the test-retest reliability of Stroop interference is lower than expected. Measuring several times with 48–72 trials and looking at the trend is better than relying on a single number.

What affects measurement accuracy

  • Input device — Wired keyboards usually have short latency, while wireless or Bluetooth keyboards and touchscreens can add longer delays.
  • Screen refresh rate — A 60 Hz screen redraws every 16.7 ms, so there is an error of about one frame in when the word actually became visible.
  • Browser and device load — If other tabs or apps are heavy, event handling can be delayed. Switching to another tab during the test stops it automatically.

Most of these delays are added equally to every condition, so the interference effect, being a difference between conditions, is less affected than absolute reaction times. Switching devices can change the absolute values, though, so compare results on the same device.

This is not a diagnosis

This test is not a standardized neuropsychological test and does not provide a medical or psychological diagnosis. Use the results only for fun and self-observation.

References

Hedge, C., Powell, G., & Sumner, P. (2018). The reliability paradox: Why robust cognitive tasks do not produce reliable individual differences. Behavior Research Methods, 50(3), 1166–1186.

MacLeod, C. M. (1991). Half a century of research on the Stroop effect: An integrative review. Psychological Bulletin, 109(2), 163–203.

Try the test →