Simpson’s Paradox, explained.
Simpson's Paradox occurs when a relationship seen within subgroups changes direction after the groups are combined, because the groups contribute different weights.
Why it happens
A total rate mixes performance with case composition. If one option handles mostly difficult cases and another mostly easy ones, the aggregate can reverse the comparison within each difficulty level. The appropriate grouping depends on the causal question.
An aggregate comparison can tell a different story from every subgroup comparison when the groups have different weights. Edward Simpson analyzed how combining contingency tables changes interpretation.
Read the result
Balance the case mix while holding subgroup rates fixed. If the overall winner changes, the comparison was partly driven by weights rather than a change in either program's subgroup performance.
A worked example
Two programs handle different cases
Program A succeeds at 90% of easy cases and 30% of hard cases; B succeeds at 80% and 20% respectively.
If A has only 10% easy cases, its total is 36%. If B has 90% easy cases, its total is 74%.
A leads in each subgroup but trails overall. The different case mixes explain the reversal.
OPTIONAL DEEPER DETAILGo deeper: inside the model
Inside this model
In each difficulty group A succeeds 10 percentage points more often: easy 90% versus 80%; hard 30% versus 20%. Overall A = 0.30+0.60×easyA and B = 0.20+0.60×easyB, with shares expressed as fractions. These are illustrative expected rates for equal-sized programs.
Where this idea is useful
A practical use
Before judging two hospitals or schools by one headline rate, compare the difficulty of the cases each accepts.
A common misconception
“The subgroup result is always the correct result.”
Aggregation and conditioning answer different questions. Choosing what to control requires context and, for causal claims, a causal model.
What this explanation leaves out
- This toy comparison is not a causal estimate. Whether to pool or split real data depends on how groups arise, what was measured, and the question being asked.
How can I recognize a misleading aggregate?
Look for groups with different baseline difficulty and different proportions across the compared options. Compute rates within groups and then check how the weights affect the combined result.
Is the headline rate comparing performance, case mix, or both?
Associated thinkers
Further reading
Explore the original research or the teaching reference behind this experiment.