Simpson's Paradox: How Combining Groups Can Flip Every Conclusion
In 1973, UC Berkeley looked like it was running a textbook case of gender discrimination. Men applying to graduate school got in at a rate of 44.2%; women were admitted at only 34.6%. The gap was big enough that, statistically, it was very unlikely to be chance. Lawyers could smell a case.
Then someone actually checked the individual departments. In most of them, women were admitted at equal or higher rates than men. Somehow a bias in favor of women at the department level had assembled itself into a bias against women at the university level. Nobody had cooked the books. No dean had lied. The math itself had produced a mirage — and it has a name: Simpson's Paradox.
The Concept
Simpson's Paradox happens when a trend that shows up in every subgroup of your data reverses — or simply vanishes — when you combine those subgroups into one big pool. It sounds like it should be mathematically impossible. It isn't. It's just a consequence of unequal group sizes hiding inside an average.
Here's the mechanism in miniature. Imagine two departments at a university:
- Engineering (hard to get into): 20 women apply, 2 get in (10%). 100 men apply, 15 get in (15%).
- English (easier to get into): 100 women apply, 80 get in (80%). 20 men apply, 17 get in (85%).
Notice: in both departments, men are admitted at a slightly higher rate than women. Look closely, though, at who's applying where. Women are disproportionately applying to the competitive department (Engineering), and men are disproportionately applying to the easier one (English).
Now pool everyone together. Women: 82 admitted out of 120 applicants = 68%. Men: 32 admitted out of 120 applicants = 27%. The combined numbers show women getting in at a dramatically higher rate — even though men had the edge in every single department. The "hidden variable" — which department someone applied to — was doing all the work, and it vanished the moment the data got pooled.
This is the whole trick of Simpson's Paradox: an aggregate statistic quietly averages over a lurking variable that's unevenly distributed across your groups. If that variable correlates with both the thing you're measuring (admission rate) and the group sizes (how many people are in each subgroup), the pooled number can say the opposite of what every individual piece says.
Why It Matters
The paradox isn't a cute math trivia fact — it has misled real institutions, courts, and hospitals.
The actual Berkeley case. The 1975 study by statisticians Peter Bickel, Eugene Hammel, and J. William O'Connell, published in Science, looked at 12,763 applicants across 101 departments. Aggregate: 44.2% of men admitted vs. 34.6% of women. But when they broke it down by department, most departments showed either no bias or a slight bias in favor of women. The explanation: women disproportionately applied to the university's most competitive, lowest-acceptance-rate departments (humanities, with rates sometimes below 20%), while men disproportionately applied to less competitive, higher-acceptance departments (like engineering, at Berkeley in that era). The pooled number wasn't measuring discrimination — it was measuring which departments people chose to apply to.
A real medical example. A 1986 study in the British Medical Journal compared two treatments for kidney stones: open surgery (Treatment A) and a less invasive procedure, percutaneous nephrolithotomy (Treatment B). Broken down by stone size, Treatment A won both times — 93% success vs. 87% for small stones, and 73% vs. 69% for large stones. But pooled across all patients, Treatment B appeared to win, 83% to 78%. Why? Doctors tended to reserve the more invasive Treatment A for the tougher, large-stone cases, and gave the gentler Treatment B to easier, small-stone cases. Stone size — the "severity" variable — was unevenly split between the two treatment groups, and it flipped the overall verdict. A doctor reading only the pooled statistic would recommend the objectively worse treatment for a given patient.
Sports and pay statistics. Simpson's Paradox regularly appears in batting averages — a player can have a higher batting average than a teammate in both of two seasons, yet a lower combined average across both seasons, if their at-bat counts were split unevenly. It also shows up in wage-gap analyses, vaccine efficacy comparisons across age groups, and even COVID-19 case-fatality-rate comparisons between countries, where age distribution acted as the hidden confounder.
The Details
The formal name traces to a 1951 paper, "The Interpretation of Interaction in Contingency Tables," by the British statistician Edward H. Simpson — though the underlying effect had already been noticed decades earlier by Karl Pearson in 1899 and by Udny Yule in 1903. The phenomenon only picked up Simpson's name in 1972, when statistician Colin R. Blyth labeled it "Simpson's Paradox" in a paper on causal reasoning — so, fittingly, even the name of the paradox has a slightly confounded history.
Mathematically, the paradox comes down to comparing weighted averages. When you combine group A and group B, the pooled rate is a weighted average of the two group rates, and the weights are the group sizes. If the group with the higher rate (say, men in English) is heavily represented in the group with generally higher outcomes (English overall), while the group with the higher rate on the other side (women in Engineering) is a small slice of the group with generally lower outcomes, the weights can outrun the individual rates entirely. You can construct this with any set of numbers you like — pick two fractions where a/b < c/d and e/f < g/h, but (a+e)/(b+f) > (c+g)/(d+h). It's not a coincidence or a computational error; it's baked into how fractions of unevenly-sized groups combine.
Picture it as a scatter of points on two separate trend lines, each sloping downward. If you shift one entire cluster of points to the right and up relative to the other, then squash both clusters together and draw a single trend line through the combined cloud, that overall line can slope upward — even though each individual cluster's line sloped down. Graphically, this is exactly what's called a "reversal paradox": two negatively-sloped clusters, positioned so their combination reads as positive.
Statisticians today handle this with causal graphs — diagrams (popularized by computer scientist Judea Pearl) that map out which variables influence which others. The tool tells you whether to control for a variable (like department, or stone size) or whether doing so would actually introduce bias rather than remove it. This matters because Simpson's Paradox has an evil twin: sometimes stratifying the data is the right move (as in Berkeley and kidney stones), but in other structures, splitting into subgroups is what introduces a spurious reversal, and the pooled number is actually the correct one to trust. Telling these two situations apart requires knowing the causal story behind the data, not just running the arithmetic.
Takeaways
- Simpson's Paradox occurs when a trend present in every subgroup of a dataset reverses or disappears when the subgroups are combined — caused by a hidden variable unevenly distributed across the groups.
- It's not a rare curiosity: it has shown up in real gender-discrimination allegations (UC Berkeley, 1973), real medical treatment comparisons (kidney stone surgery, 1986), and routinely appears in sports statistics and wage-gap studies.
- The fix is not "always trust the pooled data" or "always trust the subgroup data" — it depends on the causal structure connecting the hidden variable to the outcome, which is why modern statisticians lean on causal diagrams rather than raw arithmetic.
- Before trusting any aggregate statistic, ask: is there a variable that differs a lot between my groups and also affects the outcome? If so, the headline number might be averaging over exactly the thing that matters.
- The paradox is a humbling reminder that "the numbers don't lie" is itself a claim that needs scrutiny — numbers report exactly what you fed them, and pooling can quietly erase the story hiding inside.
Resources: - Bickel, P. J., Hammel, E. A., & O'Connell, J. W. (1975). "Sex Bias in Graduate Admissions: Data from Berkeley." Science, 187(4175), 398–404. - Simpson, E. H. (1951). "The Interpretation of Interaction in Contingency Tables." Journal of the Royal Statistical Society, Series B. - Charig, C. R., et al. (1986). "Comparison of treatment of renal calculi by open surgery, percutaneous nephrolithotomy, and extracorporeal shockwave lithotripsy." British Medical Journal, 292(6524), 879–882.