Treatment A has a higher success rate than Treatment B among mild cases, and also among severe cases. And yet, when you combine the two groups into one, Treatment B ends up with the higher overall success rate. It sounds impossible, but it really happens — this is Simpson's Paradox.

Mild A 93% B 87% Severe A 73% B 69% Overall A 78% B 83%
A leads in both mild and severe cases separately, but B comes out ahead once they're combined

The numbers above come from a famous real study comparing kidney stone treatments (UK, 1986). Split by mild and severe cases, Treatment A (open surgery) is always better, but because the two groups have very different numbers of patients, the overall result flips once you combine them.

Why does this happen? The overall success rate isn't a simple average of the two groups' rates — it's a weighted average, weighted by how many patients were in each group. Treatment A was used on far more severe cases, which have a lower success rate, while Treatment B was used on far more mild cases, which have a higher success rate. So A's overall average gets pulled way down toward the low success rate of severe cases, while B's overall average gets pulled up toward the high success rate of mild cases. In other words, it's not a difference in how good A and B actually are — it's a difference in "who got treated more," a difference in sample composition, that flips the overall ranking. In statistics, this is called the effect of a "confounding variable," and here the hidden confounding variable is "how severe the patient's case was."

Something similar happened with 1973 graduate admissions data from UC Berkeley. Looking only at the overall acceptance rate, male applicants were accepted at a higher rate than female applicants, which raised suspicions of gender discrimination. But once the data was broken down by department, women's acceptance rates were about the same or even higher in most departments. The cause was that female applicants tended to apply more to popular, highly competitive departments with lower acceptance rates overall.

Simpson's Paradox takes its name from statistician Edward Simpson's 1951 paper, but Karl Pearson had actually noted a similar phenomenon back in 1899. It's still considered a trap to watch out for in every field that works with data — medicine, education, sports statistics, policy evaluation. That's why, whenever you look at statistics, it's important to stay in the habit of asking whether there's a subgroup you should be splitting the data by, rather than trusting the overall number alone. On our activity page, you can switch between the mild, severe, and overall views with buttons to compare them, and change the group ratios yourself with a slider to see exactly how the paradox appears and disappears.