← Back to MasteryLoop HQ
MasteryLoop HQ · Cliff Notes preview

The Two-Sigma Problem: What Bloom Actually Found

A free preview, generated by MasteryLoop HQ.

1 — The number everyone quotes

In 1984, educational psychologist Benjamin Bloom published a finding in Educational Researcher (13(6)) that has anchored education policy ever since: students tutored one-on-one scored about two standard deviations above students in a conventional classroom — the average tutored student outperformed roughly 98% of the untutored control group. Bloom called it "the 2 sigma problem": not because tutoring failed, but because it worked too well to ever be affordable at scale. For four decades, "two sigma" has been the number cited to justify everything from voucher programs to today's AI tutoring products.

Why it matters: This is the number most people repeat, and it sits well above what the broader body of later research found — which is the story this piece exists to untangle.

2 — What Bloom's studies actually tested

Bloom's team compared three conditions, not two: conventional instruction; "mastery learning," meaning frequent testing plus corrective feedback with no individual tutor; and one-on-one tutoring that also included mastery learning. The famous 2.0 standard deviation result came from that third group. The second group — the testing loop alone, no tutor — landed around one standard deviation on its own. The studies behind these figures were also small, a handful of classrooms rather than a national sample.

Why it matters: Roughly half the effect showed up without any tutor at all, from the testing loop by itself — which is the detail the popular one-sentence version of this finding almost always drops.

3 — The broader evidence that came later

In 2011, Kurt VanLehn published a review in Educational Psychologist (46(4)) — a meta-analysis pooling results across dozens of separate tutoring studies, covering human tutors, intelligent tutoring systems, and other formats, all measured against no-tutoring instruction. Human tutoring produced an effect size around 0.79; intelligent tutoring systems came in around 0.76. Both are substantial, but neither is close to 2.0. An effect size of 0.79 moves a student from the 50th percentile to roughly the 79th — meaningful, but not "better than 98% of peers."

Why it matters: This was never a failed attempt to redo Bloom's experiment. It was a much larger and more varied body of evidence arriving at smaller numbers — a different, more honest story than either "it's true" or "it's a myth."

4 — Why the gap matters more than it sounds

An effect size of 2.0 versus 0.79 isn't a rounding error — it changes what tutoring can credibly promise. At 2.0, one-on-one attention alone would be close to a silver bullet. At 0.79, tutoring is one of the most effective interventions education research has documented, but it's not magic, and it doesn't work identically for every subject, student, or format.

Why it matters: For anyone building or marketing an AI tutoring product, this distinction isn't academic — claiming the inflated number sets an expectation the product then has to live up to, and probably can't.

5 — What was actually driving the results

Neither Bloom's original work nor VanLehn's review credits "attention" alone as the mechanism. Tutored students weren't just getting a person in the room — they were tested frequently on material they'd just covered, given specific feedback on what they'd missed, and required to fix gaps before moving forward. That loop — teach, test, correct, retest — is mastery learning, and Bloom's own no-tutor condition suggests it carries a large share of the effect by itself.

Why it matters: The delivery mechanism (a human) and the pedagogical mechanism (forced correction of gaps) are two different things. And because Bloom's tutored students had to clear a 90% mastery bar while the classroom mastery-learning group only had to clear 80%, some of the remaining gap may be the standard being enforced, not the tutor at all.

6 — Why this became the pitch for AI tutoring

Once large language models could hold a coherent tutoring conversation, "solve the two sigma problem" became a natural pitch — software delivering Bloom's result at near-zero marginal cost. But a chatbot that adapts explanations to a student's level is only doing the easy half of tutoring, and it does that by default. The harder half — testing comprehension, refusing to let a gap slide, requiring demonstrated understanding before moving on — is a design choice, not a byproduct of using AI at all.

Why it matters: A product can use AI and still fail to reproduce what made Bloom's tutored students outperform their peers, if it skips the testing-and-correction loop.

7 — What the corrected number still supports

None of this makes tutoring — human or AI — overhyped as a category. A 0.76–0.79 effect size, found consistently across dozens of studies and across both human and software tutors, is still large by education-research standards. The honest pitch isn't "AI will make every student outperform 98% of peers." It's closer to: the mechanism behind the two sigma finding — frequent low-stakes testing, immediate corrective feedback, refusing to let students advance on a guess — is a real, well-evidenced lever AI can now implement systematically at low cost.

Why it matters: The corrected number is still a strong case for building tutoring tools — it's just a different, more defensible case than the original headline figure.

8 — The question worth asking about any tool that cites this

Whenever a learning product invokes Bloom's two sigma problem, the useful follow-up question isn't whether the AI talks like a tutor — most do. It's whether the product implements the mastery-learning loop that produced the effect: does it actually test comprehension rather than just explain, does it require correction rather than let errors pass, and does it gate progress on demonstrated understanding rather than time spent. The sharpest version of that test is the one Richard Feynman was known for — if you can't explain an idea in plain language, you don't yet understand it — because an explanation has to be constructed, while a multiple-choice answer only has to be recognized.

Why it matters: That loop, not the presence of a chat window, is what the research actually supports — and it's the standard any tutoring product, AI or human, should be measured against.

The big picture

Bloom's 1984 two-sigma finding is real, but later evidence puts tutoring well below it. A 2011 review in Educational Psychologist found human tutors around 0.79 and software tutors around 0.76 — still substantial effects, and both driven less by personal attention than by a loop of frequent testing and forced correction of gaps.

A headline figure later evidence didn't supportTesting loop, not personal attentionJudge the loop, not the citation

Ask someone which number they'd heard before reading this, then ask what they think an AI tutor would actually need to do — beyond talking clearly — to reproduce Bloom's result rather than just gesture at it.

You just read about it. Could you teach it?

Reading a good explanation feels like understanding it. Usually it isn't the same thing — and you don't find out which one you've got until someone asks you to explain it back.

Deep Study: The Two-Sigma Problem: What Bloom Actually Found Start → Free — ten minutes and you'll know exactly what stuck.
Get the quick version → Free — takes under a minute to set up
More Cliff Notes →