A free preview, generated by MasteryLoop HQ.
In 1984, educational psychologist Benjamin Bloom published a finding in Educational Researcher (13(6)) that has anchored education policy ever since: students tutored one-on-one scored about two standard deviations above students in a conventional classroom — the average tutored student outperformed roughly 98% of the untutored control group. Bloom called it "the 2 sigma problem": not because tutoring failed, but because it worked too well to ever be affordable at scale. For four decades, "two sigma" has been the number cited to justify everything from voucher programs to today's AI tutoring products.
Bloom's team compared three conditions, not two: conventional instruction; "mastery learning," meaning frequent testing plus corrective feedback with no individual tutor; and one-on-one tutoring that also included mastery learning. The famous 2.0 standard deviation result came from that third group. The second group — the testing loop alone, no tutor — landed around one standard deviation on its own. The studies behind these figures were also small, a handful of classrooms rather than a national sample.
In 2011, Kurt VanLehn published a review in Educational Psychologist (46(4)) — a meta-analysis pooling results across dozens of separate tutoring studies, covering human tutors, intelligent tutoring systems, and other formats, all measured against no-tutoring instruction. Human tutoring produced an effect size around 0.79; intelligent tutoring systems came in around 0.76. Both are substantial, but neither is close to 2.0. An effect size of 0.79 moves a student from the 50th percentile to roughly the 79th — meaningful, but not "better than 98% of peers."
An effect size of 2.0 versus 0.79 isn't a rounding error — it changes what tutoring can credibly promise. At 2.0, one-on-one attention alone would be close to a silver bullet. At 0.79, tutoring is one of the most effective interventions education research has documented, but it's not magic, and it doesn't work identically for every subject, student, or format.
Neither Bloom's original work nor VanLehn's review credits "attention" alone as the mechanism. Tutored students weren't just getting a person in the room — they were tested frequently on material they'd just covered, given specific feedback on what they'd missed, and required to fix gaps before moving forward. That loop — teach, test, correct, retest — is mastery learning, and Bloom's own no-tutor condition suggests it carries a large share of the effect by itself.
Once large language models could hold a coherent tutoring conversation, "solve the two sigma problem" became a natural pitch — software delivering Bloom's result at near-zero marginal cost. But a chatbot that adapts explanations to a student's level is only doing the easy half of tutoring, and it does that by default. The harder half — testing comprehension, refusing to let a gap slide, requiring demonstrated understanding before moving on — is a design choice, not a byproduct of using AI at all.
None of this makes tutoring — human or AI — overhyped as a category. A 0.76–0.79 effect size, found consistently across dozens of studies and across both human and software tutors, is still large by education-research standards. The honest pitch isn't "AI will make every student outperform 98% of peers." It's closer to: the mechanism behind the two sigma finding — frequent low-stakes testing, immediate corrective feedback, refusing to let students advance on a guess — is a real, well-evidenced lever AI can now implement systematically at low cost.
Whenever a learning product invokes Bloom's two sigma problem, the useful follow-up question isn't whether the AI talks like a tutor — most do. It's whether the product implements the mastery-learning loop that produced the effect: does it actually test comprehension rather than just explain, does it require correction rather than let errors pass, and does it gate progress on demonstrated understanding rather than time spent. The sharpest version of that test is the one Richard Feynman was known for — if you can't explain an idea in plain language, you don't yet understand it — because an explanation has to be constructed, while a multiple-choice answer only has to be recognized.
Bloom's 1984 two-sigma finding is real, but later evidence puts tutoring well below it. A 2011 review in Educational Psychologist found human tutors around 0.79 and software tutors around 0.76 — still substantial effects, and both driven less by personal attention than by a loop of frequent testing and forced correction of gaps.
Ask someone which number they'd heard before reading this, then ask what they think an AI tutor would actually need to do — beyond talking clearly — to reproduce Bloom's result rather than just gesture at it.
Reading a good explanation feels like understanding it. Usually it isn't the same thing — and you don't find out which one you've got until someone asks you to explain it back.