← Back to MasteryLoop HQ
MasteryLoop HQ

Why This Works

What actually makes learning stick — and why most "AI tutor" headlines miss it.

When "Personalized" Isn't "Proven"

A country just made headlines for its national AI-tutor rollout. Big pilot, promising test scores, comparisons to Germany and Sweden. It's a good story, and it's not really about AI tutoring — not on its own.

The pilot schools also got rebuilt buildings, new devices, better connectivity, and extra teacher training, all bundled together. Nobody's published a study isolating what the AI tutor itself contributed. Even the World Bank's own statement on the results was careful to call them "preliminary findings, drawn from a group of schools" — not a verdict. That's not a knock on the program — it's a reminder of something easy to miss in every AI-education headline right now: most of these claims are about delivery, not understanding.

An AI that adapts to your pace, explains things patiently, meets you where you are — that's real and valuable. But personalized delivery and verified comprehension are two different things. A test score tells you a group did better. It doesn't tell you which of them could actually explain why, or would still remember it next month, or just learned to pattern-match the test.

That gap is why MasteryLoop exists. We don't just deliver content adapted to you — we make you prove you understand it before you move on. Comprehension checks are built around why, not what, because that's the real line between being taught something and actually knowing it.

That's why a missed concept isn't a dead end. When you're stuck, Arti helps you find where the understanding broke, nudges you with hints that never hand over the answer, and explains the why. Then you try again in your own words, against the same standard.

The Two Sigma Problem

In 1984, educational psychologist Benjamin Bloom published a finding that has shaped the field ever since: students taught one-to-one performed about two standard deviations better than students in a conventional classroom. The average tutored student scored above roughly 98% of the students in the control group. Bloom called it the "2 sigma problem" — not because tutoring failed, but because it worked so well and scaled so badly. One tutor per student was never going to be affordable.

Two things usually get left out when that number gets quoted. The first is that it hasn't held up at that size. Kurt VanLehn's 2011 review in Educational Psychologist found human tutoring produced an effect size of about d = 0.79 against no-tutoring instruction, and intelligent tutoring systems about d = 0.76 — against the 2.0 and 1.0 that had been widely assumed for each. Both are real, substantial effects. Neither is two sigma. Treating the 1984 headline figure as settled science would be the same mistake as treating a press release about a national pilot as a verdict.

The second omission is more useful, and it's the part that matters here. Bloom's tutored students weren't just getting personal attention. They were getting mastery learning — tested frequently on what they'd just covered, given feedback on what they'd missed, and required to fix it before moving on. The tutoring was the delivery. The testing loop was the mechanism.

Bloom's own data is what supports that reading. He ran mastery learning as a separate condition, with no tutor at all, across all six of the original studies — and on its own it produced about one standard deviation over conventional instruction, roughly half the tutoring effect, from the testing loop alone. The tutored students were also held to a stricter bar: 90% mastery before advancing, against 80% for the classroom group. Some of the gap that remains may be the standard being enforced, not the tutor's presence. We go through the full research story, including where the famous number came from and what later evidence found, in The Two-Sigma Problem: What Bloom Actually Found.

The form that testing takes matters too. The Feynman method — named for the physicist, who was famous for insisting that if you can't explain something in plain language, you don't actually understand it — says the fastest way to find a gap in your own knowledge is to try teaching the idea to someone else. Where you go vague, hedge, or reach for jargon is exactly where the understanding isn't. That's why MasteryLoop asks you to explain concepts back in your own words rather than pick from four options: a multiple-choice answer can be recognised, but an explanation has to be constructed.

Put together, that's the whole design. Adapting an explanation to you is the easy half, and it's what most AI tutors sell. Asking you to explain it back, telling you what's still fuzzy, and not letting you advance until it isn't — that's the half Bloom's students actually had, and it's the half software can genuinely scale.

That loop doesn't stop when a course ends. Concepts you've passed come back on a spaced schedule, and the exact rules behind it — the 0-5 scoring scale, how each interval is calculated, and one disclosed deviation from the standard algorithm — are written out on the Review Scheduling page.

Sources: Bloom, B.S. (1984), Educational Researcher, 13(6), reporting Anania (1981). VanLehn, K. (2011), Educational Psychologist, 46(4).