
Chiranjeevi Maddala
August 20, 2026
Every AI-in-education pitch now cites a study. Fewer mention that the same body of research also documents, in controlled trials, students getting measurably worse at exactly the thing the tool was supposed to help with. Both things are true, and the difference between them is not the technology. It is a design choice — one that turns out to be checkable before a school ever signs a contract.
In 1984, educational psychologist Benjamin Bloom published a finding that has quietly shaped every tutoring product built since, including the AI ones. Comparing conventional classrooms, classrooms with structured feedback, and one-to-one human tutoring, Bloom's students found that the average tutored student outperformed roughly 98 percent of students in a conventional classroom — a gap of about two standard deviations, or "two sigma." Bloom called it a problem, not a triumph, because individual tutoring for every child is too expensive to deliver at any scale a public school system can afford.
That number — two sigma — has been the explicit target of education technology ever since, from Intelligent Tutoring Systems in the 1990s through today's AI tutoring platforms. It is also the yardstick this brief uses, because it is the one figure that lets very different studies, from different decades and different countries, be compared on the same scale.

Figure 1. Bloom's two-sigma effect: average student's percentile rank by instructional method (Bloom, 1984).
For most of AI tutoring's history, the evidence was observational — usage correlated with better grades, but usage was also chosen by more motivated students and better-resourced schools, which made the causal claim shaky. That changed with a run of randomized controlled trials, the design that lets researchers actually isolate the effect of the tool itself.
The best-known is Kestin and colleagues' 2025 Harvard study, published in Scientific Reports, which randomly assigned physics students to a custom AI tutor or to an active-learning classroom taught by experienced instructors — previously the strongest known classroom method. The AI-tutored group learned more, and did so in less time, than the active-learning group. A separate 2025 review by Létourneau and colleagues, covering 28 quasi-experimental studies and roughly 4,600 students, found generally positive effects from intelligent tutoring systems on learning outcomes, with the strongest gains when the system replaced whole-class instruction rather than supplementing it. An exploratory randomized trial run directly in UK classrooms reached a similar conclusion: AI tutoring, deployed carefully, supported students safely and effectively. A separate randomized trial embedding a Socratic-style AI tutor inside a science-inquiry model for K-12 students found gains in critical thinking and self-regulation specifically — not just test scores.
Four independently run trials, four different countries and age groups, one consistent finding: a well-designed AI tutor can match or exceed strong human instruction. The open question was never whether AI tutoring could work. It was which designs make it work.

Figure 2. Randomized controlled trial design used in Kestin et al. (2025): students randomly assigned to an AI tutor or an active-learning classroom, then compared on a common post-test.
It is worth sitting with the Harvard result for a moment longer, because the comparison group is the part vendors usually skip past. The researchers did not compare AI tutoring to a lecture, or to no intervention — they compared it to active-learning instruction, the method that decades of physics-education research already considered the strongest classroom alternative to one-to-one tutoring. Beating that bar, in a randomized design, with less time spent, is a materially different claim than beating a lecture. It is also, notably, a study run on college undergraduates in an introductory physics course — a population and a subject a considerable distance from a Grade 6 classroom in Raipur or Dehradun, which is exactly why the K-12-specific trials matter as a separate line of evidence rather than an afterthought.
The Socratic-AI trial fills part of that gap directly. Run inside an Argument-Driven Inquiry framework — a model where students construct a claim, defend it with evidence, and revise it under challenge — the AI was deliberately restricted to asking questions during argument construction and evidence evaluation, rather than supplying conclusions. Comparing three conditions (AI-supported inquiry, teacher-led inquiry without AI, and no intervention), the trial measured scientific reasoning and self-regulation, not just a test score, and found the AI-supported group ahead on both. The design constraint — question-asking, not answer-giving — is doing real work in that result, and it previews the mechanism this brief returns to in Section 4.
The same 2025–26 wave of trials also produced a less comfortable finding, and it comes from a study run at a scale most vendors would want on their homepage: Bastani and colleagues' randomized trial across roughly 1,000 high-school students in a real mathematics classroom, published in 2025. Students given unrestricted access to a GPT-4-based assistant during practice problems performed 48 percent better while assisted than a no-access control group. On a later exam, taken without the assistant, the same students performed 17 percent worse than the control group. The tool did not fail to help — it helped enormously in the moment, and that same help appears to have substituted for the cognitive work that makes learning stick, an effect consistent with decades of cognitive-science research on why active engagement with material produces more durable learning than passively receiving a correct answer.
A second, 2026 trial by Liu and colleagues, run across three separate experiments, reached a compatible conclusion from a different angle: AI assistance reliably improved in-the-moment math and reading performance, but reduced students' persistence and performance on subsequent tasks attempted without assistance. Put plainly — the trials that measure only assisted performance and the trials that also measure unassisted performance afterward are measuring two different things, and a tool can score well on the first while quietly costing a school on the second.

Figure 3. The assist-versus-test reversal: unrestricted GPT-4 access during practice raised assisted performance but lowered later unassisted exam performance, relative to a no-access control (Bastani et al., 2025).
Before AI enters the picture at all, tutoring research has already documented a second, separate warning that any school scaling a pilot to a full campus rollout needs to sit with. Nickow, Oreopoulos, and Quan's 2020 meta-analysis of PreK-12 tutoring experiments — one of the largest such reviews ever conducted — found a pooled effect of roughly 0.37 standard deviations across the studies it covered; a later, expanded version of the same review put the figure closer to 0.29 SD. Both are substantial. Neither is what shows up once a program leaves the hands of the research team that designed it.
A 2024 follow-up by Kraft, Schueler, and Falken, examining 265 to 282 randomized trials, deliberately re-weighted the evidence toward the question school leaders actually face: not "does tutoring work in a tightly run study," but "what should we expect once a program runs at the scale of a real district." Pooled effects that better matched large-scale, standardized-test-focused programs came out to only a third to a half the size of the full-sample average — and independent, real-world evaluations of post-pandemic U.S. tutoring initiatives, examined using a value-added framework, found first-year effects that were often statistically indistinguishable from zero, below 0.04 SD. The same paper notes that Bloom's own original 2-sigma studies have separately come under scrutiny from other researchers for how well two small 1980s doctoral experiments generalize to anything run at scale today.
This is not a reason to distrust the RCTs in Section 2 — it is a reason to ask a different question at the pilot stage. Not only "did this work for the class that tried it," but "what, specifically, about how it was run would still be true at ten times the size."

Figure 4. Pooled tutoring effect sizes shrink as the evidence base moves from small-scale RCT meta-analysis to real-world, district-scale rollout (Nickow et al., 2020/2024; Kraft, Schueler & Falken, 2024).
Kraft and colleagues go on to identify a bundled set of design features — high tutor-to-student ratios, tight alignment between tutoring content and classroom curriculum, consistent scheduling, and strong quality monitoring — that appeared to partially protect programs from this attenuation. None of the four is specific to AI, and none of them is automatically present just because a tool is AI-powered. If anything, they are a second, independent checklist layered on top of the design questions in Section 6.
The mechanism behind the assist-versus-test reversal is not mysterious, and recent technical research names it directly. Large language models are trained, above all, to be helpful — which in a tutoring context means they default to producing the full worked solution as soon as a student asks, rather than working through a hint sequence first. A 2025 study of an LLM tutor supporting data-science exercises measured this directly across more than 500 tutoring turns and found the model routinely disclosed complete solutions before the student had made a first attempt at the problem, undermining the productive struggle the exercise was designed to create. The researchers' proposed fix — a scaffolding-first constraint layer that enforces a hint ladder and delays full solutions until an attempt is observed — is a design decision, not a capability limitation. Nothing about the underlying model prevents good tutoring design. Most default configurations simply do not choose it.
The scaffolding-first design referenced above is not an abstraction — it is a specific, describable sequence, and it is worth spelling out because it is also the clearest thing a school can ask a vendor to demonstrate live. A student asks a question. A well-scaffolded tutor's first move is a clarifying or diagnostic question of its own — what have you tried, or what part is confusing — rather than a solution. Its second move, if the student is still stuck, is a partial hint that narrows the problem without resolving it: a relevant formula named but not applied, a similar worked example without the final answer, a question that isolates which step is the actual sticking point. Only after the student has made a genuine attempt at that narrowed step does a full explanation become the appropriate next move — and at that point, research on worked examples suggests it usually is appropriate; the goal is not struggle for its own sake, but struggle sequenced ahead of explanation rather than instead of it.
A poorly-scaffolded tutor collapses this into one step: question in, full solution out. It will still test well on satisfaction surveys, because students — like anyone — generally prefer the version that removes friction immediately. That preference is precisely why assisted-performance metrics and satisfaction scores are the wrong instruments for catching this failure mode, and why Section 3's assist-versus-test reversal had to be measured with a delayed, unassisted exam to show up at all.

Figure 5. A scaffolding-first hint ladder: full explanations are withheld until the student has made an attempt at a narrowed step (after Puech et al., 2025).
None of this should be surprising to anyone who has followed Manu Kapur's research on what he termed productive failure, built over more than fifteen years and dozens of classroom studies, mostly in Singapore. The consistent finding: students who first struggle with a problem before receiving instruction — provided that struggle is scaffolded, not abandoned — develop deeper and more transferable understanding than students given the instruction up front. A 2024 synthesis of 53 such studies, reported in Education Week, found the benefit held broadly across grade levels and subjects, with one honest exception worth naming directly: students in grades 2 through 5 did not show the same advantage, largely because the instructional materials used at that age rarely built early problem-solving into a coherent sequence. Struggle helps learning only when it is designed for — thrown at students without scaffolding, it simply produces failure, full stop.
The AI tutoring trials of 2025 and the classroom research of the previous two decades are describing the same underlying rule from two different directions: help that arrives too early removes the exact difficulty that produces learning.
Taken together, the research points to a short, concrete checklist — one that does not require a data-science background to apply during a vendor demo.
Does it show its work, or does it show the answer? Ask the tool a curriculum question live, and watch whether it asks a clarifying or diagnostic question back before it answers, or goes straight to the solution.
Has assisted performance ever been tested against unassisted performance? A tool that only reports engagement or in-app scores has not answered the question that matters. Ask directly whether the vendor has measured what happens on an unassisted exam afterward.
Was the pedagogy designed by learning scientists, or bolted on by engineers after the fact? Scaffolding-first behavior — hint ladders, delayed solutions, questions instead of answers — has to be a deliberate constraint on the system. It rarely happens by default.
Does the evidence come from a controlled trial, or from a testimonial? Usage statistics and satisfaction quotes are not the same claim as a randomized trial, and the 2025–26 literature shows the two can point in opposite directions for the same underlying tool.

Figure 6. Not all evidence carries equal weight — a testimonial and a meta-analysis of randomized trials are different classes of claim.
Is there a published account of where it did not work? Kapur's own research found the grades-2-to-5 exception because his team looked for it. A vendor, or a school, willing to publish a disappointing pilot alongside the good ones is a stronger signal than a vendor who only publishes wins.
Does the pitch describe a pilot, or a design built to survive scale? A tutor-to-student ratio, a curriculum-alignment process, and a quality-monitoring routine are boring compared to a feature list, but they are the specific ingredients Kraft, Schueler, and Falken found separated programs that held their effect size at scale from programs that didn't.
A research brief that only lists favorable findings is doing marketing with footnotes, so it is worth being explicit about what this body of evidence has not yet established. Most of the strongest randomized trials cited here — Kestin et al.'s physics study chief among them — were run over days or weeks, not full academic years, so what they show is a real, causal learning effect over a short window, not proof that the gain holds up or compounds over a full term. Several were also run in single-subject settings — physics, high-school mathematics, data-science exercises — and effects that hold in a subject built on procedural, checkable steps do not automatically transfer to a subject like history or literature, where correctness is a matter of argument quality rather than a single right answer.
There is also, as Section 3's own scale-effects discussion shows, a wide and well-documented gap between what a research team achieves running a tightly monitored pilot and what a school administration achieves running the same tool across an entire campus with ordinary staffing. None of this undermines the design principles this brief draws out — a hint ladder is a hint ladder regardless of subject or duration — but it is the honest reason a single successful pilot, ours included, is evidence of a promising design, not proof of a solved problem.
We built Cypher around the specific failure mode this brief describes, not as an afterthought — it is designed to ask questions back and surface a student's own reasoning gaps, rather than complete the work for them, for exactly the reasons Bastani et al. and Liu et al. found in 2025 and 2026. Morpheus is built on the same premise for teachers: it should return time, not replace the professional judgment that decides when a student needs more struggle rather than less.
We will not claim a published randomized trial of our own yet — we do not have one, and the honest answer is that almost no school-vendor in this category does. What we do publish, deliberately, is every pilot report as it actually came out, including the ones with mixed results, on our own blog. A three-day pilot is not a randomized trial and we don't present it as one — but it is the same instinct that runs through this brief: the evidence should be checkable, not just claimed.
AI tutoring is not, on the current evidence, a settled win or a settled harm — it is a design-dependent technology, and the 2025–26 research finally has the controlled trials to show exactly where the line sits. A school evaluating any AI tool, ours included, is better served asking about hint ladders and unassisted-performance data than about feature lists. Bloom's two-sigma gap has not been closed by a press release. The research suggests it can be narrowed — by the tools that were built to make students struggle a little longer, not a little less.
See the design principles in this brief applied in your own classroom.
AI Ready School runs a free, structured 3-day pilot — one class, one teacher, zero cost — that ends with a measurable Impact Report for your school leadership.
Apply for the Free 3-Day Pilot