Research/AI Research/Position paper

AI tutoring is a motivation problem, not an intelligence problem

Tutors that give answers make students worse. Tutors that refuse get ignored. Smarter models fix neither.

Supersafe Research · September 2, 2026 · 9 min read

Summary

The case for AI tutors rests on a famous promise: one-to-one tutoring dramatically outperforms classroom teaching, and models can now tutor everyone. Two recent field experiments show the promise breaking in opposite directions. A tutor that hands out answers improved practice and hurt learning. A tutor that refused to hand out answers was mostly ignored. We argue that the limiting factor is not how intelligent the tutor is but whether the learner does the work, and we describe how Teach, built into every Supersafe product, is designed around that.

Two failures with one cause

Bastani and colleagues gave nearly a thousand high school math students in Turkey access to GPT-4 during practice [1]. Students with an unrestricted chat interface did 48% better on practice problems than a control group, then 17% worse on the exam taken without it. A second group used a version with teacher-written guardrails that gave hints instead of answers. They did 127% better in practice and about the same as the control group on the exam. The unrestricted tutor worked as a crutch.

Oreopoulos and Low ran the opposite design: a two-year randomized trial in 18 Tennessee middle schools, with Khan Academy's Khanmigo configured to coach rather than answer [2]. Nearly every student tried it. But the median student messaged the tutor on only about a third of the days they practiced, and students reached out in only 17% of the sessions where they made a mistake. Many messages were bare answers or clicks on suggested prompts. The program raised math achievement modestly, by about as much as Khan Academy practice without AI had in earlier studies. The authors conclude that engagement, not the technology, is the binding constraint.

Put together: give students answers and they stop learning; withhold answers and they stop asking. Neither problem gets better with a smarter model.

PERFORMANCE VS. CONTROL GROUP · ~1,000 STUDENTS · BASTANI ET AL.-25%0%+25%+50%+75%+100%+125%PRACTICE (AI AVAILABLE)Unrestricted GPT-4Unrestricted GPT-4, Practice (AI available): +48%+48%Guardrailed tutorGuardrailed tutor, same: +127%+127%EXAM (AI REMOVED)Unrestricted GPT-4Unrestricted GPT-4, Exam (AI removed): -17%-17%Guardrailed tutorAbout the same as control≈ 0 (no harm)The unrestricted assistant helped on practice and hurt once it was taken away.
Show the numbers
ConditionPracticeExam
Unrestricted GPT-4+48%−17%
Guardrailed tutor+127%≈ 0%
Figure 1 Performance relative to a control group without AI. The unrestricted assistant improved practice and hurt the exam taken without it; the guardrailed tutor avoided the harm.
TWO-YEAR RANDOMIZED TRIAL · 18 MIDDLE SCHOOLS · OREOPOULOS & LOW96%of students tried the AI tutorat least once~1/3of practice days, for themedian student, included amessage to it17%of sessions with a mistake ledthe student to ask it for help
Show the numbers
MeasureValue
Tried the tutor at least once96%
Practice days with a message (median student)about one third
Mistake sessions with a message17%
Figure 2 A tutor that refused to give answers was tried by nearly everyone and then mostly left alone.

What tutoring actually does

Bloom's "two sigma" finding, that tutored students outperformed about 98% of conventionally taught students, is the founding myth of the field [3]. VanLehn's later review of tutoring studies found much smaller effects for human tutors, about 0.8 standard deviations, and found that well-designed step-by-step tutoring systems came close to that [4]. The effective ingredients are structural: mastery before moving on, immediate feedback, and practice that makes the learner retrieve what they know. Testing yourself beats rereading for long-term retention [5]. None of these require conversation.

EFFECT ON LEARNING, IN STANDARD DEVIATIONS0.00.51.01.52.0Bloom (1984): human tutoring, claimedBloom (1984): human tutoring, claimed: 2.0 SD2.0 SDVanLehn (2011): human tutoringVanLehn (2011): human tutoring: 0.79 SD0.79 SDVanLehn (2011): step-based tutoring softwareVanLehn (2011): step-based tutoring software: 0.76 SD0.76 SDKhanmigo trial: per school yearKhanmigo trial: per school year: 0.06–0.08 SD0.06–0.08 SDDifferent studies, outcomes and durations; read as orders of magnitude, not a ranking.
Show the numbers
SourceEffect size
Bloom (1984): human tutoring, claimed2.0 SD
VanLehn (2011): human tutoring0.79 SD
VanLehn (2011): step-based tutoring software0.76 SD
Khanmigo trial: per school year0.06–0.08 SD
Figure 3 How large tutoring effects really are. The famous two-sigma result has not held up at that size; well-designed tutoring software gets close to human tutors; the AI tutor trial produced a small effect.

Why chat is the wrong default

Learning research has a name for practice that feels harder but works better: desirable difficulties [9]. Retrieving an answer from memory, spacing practice out, and struggling briefly before getting help all slow learners down in the moment and improve what they keep. Learners consistently prefer the opposite, because fluency feels like progress.

A chat box makes the easy path one sentence away: ask for the answer. That is why the unrestricted tutor hurt learning and why students abandoned the one that refused. The interface, not the intelligence behind it, sets the default behavior.

The Alpha School case

Alpha School is often cited as proof that AI can compress a school day into two hours. A detailed parent's review, published as a contest finalist on Astral Codex Ten, complicates that story [6]. By the reviewer's account, the student-facing platform contains no generative AI at all: it is mastery learning, spaced repetition and immediate feedback assembled mostly from existing apps, wrapped in an elaborate incentive system with its own currency and prize store. The reported results are strong, but they come from an expensive, selective school without a control group. Most telling, the reviewer reports that a home pilot using the same software without the surrounding system showed roughly ordinary growth.

The reviewer's own conclusion is that the motivation system, not the software, is doing much of the work. That matches the field experiments.

What Alpha gets right that chatbots miss

Set aside the marketing and the incentives, and the Alpha School review describes design choices any learning product can copy. Work is time-boxed, so effort has a clear end. Students must show mastery before moving on, so gaps don't compound. Progress is visible every day. And a person checks in regularly, which turns a private struggle into something someone notices.

None of these depend on the tutor being articulate. They depend on structure: what the learner is asked to do, when, and what happens next. That is where we spend Teach's design effort, and where we think most AI tutors underinvest.

How Teach is designed

Teach borrows its shape from Matt Pocock's open-source teach skill for coding agents [7], which treats learning as a persistent workspace rather than a chat:

  • A written mission: what the learner is trying to be able to do, and why.
  • Short lessons, around ten minutes, each ending in something to do that same day.
  • Retrieval practice: questions answered from memory before the explanation is shown.
  • Learning records that persist between sessions, including corrections when an earlier lesson got something wrong.
  • Reference sheets for what the learner will come back to, separate from disposable lessons.

Inside an Island, Teach adds one thing a standalone tutor can't: relevance. It knows what is actually in front of the learner, whether that's tonight's homework, a new hire's first real task, or a skill someone said they wanted. Lessons attach to work the learner already has to do. Hints come before answers, and answers come after an attempt, not instead of one.

AI tutoring is a motivation problem wearing an intelligence costume.

Teach in practice

  • A child stuck on fractions gets a question about the pizza they shared at dinner, not a definition. After they try, Teach shows one step. The answer comes only after a second attempt.
  • A new employee learning to quote a job gets a ten-minute lesson built from three of the business's own past quotes, then quotes a real job with Teach checking the work.
  • An adult learning statistics to read a research paper gets five retrieval questions the next morning on yesterday's lesson, before anything new.

In each case the motivation comes from the work itself: homework due tomorrow, a quote the customer is waiting for, a paper the learner chose to read.

Where we might be wrong

  • Relevance may not be enough, especially for children. Alpha suggests explicit incentives matter, and Teach doesn't pay anyone.
  • Incentives have their own costs. A large meta-analysis found that tangible rewards can undermine intrinsic motivation for interesting tasks [8]. We're wary of building a prize store.
  • Engagement design can slide into manipulation. The line between a nudge to keep learning and a dark pattern is thin, and we've promised not to cross it.
  • The measured effects are small. Even the best randomized results here are modest, and expectations for any AI tutor, ours included, should be too.

What this means for parents and employers

The practical advice that follows is unglamorous. Don't judge an AI tutor by how well it explains; judge it by what the learner can do a week later without it. Be suspicious of any tool that makes practice feel easy. Expect learners to resist being made to try first, and plan for that resistance instead of hoping a friendlier chatbot will dissolve it.

For parents, that means sitting with a child for the first few sessions, the same way you would with a human tutor. For employers, it means tying lessons to real work with a real deadline, and checking the work, not the lesson completion rate.

What we're measuring

  • Delayed retention without the tutor: a short check a week later with Teach switched off. This is the lesson of the Bastani study.
  • Engagement after the first refusal: whether learners keep using Teach after it declines to give an answer.
  • Transfer to real work: whether a new hire completes the task alone, or a child finishes the next assignment without help.

References

  1. [1]Bastani, H., Bastani, O., Sungu, A., Ge, H., Kabakcı, Ö., & Mariman, R. (2025). Generative AI without guardrails can harm learning: Evidence from high school mathematics. PNAS, 122. Link ↗
  2. [2]Oreopoulos, P., & Low, N. (2026). One click away: AI tutoring with Khanmigo in a two-year school experiment. NBER Working Paper 35620. Link ↗
  3. [3]Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6).
  4. [4]VanLehn, K. (2011). The relative effectiveness of human tutoring, intelligent tutoring systems, and other tutoring systems. Educational Psychologist, 46(4).
  5. [5]Roediger, H. L., & Karpicke, J. D. (2006). Test-enhanced learning: Taking memory tests improves long-term retention. Psychological Science, 17(3).
  6. [6]Anonymous reader review (2025). Your review: Alpha School. Astral Codex Ten. Link ↗
  7. [7]Pocock, M. Skills for real engineers: teach. Open-source agent skills. Link ↗
  8. [8]Deci, E. L., Koestner, R., & Ryan, R. M. (1999). A meta-analytic review of experiments examining the effects of extrinsic rewards on intrinsic motivation. Psychological Bulletin, 125(6).
  9. [9]Bjork, R. A. (1994). Memory and metamemory considerations in the training of human beings. In Metacognition: Knowing about Knowing. MIT Press.