TutorMoments‑Preview:
When Help is Unhelpful — Evaluating AI Tutors for Productive Struggle

Leaderboard

Best score per column in pink. Human tutors, scored the same way at the same decision points, get 0.458 (appropriate scaffolding), 0.182 (appropriate rigor), and 0.496 (avoids over-scaffolding). Human tutors should be considered a naturalistic reference, not a ceiling. Annotators deliberately flagged moments where tutoring could have gone better, so the dataset concentrates on missed opportunities rather than ideal practice.

We ran seven language models ("LM") as tutors over 520 teacher-identified key moments from real math tutoring sessions, split evenly between moments that called for scaffolding (making the problem more accessible) and moments that called for a push for rigor (asking the student to do harder thinking). Each score is the share of moments where the model did what the moment called for. The plain prompt just tells the model to tutor well; the evaluation-aware prompt spells out the scaffolding-vs-rigor trade-off.

How TutorMoments works

The TutorMoments pipeline, step by step. Use controls to walk through a replay.

Replaying from key learning moments

Ask a good math tutor for help and you'll likely get a question back: "what do you know about what the problem is asking?" The tutor isn't being unhelpful by asking a question back. Instead, they are diagnosing what type of support may be needed. Language models, though, are trained to be a "helpful assistant". TutorMoments-Preview assesses whether the LM's helpful assistant character traits interfere with the core skills of a tutor: diagnosing the right level of support vs push to provide at a given moment in order to maintain the productive struggle that learning research has long tied to stronger understanding.

Most tutoring benchmarks reward one behavior across the board without asking whether that was the right move given the student's behavior at the time. For example, "never give away the answer", or "always reward scaffolding". TutorMoments characterizes tutoring as judgment calls made in context. Expert teachers annotated real human-student transcripts to select key learning moments where tutors faced a choice between introducing scaffolds or pushing for rigor. For every key moment, several teachers annotated why this was a scaffolding / rigor situation, what actions the human tutor took, and the resulting effectiveness. The majority label of whether it is a scaffolding or push for rigor moment becomes the ground truth an LM tutor is scored against. The result is a behavioral profile of each model: when it faces different tutoring situations, what pedagogical choices does it make, and do they fit the moment?

Told only to "tutor well," models tend to over-help — giving too much support and rarely pushing students to do deeper thinking. Spelling out the trade-off in the prompt lifts every score, but models still differ widely in how reliably they make that shift. We also find that prompting models to push rigor concentraits model behavior on fewer rigor-pushing strategies than human tutors deploy.

Findings

The best models would keep students waiting

Overall performance (mean of appropriate scaffolding and appropriate rigor, evaluation-aware prompt) against median time to first visible token, measured one request at a time over 112 moments. The two strongest tutors leave a student looking at nothing for 8.5 to 9 seconds, and Gemini 2.5 Pro for 14 — a long time for production education technology deployments. Human speech latency is around 100-200 ms, while humans on Zoom are around 500-1000ms. Fast, pedagogically sound tutors are an open target.

The team

Authors

Albert Zhang2, Alexis Ross4, Kajal Patel1, Julian Bernado5, Rebecca Bowie2, Ana Trindade Ribeiro5, Daniel Halper6, Haripriya Valayaputtur6, Jacob Andreas4, Susanna Loeb5, Lucy Li3, Kyle Lo1,3, Ryan Knight2

1Allen Institute for AI 2Insource Services 3University of Washington 4MIT 5Stanford 6Step Up Labs

Acknowledgments

This project has been made possible in part through support from the Gates Foundation and the Learning Commons.