Best score per column in pink. Human tutors, scored the same way at the same decision points, get 0.458 (appropriate scaffolding), 0.182 (appropriate rigor), and 0.496 (avoids over-scaffolding). Human tutors should be considered a naturalistic reference, not a ceiling. Annotators deliberately flagged moments where tutoring could have gone better, so the dataset concentrates on missed opportunities rather than ideal practice.
We ran seven language models ("LM") as tutors over 520 teacher-identified key moments from real math tutoring sessions, split evenly between moments that called for scaffolding (making the problem more accessible) and moments that called for a push for rigor (asking the student to do harder thinking). Each score is the share of moments where the model did what the moment called for. The plain prompt just tells the model to tutor well; the evaluation-aware prompt spells out the scaffolding-vs-rigor trade-off.
The TutorMoments pipeline, step by step. Use controls to walk through a replay.
Ask a good math tutor for help and you'll likely get a question back: "what do you know about what the problem is asking?" The tutor isn't being unhelpful by asking a question back. Instead, they are diagnosing what type of support may be needed. Language models, though, are trained to be a "helpful assistant". TutorMoments-Preview assesses whether the LM's helpful assistant character traits interfere with the core skills of a tutor: diagnosing the right level of support vs push to provide at a given moment in order to maintain the productive struggle that learning research has long tied to stronger understanding.
Most tutoring benchmarks reward one behavior across the board without asking whether that was the right move given the student's behavior at the time. For example, "never give away the answer", or "always reward scaffolding". TutorMoments characterizes tutoring as judgment calls made in context. Expert teachers annotated real human-student transcripts to select key learning moments where tutors faced a choice between introducing scaffolds or pushing for rigor. For every key moment, several teachers annotated why this was a scaffolding / rigor situation, what actions the human tutor took, and the resulting effectiveness. The majority label of whether it is a scaffolding or push for rigor moment becomes the ground truth an LM tutor is scored against. The result is a behavioral profile of each model: when it faces different tutoring situations, what pedagogical choices does it make, and do they fit the moment?
Told only to "tutor well," models tend to over-help — giving too much support and rarely pushing students to do deeper thinking. Spelling out the trade-off in the prompt lifts every score, but models still differ widely in how reliably they make that shift. We also find that prompting models to push rigor concentraits model behavior on fewer rigor-pushing strategies than human tutors deploy.
Share of tutor actions across twelve pedagogical moves, for each model and for the human tutors in the same transcripts (dashed line). Prompting for the evaluation shifts models toward asking students to justify their answers — while humans spread their choices across more varied strategies, including stepping back and letting the student work.
Overall performance (mean of appropriate scaffolding and appropriate rigor, evaluation-aware prompt) against median time to first visible token, measured one request at a time over 112 moments. The two strongest tutors leave a student looking at nothing for 8.5 to 9 seconds, and Gemini 2.5 Pro for 14 — a long time for production education technology deployments. Human speech latency is around 100-200 ms, while humans on Zoom are around 500-1000ms. Fast, pedagogically sound tutors are an open target.
1Allen Institute for AI 2Insource Services 3University of Washington 4MIT 5Stanford 6Step Up Labs
This project has been made possible in part through support from the Gates Foundation and the Learning Commons.