TutorMoments: Do AI tutors know when to help and when to hold back?
TutorMoments: Do AI tutors know when to help and when to hold back?
TutorMoments:AI 导师知道何时该提供帮助,何时该适可而止吗?
Today we’re introducing a preview of TutorMoments, a framework to measure whether cutting-edge LLMs can balance one of the hardest trade-offs in education: when to step in and help a student and when to hold back and let the student do more of the work. 今天,我们推出了 TutorMoments 的预览版。这是一个旨在评估尖端大语言模型(LLM)能否平衡教育中最难权衡的问题之一的框架:即何时介入帮助学生,以及何时该适可而止,让学生独立完成更多思考。
TutorMoments is a replay-based evaluation built off real one-on-one math tutoring sessions. Experienced math teachers go through transcripts collected from a U.S. tutoring program and flag the moments where a tutor had to choose between making a problem easier to get started on and pushing the student to do more of the reasoning themselves. TutorMoments 是一项基于真实一对一数学辅导课程的“重演式”评估。经验丰富的数学老师会审阅从美国某辅导项目中收集的教学记录,并标记出导师必须做出抉择的关键时刻:是降低题目难度以帮助学生入门,还是推动学生进行更多的自主推理。
TutorMoments then takes the transcript up to that decision point, hands it to a language model, and has the model take over as the tutor in a simulated session – with the student played by another language model – to see what the LLM tutor does. TutorMoments 会将教学记录截取至该决策点,交给大语言模型,让模型在模拟课程中接管导师的角色(学生角色则由另一个大语言模型扮演),以观察该 AI 导师的表现。
Told only to “tutor well,” we find that models tend to over-help by giving too much support and rarely pushing students to do deeper thinking. Spelling out the trade-off (when to help versus when to hold back) in the tutor’s prompt improves performance, but it doesn’t close the gap to human tutoring that consistently fits the moment, and LLMs still differ widely in how reliably they make that call. 在仅被要求“做好辅导”的情况下,我们发现模型往往会提供过多的支持,从而导致“过度帮助”,却很少推动学生进行更深层次的思考。在导师的提示词(Prompt)中明确指出这种权衡(何时帮助与何时克制)可以提升表现,但仍无法弥补与人类导师在精准把握时机方面的差距,且不同大语言模型在做出此类判断的可靠性上仍存在巨大差异。
As part of our commitment to open research, we’re releasing a dataset of de-identified tutoring transcripts, the code for running our replay pipeline, and the model tutor replays of the key moments we evaluated in those transcripts for reproducibility. We hope TutorMoments gives educators, researchers, and the teams building AI tutors a sharper way to ask how a model handles the pedagogical decisions that matter most—and helps the field build tutors that adapt to each student instead of doing the work for them. 作为我们致力于开放研究的一部分,我们发布了去标识化的辅导记录数据集、运行重演流程的代码,以及我们在这些记录中评估的关键时刻的模型重演数据,以供复现。我们希望 TutorMoments 能为教育工作者、研究人员和构建 AI 导师的团队提供一种更敏锐的方法,来审视模型如何处理最关键的教学决策,并帮助该领域构建出能够适应每位学生、而非代劳工作的 AI 导师。
What makes a good tutor? Ask a good math tutor for help and you’ll likely get a question back like, “What do you know about what the problem is asking?” That isn’t unhelpfulness–part of strong teaching is diagnosing what students do know and providing the right support for them in the moment. Immediately volunteering support would rob a student of the intellectual work that helps them learn. 什么才是好的导师?向优秀的数学老师求助时,你很可能会得到这样的反问:“关于题目要求,你了解到了什么?”这并非不乐于助人——优秀教学的一部分在于诊断学生已掌握的知识,并在当下提供恰当的支持。如果立即提供帮助,反而会剥夺学生通过智力劳动进行学习的机会。
Sometimes support is needed; other times what’s most effective is a push to solidify understanding by explaining a correct answer. Language models, though, are trained to be helpful, and a helpful assistant tends to do the hard part for you—explaining the concept, laying out the steps, and guiding you to the answer. In a tutoring session, that can cut short the productive struggle—the effortful, sometimes frustrating problem-solving that learning research has long tied to stronger understanding. 有时需要支持;而另一些时候,最有效的方法是通过解释正确答案来巩固理解。然而,大语言模型被训练为“乐于助人”,而一个乐于助人的助手往往会替你完成困难的部分——解释概念、列出步骤并引导你得出答案。在辅导课程中,这可能会缩短“富有成效的挣扎”(productive struggle)——即学习研究中长期认为有助于加深理解的、费力且有时令人沮丧的问题解决过程。
Most benchmarks for language models acting as tutors don’t capture this tension. They tend to reward one behavior in particular – never giving away the answer to a problem, say, or always offering a hint – without accounting for whether that was the right move for where the student actually was in their understanding. But good tutoring isn’t a single fixed behavior you can identify across the board. It’s a judgment call: what does this student need, right now, on this problem? 大多数针对 AI 导师的基准测试并未捕捉到这种张力。它们往往倾向于奖励某种特定的行为——比如从不直接给出答案,或者总是提供提示——却不考虑这对于学生当前的理解水平而言是否是正确的举措。但好的辅导并非一种可以一概而论的固定行为,它是一种判断:此时此刻,针对这道题,这位学生真正需要什么?
How TutorMoments works: TutorMoments is built on real tutoring data. The dataset we’re releasing, TutorMoments-Preview, is 462 de-identified, text-only transcripts of real one-on-one math tutoring with U.S. students in grades 2-7, with more than 1,500 teacher-annotated key moments and several thousand free-text annotations from 27 U.S.-based teacher annotators. TutorMoments 的工作原理:TutorMoments 基于真实的辅导数据构建。我们发布的 TutorMoments-Preview 数据集包含 462 份去标识化的纯文本记录,来自美国 2-7 年级学生的一对一数学辅导,其中包含超过 1,500 个由教师标注的关键时刻,以及来自 27 位美国教师标注员的数千条自由文本注释。
The transcripts come from a high-dosage tutoring program whose students mostly attend Title I schools, shared under a research clause agreed to by parents and guardians; all data was stripped of identifying details, first by the provider and then through an additional math-aware pipeline. All annotations came from experienced math teachers, whom we asked to read the transcripts and mark key learning moments—noting what was going on, what the tutor did, and how it landed for the student. 这些记录来自一个高强度辅导项目,学生大多就读于 Title I 学校,并在家长和监护人同意的研究条款下共享;所有数据均已剔除身份识别信息,先由提供方处理,随后通过额外的数学感知流程进行二次清洗。所有注释均来自经验丰富的数学老师,我们邀请他们阅读记录并标记关键学习时刻——记录当时的情况、导师的行为以及学生对此的反应。
Each key moment is a decision point where the tutor had to weigh scaffolding (making a problem more accessible) against pushing for rigor (encouraging the student to do harder thinking). TutorMoments runs by pausing a transcript at one of those key moments and handing the session to a language model, which takes over as the tutor for five turns with a simulated student. We call each of these model-generated continuations a replay. 每个关键时刻都是一个决策点,导师必须在“脚手架”(使问题更易于理解)和“推动严谨性”(鼓励学生进行更深入的思考)之间进行权衡。TutorMoments 的运行方式是在这些关键时刻暂停记录,并将课程交给大语言模型,由其接管导师角色,与模拟学生进行五个回合的互动。我们将这些由模型生成的后续对话称为“重演”。
An LLM-based scoring pipeline then rates each replay on three things: whether the model (1) scaffolded when the student needed support, (2) pushed for rigor when the student was ready for more challenge, and (3) avoided over-scaffolding (reducing the challenge more than the moment called for). The scoring pipeline starts from a teacher-defined ground truth: for each key moment, whether it called for scaffolding or for a push for rigor. 随后,一个基于 LLM 的评分流程会对每次重演进行三项评估:模型是否 (1) 在学生需要支持时提供了脚手架,(2) 在学生准备好迎接挑战时推动了严谨性,以及 (3) 避免了过度脚手架(即降低难度超过了当时所需的程度)。评分流程始于教师定义的“基准真值”:即针对每个关键时刻,它究竟需要脚手架还是推动严谨性。
Several teachers annotated each moment, and when they disagreed we took the majority label—if three teachers annotated a moment and two called for rigor while one called for scaffolding, the ground truth is rigor. A separate LM classifier validated against teacher annotations then decides whether the tutor’s actual move matches what the moment called for—an “appropriate” turn means the tutor’s classified action (scaffold, push for rigor, or over-scaffold) lines up with what teachers judged the moment to call for. 每个时刻均由多位教师标注,若意见不一致,我们采用多数原则——如果三位教师标注某时刻,两位要求严谨性,一位要求脚手架,则基准真值为严谨性。随后,一个经过教师标注验证的独立 LM 分类器会判定导师的实际行为是否符合该时刻的需求——“适当”的回合意味着导师的分类行为(脚手架、推动严谨性或过度脚手架)与教师对该时刻的判断相符。
Preliminary results: We ran seven LLMs through TutorMoments using two prompts: a plain prompt that gives no real guidance – it only tells the model to use what it knows about good tutoring to respond to the student – and an evalua… 初步结果:我们使用两个提示词对七个大语言模型进行了 TutorMoments 测试:一个是没有任何实际指导的简单提示词——仅告诉模型利用其所知的良好辅导知识来回应学生——以及一个评估……