Examining Variation in How Guided AI Tutors Resolve Student Impasses
Examining Variation in How Guided AI Tutors Resolve Student Impasses
探究引导式 AI 导师解决学生困境方式的差异
Abstract: When a student is stuck, a tutor faces the assistance dilemma: help given too early can hinder productive struggle, while help withheld too long leaves the student in a frustrating, persistent impasse (i.e., wheel spinning). Generative AI tutors increasingly use guardrails restricting answer-giving, yet little is known about how such tutors behave once an impasse persists.
摘要: 当学生遇到困难时,导师面临着“辅助困境”:过早提供帮助可能会阻碍学生进行富有成效的思考,而过晚提供帮助则会让学生陷入令人沮丧的持续性困境(即“原地打转”)。生成式 AI 导师越来越多地使用限制直接给出答案的护栏机制,但人们对于此类导师在困境持续时如何表现知之甚少。
We analyze 20,462 student turns from 1,260 authentic sessions with a guided LLM chemistry tutor, identifying 6,630 impasse turns of three major types: conceptual errors, expressed uncertainty, or help-seeking. We then used these impasses to simulate three tutoring conditions to study variation in AI tutor guidance through impasses: baseline, no-direct-answer, and guided tutor.
我们分析了 1,260 次真实教学会话中 20,462 条学生对话轮次,这些会话由一个引导式大语言模型(LLM)化学导师辅助。我们识别出 6,630 条困境轮次,主要分为三类:概念错误、表达不确定性或寻求帮助。随后,我们利用这些困境模拟了三种辅导条件,以研究 AI 导师在处理困境时引导方式的差异:基准组、禁止直接回答组和引导式导师组。
For a sample of 150 impasses, prompt specificity changed pedagogy: a baseline tutor provided the answer directly in 50.7% of responses, a no-direct-answer tutor asked a follow-up question every time, and the guided tutor responded in a wide variety of ways depending on the context.
针对 150 个困境样本的研究发现,提示词的特异性改变了教学法:基准组导师在 50.7% 的回复中直接给出了答案;禁止直接回答组导师每次都会提出后续问题;而引导式导师则根据语境以多种方式进行回应。
We then analyzed impasse trajectories in authentic interactions, finding that each additional impasse turn lowered the odds of next-turn recovery by 12.7% (AOR = 0.873, p < .001), and early dropouts were caught in recursive concept elicitation before reaching execution. The benefit of questioning decayed as impasses persisted (scripted question x depth AOR = 0.78; follow-up x depth AOR = 0.83), whereas addressing the student’s error grew more beneficial (AOR = 1.14); after a failed scripted question, repeating it was followed by recovery in 28.1% of cases, compared with 39.8% when the tutor addressed the error instead.
我们随后分析了真实互动中的困境轨迹,发现困境轮次每增加一次,下一轮次恢复正常的几率就会降低 12.7% (AOR = 0.873, p < .001);早期放弃的学生往往在进入执行阶段前就陷入了递归式的概念诱导中。随着困境持续,提问的益处会衰减(脚本化提问 x 深度 AOR = 0.78;后续追问 x 深度 AOR = 0.83),而直接指出学生错误的效果则变得更好(AOR = 1.14);在脚本化提问失败后,重复该提问的恢复率为 28.1%,而如果导师转为指出错误,恢复率则提升至 39.8%。
For learning analytics, these findings identify impasse depth and type as observable, turn-level dialogue signals that analytics can use to trigger graduated, state-sensitive assistance in real time.
对于学习分析而言,这些发现将困境的深度和类型确定为可观察的、对话轮次级别的信号,分析系统可以利用这些信号实时触发分阶段的、针对状态的辅助干预。