Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake
Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake
基于临床医生视角的 AI 辅助精神科问诊质量保证
Abstract: Before patients can use AI-assisted psychiatric intake systems, health systems need practical ways to routinely evaluate these tools against their clinical standards for quality assurance. Because clinicians may use different intake styles, evaluation for this task must (1) support comparison across interviewing approaches, (2) minimize clinician burden, and (3) measure clinically relevant performance for health systems deploying these technologies.
摘要: 在患者使用 AI 辅助精神科问诊系统之前,医疗系统需要切实可行的方法,根据临床标准对这些工具进行常规的质量保证评估。由于临床医生可能采用不同的问诊风格,针对该任务的评估必须满足以下要求:(1)支持不同问诊方式之间的比较;(2)最大限度地减少临床医生的负担;(3)为部署这些技术的医疗系统衡量临床相关的性能指标。
We present a clinician-grounded evaluation platform built around a memory-augmented patient simulator for open-ended AI interviewing, InterviewPlayground. We created interactive patients using InterviewPlayground with our expert-authored vignettes, constructed a simulated intake platform for the interviews, and designed evaluation modalities relevant to intake.
我们提出了一个基于临床医生视角的评估平台,该平台围绕一个用于开放式 AI 问诊的记忆增强型患者模拟器——InterviewPlayground 构建。我们利用专家编写的病例小传(vignettes)通过 InterviewPlayground 创建了交互式患者,构建了一个模拟问诊平台,并设计了与问诊相关的评估模式。
In a pilot of 6 clinicians in a 25-minute assessment compared to a GPT-based LLM intake interviewer, the LLM recovered more of the clinically relevant items embedded in the patient vignettes (88.0% vs. 38.9%), but made more clinical inferences not based on the interview (56.8% vs. 27.8%), and characterized identified safety concerns less often (33.3% vs. 66.7%), setting the stage for deployed quality assurance for this task.
在一项涉及 6 名临床医生的 25 分钟评估试点中,研究人员将该平台与基于 GPT 的大语言模型(LLM)问诊员进行了对比。结果显示,LLM 能够提取出更多嵌入在患者病例中的临床相关信息(88.0% 对比 38.9%),但同时也做出了更多缺乏问诊依据的临床推断(56.8% 对比 27.8%),且对已识别的安全隐患的描述频率较低(33.3% 对比 66.7%)。这项研究为该任务的部署质量保证奠定了基础。