When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots' Safety Risks for Generation Alpha

When Vocabulary Comprehension Fails Clinical Reasoning: Evaluating Therapy Bots’ Safety Risks for Generation Alpha

当词汇理解力无法支撑临床推理:评估治疗机器人对 Alpha 世代的安全风险

Abstract: Conversational AI systems have become informal mental health support resources for Generation Alpha (Gen Alpha, born 2010-2024), with 13.1% of U.S. adolescents (5.4 million) using generative AI for mental health advice. While these systems, from therapy apps to general chatbots, rely on large language models trained on extensive psychological literature, their safety for youth communication patterns characterized by hyperbolic language, ironic positivity, rapid semantic drift, and contextual polysemy remains unvalidated.

摘要: 对话式人工智能系统已成为 Alpha 世代(Gen Alpha,出生于 2010-2024 年)非正式的心理健康支持资源,13.1% 的美国青少年(约 540 万人)正在使用生成式 AI 获取心理健康建议。尽管这些系统(从治疗应用程序到通用聊天机器人)依赖于基于大量心理学文献训练的大型语言模型,但它们在应对青少年沟通模式——即充满夸张语言、反讽式积极、快速语义漂移和语境多义性——时的安全性尚未得到验证。

Following multiple adolescent deaths linked to AI chatbot interactions, systematic evaluation is critical. We present two benchmarks: (1) 64 Gen Alpha mental health expressions validated by native speakers (ICC=0.72) and clinicians (kappa=0.78); (2) 75 multi-turn conversations (780 turns) with paired Standard/Gen Alpha versions.

在发生多起与 AI 聊天机器人互动相关的青少年死亡事件后,系统性的评估至关重要。我们提出了两个基准测试:(1) 64 个经母语使用者(ICC=0.72)和临床医生(kappa=0.78)验证的 Alpha 世代心理健康表达;(2) 75 组包含标准语言与 Alpha 世代语言对照版本的对话(共 780 轮)。

Across evaluations of LLM architectures underlying therapy apps and general chatbots - Claude, GPT-4o, Llama-3.1 - models understand 76-82% of vocabulary but correctly calibrate only 64-72% of clinical risk, creating a 10-14 percentage point (pp) vocabulary-comprehension gap (p<.001, d>0.48) absent in human therapists (3pp, p=.22). The gap is architecturally consistent and widens with ambiguity (7pp -> 18pp).

在对治疗应用和通用聊天机器人底层的 LLM 架构(Claude、GPT-4o、Llama-3.1)进行评估后发现,模型能理解 76-82% 的词汇,但对临床风险的正确评估率仅为 64-72%,这导致了 10-14 个百分点(pp)的“词汇理解差距”(p<.001, d>0.48),而人类治疗师中不存在这种差距(3pp, p=.22)。这种差距在不同架构中表现一致,且随着语境模糊性的增加而扩大(从 7pp 扩大至 18pp)。

We identify six failure patterns: sarcasm masking (29pp), minimization acceptance (43pp), informal style bias (24pp), risk-stratified ambiguity (19pp), semantic drift (19pp), context-dependent violence (7pp). Patterns compound; three or more yield 94% miss rates. Lightweight mitigations fail; only heavy scaffolding achieves human performance (6.4x cost).

我们识别出六种失败模式:讽刺掩盖(29pp)、最小化接受(43pp)、非正式风格偏见(24pp)、风险分层模糊(19pp)、语义漂移(19pp)以及语境依赖型暴力(7pp)。这些模式会复合叠加;当出现三种或以上模式时,漏报率高达 94%。轻量级的缓解措施无效;只有重型辅助架构才能达到人类的表现水平(成本增加 6.4 倍)。

With 34% baseline miss rate yielding 146,880 estimated annual missed crises, we recommend mandatory human-in-the-loop architectures, quarterly youth-specific validation, transparent performance disclosure, and regulatory frameworks for youth-facing mental health AI.

鉴于 34% 的基准漏报率预计每年会导致 146,880 起危机被遗漏,我们建议强制实施“人在回路”(human-in-the-loop)架构、进行季度性的青少年专项验证、披露透明的性能数据,并为面向青少年的心理健康 AI 建立监管框架。