Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
Large Language Models Show Metacognitive Sensitivity in Medical Reasoning
大型语言模型在医学推理中表现出元认知敏感性
Abstract: Large language models (LLMs) are increasingly evaluated and used in medicine, but clinical usefulness depends on answer accuracy and whether confidence tracks evidence quality and uncertainty. 摘要: 大型语言模型(LLMs)正越来越多地被评估并应用于医学领域,但其临床实用性取决于回答的准确性,以及模型表现出的置信度是否能与证据质量和不确定性相匹配。
We developed a controlled, psychophysics-inspired clinical benchmark to test diagnostic choice and confidence behavior in a medical LLM. The benchmark focused on probable Alzheimer-type neurocognitive disorder (AT-NCD) versus depression-related cognitive impairment (DRCI). 我们开发了一个受心理物理学启发的受控临床基准,用于测试医学大模型在诊断选择和置信度方面的表现。该基准侧重于区分可能的阿尔茨海默型神经认知障碍(AT-NCD)与抑郁相关的认知障碍(DRCI)。
We generated 45 synthetic vignettes varying evidence strength, conflicting evidence, and missing information. Each vignette was presented under three prompt variants, yielding 135 trials. In a pilot run with gpt-4.1-nano, all trials produced valid structured outputs. 我们生成了 45 个合成病例,通过改变证据强度、冲突证据和缺失信息来设计场景。每个病例在三种不同的提示词变体下进行测试,共产生 135 次试验。在使用 gpt-4.1-nano 进行的初步运行中,所有试验均输出了有效的结构化结果。
Across forced-choice trials, diagnostic accuracy was 93.5%, mean confidence was 78.4%, and AUROC2 was 0.876. Confidence increased with evidence distance from the diagnostic boundary, decreased when information was missing, and remained higher on correct than incorrect trials after adjustment for evidence strength and prompt format. 在强制选择试验中,诊断准确率为 93.5%,平均置信度为 78.4%,AUROC2 为 0.876。当证据远离诊断边界时,置信度会增加;当信息缺失时,置信度会降低;在调整了证据强度和提示格式后,正确试验的置信度始终高于错误试验。
These findings indicate partial metacognitive sensitivity rather than globally uninformative confidence. However, errors clustered in moderate, conflicting AT-NCD cases, where the model shifted toward DRCI and retained more confidence than empirical accuracy justified. 这些发现表明模型具备部分元认知敏感性,而非完全缺乏参考价值的置信度。然而,错误主要集中在证据中等且存在冲突的 AT-NCD 病例中,此时模型倾向于诊断为 DRCI,且其表现出的置信度超过了实际准确率所能支撑的水平。
Model comparison suggested that confidence quality should be measured directly rather than inferred from benchmark accuracy or model capability alone. This study establishes a reproducible framework for evaluating evidence sensitivity, metacognitive sensitivity, and localized calibration failure in medical LLMs. 模型对比研究表明,置信度质量应直接进行测量,而不应仅根据基准测试准确率或模型能力进行推断。本研究为评估医学大模型中的证据敏感性、元认知敏感性以及局部校准失败问题建立了一个可复现的框架。