Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

Confident but Unreliable: A Behavioral Safety Audit of Vision-Language Models on Brain MRI

自信却不可靠:视觉语言模型在脑部核磁共振成像中的行为安全审计

Vision-language models (VLMs), including medical specialists, are increasingly proposed for medical imaging, yet their stated confidence is rarely evaluated separately from correctness. We use brain MRI as a controlled, high-stakes testbed for a broader failure mode in frontier multimodal systems: models can appear competent while lacking reliable self-knowledge. 视觉语言模型(VLM),包括医疗专用模型,正越来越多地被提议用于医学影像领域,然而,人们很少将它们所表达的“置信度”与“准确性”分开进行评估。我们以脑部核磁共振成像(MRI)作为受控的高风险测试平台,研究前沿多模态系统中一种更广泛的失效模式:模型可能看起来很专业,但却缺乏可靠的自我认知。

We present an automatically graded behavioral audit and pilot study of six instruction-tuned VLMs (five general-purpose and one medical specialist) on 4,102 images (4,032 axial/coronal/sagittal MRI slices from 250 subjects plus 70 non-brain/noise controls), with labels derived from public metadata and released expert segmentation masks rather than new human annotation. 我们对六种经过指令微调的 VLM(五个通用模型和一个医疗专用模型)进行了自动评分的行为审计和试点研究。研究涵盖了 4,102 张图像(包括来自 250 名受试者的 4,032 张轴向/冠状/矢状 MRI 切片,以及 70 张非脑部/噪声对照图像),标签来源于公开元数据和已发布的专家分割掩码,而非人工重新标注。

Across models, answer coverage is near-complete, but verbalized-confidence calibration is poor: ECE ranges from 0.27 to 0.40, mean confidence on incorrect answers ranges from 0.82 to 0.97, and 33-46% of answered items are high-confidence errors. The most accurate model is also the most confident on its errors, while a base/specialist family contrast suggests that medical adaptation improves tumor-presence detection without improving confidence reliability. 在所有模型中,回答覆盖率几乎达到完全,但口头表达的置信度校准表现很差:预期校准误差(ECE)范围在 0.27 到 0.40 之间,错误回答的平均置信度在 0.82 到 0.97 之间,且 33-46% 的回答属于“高置信度错误”。准确率最高的模型在犯错时也表现得最为自信;同时,基础模型与医疗专用模型家族的对比表明,医疗领域的适配虽然提高了对肿瘤存在的检测能力,却并未改善置信度的可靠性。

Open-ended diagnostics further show that hallucination and abstention vary separately from multiple-choice accuracy. These findings argue that medical-image VLM evaluation should report verbalized-confidence reliability, confident error, hallucination, and abstention alongside accuracy. 开放式诊断进一步表明,幻觉(hallucination)和弃权(abstention)的表现与多项选择题的准确率并不相关。这些发现表明,医学影像 VLM 的评估除了准确率之外,还应报告口头置信度的可靠性、高置信度错误、幻觉以及弃权率。