Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
Spatially Grounded Concept Bottleneck Models for Trustworthy Breast Ultrasound Diagnosis
用于可信乳腺超声诊断的空间基础概念瓶颈模型
Abstract: Concept Bottleneck Models provide interpretable-by-design predictions by mediating diagnosis through human-understandable concepts, but in medical imaging, their trustworthiness is often limited by the quality and granularity of available supervision. In particular, predicted concept activations can be driven by irrelevant regions, leading to spatially unfaithful explanations.
摘要: 概念瓶颈模型(Concept Bottleneck Models)通过人类可理解的概念进行诊断,从而提供设计上的可解释性预测。然而,在医学影像领域,其可信度往往受到可用监督信息质量和粒度的限制。特别是,预测的概念激活可能由无关区域驱动,导致空间上不准确的解释。
We study a data-centric spatially grounded Concept Bottleneck Model (SG-CBM) that leverages coarse lesion delineations as weak supervision to encourage anatomically plausible concept evidence. For breast ultrasound, we derive two clinically motivated zones from each lesion mask: (i) an in-lesion region of interest for morphology-related concepts and (ii) a posterior acoustic band for posterior phenomena.
我们研究了一种以数据为中心的空间基础概念瓶颈模型(SG-CBM),该模型利用粗略的病灶轮廓作为弱监督,以促进解剖学上合理的概念证据。针对乳腺超声,我们从每个病灶掩模中推导出两个具有临床意义的区域:(i) 用于形态相关概念的病灶内感兴趣区域,以及 (ii) 用于后方回声现象的后方声带区域。
We train concept maps using a grouped spatial grounding objective and preserve semantic faithfulness with a linear bottleneck classifier. Across five-fold stratified group cross-validation, the proposed SG-CBM improves diagnostic AUROC and concept macro-AUROC while markedly increasing spatial alignment of concept evidence.
我们使用分组空间基础目标来训练概念图,并通过线性瓶颈分类器保持语义忠实度。在五折分层分组交叉验证中,所提出的 SG-CBM 提高了诊断 AUROC 和概念宏观 AUROC,同时显著增强了概念证据的空间对齐度。
We also perform a Train-corrupt/Test-clean annotation-quality stress test to quantify the impact of supervision quality on diagnosis and spatial faithfulness. Overall, the results underscore the need for data-quality-aware supervision design and systematic trustworthiness validation for deployable healthcare AI systems.
我们还进行了“训练损坏/测试清洁”的标注质量压力测试,以量化监督质量对诊断和空间忠实度的影响。总体而言,研究结果强调了在可部署的医疗 AI 系统中,进行数据质量感知监督设计和系统性可信度验证的必要性。