Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening
Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening
性能与一致性:评估基础模型在 Lung-RADS 筛查中的表现
Abstract: Foundation models have recently demonstrated strong capabilities across a wide range of medical imaging tasks. However, their performance in structured clinical interpretation settings remains insufficiently explored. In lung cancer screening, interpretative variability persists despite standardized frameworks such as Lung-RADS.
摘要: 基础模型最近在广泛的医学影像任务中展现出了强大的能力。然而,它们在结构化临床解读环境中的表现仍未得到充分探索。在肺癌筛查中,尽管存在 Lung-RADS 等标准化框架,但解读差异性依然存在。
In this study, we evaluate MedGemma, a medical general-purpose foundation model derived from Gemini and its fine-tuned version adapted for lung cancer detection and diagnosis, compared against radiologists performing Lung-RADS v2022 assessment on the NLST dataset. Twelve radiologists independently evaluated each case in a multi-reader design, enabling quantification of inter-reader variability.
在本研究中,我们评估了 MedGemma(一种源自 Gemini 的通用医学基础模型)及其针对肺癌检测和诊断进行微调的版本,并将其与在 NLST 数据集上执行 Lung-RADS v2022 评估的放射科医生进行了对比。通过多阅片者设计,12 名放射科医生独立评估了每个病例,从而实现了对阅片者间差异性的量化。
Radiologists achieved a mean AUC of 0.90, with substantial variability across readers (range: 0.80-0.94). The native foundation model achieved an AUC of 0.70, failing to reach clinically relevant performance. In contrast, fine-tuning significantly improved performance to an AUC of 0.83, placing the model within the lower range of individual radiologists performance.
放射科医生的平均 AUC 为 0.90,但阅片者之间存在显著差异(范围:0.80-0.94)。原生基础模型的 AUC 为 0.70,未能达到临床相关的性能水平。相比之下,微调显著提升了性能,使 AUC 达到 0.83,处于个体放射科医生表现的较低区间。
These findings highlight a trade-off between peak accuracy and prediction consistency. Unlike radiologists, under fixed conditions, the model produces deterministic outputs, removing inter-run variability under identical inputs, in contrast to inter-reader variability observed among radiologists. This supports the role of fine-tuned foundation models potential complementary tools for clinical decision support, particularly in settings with limited expertise.
这些发现突显了峰值准确性与预测一致性之间的权衡。与放射科医生不同,该模型在固定条件下产生确定性输出,消除了相同输入下的运行间差异,这与放射科医生之间观察到的阅片者间差异形成了对比。这支持了微调后的基础模型作为临床决策支持潜在补充工具的作用,特别是在专业知识有限的环境中。
However, evaluation is performed on a case-enriched cohort from NLST and does not account for real-world prevalence or external validation, limiting direct clinical generalization.
然而,本次评估是在 NLST 的病例富集队列上进行的,并未考虑现实世界中的患病率或外部验证,这限制了其直接的临床推广能力。