Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

Didactic knowledge or Clinical Cases? How Data Types Shape Medical Large Language Models

教学知识还是临床案例?数据类型如何塑造医学大语言模型

Abstract: Medical large language models are commonly trained on mixtures of didactic data (e.g., textbooks) and clinical data (e.g., patient records), yet how these data types differentially shape model capabilities remains unclear.

摘要: 医学大语言模型通常使用教学数据(如教科书)和临床数据(如患者记录)的混合体进行训练,但这些数据类型如何以不同方式塑造模型能力尚不明确。

We address this issue with token-matched experiments that vary the didactic-to-clinical ratio and analyze how data composition affects performance, capability profiles, and error patterns across knowledge-intensive and clinic-oriented tasks.

我们通过词元匹配(token-matched)实验解决了这一问题,通过改变教学数据与临床数据的比例,分析了数据构成如何影响模型在知识密集型任务和临床导向型任务中的表现、能力概况及错误模式。

We uncover an asymmetric transfer across task types: clinical data improves clinic-oriented tasks while remaining competitive on knowledge-intensive ones, whereas didactic data mainly improves knowledge-intensive tasks.

我们发现不同任务类型之间存在非对称迁移:临床数据在提升临床导向型任务表现的同时,在知识密集型任务上仍保持竞争力;而教学数据则主要提升知识密集型任务的表现。

Error analysis suggests a knowing-doing gap, where improvements in knowledge recall do not reliably generalize to clinical reasoning.

错误分析表明存在“知行差距”(knowing-doing gap),即知识回忆能力的提升并不能可靠地泛化到临床推理中。

We further observe that modest amounts of clinical data yield most of the gains on EHR-grounded tasks, while the optimal mixture ratio varies with the knowledge and clinical reasoning demands of downstream tasks.

我们进一步观察到,少量的临床数据即可在基于电子健康记录(EHR)的任务中获得大部分性能增益,而最佳的混合比例则取决于下游任务对知识和临床推理的需求。

These findings suggest that medical LLM data curation should be application-driven, with higher proportions of clinical data preferred for reasoning-intensive use cases.

这些发现表明,医学大语言模型的数据整理应以应用为导向,对于推理密集型的应用场景,应优先考虑更高比例的临床数据。