Large Language Models are Approximate Survival Estimators
Large Language Models are Approximate Survival Estimators
大型语言模型是近似生存估计器
Abstract: Survival analysis estimates time-to-event outcomes from patient covariates and is widely used for medical risk assessment. Patients seeking prognostic information after a diagnosis may turn to large language models (LLMs), now readily accessible through consumer applications. However, whether LLMs can provide accurate survival predictions has not been rigorously evaluated.
摘要: 生存分析通过患者协变量来估计事件发生的时间结果,并被广泛用于医疗风险评估。患者在确诊后寻求预后信息时,可能会转向如今通过消费级应用即可轻松访问的大型语言模型(LLM)。然而,LLM 是否能够提供准确的生存预测尚未经过严格评估。
We introduce Survprompt, a framework that converts structured patient covariates into free-text clinical vignettes and prompts pre-trained LLMs to predict survival zero-shot. We benchmark Survprompt against conventional survival models, including random survival forests (RSF), across two multi-institutional pan-cancer cohorts: the publicly available MSK-CHORD cohort and a newly curated cohort from the Providence St. Joseph Health Network constructed using an LLM-based medical abstraction framework.
我们引入了 Survprompt,这是一个将结构化患者协变量转换为自由文本临床小传,并提示预训练 LLM 进行零样本(zero-shot)生存预测的框架。我们在两个多机构泛癌队列中,将 Survprompt 与包括随机生存森林(RSF)在内的传统生存模型进行了基准测试:一个是公开的 MSK-CHORD 队列,另一个是使用基于 LLM 的医学抽象框架构建的 Providence St. Joseph 医疗网络的新整理队列。
We report censored mean absolute error (cMAE) and concordance index (c-index) and conduct feature ablations to identify variables influencing LLM predictions. Frontier LLMs achieved surprisingly competitive cMAE for individual survival times. For example, GPT-5.6-Sol achieved cMAE within 10% of state-of-the-art RSF models specifically trained for survival prediction for several cancer types and lower cMAE than RSF for prostate cancer in MSK-CHORD.
我们报告了删失平均绝对误差(cMAE)和一致性指数(c-index),并进行了特征消融研究,以确定影响 LLM 预测的变量。前沿 LLM 在个体生存时间的 cMAE 上表现出了令人惊讶的竞争力。例如,GPT-5.6-Sol 在多种癌症类型中实现的 cMAE 与专门为生存预测训练的最先进 RSF 模型相比差距在 10% 以内,且在 MSK-CHORD 的前列腺癌预测中,其 cMAE 低于 RSF。
Feature ablations revealed that LLMs prioritized clinical variables similarly to specialized survival models. However, LLMs showed inconsistent accuracy across cancer types and institutions and poorly discriminated between high- and low-risk patients (lower c-index). Zero-shot LLMs can generate surprisingly accurate prognostic estimates without specialized training, but their variable performance across cancer types and institutions remains an important limitation for clinical use.
特征消融显示,LLM 对临床变量的优先排序与专业生存模型相似。然而,LLM 在不同癌症类型和机构间的准确性表现不一致,且在高风险和低风险患者之间的区分能力较差(c-index 较低)。零样本 LLM 无需专门训练即可生成令人惊讶的准确预后估计,但其在不同癌症类型和机构间不稳定的表现,仍然是其临床应用的一个重要局限。