Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

Marking the Wrong Symptoms: Evaluating LLM Watermarks in Medical Texts

标记错误的症状:评估医疗文本中的大模型水印

Abstract: Large language models (LLMs) are increasingly integrated into clinical workflows, stressing the need for reliable traceability of model-generated output with watermarking. Yet, most watermarks are evaluated on general-purpose benchmarks, leaving domains like medicine, where small token-level perturbations can result in significant semantic changes, underexplored.

摘要: 大型语言模型(LLMs)正日益融入临床工作流程,这凸显了通过水印技术实现模型生成内容可靠溯源的必要性。然而,大多数水印技术仅在通用基准测试中进行评估,而医疗等领域却鲜有研究——在这些领域中,微小的标记级扰动就可能导致重大的语义变化。

In this work, we present the first rigorous study of how LLM watermarks affect medical performance, benchmarking 5 watermarking schemes across 11 LLMs and 7 VLMs on various tasks spanning unimodal and multimodal clinical reasoning. Importantly, we complement existing evaluations by introducing a human-expert-validated pipeline for systematically auditing medical reasoning quality, terminological precision, and induced hallucinations.

在这项工作中,我们首次对大模型水印如何影响医疗性能进行了严谨的研究,针对跨越单模态和多模态临床推理的各项任务,对 11 个大语言模型(LLMs)和 7 个视觉语言模型(VLMs)中的 5 种水印方案进行了基准测试。重要的是,我们引入了一套经人类专家验证的流程,用于系统地审计医疗推理质量、术语精确度以及诱导性幻觉,从而补充了现有的评估体系。

Our results reveal that watermarking can induce substantial degradation across multiple failure modes, including lexical corruption, hallucinated terminology, and amplified misattribution or omission of image findings. Notably, we find that the absence of domain-specific analyses, combined with aggregate metrics that miss failures inherent to临床 text, can systematically obscure practical watermark-induced degradations.

我们的研究结果表明,水印会在多种故障模式下导致性能大幅下降,包括词汇损坏、术语幻觉,以及对图像发现的错误归因或遗漏加剧。值得注意的是,我们发现缺乏领域特定的分析,加之那些忽略临床文本固有故障的汇总指标,可能会系统性地掩盖水印所带来的实际性能退化。

Our findings establish domain-specific evaluation as a prerequisite for the safe deployment of watermarked models in medicine, where current benchmarks can otherwise mask clinically consequential failures.

我们的研究结果确立了“领域特定评估”是医疗领域安全部署水印模型的先决条件,否则当前的基准测试可能会掩盖具有临床严重后果的故障。