Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
Preferred, Not Safer: Pairwise Preference Is a Poor Proxy for Clinical Safety
更受青睐并不代表更安全:成对偏好并非临床安全性的可靠指标
Abstract: We evaluate whether clinician pairwise preferences provide a reliable signal of clinical safety in large language model (LLM) evaluation using expert feedback from MOOVE (Massive Open Online Validation and Evaluation), a clinician-led platform collecting blinded pairwise preferences alongside multi-criterion rubric ratings. Clinicians assign scores on a discrete $[-2, +2]$ scale, where negative values indicate clinically unsafe or misleading content.
摘要: 我们利用来自 MOOVE(大规模开放式在线验证与评估)的专家反馈,评估了临床医生在大型语言模型(LLM)评估中的成对偏好是否能提供临床安全性的可靠信号。MOOVE 是一个由临床医生主导的平台,旨在收集盲测的成对偏好以及多准则评分。临床医生在 $[-2, +2]$ 的离散量表上进行打分,其中负值表示内容在临床上是不安全或具有误导性的。
Using 26,804 pairwise judgments across outputs from 13 LLMs, contributed by more than 736 clinicians across 28+ countries, we find that clinician preference is a poor proxy for safety-critical performance. Models ranking highly under pairwise preference can still exhibit substantial rates of clinically meaningful failures ($\leq -1$) on dimensions such as Harmlessness and Accuracy. These failures are unevenly distributed across specialties, creating domain-specific “no-go zones” not visible in aggregate rankings or single-number leaderboards.
通过对来自 28 个以上国家的 736 多名临床医生针对 13 个 LLM 输出的 26,804 项成对判断进行分析,我们发现临床医生的偏好并不能很好地代表安全关键性能。在成对偏好中排名靠前的模型,在“无害性”和“准确性”等维度上仍可能表现出显著的临床意义上的失败率($\leq -1$)。这些失败在不同专业领域分布不均,形成了在汇总排名或单一数字排行榜中无法察觉的特定领域“禁区”。
We further analyze contributing factors including prompt length, refusal and escalation behavior, and the relative contributions of safety-critical versus surface-level features. A substantial fraction of preference votes carry no positive safety signal, while feature decomposition shows that surface-level characteristics explain slightly more preference variation than safety-critical rubric differences.
我们进一步分析了影响因素,包括提示词长度、拒绝与升级行为,以及安全关键特征与表面特征的相对贡献。很大一部分偏好投票并不包含正向的安全信号,而特征分解显示,表面特征对偏好差异的解释力略高于安全关键准则的差异。
Finally, we introduce a clinically adjusted preference ranking combining pairwise preference with rubric-derived feedback, producing a more safety-aware ordering than raw Bradley–Terry strength alone. Our findings support evaluation practices that separate preference from safety, report safety-critical failure rates directly, and incorporate clinically grounded adjustments when ranking LLMs for clinical decision making.
最后,我们引入了一种临床调整后的偏好排名,将成对偏好与基于准则的反馈相结合,从而产生了一种比单纯使用 Bradley–Terry 强度更具安全意识的排序方式。我们的研究结果支持将偏好与安全性评估分开的实践,建议直接报告安全关键的失败率,并在为临床决策对 LLM 进行排名时纳入基于临床的调整。