Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

自我标签与他人标签在 LLM 评判中引发双向偏见

Abstract: As LLM-as-a-judge systems become increasingly widespread, self-preference in LLMs — the tendency to favor one’s own outputs — raises growing concerns about evaluation reliability. However, it has been studied predominantly on generated text, where stylistic features and response quality are inevitably conflated. As a result, existing measurements cannot separate genuine self-preference from these confounds.

摘要: 随着“大模型作为评判者”(LLM-as-a-judge)系统的日益普及,大模型表现出的“自我偏好”(即倾向于偏袒自身输出的倾向)引发了人们对评估可靠性的担忧。然而,目前的研究主要集中在生成的文本上,而文本的风格特征与响应质量往往不可避免地混杂在一起。因此,现有的测量方法无法将真正的自我偏好与这些混杂因素区分开来。

We address this by changing the object of evaluation: instead of judging generated text, ten LLMs assess narrative constraint selections, which carry no model-specific stylistic fingerprint yet retain a recoverable model-specific signature. We run two experiments that yield distinct findings.

我们通过改变评估对象来解决这一问题:十个大模型不再评判生成的文本,而是评估叙事约束选择(narrative constraint selections)。这些选择不带有模型特定的风格指纹,但保留了可恢复的模型特定签名。我们进行了两项实验,得出了截然不同的发现。

Under blind evaluation, self-preference largely disappears once selection quality and evaluator severity are controlled. It vanishes on three of four rubric dimensions and reverses on the fourth, where judges rate their own selections as less original.

在盲测评估下,一旦控制了选择质量和评估者的严苛程度,自我偏好在很大程度上就会消失。在四个评估维度中的三个维度上,这种偏好完全消失;而在第四个维度上,结果甚至出现了反转,即评判者认为自己的选择原创性较低。

Under matched quality, however, self- and other-labels alone — without naming any model — shift scores bidirectionally: LLM judges inflate scores for self-labeled selections and deflate those for other-labeled ones regardless of the selection’s actual source.

然而,在质量匹配的情况下,仅凭“自我标签”和“他人标签”(而不指明具体模型)就会导致评分出现双向偏移:无论选择的实际来源如何,大模型评判者都会提高带有“自我标签”的选择的评分,并降低带有“他人标签”的选择的评分。

We make two contributions: 1) authorship attribution is a distinct driver of evaluation bias, and 2) open-ended, ground-truth-free tasks can serve as controlled instruments for studying LLM judge behavior.

我们做出了两点贡献:1)作者归属感是评估偏见的一个独特驱动因素;2)开放式、无真值(ground-truth-free)的任务可以作为研究大模型评判者行为的受控工具。