Binarization Flattens the Score Space
Binarization Flattens the Score Space
二值化抹平了评分空间
Abstract: Large language model (LLM) judges are often used as rewards to train policies on objectives that deterministic verifiers cannot capture. However, these rewards are often collapsed to pass/fail ({0, 1}), which reports the verdict but not how well a response met each criterion. We model each pass/fail verdict as a score on an unreported scale, compared with one cutoff. A stretch of that scale moves every score proportionally toward or away from the cutoff, but never across it, so no verdict changes. A policy is therefore free to apply any stretch without changing anything the panel reports.
摘要: 大型语言模型(LLM)评估者常被用作奖励机制,以训练确定性验证器无法捕捉的目标策略。然而,这些奖励往往被简化为通过/失败({0, 1})的二值结果,这仅报告了结论,却无法反映回答在多大程度上满足了各项标准。我们将每个通过/失败的判定建模为未公开量表上的分数,并与一个临界值进行比较。对该量表的拉伸会使所有分数按比例向临界值靠近或远离,但绝不会跨越临界值,因此判定结果不会改变。因此,策略可以自由地应用任何拉伸操作,而不会改变评估面板报告的任何内容。
Under a joint-Gaussian model, a third grade adds a second threshold and removes this affine stretch ambiguity. On MATH and SciBench outputs from one seven-criterion judge, all 14 constructed criterionwise stretches were invisible after binarization but visible with three grades. At $n=1{,}024$, a test given both population laws had at least 96.5% power at a $1.5\times$ stress.
在联合高斯模型下,增加第三个等级相当于引入了第二个阈值,从而消除了这种仿射拉伸的不确定性。在针对 MATH 和 SciBench 数据集的七项标准评估中,所有 14 种构建的准则级拉伸在二值化后均不可见,但在使用三个等级时则清晰可见。在 $n=1{,}024$ 的样本量下,针对两种总体分布的测试在 $1.5\times$ 的拉伸强度下至少具有 96.5% 的统计功效。
Retaining grades closes one blind spot created by binarization, but verdicts alone remain insufficient as some changes are still indistinguishable from genuine improvement. These include arbitrary within-grade changes and fixed-covariance, loading-aligned mean shifts — the signature of a sycophancy-shaped lift the panel reads as competence. The shared-factor reference approximation fit MATH and SciBench but not HealthBench, delineating its empirical scope.
保留等级可以消除二值化带来的一个盲点,但仅凭判定结果仍然不足,因为某些变化与真正的能力提升无法区分。这些变化包括等级内的任意变动以及固定协方差、载荷对齐的均值偏移——这正是评估面板误以为是能力提升的“谄媚式”表现特征。共享因子参考近似法适用于 MATH 和 SciBench,但不适用于 HealthBench,这界定了其经验适用范围。
We recommend keeping at least three grades (for example, asking the judge whether each criterion is fully, partially, or not met and rewarding {0, 0.5, 1}), and externally validating gains along the remaining direction, which no finer scale removes.
我们建议至少保留三个等级(例如,询问评估者每项标准是完全满足、部分满足还是未满足,并分别给予 {0, 0.5, 1} 的奖励),并针对剩余方向上的增益进行外部验证,因为任何更精细的量表都无法消除该方向上的模糊性。