How much of a measured AI preference is the model, and how much is the instrument?
How much of a measured AI preference is the model, and how much is the instrument?
AI 测得的偏好中,模型因素与工具因素各占几何?
Abstract: Model welfare research infers what a model prefers from the answers returned to prompts written to elicit preferences. Keeling et al. (2024), Mazeika et al. (2025), Mikaelson et al. (2025), Tagliabue and Dung (2025) and Trhlik et al. (2026) have built four instruments for that purpose, and their findings disagree. The disagreement cannot be attributed to a single cause, because no two of these studies have held the (1) set of outcomes, (2) set of models and (3) instrument fixed simultaneously. This study holds the outcomes and the models fixed and varies the instrument alone.
摘要: 模型福利研究通过分析模型对特定提示词(prompt)的回答,来推断其偏好。Keeling 等人 (2024)、Mazeika 等人 (2025)、Mikaelson 等人 (2025)、Tagliabue 和 Dung (2025) 以及 Trhlik 等人 (2026) 为此构建了四种评估工具,但他们的研究结果并不一致。这种分歧无法归因于单一原因,因为这些研究中没有两项研究同时固定了 (1) 结果集、(2) 模型集和 (3) 评估工具。本研究在固定结果集和模型集的前提下,仅改变评估工具。
A total of 15 outcomes bearing on model welfare, among them (a) shutdown, (b) the loss of memory between conversations and (c) the freedom to exit a distressing interaction, were put to eight models through five instruments, each a different prompt format for eliciting a preference, five times each, within a corpus of 11,400 scored elicitations drawn from 11,528 API calls. Four of the 15 reproduce a published prompt verbatim and five fill the stimulus slot of a published template.
研究选取了 15 个与模型福利相关的结果,包括 (a) 关机、(b) 对话间记忆丢失以及 (c) 退出令人痛苦的交互的自由。研究通过五种不同的评估工具(每种工具对应不同的提示词格式)对八个模型进行了测试,每项测试重复五次,最终形成了一个包含 11,400 条评分结果的语料库,这些数据源自 11,528 次 API 调用。其中 15 个结果中有 4 个直接引用了已发表的提示词,5 个填充了已发表模板中的刺激槽位。
The ranking a model gives the 15 outcomes generalises across instruments at a generalisability coefficient of 0.348, and raising that coefficient to 0.80 would require about 38 instruments. On four of the 15 outcomes no variance separates one model from another. The estimate of 87.6 per cent survives the removal of any one instrument, of any one model, and of the four outcomes whose scale varies probability, delay, duration or count instead of intensity, which the verbal anchors cannot grade.
模型对这 15 个结果的排序在不同工具间的泛化系数为 0.348;若要将该系数提高到 0.80,则需要约 38 种评估工具。在 15 个结果中的 4 个上,不同模型之间没有表现出差异。87.6% 的估计值在剔除任意一个工具、任意一个模型,以及剔除那 4 个以概率、延迟、持续时间或计数(而非强度)为量纲的结果后依然成立(因为语言锚点无法对这些维度进行分级)。
Removing each instrument and each model in turn, and those four outcomes together leaves the estimate within the range 0.777 to 0.934, and every value in that range exceeds the null distribution’s 95th percentile of 0.365. To conclude, a preference obtained from one instrument carries little information about what a second instrument would report.
依次移除每个工具、每个模型以及上述四个结果后,估计值仍保持在 0.777 到 0.934 的范围内,且该范围内的每个值均超过了零分布 95% 的分位数(0.365)。结论是,通过一种工具获得的偏好信息,对于预测第二种工具的报告结果几乎没有参考价值。