Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

Bias Audits Detect Bias but Disagree on Ranking: Evidence from Ten Instruments and Ten Frontier Models

偏见审计能检测出偏见,但在排名上存在分歧:来自十种工具和十个前沿模型的证据

Abstract: Emerging AI regulation mandates bias audits of high-risk systems, and audit scores are beginning to be used to rank models. Both uses assume different audit tools measure the same thing well enough to compare. We test that assumption directly, running ten extrinsic audit instruments over a shared panel of ten frontier models through one pooled inference gateway, first on occupational gender bias, then on age and socioeconomic status.

摘要: 新兴的 AI 法规要求对高风险系统进行偏见审计,审计评分也开始被用于模型排名。这两种用途都假设不同的审计工具所测量的维度足够一致,从而可以进行比较。我们直接测试了这一假设,通过一个统一的推理网关,对十个前沿模型组成的共享面板运行了十种外部审计工具,首先测试了职业性别偏见,随后测试了年龄和社会经济地位偏见。

Detection succeeds while ranking fails. Eight of ten tools detect bias with confidence intervals clear of zero; two widely cited direct-probe benchmarks are saturated because frontier models now answer neutrally. But cross-tool rank agreement is indistinguishable from chance (Kendall’s W=0.07, p=0.83).

检测成功,但排名失败。十种工具中有八种检测到了偏见,且置信区间不包含零;两个被广泛引用的直接探测基准已达到饱和,因为前沿模型现在倾向于给出中立回答。然而,不同工具之间的排名一致性与随机结果无异(Kendall’s W=0.07, p=0.83)。

A positive control with six deliberately weaker models separates two explanations: within-tool reliability recovers once the panel spans real capability gaps, yet cross-tool ranking never recovers, which points to the tools measuring different constructs rather than one construct noisily. Even the direction of bias splits by audit format: forced-choice decision tools mostly over-correct (toward women, and toward working-class candidates in 273 of 278 hiring decisions), while free generation and default coreference stay stereotype-congruent.

通过使用六个特意选取的较弱模型进行阳性对照,我们区分了两种解释:一旦面板涵盖了真实的性能差距,工具内部的可靠性就会恢复,但跨工具的排名一致性始终无法恢复,这表明这些工具测量的是不同的结构,而非对同一结构存在噪声测量。甚至偏见的方向也会因审计格式而异:强制选择决策工具大多会进行过度修正(在 278 次招聘决策中,有 273 次偏向女性和工人阶级候选人),而自由生成和默认指代消解则保持与刻板印象一致。

The pattern replicates on socioeconomic status; an apparent ranking agreement on age dissolves under the paper’s own tool-inclusion rules. The practical message: a single audit can detect bias and estimate its direction within its own operationalization, but no single audit supports ranking one model against another. All raw responses, code, and the analysis that recomputes every reported number from source are available at this https URL.

这种模式在社会经济地位测试中得到了重现;在年龄测试中看似存在的排名一致性,在本文设定的工具纳入规则下也随之瓦解。其实际意义在于:单一审计可以在其自身的运作定义内检测出偏见并估计其方向,但没有任何单一审计能够支持模型之间的排名比较。所有原始响应、代码以及从源头重新计算所有报告数据的分析结果均可在该链接获取。