When benchmark inferences do not compose: Projectibility in AI evaluation

When benchmark inferences do not compose: Projectibility in AI evaluation

当基准测试推论无法组合时:人工智能评估中的可投射性问题

Abstract: An AI benchmark result rarely reaches a consequential claim in one step. Evaluators generalize it to further cases, interpret it as evidence of capability, extrapolate it to new tasks, transport it to another system or site, and combine it with assumptions about human review and downstream consequences. Validity-centred approaches require evidence for each claim. This paper identifies a further epistemic problem: warranted links don’t automatically make a warranted chain. The target of one study may not be the source of the next; system, population, outcome, or conditions may change at the interface; and shared data or model lineage may make apparently independent support dependent.

摘要: 人工智能(AI)的基准测试结果很少能一步到位地得出具有实质意义的结论。评估者通常需要将其推广到更多案例,将其解读为能力的证据,推演至新任务,迁移到其他系统或场景,并结合关于人工审核及下游影响的假设。以有效性为中心的方法要求对每一项主张都提供证据。本文指出了一个深层的认识论问题:有根据的环节并不自动构成有根据的链条。一项研究的目标可能并非下一项研究的来源;系统、群体、结果或条件可能会在接口处发生变化;而共享数据或模型血缘可能会使看似独立的支撑证据变得相互依赖。

Projectibility concerns whether a bounded extension from observed to unobserved cases is warranted. Goodman supplies the problem of rival extensions; argument-based validity supplies an architecture for testing them. The paper’s distinctive claim is a non-composition principle: support for adjacent projections warrants their composition only when endpoints and assumptions align and dependence and uncertainty are carried through. A legal-research case shows how benchmark evidence and a deployment study can each be sound while remaining parallel. A reanalysis and simulation show why aggregate stability can erase distinctions a later projection requires. The resulting projectibility audit diagnoses unsupported joins in benchmark-to-use arguments.

可投射性(Projectibility)关注的是:从已观察案例到未观察案例的有限扩展是否具有正当性。 古德曼(Goodman)提出了竞争性扩展的问题;而基于论证的有效性则提供了一种测试这些扩展的架构。本文的核心主张是一个“非组合原则”:只有当端点和假设保持一致,且依赖关系和不确定性得到有效传递时,对相邻预测的支撑才能证明其组合的合理性。一个法律研究案例展示了基准测试证据与部署研究如何各自稳健却又互不相关。通过重新分析和模拟,本文揭示了为何总体稳定性可能会抹除后续预测所必需的区分度。最终提出的“可投射性审计”旨在诊断从基准测试到实际应用论证中那些缺乏支撑的连接点。