How UK AISI and EvalEval Are Making Benchmark Results Reproducible
How UK AISI and EvalEval Are Making Benchmark Results Reproducible
英国人工智能安全研究所(AISI)与 EvalEval 如何实现基准测试结果的可复现性
The EvalEval Coalition is thrilled to share that the UK AI Security Institute (AISI) is using EvalEval’s infrastructure to openly share evaluation results, supporting more reproducible and verifiable evaluation science. AISI and EvalEval have previously collaborated on research that began at a joint workshop alongside NeurIPS 2025, and feedback from the Institute has helped shape the Every Eval Ever (EEE) schema. This next phase of the collaboration puts that shared infrastructure into practice.
EvalEval 联盟很高兴地宣布,英国人工智能安全研究所(AISI)正在使用 EvalEval 的基础设施公开分享评估结果,以支持更具可复现性和可验证性的评估科学。AISI 和 EvalEval 此前曾在 NeurIPS 2025 的联合研讨会上开展过研究合作,来自该研究所的反馈意见帮助塑造了“Every Eval Ever”(EEE)架构。此次合作的下一阶段将把这些共享基础设施付诸实践。
Why reproducible evaluation reporting matters
为什么可复现的评估报告至关重要
As AI deployment accelerates, evaluations are becoming increasingly important sources of evidence about model and system performance. Yet results are reported across many formats, platforms, and outlets, often without enough information to reproduce them. Running the evaluations again may itself be prohibitively expensive.
随着人工智能部署的加速,评估正成为衡量模型和系统性能越来越重要的证据来源。然而,评估结果通过多种格式、平台和渠道发布,往往缺乏足够的信息来复现这些结果。而重新运行这些评估本身可能成本高昂,令人望而却步。
EvalEval’s mission is to improve this ecosystem through a shared reporting schema, Every Eval Ever, and an open platform, Evaluation Cards, that brings evaluation results and the information needed to interpret them into a common structure. This builds naturally on AISI’s work to make evaluation more efficient through OptStop, more statistically rigorous through HiBayES, and more standardised in areas including transcript analysis and capability elicitation. Together, AISI and EvalEval are working to diagnose gaps in evaluation reporting and build shared infrastructure to close them.
EvalEval 的使命是通过共享报告架构“Every Eval Ever”以及开放平台“Evaluation Cards”来改善这一生态系统,将评估结果及解读所需的信息整合为统一的结构。这建立在 AISI 的工作基础之上,即通过 OptStop 提高评估效率,通过 HiBayES 增强统计严谨性,并在转录分析和能力诱导等领域实现标准化。AISI 和 EvalEval 正共同努力诊断评估报告中的差距,并构建共享基础设施以弥补这些不足。
What AISI is sharing
AISI 正在分享的内容
Transcript-level transparency matters not only for reproducibility, but also for analysis and diagnosis. In this new phase of the collaboration, AISI is making publicly reported evaluation methods and findings available through Evaluation Cards where appropriate. The release includes verified results, context, and configuration information for the five benchmarks in the paper’s main experiment: HealthBench, FrontierMath, Humanity’s Last Exam, SWE-Bench Pro, and Terminal-Bench 2.0.
转录层面的透明度不仅对可复现性至关重要,对于分析和诊断也同样重要。在合作的这一新阶段,AISI 正在通过 Evaluation Cards 公开其评估方法和研究发现。此次发布的内容包括该论文主要实验中五个基准测试的验证结果、背景信息和配置信息,这些基准测试包括:HealthBench、FrontierMath、Humanity’s Last Exam、SWE-Bench Pro 和 Terminal-Bench 2.0。
These results cover six frontier models: Claude Opus 4, Claude Opus 4.5, Claude Opus 4.6, GPT-5, GPT-5.2, and GPT-5.4. The release also includes results from two related cyber evaluations—Cyber CTFs and The Last Ones—which use a different, partially overlapping set of models. The data accompany AISI’s paper, How Inference Compute Shapes Frontier LLM Evaluation, which studies how benchmark performance depends on inference-time compute and evaluation protocol.
这些结果涵盖了六个前沿模型:Claude Opus 4、Claude Opus 4.5、Claude Opus 4.6、GPT-5、GPT-5.2 和 GPT-5.4。此次发布还包括两项相关网络安全评估的结果——Cyber CTFs 和 The Last Ones,它们使用了另一组部分重叠的模型。这些数据随 AISI 的论文《推理计算如何塑造前沿大语言模型评估》(How Inference Compute Shapes Frontier LLM Evaluation)一同发布,该论文研究了基准测试性能如何依赖于推理时计算量和评估协议。
When results are openly released with setup information, researchers and practitioners can examine individual studies more closely and compare findings across the wider ecosystem. Where other reports lack these details, releases like AISI’s provide verified reference points for interpreting evaluations in context—for example, by helping researchers understand how setup choices may influence reported performance. As more evaluators adopt EEE, open comparisons like these can support broader and more reliable meta-research.
当评估结果连同设置信息一并公开时,研究人员和从业者可以更深入地审查各项研究,并在更广泛的生态系统中比较研究发现。在其他报告缺乏这些细节的情况下,像 AISI 这样的发布提供了经过验证的参考点,有助于在特定背景下解读评估结果——例如,帮助研究人员理解设置选择如何影响报告的性能。随着越来越多的评估者采用 EEE,这类开放式比较将支持更广泛、更可靠的元研究。
Contribute to the shared mission
为共同使命做出贡献
-
Model developers: Report verified evaluation results.
-
Evaluation developers: Report benchmarks and run data using the Every Eval Ever schema.
-
Evaluation, governance, and policy researchers: Explore Evaluation Cards by benchmark or model, or use it to examine the state of evaluation reporting as a whole.
-
模型开发者: 报告经过验证的评估结果。
-
评估开发者: 使用 Every Eval Ever 架构报告基准测试和运行数据。
-
评估、治理和政策研究人员: 按基准测试或模型探索 Evaluation Cards,或利用它来审视评估报告的整体状况。
About the EvalEval Coalition
关于 EvalEval 联盟
The EvalEval Coalition is a research community developing scientifically grounded research and robust deployment infrastructure for the evaluation ecosystem. Its goal is to improve evaluation science, address the lack of consensus around documenting evaluation applicability and utility, and broaden coverage of the impacts that matter for scientific research and policy analysis. The coalition’s flagship projects include Every Eval Ever, a shared schema and repository for evaluation results, and Evaluation Cards, which combines benchmark metadata, evaluation-run data, and model metadata into interpretable records. Together, they make it easier to understand when apparently similar scores were produced under meaningfully different conditions.
EvalEval 联盟是一个研究社区,致力于为评估生态系统开发基于科学的研究和稳健的部署基础设施。其目标是改进评估科学,解决在记录评估适用性和效用方面缺乏共识的问题,并扩大对科学研究和政策分析至关重要的影响力的覆盖范围。该联盟的旗舰项目包括“Every Eval Ever”(一个用于评估结果的共享架构和存储库)以及“Evaluation Cards”(将基准测试元数据、评估运行数据和模型元数据整合为可解读的记录)。它们共同使得人们更容易理解为何看似相似的分数是在截然不同的条件下产生的。
About the UK AI Security Institute
关于英国人工智能安全研究所 (AISI)
The UK AI Security Institute is a research organisation within the UK government’s Department for Science, Innovation and Technology. Its mission is to equip governments with a scientific understanding of the risks posed by advanced AI. AISI conducts research and builds infrastructure to understand advanced AI capabilities and impacts, develop and test mitigations, and inform policy.
英国人工智能安全研究所是英国政府科学、创新和技术部下属的研究机构。其使命是为政府提供对先进人工智能所带来风险的科学理解。AISI 开展研究并构建基础设施,以了解先进人工智能的能力和影响,开发并测试缓解措施,并为政策制定提供参考。