Artificial Analysis Intelligence Index v4.2

Artificial Analysis Intelligence Index v4.2

Announcing Artificial Analysis Intelligence Index v4.2 We are accelerating elements of our upcoming v5 release with interim updates to keep pace with the frontier. Index v4.2 has more complex and realistic tasks, and more private test sets to prevent gaming.

发布 Artificial Analysis Intelligence Index v4.2 我们正在加速 v5 版本的部分功能开发,通过中期更新来跟上技术前沿的步伐。Index v4.2 包含了更复杂、更真实的测试任务,并增加了更多的私有测试集以防止刷榜。

Intelligence Index v4.2 changelog:

  • AA-Briefcase, our agentic knowledge work evaluation with a private test set
  • Surge’s GDP.pdf, long context document reasoning across 4,592 PDF pages
  • GPQA Diamond, an exceptional scientific reasoning evaluation that has now been saturated … plus greater weighting on held-out test sets to prevent gaming, and grading infrastructure upgrades to increase robustness.

Intelligence Index v4.2 更新日志:

  • AA-Briefcase:我们针对智能体知识工作的评估,包含私有测试集。
  • Surge 的 GDP.pdf:跨越 4,592 页 PDF 文档的长上下文推理测试。
  • GPQA Diamond:一项卓越的科学推理评估,目前已达到饱和状态。 ……此外,增加了对留存测试集(held-out test sets)的权重以防止刷榜,并升级了评分基础设施以提高稳健性。

This update brings the Index closer to real-world use cases with more challenging, complex and realistic tasks and private test sets to prevent gaming. We have been planning and building elements of Index v5 for months - it’s been 8 months since we launched Index v4 in January. We have deliberately held back updates to keep the Index stable through recent major model launches. However, with the frontier moving so quickly in the past weeks, we feel it is important to deliver an immediate interim update to ensure our Index remains as relevant and useful as ever to users. Beyond this interim update, our team is hard at work on v5 of the Index. We are planning more incremental releases in the near future. Stay tuned!

此次更新通过更具挑战性、复杂且真实的测试任务以及防止刷榜的私有测试集,使该指数更贴近现实应用场景。我们几个月来一直在规划和构建 Index v5 的要素——自 1 月份发布 Index v4 以来已经过去了 8 个月。为了在近期各大模型发布期间保持指数的稳定性,我们特意推迟了更新。然而,鉴于过去几周技术前沿发展如此迅速,我们认为有必要立即进行一次中期更新,以确保我们的指数对用户而言依然具有相关性和实用性。除了这次中期更新外,我们的团队正在全力开发 v5 版本。我们计划在不久的将来发布更多增量更新。敬请期待!

Intelligence Index v4.2 changes in detail:

Intelligence Index v4.2 详细变更:

Adding AA-Briefcase: Our in-house evaluation with a private held-out test set, AA-Briefcase tests models on realistic agentic knowledge work tasks in complex projects built by industry experts. Models are evaluated on multi-week knowledge work projects, each with many linked tasks and thousands of input source files. AA-Briefcase combines rubric and pairwise grading to evaluate verifiable task success, analytical quality, and presentation quality, giving a holistic view of overall agentic capability in knowledge work.

新增 AA-Briefcase: 这是我们内部开发的评估工具,包含私有留存测试集。AA-Briefcase 在由行业专家构建的复杂项目中,测试模型处理现实智能体知识工作任务的能力。模型需完成为期数周的知识工作项目,每个项目包含许多关联任务和数千个输入源文件。AA-Briefcase 结合了评分标准和成对比较评分,以评估可验证的任务成功率、分析质量和呈现质量,从而全面展示模型在知识工作中的智能体能力。

Adding GDP.pdf: Created by Surge AI, GDP.pdf evaluates single-turn professional document reasoning across 100 PDFs and ten domains. Models must synthesize evidence distributed across 4,592 pages, including text, tables, charts, footnotes, and exclusions. Responses are graded against 1,275 expert-authored atomic criteria; the headline All-pass Rate credits a task only when every criterion is satisfied.

新增 GDP.pdf: 由 Surge AI 创建,GDP.pdf 评估跨越 100 份 PDF 文档和十个领域的单轮专业文档推理能力。模型必须综合分布在 4,592 页中的证据,包括文本、表格、图表、脚注和排除项。回答将根据 1,275 条专家编写的原子标准进行评分;只有当所有标准都满足时,才会计入“全通过率”(All-pass Rate)。

Weighting to measure real-world use and prevent gaming: 40% of our Index weighting is now private, held-out test sets - double the figure from v4.1. Held-out data includes AA-Briefcase, AA-Omniscience, and solutions for CritPt. This reduces the ability for labs to game evaluations. The held-out percentage will increase further in Index v5.

衡量现实使用情况并防止刷榜的权重调整: 我们指数中 40% 的权重现在来自私有留存测试集,是 v4.1 版本的两倍。留存数据包括 AA-Briefcase、AA-Omniscience 和 CritPt 的解决方案。这降低了实验室通过刷榜来操纵评估结果的能力。在 Index v5 中,留存测试集的比例将进一步提高。

Improving our grading infrastructure: In AA-LCR v1.1, we have added a grading system prompt and corrected errors and ambiguities in answer keys, improving scoring accuracy. For GDPval-AA v2 and AA-Briefcase, we have improved our sampling and re-anchored the Elo scale, making ratings more stable as new models are added. For SciCode we have improved robustness of grading sandboxes to ensure slow but correct code does not count as a failure.

改进评分基础设施: 在 AA-LCR v1.1 中,我们添加了评分系统提示词,并纠正了答案键中的错误和歧义,提高了评分准确性。对于 GDPval-AA v2 和 AA-Briefcase,我们改进了采样方式并重新锚定了 Elo 等级分,使得在添加新模型时评分更加稳定。对于 SciCode,我们增强了评分沙箱的稳健性,以确保运行缓慢但正确的代码不会被判定为失败。

Key results:

关键结果:

Anthropic and OpenAI lead the Index: Anthropic’s Claude Fable 5.1 leads the Index, followed by OpenAI’s GPT-6 Astra, which shows a 4pt gain over GPT-5.6 Sol. Meta is the third-ranked lab on the leaderboard, followed by SpaceXAI, Moonshot/Kimi, Z.AI, and Google.

Anthropic 和 OpenAI 领跑指数: Anthropic 的 Claude Fable 5.1 位居指数榜首,其次是 OpenAI 的 GPT-6 Astra,其表现比 GPT-5.6 Sol 提升了 4 个点。Meta 在排行榜上排名第三,紧随其后的是 SpaceXAI、月之暗面/Kimi、Z.AI 和 Google。

Cost per Task Pareto frontier shared by four labs: Anthropic, OpenAI, Meta and Z.AI occupy the updated Cost per Task frontier.

四家实验室共享“单任务成本”帕累托前沿: Anthropic、OpenAI、Meta 和 Z.AI 占据了更新后的“单任务成本”前沿位置。

GPT-6 Astra dominates the output token frontier: GPT-6 Astra is more token efficient than almost every other model near the intelligence frontier, with Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite at either end of the curve (excludes models below 25 on the Index).

GPT-6 Astra 在输出 Token 效率前沿占据主导地位: 在智能前沿附近,GPT-6 Astra 的 Token 效率高于几乎所有其他模型,Claude Fable 5.1、Grok 4.5 和 Gemini 3.5 Flash-Lite 分别位于曲线的两端(不包括指数低于 25 的模型)。

Anthropic’s Claude Fable 5.1 and Opus 5 lead AA-Briefcase, followed by GPT-6 Astra and Muse Spark 1.3. GPT-6 Astra shows a substantial gain above GPT-5.6 Sol of ~85 Elo points.

Anthropic 的 Claude Fable 5.1 和 Opus 5 在 AA-Briefcase 中领先,其次是 GPT-6 Astra 和 Muse Spark 1.3。GPT-6 Astra 相比 GPT-5.6 Sol 有显著提升,Elo 分数提高了约 85 分。

OpenAI leads GDP.pdf with GPT-6 Astra at 33.2% and GPT-5.6 Sol at 28.2%, followed by Claude Fable 5.1 at 26.2%.

OpenAI 在 GDP.pdf 中处于领先地位,GPT-6 Astra 得分为 33.2%,GPT-5.6 Sol 为 28.2%,随后是 Claude Fable 5.1,得分为 26.2%。