ARC-AGI Leaderboard

ARC-AGI Leaderboard

Understanding the Leaderboard

ARC-AGI has evolved from its first versions (ARC-AGI-1 and 2) which measured passive fluid intelligence, to ARC-AGI-3 which challenges AI agents to adapt on the fly to novel interactive environments.

理解排行榜

ARC-AGI 已从衡量被动流体智力的早期版本(ARC-AGI-1 和 2)演进至 ARC-AGI-3,后者旨在挑战 AI 智能体在全新的交互式环境中进行实时适应的能力。

The scatter plot above visualizes the critical relationship between cost-per-task and performance - a key measure of efficiency. True intelligence isn’t just about solving problems, but solving them efficiently with minimal resources.

上方的散点图直观地展示了单任务成本与性能之间的关键关系——这是衡量效率的核心指标。真正的智能不仅在于解决问题,更在于以最少的资源高效地解决问题。

Interpreting the data

Reasoning Systems Trend Line solutions display connected points representing the same model at different reasoning levels. These trend lines illustrate how increased reasoning time affects performance, typically showing asymptotic behavior as thinking time increases.

数据解读

“推理系统趋势线”(Reasoning Systems Trend Line)方案展示了同一模型在不同推理水平下的连接点。这些趋势线说明了增加推理时间如何影响性能,通常表现为随着思考时间的增加,性能趋于渐近线。

Base LLMs solutions represent single-shot inference from standard language models like GPT-4.5 and Claude 3.7, without extended reasoning capabilities. These points demonstrate raw model performance without additional reasoning enhancements.

“基础大模型”(Base LLMs)方案代表了来自 GPT-4.5 和 Claude 3.7 等标准语言模型的单次推理(single-shot inference),且不具备扩展的推理能力。这些点展示了模型在没有额外推理增强情况下的原始性能。

Kaggle Systems solutions showcase competition-grade submissions from the Kaggle challenge, operating under strict computational constraints ($50 compute budget for 120 evaluation tasks). These represent purpose-built, efficient methods specifically designed for the ARC Prize.

“Kaggle 系统”(Kaggle Systems)方案展示了来自 Kaggle 竞赛的竞技级提交作品,它们在严格的计算约束下运行(120 个评估任务的计算预算为 50 美元)。这些代表了专为 ARC Prize 设计的高效专用方法。

Verification Policy

For more information, see our testing policy.

验证政策

欲了解更多信息,请参阅我们的测试政策。

Leaderboard Breakdown

Notes:

  • Only systems which required less than $10,000 to run are shown.
  • For models that were not able to produce full test outputs, remaining tasks were marked as incorrect.
  • Results marked as “preview” are unofficial and may be based on incomplete testing.
  • 1 ARC-AGI-2 score estimate based on partial testing results and o1-pro pricing.
  • 2 Provisional cost estimates based on Gemini 3 Pro pricing. Model to be retested once released.

排行榜细则

备注:

  • 仅展示运行成本低于 10,000 美元的系统。
  • 对于未能生成完整测试输出的模型,剩余任务被标记为错误。
  • 标记为“预览”(preview)的结果为非官方数据,可能基于不完整的测试。
  • 1 ARC-AGI-2 分数估算基于部分测试结果及 o1-pro 定价。
  • 2 临时成本估算基于 Gemini 3 Pro 定价。模型发布后将进行重新测试。