Claude Opus 5 leads on agentic work — and undercuts Fable 5 on cost

Claude Opus 5 leads on agentic work — and undercuts Fable 5 on cost

Claude Opus 5 正式发布,曾协助 Anthropic 进行预发布评估的 Artificial Analysis 刚刚公布了其完整的基准测试报告。核心结论是:它是目前处理代理(Agentic)知识工作的顶级模型,且单任务成本低于 Fable 5。这种在模型前沿领域兼具性能与性价比的情况并不多见。

“Opus 5 (max) scores 61 on the Artificial Analysis Intelligence Index, effectively tied with Claude Fable 5 (max, 60), and ahead of GPT-5.6 Sol (max, 59)” “Opus 5 (max) 在 Artificial Analysis 智能指数上得分为 61,与 Claude Fable 5 (max, 60) 基本持平,并领先于 GPT-5.6 Sol (max, 59)。”

What actually changed

实际变化有哪些?

New agentic leader: 1861 Elo on GDPval-AA v2 — more than 100 points ahead of both Fable 5 and GPT-5.6 Sol. On AA-Briefcase (agentic knowledge work), it’s +146 Elo over Fable 5. 新的代理任务领跑者:在 GDPval-AA v2 上获得 1861 Elo 分数,比 Fable 5 和 GPT-5.6 Sol 高出 100 多分。在 AA-Briefcase(代理知识工作)测试中,其 Elo 分数比 Fable 5 高出 146 分。

Joint first on coding: Opus 5 (xhigh) with Claude Code tops the Artificial Analysis Coding Index, including the highest score on SWE-Atlas-QnA. 89% on Terminal-Bench v2.1: Roughly in line with the current terminal leader, GPT-5.6 Sol. 编程能力并列第一:Opus 5 (xhigh) 配合 Claude Code 在 Artificial Analysis 编程指数中名列前茅,包括在 SWE-Atlas-QnA 中获得最高分。在 Terminal-Bench v2.1 上得分为 89%:与目前的终端任务领跑者 GPT-5.6 Sol 基本持平。

Cost per task: $2.03 at max effort — vs Fable 5’s $2.75. That’s 26% less for equivalent or better intelligence on agentic benchmarks. 1M token context window (same as Opus 4.8), 5 effort settings (low → max), and server-side fallback support. Pricing: $5/$25 per million input/output tokens — same rate as previous Opus launches. 单任务成本:在最高努力(max effort)设置下为 2.03 美元,而 Fable 5 为 2.75 美元。这意味着在代理基准测试中,以同等或更好的智能水平,成本降低了 26%。支持 100 万 token 上下文窗口(与 Opus 4.8 相同)、5 种努力程度设置(从 low 到 max)以及服务器端回退支持。定价:每百万输入/输出 token 为 5 美元/25 美元——与之前的 Opus 发布价格相同。

The cost-intelligence shift

成本与智能的权衡转变

For agentic workloads — the things most teams are actually building on right now — Opus 5 doesn’t just match Fable 5. It beats it, and charges less to do it. Fable 5 was the “throw more at it” option. Opus 5 reframes the trade-off: better agentic outcomes and a lower bill. At mid-tier effort settings (high, xhigh), it can outperform both Opus 4.8 and Sonnet 5 on a cost-per-task basis. That’s a lot of headroom to play with before you’re even at max effort. 对于代理工作负载(目前大多数团队实际构建的方向),Opus 5 不仅仅是追平了 Fable 5,它在超越对方的同时还降低了成本。Fable 5 曾是那种“堆算力”的选择,而 Opus 5 重塑了这种权衡:它能带来更好的代理任务结果,且账单更低。在中等努力设置(high, xhigh)下,它在单任务成本上就能超越 Opus 4.8 和 Sonnet 5。这意味着在达到最高努力设置之前,你还有很大的优化空间。

The caveat worth flagging: factual knowledge still lags.

值得注意的警告:事实知识仍有欠缺。

Opus 5 improved +7 points on AA-Omniscience over Opus 4.8, but its hallucination rate climbed 14 points to 50% — it guesses more confidently when uncertain. For retrieval-heavy or factual precision tasks, Fable 5 still holds the edge. Opus 5 在 AA-Omniscience(全知指数)上比 Opus 4.8 提高了 7 分,但其幻觉率上升了 14 个百分点,达到 50%——这意味着它在不确定时会更自信地胡编乱造。对于检索密集型或要求事实精确的任务,Fable 5 仍然占据优势。

What to do

建议与行动

  • Running agentic pipelines? Opus 5 is the new default to benchmark. Start at high or xhigh effort before committing to max. 正在运行代理流水线? Opus 5 是新的基准测试默认选择。在决定使用 max 模式前,先从 high 或 xhigh 努力程度开始尝试。
  • On Claude Code? You’re already getting the benefit — joint first on the Coding Agent Index. 使用 Claude Code? 你已经从中受益了——它在编程代理指数中并列第一。
  • Cost-sensitive on frontier models? Max-effort Opus 5 undercuts Fable 5 by 26%. Re-run your cost model — this changes the calculus. 对前沿模型的成本敏感? 最高努力模式下的 Opus 5 比 Fable 5 便宜 26%。重新计算你的成本模型吧——这改变了计算逻辑。
  • Factual knowledge tasks? Hold off. A 50% hallucination rate is a hard limit for anything knowledge-intensive. Fable 5 still wins there. 事实知识类任务? 先等等。50% 的幻觉率对于任何知识密集型任务来说都是硬伤。在这方面,Fable 5 依然胜出。

Full benchmark breakdown: Artificial Analysis — Claude Opus 5 完整基准测试报告:Artificial Analysis — Claude Opus 5