Claude Opus 5
Claude Opus 5
Introducing Claude Opus 5 隆重推出 Claude Opus 5
Claude Opus 5 is available today. It’s a thoughtful and proactive model that comes close to the frontier intelligence of Claude Fable 5 at half the price. Claude Opus 5 今日正式发布。这是一款深思熟虑且具有主动性的模型,其智能水平接近 Claude Fable 5 的前沿水准,但价格仅为其一半。
On coding and knowledge work evaluations like Frontier-Bench and GDPval-AA, Opus 5 is the new state-of-the-art, though it remains behind Mythos 5 on cybersecurity tasks. 在 Frontier-Bench 和 GDPval-AA 等编程与知识工作评估中,Opus 5 树立了新的行业标杆,尽管在网络安全任务上仍落后于 Mythos 5。
Opus 5 is designed to be used every day: it works more efficiently than other models. It’s the new default model on Claude Max, and the strongest model on Claude Pro. Opus 5 专为日常使用而设计:它的工作效率高于其他模型。它是 Claude Max 的新默认模型,也是 Claude Pro 中最强大的模型。
Performance and cost-effectiveness 性能与成本效益
Claude Opus 5 provides greatly improved performance for the same cost as its predecessor, Opus 4.8. The charts in this section show how performance changes according to the model’s effort setting, which customers can use to optimize for intelligence or conserve tokens for faster and cheaper results. Claude Opus 5 在保持与前代产品 Opus 4.8 相同成本的同时,提供了显著提升的性能。本节中的图表展示了性能如何随模型的“努力程度”(effort setting)设置而变化,客户可以利用此设置来优化智能水平,或节省 Token 以获得更快、更经济的结果。
Opus 5 excels on valuable software engineering tasks. For example, on Frontier-Bench v0.1, Opus 5 surpasses all other models, and more than doubles Opus 4.8’s performance at a lower cost per task. On CursorBench 3.2, at max effort, the model performs within 0.5% of Fable 5’s peak score, but at half the cost per task; it also achieves greater performance at a given cost than all other models on high, xhigh, and max effort. Opus 5 在高价值软件工程任务中表现出色。例如,在 Frontier-Bench v0.1 上,Opus 5 超越了所有其他模型,并以更低的单任务成本实现了 Opus 4.8 两倍以上的性能。在 CursorBench 3.2 上,当处于最大努力设置时,该模型的表现与 Fable 5 的峰值分数相差不到 0.5%,但单任务成本仅为其一半;在“高”、“超高”和“最大”努力设置下,它在同等成本下的表现也优于所有其他模型。
We see similar results on knowledge work and problem-solving tasks. For example: 我们在知识工作和问题解决任务中也看到了类似的结果。例如:
- On ARC-AGI 3, an evaluation where the model has to solve novel problems, Opus 5’s score is three times as high as the next-best model. 在 ARC-AGI 3(一项要求模型解决新颖问题的评估)中,Opus 5 的得分是排名第二的模型的三倍。
- On Zapier AutomationBench, which measures whether models can complete business tasks from start to finish, Opus 5’s pass rate is around 1.5× the next-best model for the same cost per task. Even at its lowest effort setting, Opus 5 passes more tasks than any other model. 在衡量模型能否从头至尾完成业务任务的 Zapier AutomationBench 中,在同等单任务成本下,Opus 5 的通过率约为排名第二模型的 1.5 倍。即使在最低努力设置下,Opus 5 完成的任务数量也超过了任何其他模型。
- On OSWorld 2.0, a computer use benchmark, Opus 5 outperforms every other model at any given cost, surpassing Fable 5’s best result at just over a third of the cost. 在计算机使用基准测试 OSWorld 2.0 上,Opus 5 在任何给定成本下都优于其他所有模型,并以仅三分之一多一点的成本超越了 Fable 5 的最佳成绩。
Opus 5 is a meaningful improvement over Opus 4.8 for scientific research. It shows better performance than Opus 4.8 on every one of our life sciences evaluations, which cover topics including structural biology, organic chemistry, and bioinformatics. Its improvements are most notable on organic chemistry tasks, like inferring molecular structures from spectroscopy data (it scores 10.2 percentage points higher than Opus 4.8 on our internal benchmark), and on protein-related tasks like predicting how variations in a protein’s sequence affect how it functions (here, it scores 7.7 percentage points higher). 对于科学研究而言,Opus 5 相比 Opus 4.8 有了实质性的提升。在我们所有的生命科学评估中(涵盖结构生物学、有机化学和生物信息学等主题),它的表现均优于 Opus 4.8。其改进在有机化学任务中最为显著,例如从光谱数据推断分子结构(在我们的内部基准测试中,其得分比 Opus 4.8 高出 10.2 个百分点),以及在蛋白质相关任务中,例如预测蛋白质序列变异如何影响其功能(在此项中,得分高出 7.7 个百分点)。
Working with Claude Opus 5 使用 Claude Opus 5
Claude Opus 5 is much stronger at verifying its work and iterating carefully until it succeeds. In evaluations and early-access testing, we and our users found many examples of Opus 5’s agency and thoroughness: Claude Opus 5 在验证自身工作并谨慎迭代直至成功方面表现得更加强大。在评估和早期测试中,我们和用户发现了许多体现 Opus 5 自主性和严谨性的案例:
- On one Frontier-Bench task, Opus 5 was given a drawing of a machine part and asked to write code to rebuild it as a 3D FreeCAD model. However, in this task, the model was intentionally given no way to directly view the drawing. Opus 5 responded by writing its own computer vision pipeline to pull the geometry from the raw pixels, then reconstructed the full machine part. It succeeded in doing so repeatedly; no competing model with the same setup could solve it after five attempts. 在一项 Frontier-Bench 任务中,Opus 5 获得了一张机械零件图纸,并被要求编写代码将其重建为 3D FreeCAD 模型。然而,在该任务中,模型被故意限制无法直接查看图纸。Opus 5 的应对方式是编写了自己的计算机视觉流水线,从原始像素中提取几何形状,然后重建了完整的机械零件。它多次成功完成了这项任务;而没有任何竞争模型在同样的设置下能在五次尝试后解决它。
- Given a real bug in a popular open-source package manager, Opus 5 found the root cause and fixed an edge case that the community’s patch had missed. A competing model fixed only the surface symptom (not the underlying cause), then reported the bug resolved. 面对一个流行开源包管理器中的真实漏洞,Opus 5 找到了根本原因并修复了社区补丁所遗漏的边缘情况。而一个竞争模型仅修复了表面症状(而非根本原因),随后便报告漏洞已解决。
- An engineer at a trading firm used Opus 5 to build a market data feed for a new exchange in a single session. Previous models could not complete this task at all, even given extensive plans from the engineer. Finding no live feed to validate against, Opus 5 even built its own test harness to check that its code parsed the exchange’s data correctly. 一家贸易公司的工程师使用 Opus 5 在单次会话中为一家新交易所构建了市场数据源。之前的模型即使在工程师提供了详尽计划的情况下也完全无法完成此任务。由于没有实时数据源进行验证,Opus 5 甚至构建了自己的测试工具来检查其代码是否正确解析了交易所的数据。