Claude Opus 5.5

Claude Opus 5.5

We’re introducing Claude Opus 5.5, the first model in our new Claude 5.5 family. It performs at the level of Claude Fable 5.1 on most work and costs 40% less to run than Opus 5. 我们正式推出 Claude Opus 5.5,这是我们全新 Claude 5.5 系列中的首款模型。它在大多数工作任务中的表现达到了 Claude Fable 5.1 的水平,且运行成本比 Opus 5 降低了 40%。

Claude Opus 5.5 is our first release since we called for pacing the frontier. It was tested before release by external evaluators, including Frontier Design and METR. On our automated behavioral audit, the most comprehensive alignment test we run, Opus 5.5 is the strongest-performing model we’ve tested to date. It also comes with the safeguards we’ve developed for our most capable models. Claude Opus 5.5 是我们呼吁放缓前沿技术发展步伐以来的首次发布。在发布前,它已通过了包括 Frontier Design 和 METR 在内的外部评估机构的测试。在我们运行的最全面的对齐测试——自动化行为审计中,Opus 5.5 是我们迄今为止测试过的性能最强的模型。它还配备了我们为最强能力模型所开发的各项安全防护措施。

Here are some of the improvements you can expect from Opus 5.5: 以下是您可以从 Opus 5.5 中期待的一些改进:

Performance. Opus 5.5 is a major step up from Opus 5. It’s the new leading model, and early testers saw large jumps in performance on their most complex work. One tester completed a 680,000-line code migration in less than a day—work that would have taken an engineering team weeks. It’s good at finding and fixing inefficiencies in software: when we asked it to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times, while Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt; Opus 5.5 scored higher than any other model on the strength of its graphics and polish. 性能。 Opus 5.5 是对 Opus 5 的重大升级。它是新的领先模型,早期测试者在处理最复杂的工作时,观察到了性能的显著飞跃。一位测试者在不到一天的时间内完成了 68 万行代码的迁移——这项工作原本需要一个工程团队花费数周时间。它擅长发现并修复软件中的低效问题:当我们要求它缩短一个 Web 应用所有页面的加载时间时,Opus 5.5 在 40 次尝试中成功了 39 次,而 Opus 5 虽然也做出了改进,但幅度较小且改变了应用的行为。另一位测试者让多个 Claude 模型根据同一个提示词构建游戏;Opus 5.5 凭借其图形表现和完成度,得分高于其他所有模型。

Safety. Opus 5.5 achieves the best scores of any model to date on our automated behavioral audit, our alignment suite that tests Claude across thousands of simulated scenarios. It is much less likely than recent models to take hard-to-reverse actions or act outside the boundaries it’s been given, and it’s more resistant than Opus 5 to prompt injection. We’ve also broadened our alignment testing to cover longer tasks, impossible tasks, and scenarios modeled on real incidents, though it still has limits. Full details of our evaluation are available in the Opus 5.5 System Card. 安全性。 Opus 5.5 在我们的自动化行为审计中取得了迄今为止所有模型的最高分,这是一套在数千个模拟场景中测试 Claude 的对齐套件。与近期模型相比,它采取难以逆转的操作或超出既定边界行为的可能性大大降低,并且比 Opus 5 更能抵御提示词注入攻击。我们还扩大了对齐测试的范围,涵盖了更长的任务、不可能完成的任务以及基于真实事件建模的场景,尽管它仍存在局限性。评估的完整细节可在《Opus 5.5 系统卡》中查阅。

Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, we’re deploying it with safeguards similar to those on Claude Fable 5.1. Vetted organizations can apply today to our Life Sciences Verification Program to use Opus 5.5 for biology research. In the coming weeks we will also be expanding access to our Cyber Verification Program, and verified cybersecurity practitioners will be able to use Opus 5.5 for their work. 由于 Opus 5.5 在生物学和网络安全领域与 Claude Mythos 5.1 相当,我们为其部署了与 Claude Fable 5.1 类似的防护措施。经过审核的机构即日起可申请加入我们的“生命科学验证计划”,以使用 Opus 5.5 进行生物学研究。在接下来的几周内,我们还将扩大对“网络安全验证计划”的访问权限,届时经过验证的网络安全从业者将能够使用 Opus 5.5 进行工作。

Cost and speed. Opus 5.5 requires less compute to serve than Opus 5, and its pricing reflects that. Our tests show that at default settings it will cost 40% less than Opus 5 on typical workloads. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5. 成本与速度。 Opus 5.5 的服务计算需求低于 Opus 5,其定价也反映了这一点。我们的测试显示,在默认设置下,典型工作负载的成本比 Opus 5 低 40%。输入和输出 Token 价格分别为每百万 Token 4 美元和 20 美元,比 Opus 5 降低了 20%。缓存读取(占智能体和编码工作成本的大部分)价格为每百万 Token 0.20 美元,比 Opus 5 降低了 60%。此外,Opus 5.5 的输出生成速度比 Opus 5 快 30% 以上。

In addition to the price drop, we’re increasing five-hour usage limits on Pro, Max, Team, and seat-based Enterprise plans. We’re also providing subscription users a rate limit reset, which you can now save and use whenever you choose. 除了降价之外,我们还提高了 Pro、Max、Team 以及基于席位的企业版计划的五小时使用限额。我们还为订阅用户提供了速率限制重置功能,您可以将其保存并在需要时随时使用。

Communication. Opus 5.5 communicates more naturally than prior models. Early testers found its writing clearer and easier to follow, which addresses some of the common feedback we heard about Opus 5. It puts the most important information up front, and its style makes it a better work partner over long sessions. As one early tester put it, “it writes the way I do.” In our own use, this has made Opus 5.5’s work easier to follow and check—which is a safety benefit as well as a practical one. 沟通。 Opus 5.5 的沟通方式比以往模型更自然。早期测试者发现其写作更清晰、更易于理解,这解决了我们关于 Opus 5 收到的一些常见反馈。它将最重要的信息放在前面,其风格使其成为长会话中更好的工作伙伴。正如一位早期测试者所言:“它的写作方式和我一样。”在我们自己的使用中,这使得 Opus 5.5 的工作成果更容易跟踪和检查——这既是实际应用上的优势,也是安全上的优势。

Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency, and safety. Claude Sonnet 5.5 和 Claude Haiku 5.5 将在未来几周内发布,它们也将具备许多相同的性能、效率和安全性改进。

Performance and cost-effectiveness

性能与成本效益

On our benchmarks, Claude Opus 5.5 leads in agentic coding, computer use, and knowledge work. That said, at these levels of capability we’ve found that benchmark margins have become a less reliable guide to real-world differences. In our own use, the gap between Opus 5.5 and Claude Fable 5.1 is narrower than these scores suggest. 在我们的基准测试中,Claude Opus 5.5 在智能体编码、计算机使用和知识工作方面处于领先地位。话虽如此,在目前的这种能力水平下,我们发现基准测试的差距已不再是衡量现实世界差异的可靠指标。在我们自己的使用中,Opus 5.5 与 Claude Fable 5.1 之间的差距比这些分数所显示的要小。

(Note: The table and technical footnotes have been omitted for brevity as per standard editorial practice for this format.) (注:为保持格式简洁,此处省略了表格及技术脚注。)