Sonnet 5.5

Sonnet 5.5

Introducing Claude Sonnet 5.5, the second model in the Claude 5.5 family. It’s a clear upgrade over Claude Sonnet 5, runs 30%+ faster, and costs up to 30% less for most work. 隆重推出 Claude 5.5 系列的第二款模型:Claude Sonnet 5.5。相比 Claude Sonnet 5,它有着显著的提升,运行速度提高了 30% 以上,且在大多数工作负载下成本降低了多达 30%。

Sonnet 5.5 is a faster, lower-cost complement to Claude Opus 5.5. Where Opus 5.5 is built for complex work requiring careful judgment, Sonnet 5.5 is strongest at well-scoped everyday tasks, fixing bugs, and creating polished documents, slides, and spreadsheets. It’s also got a sharp eye for design. Claude Haiku 5.5, built for high-volume and cost-sensitive applications, will join the Claude 5.5 family in the coming weeks. Sonnet 5.5 是 Claude Opus 5.5 更快速、更经济的补充。如果说 Opus 5.5 是为需要审慎判断的复杂工作而生,那么 Sonnet 5.5 则在明确范围的日常任务、修复 Bug 以及创建精美的文档、幻灯片和电子表格方面表现最为出色。它还具备敏锐的设计洞察力。专为高并发和成本敏感型应用打造的 Claude Haiku 5.5,也将在未来几周内加入 Claude 5.5 家族。

Sonnet 5.5 improves over Sonnet 5 on: Sonnet 5.5 在以下方面优于 Sonnet 5:

Performance. Sonnet 5.5 scores 70.6% on Terminal-Bench 4.0, an agentic coding evaluation, compared to Sonnet 5’s 10.3%. It scores two points below Opus 5.5 on GDPval-AA, a test of real-world work across a variety of occupations. And it’s strong on long-horizon work and image understanding—it’s the first Sonnet model to beat Pokémon Red working only from screenshots. 性能:在代理编程评估 Terminal-Bench 4.0 中,Sonnet 5.5 的得分为 70.6%,而 Sonnet 5 仅为 10.3%。在涵盖多种职业的真实工作测试 GDPval-AA 中,它仅比 Opus 5.5 低两分。此外,它在长周期任务和图像理解方面表现强劲——它是首个仅通过截图就能通关《宝可梦:红》的 Sonnet 模型。

Collaboration. Like Opus 5.5, Sonnet 5.5 writes more clearly than our previous generation of models; early testers described it as a better partner for collaboration than Sonnet 5. Its speed also makes it well suited to fast iteration on less complex tasks. 协作:与 Opus 5.5 一样,Sonnet 5.5 的写作比我们上一代模型更清晰;早期测试者称它比 Sonnet 5 更适合作为协作伙伴。其速度也使其非常适合对复杂度较低的任务进行快速迭代。

Cost. Sonnet 5.5 is priced the same as Sonnet 5 at $2 per million input tokens, $10 per million output tokens, and $0.20 per million tokens for cache reads, but it typically needs far fewer tokens to do the same work. In our testing, it costs up to 30% less per task than its predecessor. 成本:Sonnet 5.5 的定价与 Sonnet 5 相同,即每百万输入 Token 2 美元,每百万输出 Token 10 美元,缓存读取每百万 Token 0.20 美元,但它通常完成同样工作所需的 Token 数量要少得多。在我们的测试中,它每项任务的成本比前代产品降低了多达 30%。

Speed. Sonnet 5.5 generates outputs 30%+ faster than Sonnet 5, making it our fastest Sonnet model to date. 速度:Sonnet 5.5 的输出生成速度比 Sonnet 5 快 30% 以上,使其成为我们迄今为止最快的 Sonnet 模型。

Alignment and safety. On our automated behavioral audit, Sonnet 5.5 improves on or matches Sonnet 5 on most measures of alignment. Because its cybersecurity capabilities are comparable to Opus 5’s, it’s the first Sonnet model to launch with cyber safeguards and fallbacks like those we’ve developed for our most capable models. Its biology safeguards are the same as Sonnet 5’s. Both safeguards target a narrow set of high-risk requests; routine software development and most life sciences work are unaffected. 对齐与安全:在我们的自动化行为审计中,Sonnet 5.5 在大多数对齐指标上均优于或持平于 Sonnet 5。由于其网络安全能力与 Opus 5 相当,它是首个配备了我们为最强模型开发的网络安全防护和回退机制的 Sonnet 模型。其生物安全防护与 Sonnet 5 相同。这两项防护措施均针对极少数高风险请求;常规软件开发和大多数生命科学工作不受影响。

Performance 性能 Sonnet 5.5 improves on Sonnet 5 across domains—in some cases dramatically. On several evaluations, Sonnet 5.5 at Max effort even performs comparably to Opus 5.5. However, benchmark scores capture only one facet of a model’s capabilities; in our own testing, and in that of external testers, Opus 5.5 remains clearly stronger at complex, open-ended work requiring sustained judgment. Sonnet 5.5 在各个领域都比 Sonnet 5 有所提升,在某些情况下提升幅度巨大。在多项评估中,处于“最大努力(Max effort)”模式下的 Sonnet 5.5 表现甚至可与 Opus 5.5 相媲美。然而,基准测试分数仅反映了模型能力的一个侧面;在我们自己以及外部测试者的测试中,Opus 5.5 在需要持续判断的复杂、开放式工作中依然明显更强。

For details on how we run our evaluations, see the Sonnet 5.5 System Card. 有关我们如何进行评估的详细信息,请参阅《Sonnet 5.5 系统卡》。

The charts below plot each model’s score against its cost per task at every effort level. As effort goes up, models typically work for longer, leading to a higher cost per task but generally also a higher score. The closer a point is to the top left of the chart, the more capability it delivers per dollar. 下方的图表绘制了每个模型在不同努力程度下的得分与单项任务成本的关系。随着努力程度的提高,模型通常会工作更长时间,导致单项任务成本增加,但通常也会获得更高的分数。数据点越靠近图表的左上角,意味着其单位美元所提供的能力越强。

On several benchmarks, Sonnet 5.5 at Low or Medium effort beats Sonnet 5’s best score for about a tenth of the cost per task. It complements Opus 5.5 best when running at lower effort settings, where it costs less per task. At higher settings, it can perform comparably at a similar cost. 在多项基准测试中,处于“低”或“中”努力程度的 Sonnet 5.5 以约十分之一的单项任务成本,超越了 Sonnet 5 的最佳得分。它在较低努力程度设置下能最好地补充 Opus 5.5,此时单项任务成本更低。在较高设置下,它能以相似的成本实现相当的性能。

Coding 编码 Sonnet 5.5’s jump in performance is particularly noticeable in coding. At High effort on FrontierCode, it scores 10 points higher than Sonnet 5 at the same setting, at about one fifteenth of the cost per task. On CursorBench, which tests models on tasks from real Cursor coding sessions, its best score is within about two points of Opus 5.5. Sonnet 5.5 在编码方面的性能飞跃尤为显著。在 FrontierCode 的“高”努力程度设置下,它的得分比同样设置下的 Sonnet 5 高出 10 分,而单项任务成本仅为后者的十五分之一左右。在测试真实 Cursor 编码会话任务的 CursorBench 上,其最佳得分与 Opus 5.5 的差距在两分以内。

Early testers appreciated how quickly Sonnet 5.5 can understand a codebase. 早期测试者非常赞赏 Sonnet 5.5 理解代码库的速度。