Claude Haiku 5.5

本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。

Introducing Claude Haiku 5.5: the cheapest, fastest, and most capable small model we’ve ever released.

隆重推出 Claude Haiku 5.5:这是我们迄今为止发布的最便宜、最快且能力最强的小型模型。

Claude Haiku 5.5 is designed for high-volume, cost-sensitive tasks. It reliably handles quick and repetitive workloads (like summaries, compactions, database queries, and classification requests). It pairs well with Opus 5.5 and Sonnet 5.5 as a subagent on coding work. And, since it’s also our fastest model to date, it works especially well for speed-sensitive tasks like live customer support and browser use.¹

Claude Haiku 5.5 专为高容量、对成本敏感的任务而设计。它能可靠地处理快速且重复的工作负载(如摘要、压缩、数据库查询和分类请求)。它非常适合与 Opus 5.5 和 Sonnet 5.5 搭配,作为编码工作中的子代理使用。此外,由于它也是我们迄今为止速度最快的模型,因此在实时客户支持和浏览器使用等对速度敏感的任务中表现尤为出色。¹

Haiku 5.5 is available at a much lower price than Haiku 4.5. On average, it now costs around 75% less to run.²

Haiku 5.5 的价格远低于 Haiku 4.5。平均而言,其运行成本现在降低了约 75%。²

Along with this launch, we’re making improvements to the value of our model range. We’re halving the price of Claude Sonnet 5.5’s cache reads, which means Sonnet 5.5 now runs around 20% cheaper on most agentic work. And we’re introducing a new monthly API credit for our Claude Max and Team subscribers, designed to support our users in building new agents and applications that run on the Claude Platform.

随着此次发布,我们正在提升模型系列的产品价值。我们将 Claude Sonnet 5.5 的缓存读取价格减半,这意味着 Sonnet 5.5 在大多数代理任务中的运行成本降低了约 20%。此外,我们还为 Claude Max 和 Team 订阅用户推出了新的每月 API 额度,旨在支持用户构建在 Claude 平台上运行的新代理和应用程序。

Here’s how Claude Haiku 5.5 performs across a range of benchmarks:

以下是 Claude Haiku 5.5 在一系列基准测试中的表现:

For details on how we run our evaluations, see the Haiku 5.5 System Card.

有关我们如何进行评估的详细信息,请参阅 Haiku 5.5 系统卡。

Haiku 5.5 is our first Haiku-class model to come with an adjustable effort setting. This means that, as with our other models, users can decide whether to optimize for cost or intelligence. The charts below show how Haiku 5.5 performs on three benchmarks at each effort setting:

Haiku 5.5 是我们首款配备可调节工作量设置的 Haiku 级模型。这意味着,与我们的其他模型一样,用户可以决定是优化成本还是优化智能。下表显示了 Haiku 5.5 在每个工作量设置下三个基准测试中的表现:

OSWorld 2.1 measures how well agents can operate a real computer to finish long, multi-step tasks.

OSWorld 2.1 用于衡量代理操作真实计算机以完成长期、多步骤任务的能力。

Artificial Analysis’s GDPval-AA v2.1 evaluates agents on real-world professional work across 44 occupations.

Artificial Analysis 的 GDPval-AA v2.1 在 44 个职业中评估代理在现实专业工作中的表现。

Humanity’s Last Exam (HLE) is a test of expert-level academic knowledge and reasoning.

“人类最后考试”(HLE)是对专家级学术知识和推理能力的测试。

In early testing, our customers reported results consistent with the performance and cost improvements shown above. Here’s what they told us about the new model:

在早期测试中,我们的客户报告的结果与上述性能和成本改进一致。以下是他们对新模型的评价:

“We’re very impressed with Claude Haiku 5.5, particularly its speed. We ran it through our eval suite for AI Teammates, our AI agent product, covering use cases like triaging bugs, setting up projects, and searching large portfolios to surface high-risk or overdue work. Compared with the model we use today, we saw over a 30% reduction in latency for task completions and up to 2.5x faster inference per agent turn. It’s a noticeably snappier experience.”

“我们对 Claude Haiku 5.5 印象深刻,尤其是它的速度。我们将其运行在我们的人工智能代理产品 AI Teammates 的评估套件中,涵盖了分类错误、设置项目以及搜索大型投资组合以发现高风险或逾期工作等用例。与我们目前使用的模型相比,任务完成延迟降低了 30% 以上,每次代理轮次的推理速度提高了 2.5 倍。这是一种明显更快捷的体验。”

“At HubSpot, we use simulated portals to evaluate new models on CRM tasks like reporting on deals. We mostly test the smaller, more efficient models, and Claude Haiku 5.5 got the best score we’ve seen on this suite yet, at 92.8% averaged over three runs. One CRM audit task asks models to identify stale but ambiguous records. Across all of the models we tested, Haiku 5.5 was fastest to complete the task, and had the highest hit rate and the lowest false positive rate.”

“在 HubSpot,我们使用模拟门户来评估新模型在 CRM 任务(如交易报告)上的表现。我们主要测试较小、更高效的模型,而 Claude Haiku 5.5 在该套件中获得了我们迄今为止见过的最高分,三次运行的平均分为 92.8%。一项 CRM 审计任务要求模型识别陈旧但模棱两可的记录。在我们测试的所有模型中,Haiku 5.5 完成任务的速度最快,命中率最高,误报率最低。”

“Ask in Document is one of our big sources of spend, doing about 8M calls a week in production. It answers very specific questions on top of one or a few documents. We ran 400 queries, and Claude Haiku 5.5 was a statistically significant improvement over Haiku 4.5: 0.84 vs. 0.76.”

“‘文档提问’(Ask in Document)是我们的一大支出来源,在生产环境中每周进行约 800 万次调用。它基于一份或几份文档回答非常具体的问题。我们运行了 400 个查询,Claude Haiku 5.5 比 Haiku 4.5 有了统计学意义上的显著改进:0.84 对比 0.76。”

“Our customers use Box AI across large volumes of their enterprise content. With widespread usage comes the need to manage efficiency and cost, and to find the best model to suit the task at hand. In early testing, Claude Haiku 5.5 scored 11 points higher than Haiku 4.5 at about half the latency. We’d put it to use on analytical work that runs at scale, from cost reports to financial summaries and weekly recurring reviews.”

“我们的客户在大量企业内容中使用 Box AI。随着广泛使用,我们需要管理效率和成本,并找到最适合当前任务的模型。在早期测试中,Claude Haiku 5.5 的得分比 Haiku 4.5 高出 11 分,而延迟仅为后者的一半左右。我们将把它用于大规模的分析工作,从成本报告到财务摘要和每周定期审查。”

“The short and high-volume work is where Claude Haiku 5.5 fits for us, like quick lookups, subagents, and summaries. While a bigger model builds the deck, a Haiku 5.5 subagent goes into the 10-K and pulls the segment revenue line the deck needs. It’s accurate enough that we’d trust it there, and fast and cheap enough that we can run it a lot.”

“Claude Haiku 5.5 非常适合我们处理简短且高容量的工作,例如快速查找、子代理和摘要。当更大的模型构建演示文稿时,Haiku 5.5 子代理会进入 10-K 文件并提取演示文稿所需的分部收入行。它的准确性足以让我们信任,而且速度快、成本低,我们可以频繁使用它。”

“Claude Haiku 5.5 joins the sidekick lineup in Devin Fusion as an excellent option. With Haiku 5.5 as the sidekick, Fusion holds a top-tier FrontierCode score of 66.2 while cutting cost and latency. You can try it today in the Devin CLI with Opus 5.5 as the lead.”

“Claude Haiku 5.5 作为绝佳选择加入了 Devin Fusion 的助手阵容。以 Haiku 5.5 作为助手,Fusion 在降低成本和延迟的同时,保持了 66.2 的顶级 FrontierCode 分数。您今天就可以在 Devin CLI 中尝试它,并以 Opus 5.5 作为主导。”

The table below shows how Claude Haiku 5.5’s pricing compares to our other models. Haiku 5.5 is especially good value when used for tasks with prompts up to 100,000 tokens, which make up around 90% of requests to our previous Haiku model.

下表显示了 Claude Haiku 5.5 的定价与其他模型的对比。当用于提示词长度不超过 100,000 个 token 的任务时,Haiku 5.5 的性价比尤为突出,这类任务占我们之前 Haiku 模型请求的约 90%。

Alignment. Claude Haiku 5.5 shows major improvements across almost all of our alignment evaluations relative to Haiku 4.5. In particular, we found far fewer instances of misaligned behavior, and a lower willingness to cooperate with misuse. The model’s system card describes our evaluation process and results in more detail.

对齐。相对于 Haiku 4.5,Claude Haiku 5.5 在我们几乎所有的对齐评估中都表现出重大改进。特别是,我们发现未对齐行为的实例大大减少,并且与滥用行为合作的意愿也降低了。该模型的系统卡更详细地描述了我们的评估过程和结果。

Safeguards. Consistent with its capabilities, Haiku 5.5’s cybersecurity safeguards are more restrictive than Haiku 4.5’s, but somewhat less restrictive than those we’ve applied to other recent models. In cybersecurity, they permit a wider range of defensive tasks than our safeguards for Sonnet 5.5, but they still block penetration testing and other techniques more likely to be used by attackers.

安全防护。与其能力相一致,Haiku 5.5 的网络安全防护比 Haiku 4.5 更严格,但比我们应用于其他近期模型的防护措施略微宽松。在网络安全方面,它们允许比 Sonnet 5.5 的防护措施更广泛的防御任务,但仍然会阻止渗透测试和其他攻击者更可能使用的技术。

Haiku 5.5’s biology safeguards are the same as for Sonnet 5, Sonnet 5.5, and Opus 5. They allow research biology questions but restrict access to requests that we judge as likely to cause harm. Organizations working on wider-ranging biology and cyber activities can apply to our Life Sciences Verification Program and Cyber Verification Program.

Haiku 5.5 的生物安全防护与 Sonnet 5、Sonnet 5.5 和 Opus 5 相同。它们允许研究性生物学问题,但限制访问我们判断可能造成伤害的请求。从事更广泛生物学和网络活动的研究机构可以申请加入我们的生命科学验证计划和网络验证计划。

Claude Haiku 5.5 is available now on all platforms, including Amazon Web Services, Google Cloud, and Microsoft Azure. On the Claude Platform, developers can get started with claude-haiku-5-5.

Claude Haiku 5.5 现已在所有平台上提供,包括 Amazon Web Services、Google Cloud 和 Microsoft Azure。在 Claude 平台上,开发人员可以开始使用 claude-haiku-5-5。

See our migration guide for details.

有关详细信息,请参阅我们的迁移指南。

Alongside our new pricing for Claude Haiku 5.5, we’re making further improvements to the value of our models and products.

除了 Claude Haiku 5.5 的新定价外,我们还在进一步提升模型和产品的价值。

First, starting today, we’re lowering the price of cache reads on Claude Sonnet 5.5. Cache reads now cost 50% less: $0.10 per million tokens rather tha

首先,从今天开始,我们降低了 Claude Sonnet 5.5 的缓存读取价格。缓存读取成本现在降低了 50%:每百万 token 0.10 美元,而不是……