Anthropic's Opus 5 is about token efficiency, not a capability leap

Anthropic’s Opus 5 is about token efficiency, not a capability leap

Anthropic 的 Opus 5 旨在提升 Token 效率,而非性能的飞跃

Today, Anthropic rolled out Opus 5, the newest update for the model that has recently become a popular choice for coding and other software development tasks, among other things. While this is a noteworthy bump for Opus, it doesn’t seem to be an Opus 4.5-level breakthrough in agentic coding performance. 今天,Anthropic 推出了 Opus 5,这是该模型系列的最新更新。该模型近期已成为编程及其他软件开发任务的热门选择。虽然这对 Opus 来说是一个值得注意的提升,但它似乎并未达到 Opus 4.5 那种在智能体编程性能上的突破水平。

A chart (made by Anthropic) with various benchmarks like Frontier-Bench and DeepSWE shows Opus 5 performing at about the same level or slightly ahead of the much-ballyhooed capabilities of Anthropic’s Fable model for coding tasks. It ostensibly beats Opus 4.8 and OpenAI’s competing GPT-5.6-Sol in just about every kind of task. The various benchmarks show an iterative increase in performance, but not a radical leap, while the main pitch is that it’s a model that offers something just shy of Fable at approximately half the cost. 一张由 Anthropic 制作的图表显示,在 Frontier-Bench 和 DeepSWE 等多项基准测试中,Opus 5 的表现与备受推崇的 Anthropic Fable 模型在编程任务上的能力相当,甚至略有领先。从表面上看,它在几乎所有类型的任务中都击败了 Opus 4.8 和 OpenAI 的竞争对手 GPT-5.6-Sol。各项基准测试显示性能呈迭代式增长,而非彻底的飞跃;其核心卖点在于,它以大约一半的成本提供了接近 Fable 的性能。

It’s also worth mentioning that Anthropic specifically avoided giving Opus 5 cutting-edge training on cybersecurity tasks, so it lags way behind Fable and Mythos in that regard. Anthropic claims it is relatively good at finding cybersecurity vulnerabilities, but because of decisions made in training the model, it is “substantially behind Mythos 5 on the exploitation of those vulnerabilities.” As such, Opus 5 doesn’t have all of the same controversial protections that Fable had, such as the policy of keeping data for review for 30 days in case of an incident. 值得一提的是,Anthropic 特意避免了让 Opus 5 接受网络安全任务方面的尖端训练,因此在这方面它远远落后于 Fable 和 Mythos。Anthropic 声称它在发现网络安全漏洞方面表现尚可,但由于模型训练时的决策,它在“利用这些漏洞方面明显落后于 Mythos 5”。因此,Opus 5 并不具备 Fable 所拥有的所有争议性保护措施,例如为应对突发事件而保留数据 30 天以供审查的政策。

But this is really a cost story. From one model update to the next subsequent one, we’ve mostly been seeing modest performance improvements—though they look far more dramatic when you look at a whole year of updates. The discourse among software developers and engineering managers right now is mostly around dealing with cost, and there’s real movement in open-weight models, local models, and other alternatives to cutting-edge frontier models. 但这本质上是一个关于成本的故事。从一个模型更新到下一个更新,我们看到的主要是适度的性能提升——尽管当你回顾一整年的更新时,这些变化看起来要显著得多。目前,软件开发人员和工程经理之间的讨论主要围绕如何控制成本,而开放权重模型、本地模型以及其他替代尖端前沿模型的方案正展现出真正的活力。

Opus 5 sits at $5 per million input tokens, or $25 per million output tokens—on par with its predecessor, but cheaper than Fable. That said, the recently announced Chinese open-weight model Kimi K3 comes in at just $15 per million output tokens at similar performance, so competition is fierce. Opus 5 的定价为每百万输入 Token 5 美元,或每百万输出 Token 25 美元——与前代产品持平,但比 Fable 更便宜。话虽如此,最近发布的中国开放权重模型 Kimi K3 在性能相当的情况下,每百万输出 Token 的价格仅为 15 美元,竞争十分激烈。

Lately, companies like Cursor and Meta have been building “model routers,” systems that automatically select models of varying size and capability from an array of options based on the nature of the prompt. The idea is that you save a lot of tokens (and therefore compute and/or money) by not using something like Fable for every task. Companies like Anthropic are going to have to keep bringing token costs down, or at least (as is the case here) offering more performance for no additional cost, to continue to see the growth they’ve seen. Otherwise, they will start seeing reduced usage as people turn to smaller or open models once those models get good enough to handle the less challenging development tasks. 最近,Cursor 和 Meta 等公司一直在构建“模型路由器”(model routers),这是一种根据提示词的性质,从一系列选项中自动选择不同规模和能力模型的系统。其理念是,通过避免在每个任务上都使用像 Fable 这样的大模型,可以节省大量的 Token(从而节省计算资源和/或资金)。像 Anthropic 这样的公司必须不断降低 Token 成本,或者至少(如本例所示)在不增加成本的情况下提供更高的性能,才能维持其增长势头。否则,一旦小型或开放模型足以处理较简单的开发任务,用户就会转向这些模型,从而导致 Anthropic 的使用量下降。