New Anthropic, OpenAI models make same promise: A little more for a lot less money

New Anthropic, OpenAI models make same promise: A little more for a lot less money

Anthropic 与 OpenAI 发布新模型,承诺以更低成本提供更强性能

OpenAI and Anthropic both recently released new models aimed at lowering costs. Anthropic announced Opus 5.5, the latest version of its main mass-market workhorse model, used for tasks like coding and other complex knowledge work. And OpenAI announced GPT-6 Sol and Luna, the latest versions of its middle-of-the-road or smaller models focused on efficiency and speed. OpenAI 和 Anthropic 近期均发布了旨在降低成本的新模型。Anthropic 推出了 Opus 5.5,这是其面向大众市场的主力模型最新版本,主要用于编程及其他复杂的知识型工作。OpenAI 则发布了 GPT-6 Sol 和 Luna,这是其专注于效率与速度的中端或小型模型系列的最新版本。

These new releases are not about groundbreaking new capabilities. Rather, they’re about efficiency. As both OpenAI and Anthropic target enterprise customers, they’re racing to compete with open-weight models as organizations have explored changing their practices and using model routers to use these pricey, frontier models less in favor of cheaper alternatives. Anthropic and OpenAI argue that these new releases push the envelope at the frontier (albeit mostly in modest ways), while bringing costs substantially down. 这些新发布的产品并非为了展示突破性的新功能,而是侧重于效率。随着 OpenAI 和 Anthropic 竞相争夺企业客户,它们正面临来自开源权重模型的竞争压力。许多机构正在探索改变工作流程,利用模型路由技术减少对昂贵的前沿模型的使用,转而采用更便宜的替代方案。Anthropic 和 OpenAI 认为,这些新模型在推动前沿技术发展的同时(尽管幅度较为温和),也大幅降低了成本。

Opus 5.5: A more affordable, more competitive flagship

Opus 5.5:更实惠、更具竞争力的旗舰产品

Opus 5.5 is at the higher end of the models announced today, but the wider context here is that Anthropic is playing a bit of catch-up in its race with OpenAI. OpenAI earlier this month released GPT-6 Astra, which has sometimes been modestly beating Opus 5 in benchmarks and user sentiment. (By price and capability, Astra is competing with both Opus and Fable.) Benchmarks by Anthropic and its partners now show Opus 5.5 performing better at coding and knowledge work than GPT-6 Astra in some cases, albeit modestly. Opus 5.5 是今天发布的高端模型,但从更广阔的背景来看,Anthropic 在与 OpenAI 的竞争中正处于追赶状态。OpenAI 本月初发布了 GPT-6 Astra,该模型在基准测试和用户评价中偶尔会小幅领先于 Opus 5。(从价格和能力来看,Astra 同时与 Opus 和 Fable 竞争。)Anthropic 及其合作伙伴的基准测试显示,Opus 5.5 在编程和知识工作方面的表现已在某些情况下优于 GPT-6 Astra,尽管优势较为微弱。

The real story here is cost. From Anthropic’s announcement: “Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads (which make up the majority of agentic and coding work costs) are $0.20 per million tokens, 60% less than Opus 5. Opus 5.5 also generates output more than 30% faster than Opus 5.” Anthropic further claims that the savings are closer to 40 percent compared to Opus 5 for typical workloads at default settings, because in addition to the cost of tokens going down, Opus 5.5 also uses fewer tokens when completing tasks. 真正的重点在于成本。根据 Anthropic 的公告:“输入和输出 token 的价格分别为每百万 4 美元和 20 美元,比 Opus 5 降低了 20%。缓存读取(占智能体和编程工作成本的大部分)价格为每百万 token 0.20 美元,比 Opus 5 降低了 60%。此外,Opus 5.5 的输出生成速度比 Opus 5 快 30% 以上。”Anthropic 进一步声称,在默认设置下的典型工作负载中,相较于 Opus 5,实际节省的成本接近 40%,因为除了 token 价格下降外,Opus 5.5 在完成任务时使用的 token 数量也更少。

Opus 5.5 is said to be notably capable in “high risk areas” like cybersecurity and biology, so the same protections that applied to Fable 5.1 will also apply here—your requests might be automatically and transparently routed to an older model if they get flagged as treading into protected territory. 据悉,Opus 5.5 在网络安全和生物学等“高风险领域”表现出色,因此适用于 Fable 5.1 的保护措施同样适用于此——如果你的请求被标记为涉及受保护领域,系统可能会自动且透明地将其路由至旧版本模型。

GPT-6 Sol and Luna: A modest, cost-focused upgrade

GPT-6 Sol 和 Luna:以成本为导向的温和升级

The release of OpenAI’s GPT-6 Sol and Luna is an iterative step forward. Not long ago, the company introduced its Sol, Terra, Luna naming convention for its GPT-5.6 family of models. And more recently (just this month), it introduced GPT-6 Astra, which has been both its most advanced and most powerful model. It may not be obvious what those names mean, so here’s the quick rundown. Astra is the most aggressively powerful (and pricey) model, meant for heavy-duty coding, research, and so on. Sol is a capable but more efficient and affordable alternative—meant to be the sort of daily driver for tasks like that. Terra is the balanced, general-use model. And Luna is the fast, cheap option. OpenAI 发布 GPT-6 Sol 和 Luna 是迭代式的前进。不久前,该公司为其 GPT-5.6 系列模型引入了 Sol、Terra 和 Luna 的命名规范。最近(就在本月),它又推出了 GPT-6 Astra,这是其迄今为止最先进、最强大的模型。这些名称的含义可能不太直观,以下是简要说明:Astra 是功能最强大(也最昂贵)的模型,专为繁重的编程、研究等任务设计;Sol 是一个功能强大但更高效、更实惠的替代方案,旨在成为处理此类任务的日常主力;Terra 是平衡的通用模型;而 Luna 则是快速、廉价的选择。

You could reasonably and roughly position Astra against Anthropic’s Fable, Sol against Opus, Terra against Sonnet, and Luna against Haiku, but it would be an imperfect mapping, especially for the higher-end models. OpenAI says GPT-6 Sol and Luna were trained with similar methods to those used to train GPT-6 Astra. Depending on the benchmark, they’re sometimes a few percentage points more capable than their predecessors at certain tasks, but they cost half as much to use. GPT-6 Sol’s API pricing is $2 per 1 million input tokens and $10 per 1 million output tokens. For Luna, it’s $0.10 and $0.50, respectively. 你可以大致将 Astra 对标 Anthropic 的 Fable,Sol 对标 Opus,Terra 对标 Sonnet,Luna 对标 Haiku,但这并非完美的映射,尤其是在高端模型方面。OpenAI 表示,GPT-6 Sol 和 Luna 的训练方法与 GPT-6 Astra 类似。根据基准测试的不同,它们在某些任务上的能力有时比前代产品高出几个百分点,但使用成本却降低了一半。GPT-6 Sol 的 API 定价为每百万输入 token 2 美元,每百万输出 token 10 美元;Luna 的价格则分别为 0.10 美元和 0.50 美元。

Opinion: The devil is in the dollars

观点:魔鬼藏在金钱里

The discourse around frontier models is chaotic. You have people on social media declaring that it’s possible to one-shot complex 3D video games with models like GPT-6 Astra, but you also have news of security breaches and other alignment issues, calls for slowdowns and regulation, and so on. Some of that is worthy of serious attention, while some of it is noise—and some of it is serious, but being spun or exaggerated for commercial positioning. 围绕前沿模型的讨论十分混乱。社交媒体上有人宣称 GPT-6 Astra 等模型可以“一键生成”复杂的 3D 电子游戏,但同时也有关于安全漏洞和其他对齐问题的报道,以及要求放缓开发和加强监管的呼声。其中一些值得认真关注,有些则是噪音,还有一些虽然严肃,却被为了商业定位而扭曲或夸大。

Further, there’s a growing recognition that the models—be they Anthropic’s, OpenAI’s, Alibaba’s, or any other big player’s—are not the only important engines of progress and results in recent months. The orchestration and harnesses for those models are at least as important. The frontier is still moving forward, and today’s models are apparently more capable and aligned than those from a few months ago. But some developers and enterprises are a little less focused on the frontier and more focused on exactly how to operationalize all this and keep it cost-effective. 此外,人们越来越认识到,无论是 Anthropic、OpenAI、阿里巴巴还是其他大厂的模型,都不是近几个月来推动进步和成果的唯一引擎。这些模型的编排和配套工具至少同样重要。前沿技术仍在不断推进,今天的模型显然比几个月前的模型能力更强、对齐更好。但一些开发者和企业对“前沿”的关注度有所下降,转而更关注如何将这些技术落地并保持成本效益。

So while we’re seeing AI leaders calling for a slowdown ostensibly or partially for safety reasons, there’s also a practical and economic reality: The models we have now are good enough to do a lot of helpful things, but they require sophisticated contextualization and operationalization by human beings—whether in the runtime and harnesses or in organizational practices—to be put to use. As a result, some of these companies’ customers may be increasingly less focused on demanding better performance. They’re looking for predictable deployments and, most of all, reasonable costs. That has the potential to be a natural slowdown of its own, and these models are both being positioned for that new, on-the-ground reality. 因此,尽管我们看到 AI 领袖们表面上或部分出于安全原因呼吁放缓开发,但背后也存在着实际的经济现实:我们现有的模型已经足以完成许多有用的工作,但它们需要人类进行复杂的语境化和操作化——无论是在运行时环境、配套工具还是组织实践中——才能真正投入使用。结果是,这些公司的一些客户可能越来越不追求极致的性能,而是寻求可预测的部署,最重要的是,合理的成本。这本身可能成为一种自然的放缓,而这些新模型正是为了适应这种新的现实而定位的。