Writer introduces new AI model and upgraded harness to contain token costs

Writer introduces new AI model and upgraded harness to contain token costs

Writer 推出全新 AI 模型及升级版架构,旨在控制 Token 使用成本

Across the AI industry, users are becoming more conscious of just how expensive their deployments can be —and feeling a new urgency to cut costs. But while open source models offer significantly lower per-token costs, it can be difficult to find the right model for a given job. 在整个 AI 行业中,用户正愈发意识到部署成本的高昂,并感受到削减成本的紧迫性。尽管开源模型提供了显著更低的单 Token 成本,但要为特定任务找到合适的模型往往并非易事。

On Thursday, Writer, which offers AI tools and agents for marketers, launched a new flagship model called Palmyra X6, aimed at solving that problem for its users. Built as a post-training variation on Z.ai’s open source model GLM-5.2, Writer says the new system should provide deployment-ready capabilities at a much lower price. 周四,为营销人员提供 AI 工具和智能体的 Writer 公司发布了名为 Palmyra X6 的全新旗舰模型,旨在为用户解决上述难题。该模型基于 Z.ai 的开源模型 GLM-5.2 进行后训练优化,Writer 表示,新系统能够以更低的价格提供可直接部署的能力。

The company estimates the new model, combined with changes to the companies harness infrastructure, will cut costs for its customers by as much as 50% for basic tasks. Together with the new model, the company also released significant upgrades to its standard agentic harness. Both features will be available to Writer clients starting Thursday. 该公司预计,新模型结合其架构基础设施的改进,将使客户在处理基础任务时的成本降低高达 50%。除了新模型,该公司还对其标准的智能体架构(agentic harness)进行了重大升级。这两项功能已于周四向 Writer 的客户开放。

“I think the enterprise is absolutely sick of chasing the next benchmark,” CEO May Habib told TechCrunch. “They want flattening cost, and it seems like nobody can deliver that.” “我认为企业已经厌倦了盲目追逐下一个基准测试,”首席执行官 May Habib 在接受 TechCrunch 采访时表示,“他们想要的是成本平稳化,但似乎没人能做到这一点。”

The new approach puts particular emphasis on complex, multi-step tasks, executed faster and with fewer tokens. And Writer sees harness optimization as a crucial lever toward making that happen. A recent paper from Writer researchers lends credence to this approach, testing small changes in harness efficiency across multiple different models. 这种新方法特别强调复杂的多步骤任务,旨在以更快的速度和更少的 Token 完成任务。Writer 将架构优化视为实现这一目标的关键杠杆。Writer 研究人员最近发表的一篇论文为这一方法提供了佐证,该论文测试了在多个不同模型中对架构效率进行微小调整的效果。

The research found that, in many cases, changes in the harness were a more reliable way to reduce costs than model choice, with costs falling an average of 40% across their testing. “The harness is the one component whose efficiency multiplies across every model an organization runs—present and future,” the researchers wrote. 研究发现,在许多情况下,调整架构比单纯选择模型更能可靠地降低成本,测试中的平均成本降幅达到 40%。研究人员写道:“架构是唯一一个其效率能够叠加到组织运行的所有模型(无论是现在还是未来)上的组件。”

For Writer’s clients, the experience is still model-agnostic: Palmyra X6 will sit alongside other Writer models or outside models imported through Azure or Amazon Bedrock. But Habib also sees the push to cut costs as driving a broader distrust toward major AI labs, which have a financial incentive to drive up token use. 对于 Writer 的客户来说,使用体验依然与模型无关:Palmyra X6 将与其他 Writer 模型或通过 Azure 或 Amazon Bedrock 导入的外部模型并存。但 Habib 也认为,削减成本的压力正在加剧人们对大型 AI 实验室的普遍不信任,因为这些实验室在经济利益驱动下倾向于提高 Token 的使用量。

“The cost explosion here is just unprecedented for customers, and so is the degree to which CIOs are giving up on the labs,” Habib told TechCrunch, adding that the AI labs “don’t deeply understand right how to help an enterprise get benefit from AI.” “对于客户而言,这种成本爆炸是前所未有的,首席信息官(CIO)们对这些实验室的失望程度也是如此,”Habib 对 TechCrunch 说道,并补充称这些 AI 实验室“并不真正了解如何帮助企业从 AI 中获益。”