The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
The Price of Thinking: Reasoning Effort as a Model-Specific API Contract
思考的代价:作为模型特定 API 合约的推理努力
Abstract: API buyers purchase a dated contract, not a model name alone: the contract includes the requested and served model, reasoning-effort term or its omission, output rail, service product, prompt, and price schedule. 摘要: API 买家购买的是一份有时效的合约,而不仅仅是一个模型名称:该合约包含了请求和提供的模型、推理努力条款(或其缺失)、输出限制、服务产品、提示词以及价格表。
We study the reasoning-effort term through a registered paired contrast of Sonnet 5 with explicit high effort against the same model with effort omitted, using 30 AIME 2026 items and five calls per item. Every paid attempt was assigned one frozen terminal category, and inference resampled items while retaining their repeated calls. 我们通过对 Sonnet 5 模型进行注册配对对比研究来探讨“推理努力”这一条款,对比了显式高努力设置与省略努力设置下的表现,使用了 30 道 AIME 2026 题目,每道题调用五次。每次付费尝试都被分配了一个固定的终止类别,并在保留重复调用的同时对推理项目进行了重采样。
Mean delivered cost was $0.01031 per call higher under the explicit-high contract than under the omitted contract [+$0.00204, +$0.01974]. The corresponding accuracy contrast was +0.0133 [-0.0267, +0.0467]; we did not detect an accuracy difference, and the interval permits a gain of up to 4.67 percentage points that this design cannot rule out. 在显式高努力合约下,平均交付成本比省略努力合约高出每调用 0.01031 美元 [+0.00204, +0.01974]。相应的准确率对比为 +0.0133 [-0.0267, +0.0467];我们未检测到准确率差异,且该区间允许存在最高 4.67 个百分点的增益,这是本设计无法排除的。
Cost per correct answer was $0.08665 under the high-effort contract and $0.07662 under the omitted contract, as registered point estimates. A dated contract census, Models-API metadata, and preregistered raw-response probes further documented model-specific omission semantics, including within a provider; claims remained at documentation grade when raw structure was indeterminate. 根据注册的点估计值,高努力合约下的正确答案成本为 0.08665 美元,而省略努力合约下为 0.07662 美元。通过有时效的合约普查、Models-API 元数据以及预注册的原始响应探测,进一步记录了模型特定的省略语义(包括同一提供商内部的情况);当原始结构不确定时,相关结论仍保持在文档记录级别。
The request registry, parser, terminal taxonomy, statistical plan, and analysis pipeline were frozen before outcomes were examined; the resulting claims are bounded to the model, task, and collection date studied. 请求注册表、解析器、终止分类法、统计计划和分析流程在检查结果之前均已锁定;所得结论仅限于所研究的模型、任务和收集日期。