Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap
Mind the Cap: Output-Budget Regimes Change the Measured Multilingual Reasoning Gap
关注上限:输出预算机制改变了多语言推理能力的衡量差距
Abstract: Multilingual evaluations report accuracy at a single output-token cap, but languages need different numbers of tokens to express the same content, so the cap is a hidden experimental variable. 摘要: 多语言评估通常在单一的输出 Token 上限下报告准确率,但不同语言表达相同内容所需的 Token 数量各不相同,因此“上限”实际上是一个隐藏的实验变量。
We test whether the native-vs-translate gap on MGSM (German, Thai, Swahili) is a token-budget artifact for Qwen3-8B and Llama-3.1-8B-Instruct under four prompting strategies. The measured gap swings by up to 57 points across budgets, length normalization moves it by up to 38.9 points where the cap binds, and at tight caps normalization can reverse which strategy scores higher. 我们测试了在四种提示策略下,Qwen3-8B 和 Llama-3.1-8B-Instruct 模型在 MGSM(德语、泰语、斯瓦希里语)任务中,“母语 vs. 翻译”的性能差距是否仅仅是 Token 预算带来的伪影。研究发现,在不同的预算下,测得的差距波动高达 57 个百分点;在上限受限的情况下,长度归一化(length normalization)可使差距变动 38.9 个百分点;而在严格的上限限制下,归一化甚至可能逆转不同策略之间的优劣排名。
We prospectively froze the sweep’s three Qwen peaks and its near-zero value at 1024 and evaluated them on 540,000 independently hard-capped decodes: a second frozen family of six Holm-corrected tests rejects every null. The frozen test at $B^=1024$ still fails to reject because native accuracy has already saturated there; above saturation, the residual difference is a strategy-performance gap, not an identified reasoning deficit. 我们预先锁定了扫描中 Qwen 的三个峰值及其在 1024 Token 处的近零值,并在 540,000 次独立硬上限解码上进行了评估:第二组包含六项 Holm 校正测试的固定实验拒绝了所有零假设。在 $B^=1024$ 的固定测试中,由于母语准确率已达到饱和,无法拒绝零假设;在饱和点之上,剩余的差异属于策略性能差距,而非推理能力的缺陷。
The same truncation channel prices a cost-ordered adaptation ladder: a cross-fitted Thai vocabulary extension closes 0.0 points of the gap at the frozen budget and 4.9 points where 19% of traces still truncate. A third frozen family varies only the announced budget at a fixed enforced cap; announcing 128 rather than 2048 tokens moves Thai native accuracy by 5.1 points, so accuracy is not a function of the enforced cap alone. 同样的截断通道也影响了成本排序的适应性阶梯:在固定预算下,交叉拟合的泰语词汇扩展对缩小差距毫无帮助(0.0 点),但在 19% 的轨迹仍被截断的情况下,差距缩小了 4.9 点。第三组固定实验仅改变了在固定强制上限下所“声明”的预算;将声明的 Token 数从 2048 改为 128,泰语母语准确率波动了 5.1 点,这表明准确率并非仅由强制上限决定。
A correct-emission timing identity computed from one long-cap run matches the three pre-specified MGSM peaks to 0.65 points and, in an exploratory Qwen-only analysis of three further benchmarks, tracks held-out items to 0.92 points, locating the peak exactly in five of seven cells. Treat the output cap as an independent variable and report accuracy across the budget regime, not at a single budget. 通过一次长上限运行计算出的“正确输出时序恒等式”,与预先设定的三个 MGSM 峰值匹配度在 0.65 点以内。在针对另外三个基准测试的 Qwen 探索性分析中,该恒等式对留出项(held-out items)的追踪误差仅为 0.92 点,并在七个单元格中的五个准确锁定了峰值。因此,研究建议将输出上限视为一个自变量,并在整个预算范围内报告准确率,而非仅在单一预算下进行报告。