Learning from the Gap Between Pass@K and Pass@1
Learning from the Gap Between Pass@K and Pass@1
从 Pass@K 与 Pass@1 的差距中学习
Large language models (LLMs) are increasingly trained with reinforcement learning from verifiable rewards (RLVR). An exact verifier can also support test-time scaling by selecting a passing response from multiple samples, while other deployments use beam search, adaptive sampling, or tools. We study single-sample decoding, where each query receives one response without search, to ask whether search-exposed behavior can be absorbed into the model.
大型语言模型(LLMs)正越来越多地通过基于可验证奖励的强化学习(RLVR)进行训练。精确的验证器还可以通过从多个样本中选择一个正确的响应来支持测试时扩展(test-time scaling),而其他部署方式则使用束搜索(beam search)、自适应采样或工具。我们研究了单样本解码(即每个查询在不进行搜索的情况下仅接收一个响应),旨在探讨搜索所带来的行为是否可以被模型吸收。
Existing verified-response post-training recipes do not generally distinguish problems already solved on the first decode from failures recovered within K samples. Under a fixed budget, this can spend examples repeating behavior the deployed policy already has. We introduce GapFT, which selects training evidence by the source checkpoint’s single-sample outcome and fine-tunes on the Pass@K-Pass@1 gap: problems the policy fails on one sample but solves within K samples.
现有的验证响应后训练方案通常无法区分在第一次解码时就已经解决的问题,以及在 K 个样本内恢复的失败案例。在固定预算下,这可能会导致模型将训练样本浪费在重复部署策略已有的行为上。我们引入了 GapFT,它根据源检查点的单样本结果选择训练证据,并针对 Pass@K 与 Pass@1 之间的差距进行微调:即那些策略在单样本中失败,但在 K 个样本内能够解决的问题。
We match training examples, processed tokens, and optimizer steps while keeping the objective unchanged. GapFT fills the matched budget with recovered failures and uses an exact decomposition to distinguish corrections of recovered and missed failures from regressions on first-decode successes.
我们在保持目标不变的同时,匹配了训练样本、处理的 Token 数量和优化器步骤。GapFT 利用恢复的失败案例填补了匹配的预算,并使用精确分解来区分“对恢复和遗漏失败的修正”与“对首次解码成功案例的回归”。
On LogiQA 2.0 and ReClor with Llama-3.1-8B, GapFT improves Pass@1 by 14.4 and 13.9 points over the source model, outperforms budget-matched uniform verified RFT at the same learning rate, and matches fine-tuning on the full verified pool using one third of the data. A single decode matches the source model’s verifier-selected Pass@4 accuracy. A randomized control attributes gains to covering distinct failures, and our analysis relates available gains to transferable failure support. A three-seed Qwen2.5-7B replication retains positive gains over uniform RFT on both logic tasks.
在 Llama-3.1-8B 模型上进行的 LogiQA 2.0 和 ReClor 测试显示,GapFT 将 Pass@1 指标较源模型分别提升了 14.4 和 13.9 个百分点;在相同学习率下,其表现优于预算匹配的统一验证 RFT,并且仅使用三分之一的数据就达到了在完整验证池上进行微调的效果。单次解码的准确率即可匹配源模型通过验证器选择后的 Pass@4 准确率。随机对照实验将这些增益归因于对不同失败案例的覆盖,我们的分析将可获得的增益与可迁移的失败支持联系起来。在 Qwen2.5-7B 模型上的三种子集重复实验表明,在两项逻辑任务中,该方法相比统一 RFT 均保持了正向增益。