When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO
Computer Science > Machine Learning arXiv:2610.06861 (cs) [Submitted on 8 Jul 2026] Title: When Does External Guidance Help LLM Reasoning? A Bias-Variance Theory of Guidance-Augmented GRPO Authors: Sofia Torres, Gabriel Almeida, Carter Adams, Camila Rocha.
计算机科学 > 机器学习 arXiv:2610.06861 (cs) [提交于 2026 年 7 月 8 日] 标题:外部指导何时有助于大语言模型推理?指导增强型 GRPO 的偏差-方差理论 作者:Sofia Torres, Gabriel Almeida, Carter Adams, Camila Rocha。
Abstract: Reinforcement learning with verifiable rewards (RLVR) has become the dominant paradigm for eliciting multi-step reasoning in large language models, and a recent wave of methods (LUFFY, ExPO, PAPO, TAPO) further augments RL with external guidance - expert traces, self-explanations, or retrieved thought patterns. Although each method reports empirical gains, none provides convergence rates, bias bounds, or an optimal weighting rule for the guidance signal.
摘要:带有可验证奖励的强化学习 (RLVR) 已成为激发大语言模型多步推理的主导范式,近期涌现的一系列方法(LUFFY、ExPO、PAPO、TAPO)进一步通过外部指导(专家轨迹、自我解释或检索到的思维模式)增强了强化学习。尽管每种方法都报告了实证收益,但均未提供收敛速度、偏差界限或指导信号的最佳加权规则。
We close this gap with Guidance-Augmented GRPO (GA-GRPO), a unified theoretical framework that casts external guidance as a stochastic guidance operator G re-writing the question distribution, and analyses the resulting policy-gradient estimator as a biased on-policy estimator whose bias is bounded by the total-variation guidance divergence delta_G between the guidance-augmented sampling distribution and the policy’s own distribution.
我们通过指导增强型 GRPO (GA-GRPO) 填补了这一空白。这是一个统一的理论框架,将外部指导视为重写问题分布的随机指导算子 G,并将由此产生的策略梯度估计器分析为一种有偏的同策略估计器,其偏差受限于指导增强采样分布与策略自身分布之间的总变差指导散度 delta_G。
The framework subsumes vanilla GRPO, LUFFY, ExPO, PAPO, and TAPO as special cases obtained by particular choices of G. Under smoothness and bounded-divergence assumptions we prove that GA-GRPO converges at rate O(1/sqrt(T)) to an O(delta sqrt(T))-neighbourhood of the GRPO stationary point, derive the closed-form MSE-optimal guidance weight lambda-star(T, delta, sigma_0 squared) = sigma_0 squared / (sigma_0 squared + R_max squared delta squared T), and prove a matching minimax lower bound showing the Omega(delta squared T) bias term is unavoidable.
该框架将原始 GRPO、LUFFY、ExPO、PAPO 和 TAPO 作为通过特定 G 选择获得的特殊情况纳入其中。在平滑性和有界散度假设下,我们证明了 GA-GRPO 以 O(1/sqrt(T)) 的速率收敛到 GRPO 平稳点的 O(delta sqrt(T)) 邻域内,推导出了闭式 MSE 最优指导权重 lambda-star(T, delta, sigma_0 squared) = sigma_0 squared / (sigma_0 squared + R_max squared delta squared T),并证明了相应的极小极大下界,表明 Omega(delta squared T) 的偏差项是不可避免的。
Experiments on Qwen2.5-Math-7B-Base across nine math and OOD benchmarks confirm that optimal-weight GA-GRPO matches or surpasses TAPO, LUFFY, ExPO, and vanilla GRPO while requiring 31% fewer GPU-hours, and eight analysis experiments validate each theoretical prediction.
在九个数学和分布外 (OOD) 基准测试中对 Qwen2.5-Math-7B-Base 进行的实验证实,最优权重 GA-GRPO 在所需 GPU 小时数减少 31% 的情况下,匹配或超越了 TAPO、LUFFY、ExPO 和原始 GRPO,且八项分析实验验证了每一项理论预测。