Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

Characterizing a Configuration Where Inference-Time PRM-Pruned Fragment Grafting Is Inert: Evidence from Three Reasoning LMs

表征推理时 PRM 剪枝片段移植失效的配置:来自三种推理大模型的证据

Abstract: Diversity collapse in parallel chain-of-thought has motivated inference-time interventions built on a natural design: when a process reward model (PRM) prunes a chain, its high-PRM prefix is extracted and grafted verbatim as an in-context demonstration into a still-decoding sibling.

摘要: 并行思维链(Parallel Chain-of-Thought)中的多样性崩溃促使研究者们基于一种直观的设计进行推理时干预:当过程奖励模型(PRM)剪枝掉一条思维链时,提取其高 PRM 分数的前缀,并将其逐字移植到仍在解码的同级思维链中作为上下文示例。

We isolate this mechanism, PRM-Pruned Fragment Grafting (PPFG), as the most cost-minimal operationalization of cross-trajectory step-level transfer, and test it at the operating point where prior fragment-grafting work reports gains only under additional compensating ingredients.

我们将这种机制称为“PRM 剪枝片段移植”(PPFG),并将其视为跨轨迹步骤级迁移成本最低的操作方式。我们在先前片段移植研究仅在额外补偿条件下才能获得收益的运行点上对其进行了测试。

On Qwen2.5-7B-Instruct with Math-Shepherd on full MATH500 (n=500, three seeds), PPFG in both stagnation- and random-targeting variants is statistically indistinguishable from an independent parallel-CoT baseline on every measured axis.

在 Qwen2.5-7B-Instruct 模型上使用 Math-Shepherd 在完整 MATH500 数据集(n=500,三个随机种子)上的实验表明,无论是针对停滞状态还是随机目标的 PPFG 变体,在所有测量维度上与独立的并行思维链基线相比,统计学上均无显著差异。

We characterize why: a four-bucket classification of 322 stagnation-rule injection events shows only 14% targeted a genuinely struggling chain; the rest landed on chains that had already succeeded, were near completion, or sat on a flat PRM plateau, states a rescue graft cannot change.

我们分析了原因:对 322 次停滞规则注入事件进行四类分类后发现,只有 14% 的注入真正针对了处于困境的思维链;其余的注入则落在了已经成功、接近完成或处于 PRM 平坦区域的思维链上,而救援性移植无法改变这些状态。

No compound-gate refinement jointly achieves well-targeted firing and adequate density, and a random control matches the same parity at 2.4x the firing rate, so the inertness is not heuristic-specific.

没有任何复合门控优化能同时实现精准触发和足够的密度,且随机对照组在触发率高出 2.4 倍的情况下表现出相同的等效性,因此这种失效并非特定于某种启发式算法。

The finding replicates across three base LMs, six benchmarks, a second PRM, and a compatibility-gate sweep; two-one-sided-tests analysis promotes the parity to positive equivalence on all twelve Qwen/LLaMA cells.

这一发现在三种基础大模型、六个基准测试、第二个 PRM 以及兼容性门控扫描中均得到复现;双单侧检验(TOST)分析将这种等效性提升为所有 12 个 Qwen/LLaMA 实验单元上的正向等价性。

A per-event spot-check finds injected chains prune at 2.75x the matched-step rate, but a surviving-sibling counterfactual finds no population-level compensation.

针对单次事件的抽查发现,被注入的思维链剪枝率是匹配步骤的 2.75 倍,但对存活同级思维链的反事实分析显示,在群体层面并没有产生补偿效应。

A hindsight oracle bounds any per-problem gain from choosing PPFG over independent at +0.13 pp. We contribute an equivalence-testing template for establishing inference-time mechanism nulls, with every claim scoped to its tested operating point.

事后预言机(Hindsight oracle)将选择 PPFG 而非独立运行带来的每个问题收益限制在 +0.13 个百分点以内。我们贡献了一个用于确立推理时机制无效性的等效性测试模板,并将每一项结论限定在其测试的运行点范围内。