Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

Counterexample Generation via Per-Theorem Symbolic Verifiers: When Imitation Hurts and Reinforcement Repairs

通过定理级符号验证器生成反例:当模仿产生负面影响而强化学习能够修复时

Large language models often solve a theorem forward yet fail to disprove a closely related false one: a falsification gap that supervised fine-tuning does not close and can actively worsen. 大型语言模型通常能够正向证明一个定理,却无法证伪一个高度相关的错误命题:这种“证伪鸿沟”不仅无法通过监督微调(SFT)弥合,甚至可能因微调而进一步恶化。

We frame counterexample generation as constrained witness emission against a deterministic per-theorem Python verifier, and release SymCE, a corpus of 4,707 false undergraduate-algebra and real-analysis conjectures, each paired with executable verifiers. 我们将反例生成定义为针对确定性定理级 Python 验证器的受限见证输出任务,并发布了 SymCE 数据集。该数据集包含 4,707 个本科代数和实分析领域的错误猜想,每个猜想都配有可执行的验证器。

The verifier also serves as the reward function, making SymCE a training environment. 该验证器同时充当奖励函数,使 SymCE 成为一个训练环境。

Training Qwen3-4B with SFT followed by GRPO under this oracle reveals an imitation trap: counterexample-only SFT collapses true-theorem recognition from 0.27 to 0.00, while RLVR with a sparse outcome-only reward repairs this and exceeds the base, to 0.66. 在这一预言机(Oracle)指导下,使用 SFT 随后进行 GRPO 训练 Qwen3-4B 模型揭示了一个“模仿陷阱”:仅针对反例的 SFT 会导致模型对正确定理的识别能力从 0.27 崩溃至 0.00;而使用稀疏结果奖励的 RLVR(基于验证的强化学习)则修复了这一问题,并将识别率提升至 0.66,超过了基座模型。

The collapse replicates across four seeds and on Gemma-3-4B. 这种性能崩溃在四个随机种子以及 Gemma-3-4B 模型上均得到了复现。

Sparse and dense rewards yield statistically indistinguishable in-domain success yet diverge by 33 points on a held-out calibration probe, a dissociation we trace to the partial-credit term. 稀疏奖励和密集奖励在域内成功率上表现出统计学上的无差异,但在留出的校准探针测试中却出现了 33 个百分点的偏差,我们将这种分离归因于部分得分项(partial-credit term)。

Our 4B model outperforms every evaluated 7B open-weights math specialist, remains competitive with six frontier commercial APIs, and transfers under unchanged prompting to GSM8K, MATH-500 and MMLU-college-math. 我们的 4B 模型表现优于所有参与评估的 7B 开源数学专家模型,与六个前沿商业 API 保持竞争力,并且在无需更改提示词的情况下,能够迁移至 GSM8K、MATH-500 和 MMLU-college-math 等任务中。

A human audit of 177 verifier decisions finds 97.7% accuracy. 对 177 个验证器决策的人工审计显示,其准确率达到 97.7%。