Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation
Improving OCR Faithfulness via Gated and Attenuated On-Policy Distillation
通过门控与衰减在线策略蒸馏提升 OCR 忠实度
Abstract: Vision-language models may rewrite anomalous text in images into linguistically plausible expressions, compromising OCR transcription faithfulness. Sequence-level task rewards and local teacher guidance are complementary, but guidance from the same teacher may not remain equally effective as the student improves.
摘要: 视觉语言模型可能会将图像中的异常文本重写为语言上合理的表达,从而损害 OCR 转录的忠实度。序列级任务奖励与局部教师指导是互补的,但随着学生模型的进步,来自同一教师的指导可能不再保持同等的有效性。
Offline analysis shows that supervision from a fixed teacher becomes progressively less favorable as the student improves, both across training checkpoints and across response groups with different task rewards. Motivated by this observation, we introduce GAD-RL, which adaptively regulates teacher supervision during joint post-training according to the student’s current task performance and local distributions.
离线分析表明,随着学生模型的进步,来自固定教师的监督作用会逐渐减弱,这在不同的训练检查点以及具有不同任务奖励的响应组中均有体现。受此观察启发,我们引入了 GAD-RL,它能够根据学生当前的性能和局部数据分布,在联合后训练过程中自适应地调节教师监督。
A frozen teacher conditions on reference transcriptions and student-generated prefixes. GAD-RL disables distillation for response groups containing an output with task reward at least 0.95 and continuously attenuates distillation strength as group-mean reward increases. It also weights forward KL by the student’s probability of the teacher’s Top-1 token, moderating local auxiliary updates when student support for that candidate is low.
一个冻结的教师模型基于参考转录和学生生成的上下文前缀进行条件化。对于包含任务奖励至少为 0.95 的输出响应组,GAD-RL 会禁用蒸馏,并随着组平均奖励的增加持续衰减蒸馏强度。此外,它还通过学生对教师 Top-1 Token 的预测概率来加权前向 KL 散度,从而在学生对该候选词支持度较低时,调节局部辅助更新的力度。
On Qwen3.5-2B, GAD-RL achieves 59.92% Micro Recall on CHAOS-Bench, surpassing GRPO and GRPO+OPD (fixed-weight) by 8.45 and 4.43 percentage points, respectively, while achieving an Overall score of 91.18 on OmniDocBench v1.6.
在 Qwen3.5-2B 模型上,GAD-RL 在 CHAOS-Bench 上实现了 59.92% 的微观召回率(Micro Recall),分别比 GRPO 和 GRPO+OPD(固定权重)高出 8.45 和 4.43 个百分点,同时在 OmniDocBench v1.6 上获得了 91.18 的总分。