GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
GRPO Beyond English: A Large-Scale Study of GRPO in Non-English and Multilingual Settings
GRPO 超越英语:非英语及多语言环境下 GRPO 的大规模研究
Reinforcement Learning with Verifiable Rewards (RLVR), often optimized with Group Relative Policy Optimization (GRPO), has become a central recipe for improving the reasoning capabilities of pretrained language models but current studies remain heavily English-centric.
基于可验证奖励的强化学习(RLVR)通常采用组相对策略优化(GRPO)进行优化,已成为提升预训练语言模型推理能力的核心方法,但目前的研究仍主要集中在英语领域。
We conduct a large-scale empirical study of multilingual and non-English GRPO across a wide range of base models, training languages, and different reasoning language rewards.
我们针对多种基础模型、训练语言以及不同的推理语言奖励,对多语言和非英语环境下的 GRPO 进行了大规模实证研究。
We find that training to reason in the native language often leaves only a small gap to training for English reasoning. We further observe strong crosslingual transfer: training in one language often improves performance in many others.
研究发现,使用母语进行推理训练与使用英语进行推理训练之间的差距往往很小。我们还观察到强大的跨语言迁移能力:在一种语言上的训练往往能提升模型在其他多种语言上的表现。
However, specific trends are highly model- and language-dependent. In some cases, training in a particular language induces severe regressions on out-of-domain capabilities in other languages.
然而,具体的趋势高度依赖于模型和语言本身。在某些情况下,针对特定语言的训练会导致模型在其他语言的域外能力上出现严重的性能倒退。
Our analysis shows that RLVR beyond English can provide broad crosslingual gains, but also requires broad evaluation to detect language-specific regressions.
我们的分析表明,在英语之外应用 RLVR 可以带来广泛的跨语言收益,但也需要进行全面的评估,以检测特定语言的性能倒退问题。