One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
本文为原文前 6,000 字符的节选翻译,完整内容请查看原文。
One Model Family, Two Gold-Level Results: Fine-Tuning Nemotron for IOI and IMO
国际信息学奥林匹克竞赛(IOI)和国际数学奥林匹克竞赛(IMO)测试的是不同的技能。IOI 要求算法和代码在严格的时间和提交限制下通过隐藏测试;IMO 则要求严谨的自然语言证明。在任何一项比赛中取得成功都很困难,而在两项比赛中都取得成功则指向了更广泛的意义。
The International Olympiad in Informatics (IOI) and the International Mathematical Olympiad (IMO) test different skills. IOI requires algorithms and code that pass hidden tests under strict time and submission limits. IMO demands rigorous natural-language proofs. Success at either competition is difficult. Success at both points to something broader.
我们最近的结果表明,Nemotron 是构建世界级专家模型的强大且适应性强的基础。从 Nemotron 3 开始,我们的团队使用了监督微调(SFT)、强化学习(RL)和反馈驱动推理,创建了在 2026 年 IMO 和 2026 年 IOI 中均达到金牌水平的系统。
Our recent results show that Nemotron is a strong, adaptable foundation for building world-class specialist models. Starting from Nemotron 3, our teams used supervised fine-tuning (SFT), reinforcement learning (RL), and feedback-driven inference to create systems that reached gold-medal level at both IMO 2026 and IOI 2026.
“易于微调”不仅仅意味着使检查点可训练。它应该意味着一个强大的基础模型可以通过清晰、可重用的方案适应苛刻的领域。在两个项目中,该方案包含四个部分:从强大的 Nemotron 基础模型开始;策划特定领域的难题和高质量的推理轨迹;应用标准的训练后方法(如 SFT,并在有用时使用 RL);将专家模型与生成、评估和改进候选答案的推理循环配对。
A reusable specialization recipe “Easy to fine-tune” should mean more than making a checkpoint trainable. It should mean that a capable foundation model can be adapted to a demanding domain with a clear, reusable recipe. Across the two projects, that recipe had four parts: Start with a strong Nemotron base model. Curate domain-specific problems and high-quality reasoning traces. Apply standard post-training methods such as SFT and, where useful, RL. Pair the specialist model with an inference loop that generates, evaluates, and improves candidate answers.
对于竞技编程,我们策划了 22,000 个问题并生成了合成推理轨迹来训练两名专家。Nemotron-3-Nano-CC(拥有 300 亿总参数和 30 亿活跃参数)接受了 SFT 和 RL 训练。Nemotron-3-Ultra-CC(拥有 5500 亿总参数和 550 亿活跃参数)接受了 SFT 训练。这些实验还表明,适应性在不同规模下不必表现得完全一样。
For competitive programming, we curated 22,000 problems and generated synthetic reasoning traces to train two specialists. Nemotron-3-Nano-CC, with 30 billion total parameters and 3 billion active parameters, received both SFT and RL. Nemotron-3-Ultra-CC, with 550 billion total parameters and 55 billion active parameters, received SFT. These experiments also showed that adaptation does not have to look the same at every scale.
IMO 项目将同样的理念应用于奥林匹克数学。从 Nemotron 3 Ultra 开始,我们用 SFT 训练了一名专家,用 RL 训练了另一名专家。SFT 语料库包含了 15,818 个独特证明问题中的 414,890 个经过质量过滤的示例。它不仅教授最终答案,数据还涵盖了证明生成、细化、验证和元验证,因此模型学会了构建论点、识别差距、回应批评并判断证明是否完整。
The IMO project applied the same idea to olympiad mathematics. Starting from Nemotron 3 Ultra, we trained one specialist with SFT and another with RL. The SFT corpus contained 414,890 quality-filtered examples across 15,818 unique proof problems. It did more than teach final answers. The data covered proof generation, refinement, verification, and meta-verification, so the model learned to construct arguments, identify gaps, respond to critiques, and judge whether a proof was complete.
微调和推理时计算共同发挥作用。我们早期的 IOI 2025 Hugging Face 文章展示了推理时计算如何将开源权重模型推向金牌水平。新的结果增加了一个重要的部分:更好的专业化为推理系统提供了更好的候选者、更好的批评者和更好的改进。在两种情况下,最好的结果都来自于将有能力的专家与能够搜索、验证和改进的系统相结合。
Fine-tuning and test-time compute work together. Our earlier IOI 2025 Hugging Face post showed how test-time compute can push open-weight models to gold-level performance. The new results add an important piece: better specialization gives the inference system better candidates, better critics, and better refinements. In both cases, the best outcome came from combining a capable specialist with a system that could search, verify, and improve.