Beam: Reflection's 501B open-weight model
Beam: Reflection’s 501B open-weight model
Beam:Reflection 的 501B 开源权重模型
We are introducing Beam, Reflection’s first open-weight model. Beam is a sparse Mixture-of-Experts model with 501 billion total parameters, 23 billion active, built for coding, reasoning, and agentic workloads. 我们隆重推出 Reflection 的首个开源权重模型 Beam。Beam 是一个稀疏混合专家(MoE)模型,总参数量为 5010 亿,激活参数量为 230 亿,专为编程、推理和智能体(Agentic)工作负载而构建。
Beam’s capabilities come from major investments in both pretraining and reinforcement learning (RL). We pretrained the model on 23.8 trillion diverse, curated, high-quality tokens from the web and proprietary licensed datasets, matching or outperforming available similar-sized open base models. Beam 的能力源于我们在预训练和强化学习(RL)方面的大量投入。我们使用来自网络和专有授权数据集的 23.8 万亿个多样化、经过筛选的高质量 Token 对模型进行了预训练,其表现与同等规模的现有开源基础模型相当或更胜一筹。
In parallel, we developed the algorithms, training environments, and infrastructure needed to sustain high-compute RL at exceptional scale. Our high-compute RL run generated over 100 million rollouts on 10.5K NVIDIA GB300 GPUs over 4 weeks of training. 与此同时,我们开发了在大规模高算力强化学习下所需的算法、训练环境和基础设施。我们进行的高算力强化学习训练在 10.5 万块 NVIDIA GB300 GPU 上运行了 4 周,生成了超过 1 亿次 Rollout(推演)。
Together, these efforts produced competitive open-weight performance with frontier inference compute efficiency. Beam is undergoing final red-teaming and evaluations. You can sign up here for early access to the model. We will release the weights, technical report, model card, and developer artifacts later this month. 这些努力共同造就了具有竞争力的开源权重性能以及前沿的推理计算效率。Beam 目前正在进行最终的红队测试和评估。您可以点击此处注册以获取模型的抢先体验资格。我们将在本月晚些时候发布权重、技术报告、模型卡和开发者工具。
Model Capability
模型能力
We trained Beam with a particular focus on coding and agentic performance. Beam advances the Western open-weight frontier and is competitive with larger open models like GLM 5.2 and approaching Qwen 3.8-Max on coding and agentic tasks. Where frontier open models like Kimi K3 remain ahead on raw capability, Beam’s advantage is efficiency at inference time. 我们在训练 Beam 时特别关注编程和智能体性能。Beam 推动了西方开源权重模型的前沿水平,在编程和智能体任务上与 GLM 5.2 等更大的开源模型具有竞争力,并正在接近 Qwen 3.8-Max 的水平。虽然 Kimi K3 等前沿开源模型在原始能力上仍保持领先,但 Beam 的优势在于推理时的效率。
Beam pairs coding and agentic capabilities with highly efficient reasoning. On advanced reasoning benchmarks, it achieves scores comparable to GLM-5.2 while using 3–4× less inference compute. Efficiency gains are even more pronounced when comparing to models in the 2T+ parameter family like Qwen 3.8-Max, which require significantly more inference compute per token. Beam 将编程和智能体能力与高效推理相结合。在高级推理基准测试中,它在推理算力消耗减少 3-4 倍的情况下,取得了与 GLM-5.2 相当的分数。与 Qwen 3.8-Max 等 2 万亿参数量级的模型相比,这种效率提升更为显著,因为后者每个 Token 需要消耗更多的推理算力。
These results translate into more intelligence per token, delivering strong model capabilities at lower cost, making Beam a powerful workhorse model for enterprise coding and agentic workloads. 这些结果转化为每个 Token 更高的智能密度,以更低的成本提供强大的模型能力,使 Beam 成为企业级编程和智能体工作负载的强大主力模型。
High-Compute Reinforcement Learning
高算力强化学习
We made high-compute reinforcement learning a central scaling axis for Beam, investing in RL science, data, and infrastructure to turn more compute into stronger capabilities. Scaling RL enables more extensive exploration of problem-solving strategies, while longer rollouts support multi-step reasoning, tool use, and adaptation to environment feedback. 我们将高算力强化学习作为 Beam 的核心扩展轴,投入于强化学习科学、数据和基础设施,旨在将更多的算力转化为更强的能力。扩展强化学习能够更广泛地探索解决问题的策略,而更长的 Rollout 则支持多步推理、工具使用以及对环境反馈的适应。
To scale reinforcement learning, we deployed 10.5K NVIDIA GB300 GPUs for four weeks generating more than 100 million rollouts with a maximum context length of 256K tokens. Training and grading used approximately 1.3 billion sandboxes. To sustain a run of this magnitude, we sourced one million high-quality coding, agentic, and STEM environments. We believe this is one of the largest scale RL runs conducted by any open lab to date. 为了扩展强化学习,我们部署了 10.5 万块 NVIDIA GB300 GPU,进行了为期四周的训练,生成了超过 1 亿次 Rollout,最大上下文长度为 256K Token。训练和评分使用了大约 13 亿个沙盒环境。为了维持如此大规模的运行,我们搜集了 100 万个高质量的编程、智能体和 STEM 环境。我们相信这是迄今为止任何开源实验室进行的最大规模强化学习训练之一。
We trained Beam with asynchronous policy gradients. At scale, policy staleness becomes a major source of instability for these methods. Long running rollouts have tokens that are generated by multiple model checkpoints, with earlier tokens becoming increasingly stale relative to the current policy. 我们使用异步策略梯度(Asynchronous Policy Gradients)训练了 Beam。在大规模环境下,策略滞后(Policy Staleness)成为这些方法不稳定的主要来源。长时间运行的 Rollout 包含由多个模型检查点生成的 Token,早期的 Token 相对于当前策略而言会变得越来越滞后。