Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

Introducing Olmo-core 3: Open, scalable training infrastructure for large MoEs

介绍 Olmo-core 3:面向大型混合专家模型(MoE)的开放式可扩展训练基础设施

Today we’re releasing Olmo-core 3, a significant upgrade to our framework for developing large language models featuring a redesigned open mixture-of-experts (MoE) training system. Olmo-core 3 is designed to scale MoE training into the trillion-parameter range while preserving computational efficiency. It’s one of the core systems behind the next generation of Olmo, and part of our ongoing commitment to open up the tools and training infrastructure behind each new model.

今天,我们发布了 Olmo-core 3,这是我们大型语言模型开发框架的一次重大升级,其特色在于重新设计的开放式混合专家(MoE)训练系统。Olmo-core 3 旨在将 MoE 训练扩展至万亿参数规模,同时保持计算效率。它是下一代 Olmo 背后的核心系统之一,也是我们致力于公开每一款新模型背后工具与训练基础设施的持续承诺的一部分。

Training large AI models takes a lot of compute, driving up costs and energy use and putting advanced model development out of reach for many academic researchers and smaller labs. MoE models offer a more efficient approach—they can contain many more learned components, or parameters, without requiring every input to use all of them. But the full model still has to be stored across GPU memory and updated during training, and directing inputs to the right experts – the specialized components within an MoE – across a cluster creates its own communication and coordination costs. As MoEs grow, those costs can erode much of the computational advantage of using only part of the model for each input. Olmo-core 3 is built to close that gap.

训练大型 AI 模型需要大量计算资源,这推高了成本和能源消耗,使得许多学术研究人员和小型实验室难以进行先进的模型开发。MoE 模型提供了一种更高效的方法——它们可以包含更多的学习组件(即参数),而无需每个输入都调用所有参数。然而,整个模型仍需存储在 GPU 内存中并在训练期间进行更新,且在集群中将输入引导至正确的专家(MoE 内部的专业化组件)会产生额外的通信和协调成本。随着 MoE 规模的扩大,这些成本可能会抵消掉“每个输入仅使用模型一部分”所带来的大部分计算优势。Olmo-core 3 正是为了弥补这一差距而构建的。

In one benchmark, we increased the expert pool from 8 to 128 while still selecting only four experts per token – the small units of text a language model processes – keeping the number of active parameters per token roughly fixed at about 3.2B. Total parameter capacity grew from 4.6B to 47B, while training throughput fell by less than 5%. The same infrastructure has been benchmarked at over one trillion total parameters.

在一项基准测试中,我们将专家池从 8 个增加到 128 个,同时保持每个 Token(语言模型处理的文本小单位)仅选择 4 个专家,从而将每个 Token 的活跃参数数量大致固定在 32 亿左右。总参数容量从 46 亿增长到 470 亿,而训练吞吐量下降不到 5%。同一基础设施已在超过万亿总参数的规模下进行了基准测试。

Building a training stack around how MoEs actually work

构建围绕 MoE 实际运作方式的训练栈

Olmo-core has evolved with each generation of Olmo. Our work on sparse models goes back to OlmoE, which used an MoE architecture with 64 routed experts. Olmo 3, by contrast, used a dense architecture, meaning nearly all of the model was active for every token and its training stack was built around that design. Olmo-core 3 extends the framework with a training system designed for much larger MoE models.

Olmo-core 随着每一代 Olmo 的演进而不断进化。我们在稀疏模型上的工作可以追溯到 OlmoE,它使用了具有 64 个路由专家的 MoE 架构。相比之下,Olmo 3 使用了稠密架构,这意味着几乎整个模型在处理每个 Token 时都是活跃的,其训练栈也是围绕该设计构建的。Olmo-core 3 通过一套专为更大规模 MoE 模型设计的训练系统扩展了该框架。

Our earlier MoE implementation in Olmo-core used fully sharded data parallelism (FSDP), configured to gather and reshard model weights for each small batch of training data. Olmo-core 3 switches to a system based on distributed data parallelism (DDP). It keeps experts resident on GPUs and routes the relevant data to them, avoiding that repeated weight gathering.

我们之前在 Olmo-core 中的 MoE 实现使用了完全分片数据并行(FSDP),配置为针对每个小批量训练数据收集并重新分片模型权重。Olmo-core 3 切换到了基于分布式数据并行(DDP)的系统。它将专家常驻在 GPU 上,并将相关数据路由给它们,从而避免了重复的权重收集过程。

NVIDIA’s Megatron-Core is an established option for training large MoEs. Olmo-core 3 brings an integrated MoE training stack to the framework behind Olmo, with a redesign that improves throughput over our earlier FSDP-based implementation. In a preliminary test on eight NVIDIA B300 GPUs, a 47-billion-parameter MoE processed 52,000 tokens per second per GPU with the new stack, compared with 19,400 using our earlier implementation—about 2.7× the throughput.

NVIDIA 的 Megatron-Core 是训练大型 MoE 的成熟方案。Olmo-core 3 为 Olmo 背后的框架带来了一个集成的 MoE 训练栈,通过重新设计,其吞吐量优于我们之前基于 FSDP 的实现。在 8 张 NVIDIA B300 GPU 的初步测试中,使用新训练栈的 470 亿参数 MoE 每张 GPU 每秒可处理 52,000 个 Token,而之前的实现仅为 19,400 个——吞吐量提升了约 2.7 倍。

Scaling and optimizing MoE training

扩展与优化 MoE 训练

Olmo-core 3 combines several techniques for distributing large MoEs across GPU clusters with optimizations that make routing and computation more efficient. Three techniques determine how the model and its training state are split across hardware:

Olmo-core 3 结合了多种将大型 MoE 分布在 GPU 集群上的技术,并辅以优化措施,使路由和计算更加高效。以下三种技术决定了模型及其训练状态如何在硬件上进行拆分:

  • Expert parallelism spreads the experts across GPUs, so each GPU stores only part of the full expert pool.
  • 专家并行将专家分散在各个 GPU 上,因此每个 GPU 仅存储完整专家池的一部分。
  • Pipeline parallelism splits the model’s layers – the successive stages that transform an input – across groups of GPUs, reducing how much of the model each GPU needs to keep in memory.
  • 流水线并行将模型的层(转换输入的连续阶段)拆分到不同的 GPU 组中,减少了每个 GPU 需要保留在内存中的模型比例。
  • A distributed optimizer spreads the optimizer state – the additional data used to calculate and apply updates during training – across GPUs instead of storing a full copy on every GPU.
  • 分布式优化器将优化器状态(训练期间用于计算和应用更新的额外数据)分散在各个 GPU 上,而不是在每个 GPU 上存储完整副本。

Together, these techniques allow an MoE to scale without requiring every GPU to keep the entire model and its training state in memory. Olmo-core 3 also reduces the cost of routing data to the right experts and running their computations. Rowwise expert parallelism places routed data directly into expert input buffers, minimizing the extra work needed to rearrange it. GPU-resident routing keeps routing metadata on the GPUs, so the CPU can queue work without waiting for that information to be copied back. And grouped GEMM combines many small expert computations so GPUs can execute them more efficiently.

这些技术共同作用,使得 MoE 能够在无需每个 GPU 都将整个模型及其训练状态保留在内存中的情况下进行扩展。Olmo-core 3 还降低了将数据路由到正确专家并运行其计算的成本。行向专家并行(Rowwise expert parallelism)将路由数据直接放入专家输入缓冲区,最大限度地减少了重新排列数据所需的额外工作。GPU 常驻路由将路由元数据保留在 GPU 上,因此 CPU 可以排队工作,而无需等待信息被复制回内存。此外,分组 GEMM(矩阵乘法)将许多小的专家计算合并,使 GPU 能够更高效地执行它们。

Finally, Olmo-core 3 supports MXFP8, a lower-precision number format that represents some values with fewer bits. This can reduce computation and the amount of data moved between GPUs, as long as those savings outweigh the cost of converting between number formats. We measured MXFP8’s effect on end-to-end training throughput in a controlled benchmark on four NVIDIA B300 GPUs, with work distributed uniformly across experts. With MXFP8 enabled across the parts of the system where it helped most, training throughput was about 21% higher than with BF16, the higher-precision format we used as our baseline, while peak active memory fell from 103 GiB to 95 GiB. Most of the gain came from feed-forward computation and moving data between experts rather than attention alone.

最后,Olmo-core 3 支持 MXFP8,这是一种以更少位数表示数值的低精度数字格式。只要这些节省的资源超过了在不同数字格式之间转换的成本,它就能减少计算量和 GPU 之间传输的数据量。我们在 4 张 NVIDIA B300 GPU 上进行了受控基准测试,在工作均匀分布于专家的情况下,测量了 MXFP8 对端到端训练吞吐量的影响。在系统最受益的部分启用 MXFP8 后,训练吞吐量比我们作为基准的高精度格式 BF16 高出约 21%,同时峰值活跃内存从 103 GiB 降至 95 GiB。大部分收益来自前馈计算和专家间的数据移动,而非仅仅是注意力机制。

These techniques and optimizations have to work together. Speeding up one part of training can create costs elsewhere; faster computation may require more data movement, while moving fewer bits may not help if converting the data takes too long. Olmo-core 3 is built around those trade-offs across the full training process, giving us – and researchers using the open stack – control over how the pieces fit together.

这些技术和优化必须协同工作。加速训练的某一部分可能会在其他地方产生额外成本;更快的计算可能需要更多的数据移动,而如果数据转换耗时过长,减少传输位数可能也无济于事。Olmo-core 3 正是围绕整个训练过程中的这些权衡而构建的,使我们以及使用该开放栈的研究人员能够掌控各部分如何有机结合。

Explore our interactive walkthrough to see how data, expert, and pipeline parallelism work together to scale MoE training—from a single GPU to many. 探索我们的交互式指南,了解数据并行、专家并行和流水线并行如何协同工作,从而实现从单 GPU 到多 GPU 的 MoE 训练扩展。

Scaling into the trillion-parameter range

扩展至万亿参数规模

We’ve benchmarked Olmo-core 3 across a range of configurations on NVIDIA B3… 我们已经在 NVIDIA B3 上对 Olmo-core 3 进行了多种配置的基准测试……