K2 Horizon: A connected fleet of six open models

K2 Horizon: A connected fleet of six open models

K2 Horizon:六款开源模型组成的互联模型群

Today IFM is releasing K2 Horizon, a connected fleet of six models: 375B-A23B, 36B-A4B, 32B, 7B, 3.7B, and 0.9B. Across reasoning, mathematics, coding, agentic tasks, and general capabilities, K2 Horizon delivers top-tier performance in every size class—with the 0.9B, 3.7B, and 7B models setting new state of the art at their respective scales. 今天,IFM 正式发布 K2 Horizon,这是一个由六款模型组成的互联模型群:375B-A23B、36B-A4B、32B、7B、3.7B 和 0.9B。在推理、数学、编程、智能体任务及通用能力方面,K2 Horizon 在每一个尺寸级别都提供了顶尖的性能,其中 0.9B、3.7B 和 7B 模型在各自的规模上树立了新的行业标杆。

K2 Horizon is also our most comprehensive open release to date. For every model, we are opening the training lifecycle from pretraining through reasoning and agentic post-training. We are releasing intermediate checkpoints, training data or detailed data-construction recipes, open architecture, mixture compositions, training code, configurations, fine-grained logs, evaluation results, and final weights. The models and code are released under the Apache 2.0 license. Datasets are released under their applicable licenses, such as ODC-BY; We disclose how the data was constructed and mixed when redistribution is not possible. K2 Horizon 也是我们迄今为止最全面的开源发布。对于每一款模型,我们都公开了从预训练到推理及智能体后训练的完整生命周期。我们发布了中间检查点、训练数据或详细的数据构建配方、开放架构、混合成分、训练代码、配置、细粒度日志、评估结果以及最终权重。模型和代码均在 Apache 2.0 许可下发布。数据集则根据其适用的许可(如 ODC-BY)发布;在无法重新分发数据的情况下,我们披露了数据的构建和混合方式。

Together, K2 Horizon represents the most comprehensive open model release to date: A new performance frontier across scales. The 0.9B, 3.7B, and 7B models achieve world-leading performance in their size classes across widely used evaluations. The 36B-A4B model, equipped with our new Mixture-of-Value-Attention (MoVA) mechanism, delivers exceptional capability per active parameter, outperforming some much larger models. The 32B and 375B-A23B models rank among the top models in their respective classes. Together, the six models provide competitive performance across deployment environments ranging from edge devices to the enterprise. 总之,K2 Horizon 代表了迄今为止最全面的开源模型发布:在各个规模上都开辟了性能新前沿。0.9B、3.7B 和 7B 模型在广泛使用的评估中,于各自的尺寸级别实现了世界领先的性能。配备了我们全新“价值注意力混合”(MoVA)机制的 36B-A4B 模型,在每个激活参数上提供了卓越的能力,表现优于一些规模大得多的模型。32B 和 375B-A23B 模型在各自的类别中均名列前茅。这六款模型共同为从边缘设备到企业级的各种部署环境提供了极具竞争力的性能。

The first fully open model fleet for agents. K2 Horizon is the first open model family to expose the complete development process through agentic post-training. By releasing checkpoints, data (or data recipe), code, configurations, and training logs across every stage, K2 Horizon makes it possible to study how reasoning, tool use, planning, and agentic capabilities emerge; reproduce the methods that create them; and adapt those methods to new tools, environments, and domains. 首个面向智能体的全开源模型群。K2 Horizon 是首个通过智能体后训练公开完整开发过程的开源模型家族。通过发布各个阶段的检查点、数据(或数据配方)、代码、配置和训练日志,K2 Horizon 使研究推理、工具使用、规划和智能体能力如何涌现成为可能;同时也让复现这些能力的创建方法,并将这些方法适配到新工具、新环境和新领域成为现实。

Six models spanning edge to enterprise. The 0.9B model is designed for highly constrained environments such as watches and glasses, while the 3.7B and 7B models bring advanced capabilities to phones and other on-device applications. The dense 32B model and sparse 36B-A4B model provide powerful options for local workstations and efficient serving. The 375B-A23B model brings the fleet’s strongest capabilities to demanding enterprise deployments. All six models include quantization support. 六款模型覆盖从边缘到企业级应用。0.9B 模型专为手表和眼镜等高度受限的环境设计,而 3.7B 和 7B 模型则为手机及其他端侧应用带来了先进能力。稠密的 32B 模型和稀疏的 36B-A4B 模型为本地工作站和高效服务提供了强大的选择。375B-A23B 模型则将该系列最强的能力带到了要求严苛的企业级部署中。所有六款模型均包含量化支持。

One connected fleet. The six models share core architecture, vocabulary, training methodology, interfaces, evaluation infrastructure, and deployment tooling, with a smaller vocabulary for the 0.9B model. This consistency also makes it easier to move between sizes, route work dynamically, and study capability and efficiency across scale. 一个互联的模型群。这六款模型共享核心架构、词汇表、训练方法、接口、评估基础设施和部署工具,其中 0.9B 模型使用较小的词汇表。这种一致性也使得在不同尺寸模型间切换、动态路由任务以及研究跨规模的能力与效率变得更加容易。

World-leading performance across the scales. The 0.9B, 3.7B, and 7B models achieve state-of-the-art results in their respective classes across mathematics, reasoning, general capability, coding, and agentic tasks. The 36B-A4B model performs beyond the level normally expected from its active parameter count, demonstrating the efficiency of our unique Mixture-of-Expert design when computing attention values. The 32B and 375B-A23B models place among the top models in their respective comparison classes. 跨规模的世界领先性能。0.9B、3.7B 和 7B 模型在数学、推理、通用能力、编程和智能体任务方面,在各自的类别中均取得了最先进的成果。36B-A4B 模型的表现超出了其激活参数数量通常预期的水平,证明了我们独特的“专家混合”(MoE)设计在计算注意力值时的效率。32B 和 375B-A23B 模型在各自的对比类别中均位居前列。

The small models are especially notable. K2 Horizon 0.9B achieves an AIME 2026 score above 48, along with strong reasoning, tool-use, and agentic capabilities. K2 Horizon 3.7B and 7B extend these capabilities to more demanding software-engineering and multi-step environments, demonstrated on strong performance in SWE-bench and BrowseComp. Although complex tasks that require extensive exploration and repeated recovery, such as those in TerminalBench, remain difficult for the smallest models, K2 Horizon moves the boundary of what is possible at every scale. 小模型尤为引人注目。K2 Horizon 0.9B 的 AIME 2026 得分超过 48 分,并具备强大的推理、工具使用和智能体能力。K2 Horizon 3.7B 和 7B 将这些能力扩展到了要求更高的软件工程和多步任务环境中,在 SWE-bench 和 BrowseComp 中的强劲表现证明了这一点。尽管对于最小的模型来说,像 TerminalBench 中那样需要大量探索和反复恢复的复杂任务仍然具有挑战性,但 K2 Horizon 在每一个规模上都推动了可能性的边界。

Why the Horizon Fleet matters. A transparent model that falls far behind the capability frontier has limited value as a foundation, even for research. At the same time, a powerful model released only as final weights allows people to run it, but provides little insight into how its capabilities were created. K2 Horizon brings these two together. The fleet provides highly competitive models and releases the recipes used to train them. Researchers can study advanced capabilities in models strong enough to exhibit them, while developers can reproduce, adapt, and extend the methods rather than treating the final checkpoint as an opaque starting point. Since introducing the fully open principle in our 2023 LLM360 paper, we have released open models every year while extending that commitment to larger scales, stronger capabilities, and now the complete lifecycle through agentic post-training. Horizon 模型群为何重要。一个远远落后于能力前沿的透明模型,即使作为研究基础,其价值也有限。与此同时,一个仅以最终权重形式发布的强大模型虽然允许人们运行,但无法提供关于其能力如何产生的洞察。K2 Horizon 将这两者结合在了一起。该系列提供了极具竞争力的模型,并发布了训练它们的配方。研究人员可以在具备足够能力的模型中研究先进功能,而开发者则可以复现、适配和扩展这些方法,而不是将最终检查点视为一个不透明的起点。自我们在 2023 年的 LLM360 论文中引入全开源原则以来,我们每年都会发布开源模型,同时将这一承诺扩展到更大的规模、更强的能力,以及现在通过智能体后训练实现的完整生命周期。

A Deep Dive into The K2 Horizon Fleet. K2 Horizon 375B-A23B: the enterprise powerhouse. K2 Horizon 375B-A23B is the fleet’s largest and most capable model. Its sparse MoE architecture provides 375 billion parameters of total capacity while activating approximately 23 billion parameters for each token, allowing it to draw on the capacity of a much larger model without using every parameter for every token. The model ranks among the top models below 400 billion parameters across general, reasoning, coding, and agentic evaluations. It is designed for demanding workloads where model quality matters most, including complex reasoning, software engineering, research, and long-horizon agentic tasks. Like every model in the Horizon fleet, 375B-A23B is released not as a single endpoint but as a development tree. Its intermediate checkpoints and post-training branches expose how the base model develops into reasoning, instruction-following, and specialized agentic variants. 深入了解 K2 Horizon 模型群。K2 Horizon 375B-A23B:企业级动力源。K2 Horizon 375B-A23B 是该系列中规模最大、能力最强的模型。其稀疏的 MoE 架构提供了 3750 亿参数的总容量,同时每个 token 仅激活约 230 亿参数,使其能够在不为每个 token 使用所有参数的情况下,利用更大模型的容量。该模型在 4000 亿参数以下的通用、推理、编程和智能体评估中名列前茅。它专为模型质量至关重要的严苛工作负载而设计,包括复杂推理、软件工程、研究和长周期智能体任务。与 Horizon 系列中的每一款模型一样,375B-A23B 的发布不是作为一个单一的终点,而是一个开发树。其中间检查点和后训练分支展示了基础模型如何演变为推理、指令遵循和专门的智能体变体。

K2 Horizon 32B and 36B-A4B: strong performance for local deployment. Horizon 32B is the fleet’s most powerful dense model, providing K2 Horizon 32B 和 36B-A4B:本地部署的强劲性能。Horizon 32B 是该系列中最强大的稠密模型,提供了……