Kolibri: A Sovereign Open-Weight Model
Kolibri: A Sovereign Open-Weight Model
Kolibri:一款主权开源权重模型
Research Aleph Alpha 03/10/2026 Kolibri Has Landed: A Sovereign Open-Weight Model. On the Day of German Reunification, we are releasing our new model: Kolibri. Kolibri is an English-German Mixture-of-Experts Transformer with 78B total parameters, 3B active. It supports context lengths of up to 1M tokens. The model can be downloaded with the full weights on Hugging Face and used under the Apache 2.0 license terms. Aleph Alpha 研究部 2026年10月3日,Kolibri 正式发布:一款主权开源权重模型。在德国统一日之际,我们发布了全新的模型:Kolibri。Kolibri 是一款英德双语混合专家(MoE)Transformer 模型,总参数量为 78B,激活参数量为 3B。它支持高达 100 万 token 的上下文长度。该模型现已在 Hugging Face 上提供完整权重下载,并遵循 Apache 2.0 许可协议。
Kolibri is the result of continuous iteration of our model training effort. We first built a model training pipeline and validated it by building Kolibri Origin, a 30B total, 3B active model with a much shorter 65k token context window. Kolibri ran through the same pipeline: from data ingestion and curation, through ablations, pre-training, and post-training, to the final evals. It enabled running hundreds of ablation experiments and stable pre-training that ran without a person having to step in when hardware failed or a data connection dropped. Kolibri 是我们模型训练工作持续迭代的成果。我们首先构建了一个模型训练流水线,并通过构建 Kolibri Origin(总参数 30B,激活参数 3B,上下文窗口为 65k token)进行了验证。Kolibri 采用了相同的流水线:从数据摄取与清洗,到消融实验、预训练、后训练,直至最终评估。这套流水线支持了数百次消融实验,并实现了稳定的预训练过程——即使在硬件故障或数据连接中断时,也无需人工干预。
We continuously monitored training metrics and standardized monitoring for custom benchmarks. The time we put into building and iterating on this pipeline was a valuable investment. We see it in how much better Kolibri is than Kolibri Origin, and in how little time separates their releases. 我们持续监控训练指标,并为自定义基准测试实现了标准化监控。我们在构建和迭代该流水线上投入的时间是一项宝贵的投资。从 Kolibri 相比 Kolibri Origin 的显著提升,以及两者发布时间间隔之短,我们看到了这一投入的价值。
Kolibri is a specialized language model built for sovereign mission-critical work in regulated areas including public administration, industrials and aerospace. We specialized Kolibri for German, reasoning, math, agentic behavior, and further capabilities our customers need in production. The aim of this specialization was to optimize performance in our customers’ specific use cases. Through specialization, customers achieve contextualized performance in their AI operations and they can monitor its economic impact, so that ROI stays measurable and grows over time. Kolibri 是一款专门为受监管领域(包括公共行政、工业和航空航天)的主权关键任务而构建的专业语言模型。我们针对德语、推理、数学、智能体行为以及客户在生产环境中所需的其他能力对 Kolibri 进行了专项优化。这种专业化的目的是为了优化客户特定用例中的性能。通过专业化,客户可以在其 AI 运营中实现情境化性能,并能够监控其经济影响,从而确保投资回报率(ROI)可衡量并随时间增长。
Specialization alone is not enough. Sovereignty is just as important. Sovereignty, for us, combines two dimensions: how we built the model, and how it transfers to our customers. We offer full supply-chain integrity and account for every decision, from data ingestion, through pre- and post-training, to the final evaluations. We provide transparency. Customers have full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property of the model. Read our tech report for full details. 仅有专业化是不够的,主权同样重要。对我们而言,主权结合了两个维度:我们如何构建模型,以及它如何交付给客户。我们提供完整的供应链完整性,并对从数据摄取、预训练、后训练到最终评估的每一个决策负责。我们提供透明度。客户拥有完全的部署自由和知识产权安全,因此合规性成为了该模型的一种固有属性。阅读我们的技术报告以获取完整详情。
What Kolibri Delivers
Kolibri 的优势
We optimized Kolibri for performance across a wide range of sectors considering their particular domain-specific language, regulatory, and procedural realities. Its small and efficient size provides our customers with flexibility to run it efficiently on-premise, without sending internal data to third-party inference services. The spotlight in this section introduces the model’s capabilities, before we describe them in section How we built Kolibri at high velocity. 我们针对广泛的行业优化了 Kolibri 的性能,充分考虑了它们特定的领域语言、监管要求和流程现实。其小巧高效的体积为客户提供了灵活性,使其能够高效地在本地运行,而无需将内部数据发送至第三方推理服务。本节重点介绍了模型的能力,随后我们将在“我们如何高速构建 Kolibri”一节中进行详细描述。
Foundational capabilities for enterprise and government
面向企业与政府的基础能力
With Kolibri we optimize the trade-off between model capability and deployment costs, using 3B active parameters out of 78B total. Kolibri sits on the Pareto frontier for quality versus serving cost, for both English and German. The Pareto frontier is a concept from economics, marking the best achievable combinations of two objectives, where improving on one means giving up some of the other. None of the compared models delivers more quality at the same serving cost, or the same quality at lower cost. 通过 Kolibri,我们在 78B 总参数中使用 3B 激活参数,优化了模型能力与部署成本之间的平衡。在英语和德语方面,Kolibri 都处于质量与服务成本的帕累托前沿(Pareto frontier)。帕累托前沿是一个经济学概念,标志着两个目标之间可达到的最佳组合,即改善其中一个目标意味着必须牺牲另一个目标。在同等服务成本下,没有其他对比模型能提供更高的质量;在同等质量下,也没有模型能实现更低的成本。
Contextualized performance for real-world applications
面向实际应用的情境化性能
Public benchmarks fail to capture specialized sector needs, so we developed our own internal evaluation suites for the verticals that matter to our customers, such as the German public sector, aviation, manufacturing and the automotive industry. Each suite mirrors the skills, workflows, and edge cases required in these sectors, and paired synthetic training environments let us improve Kolibri against these evaluations without ever training on customer data. 公共基准测试无法捕捉特定行业的需求,因此我们为客户关注的垂直领域(如德国公共部门、航空、制造和汽车工业)开发了内部评估套件。每个套件都反映了这些行业所需的技能、工作流程和边缘案例,配套的合成训练环境使我们能够在不使用客户数据进行训练的情况下,针对这些评估改进 Kolibri。