Aleph Alpha Kolibri: How the sovereign German LLM works

Aleph Alpha Kolibri: How the sovereign German LLM works

Aleph Alpha Kolibri:德国主权大语言模型的工作原理

Kolibri is an open-weight large language model (LLM) from Aleph Alpha for German and English: a mixture of experts with 78 billion parameters that only uses about 3.5 billion of them for each token it reads or writes. It came out on 3 October 2026 under the Apache 2.0 license, the weights are on Hugging Face, and it was trained from scratch on infrastructure in Germany and Finland. (Kolibri is German for hummingbird, which is cute for a model whose whole trick is being light.)

Kolibri 是 Aleph Alpha 推出的针对德语和英语的开放权重大型语言模型(LLM):它采用混合专家架构(MoE),拥有 780 亿参数,但在处理每个 token 时仅使用约 35 亿参数。该模型于 2026 年 10 月 3 日在 Apache 2.0 许可下发布,权重已上传至 Hugging Face,并完全在德国和芬兰的基础设施上从零训练而成。(Kolibri 在德语中意为“蜂鸟”,对于一个以轻量化为核心卖点的模型来说,这个名字非常贴切。)

I live in Germany, and at SmashingConf New York in 2024 I told the room what I’d heard in the US when I said where I’m based: “you regulate, you don’t innovate.” It hurt to hear, and what I wished for on that stage was the middle, “the right balance between innovation and regulation around data privacy, data stewardship, environmental constraints and energy requirements.” Kolibri is a pretty direct answer to that: a German team built it with the EU AI Act in mind “from the ground up”, and in Aleph Alpha’s own evaluation it scores above every compared model of its size in both languages.

我居住在德国。2024 年在纽约举行的 SmashingConf 大会上,当我提到自己常驻德国时,我向台下观众转述了在美国常听到的一句话:“你们只管监管,从不创新。”听到这话令人心酸,当时我在台上所呼吁的是一种中间地带,即“在数据隐私、数据管理、环境约束和能源需求方面,在创新与监管之间找到正确的平衡点。”Kolibri 正是对这一呼声的直接回应:一个德国团队在设计之初就充分考虑了《欧盟人工智能法案》(EU AI Act),而在 Aleph Alpha 的内部评估中,该模型在德语和英语的表现均优于同等规模的所有对比模型。

Huge congrats to everyone at Aleph Alpha who built it, my good friend Michael Hofmann among them! This post is about how Kolibri works, where it’s strong, where it isn’t, how to run it, and when it’s the right pick. Everything here comes from Aleph Alpha’s 189 page technical report, the model card and their launch post, plus one experiment I ran on its tokenizer.

衷心祝贺 Aleph Alpha 的所有开发人员,其中包括我的好友 Michael Hofmann!本文将探讨 Kolibri 的工作原理、优缺点、运行方式以及适用场景。文中的所有信息均来自 Aleph Alpha 长达 189 页的技术报告、模型卡片及其发布博文,此外还包含了我对该模型分词器(tokenizer)进行的一项实验。

What is Kolibri?

什么是 Kolibri?

  • Parameters: 78.1 billion in total, 3.46 billion per token (4.4%)

  • Languages: German and English

  • Context: 262,144 tokens natively, tested up to 1,048,576

  • License: Apache 2.0 for the weights and configuration files (Aleph Alpha keeps the rights to its training code and methods)

  • Memory: about 78 GB of weights in 8-bit floating point (FP8)

  • Reasoning: 4 levels: none, low, medium and high

  • Tool calling: Yes

  • Knowledge cutoff: 18 June 2026

  • Training: about 24 trillion tokens, more than a fifth of them German, on 768 NVIDIA B200 graphics processing units (GPUs)

  • 参数量: 总计 781 亿,每个 token 激活 34.6 亿(4.4%)

  • 语言: 德语和英语

  • 上下文窗口: 原生支持 262,144 个 token,测试最高可达 1,048,576

  • 许可: 权重和配置文件采用 Apache 2.0(Aleph Alpha 保留其训练代码和方法的所有权)

  • 内存需求: 约 78 GB(8 位浮点数 FP8 权重)

  • 推理能力: 4 个等级:无、低、中、高

  • 工具调用: 支持

  • 知识截止日期: 2026 年 6 月 18 日

  • 训练数据: 约 24 万亿个 token(其中超过五分之一为德语),在 768 块 NVIDIA B200 GPU 上完成训练

Aleph Alpha calls Kolibri sovereign, and in their launch post that means 2 things. The first is how it was built: “teams built the model in Germany, trained it on infrastructure in Germany and Finland, under European and German law, with no foreign control.” The second is what customers get: “full freedom of deployment and intellectual-property safety, so compliance comes as an inherited property.” In plain words, a ministry or a car supplier can run it on its own servers, with its data never leaving the building, and nobody can change or switch off the model under them. Aleph Alpha has also signed the European Union’s General-Purpose AI (GPAI) Code of Practice.

Aleph Alpha 将 Kolibri 称为“主权”模型,在发布博文中,这包含两层含义。第一层是关于构建方式:“团队在德国构建模型,在德国和芬兰的基础设施上进行训练,遵循欧洲和德国法律,且不受任何外国控制。”第二层是关于客户权益:“完全的部署自由和知识产权安全,合规性成为一种固有属性。”简单来说,政府部门或汽车供应商可以在自己的服务器上运行该模型,数据无需离开本地,且没有任何外部力量可以更改或关闭该模型。此外,Aleph Alpha 还签署了欧盟的《通用人工智能(GPAI)实践准则》。

Sovereign doesn’t mean that nothing from outside Europe went in, and the model card says so itself: English web text was rephrased with Google’s Gemma 4, German with Mistral-NeMo, and Qwen3-32B labeled data for the quality filters. They then filtered the training data for the political bias such models can have, which they’ve measured in Chinese open models themselves.

“主权”并不意味着模型完全没有使用欧洲以外的资源,模型卡片对此也直言不讳:英语网页文本使用了 Google 的 Gemma 4 进行重写,德语文本使用了 Mistral-NeMo,并使用 Qwen3-32B 的标注数据进行质量过滤。随后,他们对训练数据进行了过滤,以消除此类模型可能存在的政治偏见——这些偏见是他们此前在评估中国开源模型时自行测定的。

How Kolibri works

Kolibri 的工作原理

Kolibri is 6 ideas stacked on top of each other, and each one is there to make German cheaper, longer or more honest.

  1. 384 specialists, and each token sees 6

Kolibri 由 6 个叠加的理念构成,每一个理念都是为了让德语处理更廉价、更长效或更准确。

  1. 384 个专家,每个 token 激活 6 个

In a normal (dense) model, every token goes through every parameter. In a mixture of experts (MoE), each layer has a crowd of small sub-networks called experts and a router that picks a few of them for each token. Kolibri has 50 layers, each with 384 experts plus 1 shared expert that every token goes through, and its router sends each token to 6 of the 384. That’s how 78.1 billion parameters turn into 3.46 billion of actual work per token.

在普通的(稠密)模型中,每个 token 都要经过所有参数。而在混合专家模型(MoE)中,每一层都包含一组被称为“专家”的小型子网络,以及一个为每个 token 选择专家的路由器。Kolibri 拥有 50 层,每层有 384 个专家,外加 1 个所有 token 都会经过的共享专家。路由器会将每个 token 发送给 384 个专家中的 6 个。这就是 781 亿参数如何转化为每个 token 仅需 34.6 亿参数计算量的原理。

In 2024 I gave a talk called Why Small Language Models are the future, and I argued for “smaller language models with fewer parameters and fewer places things can go wrong that require lesser compute.” My analogy was a doctor who has read every medical book in the world against a specialist in hematology: go to the first one with a blood condition and “they may not get it right cuz they know too much.” A mixture of experts puts a hospital full of specialists inside one model, and the router is the receptionist who sends each token to the right 6.

2024 年,我曾发表过题为《为什么小型语言模型是未来》的演讲,我主张“使用参数更少、出错概率更低、计算需求更小的小型语言模型”。我当时打了个比方:一位读过世界上所有医学书籍的全科医生与一位血液学专家相比,如果你患有血液病去找前者,他可能会因为“知道得太多”而无法做出准确判断。混合专家模型就像是在一个模型里装进了一整家医院的专家,而路由器就是那个负责将每个 token 引导至正确 6 位专家面前的接待员。

The analogy breaks in 2 places though. The experts aren’t neat topics like “German law”: when researchers look inside MoE models they mostly find experts for patterns of tokens, like punctuation or proper nouns, not subjects a person would pick. And the hospital has to keep all 384 specialists on staff even if you only see 6, so Kolibri computes like a 3.5 billion parameter model but needs the memory of a 78 billion parameter one. The model card says it plainly: “the full model must be held in memory even though only part of it is active at any time.”

不过,这个比喻在两处存在偏差。首先,专家并非像“德国法律”这样清晰的主题:当研究人员深入 MoE 模型内部时,他们发现专家通常是针对 token 模式(如标点符号或专有名词)进行优化的,而不是人类所理解的学科。其次,即使你每次只用到 6 个专家,医院也必须保留所有 384 名专家的编制。因此,Kolibri 的计算量相当于 35 亿参数的模型,但内存需求却需要 780 亿参数的规模。模型卡片对此说得很清楚:“尽管在任何时候只有一部分模型处于激活状态,但整个模型必须全部驻留在内存中。”

  1. A tokenizer that reads long German words
  2. 能够读取长德语单词的分词器

A model doesn’t read letters or words, it reads tokens: chunks of text from a fixed vocabulary, picked when the tokenizer is trained. German glues words together into long compound words, and a tokenizer that learned mostly from English chops them into pieces. Here’s the German name of the Federal Constitutional Court, split by the tokenizer GPT-4o and GPT-5 use (o200k_base, through OpenAI’s tiktoken), and by Kolibri’s:

模型读取的不是字母或单词,而是 token:即在分词器训练时从固定词汇表中选出的文本片段。德语习惯将单词拼接成很长的复合词,而主要从英语学习的分词器往往会将它们切碎。以下是德国联邦宪法法院的德语名称,分别由 GPT-4o 和 GPT-5 使用的分词器(o200k_base)以及 Kolibri 的分词器进行拆分的结果:

o200k_base (GPT-5): Bund | es | ver | fass | ungs | gericht (6 tokens) Kolibri: Bundes | verfassungsgericht (2 tokens)

o200k_base (GPT-5): Bund | es | ver | fass | ungs | gericht (6 个 token) Kolibri: Bundes | verfassungsgericht (2 个 token)

Kolibri’s tokenizer has 128,000 tokens, trained with a new algorithm Aleph Alpha calls UniBPE: it keeps the bottom-up merging of byte-pair encoding (BPE) and picks each merge with a different scoring rule (the Unigram objective), which respects how German builds words. The report says it needs 11.2% fewer tokens for German text than GPT-5’s tokenizer, the best of the 9 others they measured. I wanted to see that for myself, so I ran 6 tokenizers over all of the Basic Law for the Federal Republic of Germany, the German constitution (185 KB of very German legal text), and over its official English translation:

Kolibri 的分词器拥有 12.8 万个 token,采用了一种 Aleph Alpha 称为 UniBPE 的新算法训练:它保留了字节对编码(BPE)的自底向上合并方式,并使用不同的评分规则(Unigram 目标)来选择每次合并,从而更好地适应德语的构词方式。报告称,在处理德语文本时,它比 GPT-5 的分词器(在他们测试的 9 种分词器中表现最好)所需的 token 数量减少了 11.2%。为了亲自验证这一点,我用 6 种分词器对《德意志联邦共和国基本法》(即德国宪法,约 185 KB 的纯德语法律文本)及其官方英译本进行了测试: