State of Open Models: Summer 2026 Observations

State of Open Models: Summer 2026 Observations

开放模型现状:2026 年夏季观察报告

In the AI world, time feels compressed. A few months after our spring report in our biannual analysis worked through the ecosystem, there are quite a few findings that we have observed until this summer. This report lays out these observations from January to August 2026 and presents the data behind each one. 在人工智能领域,时间仿佛被压缩了。在我们半年一度的生态系统分析报告发布几个月后,我们观察到了不少新的发现。本报告梳理了 2026 年 1 月至 8 月期间的观察结果,并提供了每一项发现背后的数据支持。

Models and datasets on HF hub are growing on a daily basis. Public model repositories grew from 2.43 to 2.96 million over the period, datasets from 711,000 to 1 million, Spaces from 1.00 to 1.44 million. The distribution underneath stays extreme, roughly 85.6% of models have fewer than 200 lifetime downloads, and 1.5% of repositories account for 99.2% of all downloads. Everything below happens inside that shape. Hugging Face 平台上的模型和数据集正以日为单位增长。在此期间,公共模型仓库从 243 万个增长到 296 万个,数据集从 71.1 万个增长到 100 万个,Spaces 应用从 100 万个增长到 144 万个。其底层的分布依然极端:约 85.6% 的模型生命周期下载量不足 200 次,而 1.5% 的仓库贡献了 99.2% 的总下载量。以下所有分析均基于这一格局展开。

1. The frontier is moving fast

1. 前沿领域发展迅速

There used to be a clear progression path: labs would start by releasing smaller models and gradually work their way toward the top end of the scale. In 2026, several Chinese labs skipped this progression entirely. In almost every month of 2026, the largest and most performant open model from a Chinese lab was larger than anything an American lab released of its own. 过去,模型发布有着清晰的演进路径:实验室通常先发布较小的模型,再逐步向更大规模迈进。但在 2026 年,几家中国实验室完全跳过了这一过程。在 2026 年的几乎每个月里,中国实验室发布的最强大、性能最好的开源模型,其规模都超过了美国实验室同期发布的任何自主模型。

China’s monthly ceiling ran between 754B and 2.78 trillion parameters; America’s own ceiling stayed under 130B in five of seven months, the exception being NVIDIA’s Nemotron 3 Ultra at 561B in May and June, and Inkling from Thinking Machines Lab. The chart splits the labs into two camps. Moonshot, MiniMax, Xiaomi and Z.ai publish almost nothing below 70B, so a developer’s first encounter with them is a model too large to run on anything they own. Tencent and Alibaba Qwen cover the whole range instead, from under 1B upward. 中国实验室的月度参数上限在 7540 亿到 2.78 万亿之间;而美国实验室的上限在七个月中有五个月保持在 1300 亿以下,例外情况是英伟达 5 月和 6 月发布的 5610 亿参数的 Nemotron 3 Ultra,以及 Thinking Machines Lab 的 Inkling 模型。图表将实验室分为两个阵营:月之暗面(Moonshot)、MiniMax、小米和 Z.ai 几乎不发布 700 亿参数以下的模型,因此开发者首次接触它们时,往往会发现模型大到无法在个人设备上运行。相比之下,腾讯和阿里通义千问(Qwen)则覆盖了从 10 亿以下到超大规模的全系列。

Two things made the first camp possible. Building large stopped being a differentiator. Xiaomi and Meituan both cleared a trillion parameters this year, and neither was a household name in open weights twelve months ago. And a lab no longer has to ship a small model to be reachable, because the community’s quantization layer will make a large one runnable within days, a dependency we return to below. That leaves the size profile as a statement of intent rather than of capability. 第一阵营的出现得益于两点:首先,构建大模型不再是核心差异化优势。小米和美团今年都突破了万亿参数大关,而十二个月前它们在开源权重领域还名不见经传。其次,实验室不再需要发布小模型来获取用户,因为社区的量化层能在几天内让大模型变得可运行(我们稍后会讨论这种依赖关系)。这使得模型规模不再是能力的体现,而是一种意图的声明。

A frontier only portfolio stakes everything on benchmark position and API demand. A full spectrum portfolio is a bid to be the family developers standardise on. Both are rational, they are playing for different prizes. “前沿导向”的产品组合将赌注全押在基准测试排名和 API 需求上;而“全谱系”的产品组合则是为了成为开发者标准化的首选。两者都是理性的,只是追求的目标不同。

The United States, meanwhile, is not absent from open source. The two organizations publishing the most new open models this year are also the companies making the hardware: AMD and NVIDIA. Each released more than 200 new model repositories, far ahead of the rest of the field, with LiquidAI ranking third at around 100. Hardware vendors have realized that open models are a way to sell chips: a model optimized for your hardware and freely available is the clearest proof that the hardware works. 与此同时,美国在开源领域并未缺席。今年发布新开源模型最多的两个机构正是硬件制造商:AMD 和英伟达。它们各自发布了超过 200 个新模型仓库,遥遥领先于其他竞争对手,LiquidAI 以约 100 个位居第三。硬件厂商已经意识到,开源模型是销售芯片的利器:一个针对特定硬件优化且免费提供的模型,是证明该硬件性能最直观的证据。

When smaller models and embedding models are included, where Google, Microsoft, IBM Granite, and OpenAI’s older vision and speech models generate hundreds of millions of downloads annually, U.S. participation in open source AI is still growing. However, the center of gravity has shifted. Google and Meta now rank well below NVIDIA in new model releases, despite being the companies that defined open model publishing in previous years. Meta’s move toward closed flagship models further highlights this change. Open source has moved from model labs to hardware and infrastructure companies. 如果算上小型模型和嵌入模型(Embedding models),考虑到谷歌、微软、IBM Granite 以及 OpenAI 的旧版视觉和语音模型每年产生数亿次下载,美国在开源 AI 领域的参与度仍在增长。然而,重心已经发生了转移。尽管谷歌和 Meta 在过去几年定义了开源模型发布,但如今它们在新模型发布数量上已远落后于英伟达。Meta 向闭源旗舰模型转型的举动进一步凸显了这一变化。开源的重心已从模型实验室转移到了硬件和基础设施公司。

At the frontier scale, the picture is very different. Most U.S. releases above 100B parameters this year are not new models, but built on top of Chinese models. Only a few major original American models appear at this scale: Thinking Machines’ Inkling (952B), NVIDIA’s Nemotron 3 Ultra (561B), Nemotron 3 Super (124B), and Arcee AI’s Trinity-Large (399B). AMD contributed many conversions but no original model at this scale. This work is still important: it enables trillion-parameter Chinese models to run efficiently on American hardware. But it represents a distribution and optimization layer rather than model creation. Meanwhile, Chinese open models are increasingly optimized for domestic chips, the same competition in reverse, where models are designed around specific hardware ecosystems. 在“前沿规模”上,情况则大不相同。今年美国发布的大多数超过 1000 亿参数的模型并非原创,而是基于中国模型构建的。在此规模下,仅有少数主要的美国原创模型:Thinking Machines 的 Inkling(9520 亿)、英伟达的 Nemotron 3 Ultra(5610 亿)、Nemotron 3 Super(1240 亿)以及 Arcee AI 的 Trinity-Large(3990 亿)。AMD 贡献了许多转换版本,但没有在此规模下发布原创模型。这项工作依然重要:它使万亿参数的中国模型能够在美制硬件上高效运行。但这代表的是分发和优化层,而非模型创造。与此同时,中国的开源模型正日益针对国产芯片进行优化,这是一种反向的竞争,即模型围绕特定的硬件生态系统进行设计。

2. Attention ≠ Adoption

2. 关注度 ≠ 采用率

We took the top 25 model repositories by downloads accumulated this year and the top 25 by likes. Exactly one repository appears in both lists. We counted downloads inside the window rather than lifetime, so nothing is credited for merely having existed longer, and controlling for age makes the split sharper. Not one model published in 2026 reaches the download top 25, while thirteen of the twenty-five date from 2022. 我们选取了今年累计下载量排名前 25 的模型仓库,以及点赞数排名前 25 的仓库。结果发现,只有一个仓库同时出现在两个榜单中。我们统计的是窗口期内的下载量而非生命周期总量,因此模型不会仅仅因为发布时间长而获得加分,这种对时间因素的控制使得差异更加明显。2026 年发布的模型中,没有一个进入下载量前 25 名,而前 25 名中有 13 个模型是 2022 年发布的。

all-MiniLM-L6-v2 was pulled 1.55 billion times in seven months against 5,156 likes; Kimi-K3 was pulled about 60 times per like it received. The two numbers record different acts. A like says a release matters, and goes to frontier models in the weeks after they ship. A download says something is wired into a pipeline that runs on a schedule, and accrues to small, stable models over years. Likes are the right instrument for reading what the field is excited about, downloads for reading what it currently depends on. Treating either as a proxy for the other is the most common mistake we see in coverage of the Hub, including our own earlier work. all-MiniLM-L6-v2 在七个月内被下载了 15.5 亿次,却仅有 5156 个点赞;而 Kimi-K3 平均每个点赞对应约 60 次下载。这两个数字记录了不同的行为:点赞代表一个发布备受关注,通常发生在模型发布后的几周内,集中在“前沿模型”上;而下载则意味着模型已被整合进定期运行的流水线中,这通常属于那些小型、稳定的模型,且随时间推移而积累。点赞是衡量领域内兴奋点的正确工具,而下载量则是衡量当前依赖程度的指标。将两者互为代理指标,是我们观察到的关于 Hugging Face 报道中最常见的错误,包括我们自己早期的工作。

The same split appears at the level of the publisher. Chinese frontier labs are the only accounts on the Hub where the heavy band carries the volume. Effectively all of MiniMax’s 2026 downloads are of models above 70B, along with 88% of Moonshot’s, 55% of DeepSeek’s and 39% of Z.ai’s. No large American account looks like this: Google, Microsoft and IBM Granite record essentially none of their 2026 downloads above… 这种分化在发布者层面同样存在。中国的前沿实验室是 Hugging Face 上唯一依靠“重型模型”支撑下载量的账户。MiniMax 2026 年的下载量几乎全部来自 700 亿参数以上的模型,月之暗面(Moonshot)为 88%,深度求索(DeepSeek)为 55%,Z.ai 为 39%。没有任何大型美国账户呈现这种特征:谷歌、微软和 IBM Granite 在 2026 年的下载量中,几乎没有来自超大规模模型……