The arguments against open source AI are bad

The arguments against open source AI are bad

反对开源人工智能的论点是站不住脚的

The release of Kimi K3 has opened a fresh round of angst and confused discourse. There’s a loud cohort of journalists, business leaders, and politicians arguing that open source AI is a dangerous threat. OpenAI’s Dean Ball: One probable outcome of an open-weight-model-dominant world is full AI communism… rather than a market product, AI is a “public good” Freely available AI for anyone? The horror! Kimi K3 的发布引发了新一轮的焦虑和混乱讨论。有一群声音很大的记者、商业领袖和政客声称,开源人工智能是一种危险的威胁。OpenAI 的 Dean Ball 曾表示:一个以开放权重模型为主导的世界,其可能的结果是彻底的“人工智能共产主义”……人工智能不再是市场产品,而是一种“公共产品”。让任何人都能免费使用人工智能?太可怕了!

Frontier labs’ case against open source AI is essentially: Open source models are dangerous (and un-American!). We should open the AI Pandora’s Box, but only with responsible gatekeepers (toll collectors, preferably us!). Only trusted users (our most profitable customers) should be able to use it. I want to address some bad arguments against open source AI, but some corrections on how the argument is being framed are in order: 前沿实验室(Frontier labs)反对开源人工智能的理由本质上是:开源模型是危险的(而且是不符合美国精神的!)。我们应该打开人工智能的潘多拉魔盒,但必须由负责任的守门人(最好是我们这些收费站管理员!)来把关。只有受信任的用户(我们最赚钱的客户)才应该被允许使用它。我想反驳一些针对开源人工智能的错误论点,但首先需要纠正一下目前讨论这些问题的方式:

Open source software is the foundation for commercial software

开源软件是商业软件的基石

Ball’s framing strolls past the fact that open source software is the foundation of all proprietary software. This includes frontier models, which at the end of the day are software products. Open source software is counterintuitive to people outside of the software industry. Why work hard on a product, and give it away for free? A software program is a stack of programs, with each layer built on top of another. To build Uber, you need programming language frameworks, software to send and receive web traffic, data analysis tools, and countless other components. Most of these are not differentiators for a commercial enterprise, so it serves commercial actors to cooperate on lower components in the stack and compete on the higher level pieces that actually differentiate their products. Frontier labs would very much like AI models to not fall into the category of “so commonplace that it doesn’t make sense to compete on”. Whether that happens remains to be seen. Ball 的论调忽略了一个事实:开源软件是所有专有软件的基石。这当然也包括前沿模型,因为归根结底,它们也是软件产品。对于软件行业以外的人来说,开源软件的逻辑是反直觉的:为什么要辛苦开发一个产品,然后免费送人呢?一个软件程序是由一系列程序堆叠而成的,每一层都构建在另一层之上。要构建像 Uber 这样的应用,你需要编程语言框架、收发网络流量的软件、数据分析工具以及无数其他组件。其中大多数组件对于商业企业来说并不构成差异化优势,因此,商业参与者在底层组件上进行合作,而在真正能体现产品差异化的高层组件上进行竞争,才符合他们的利益。前沿实验室非常希望人工智能模型不要沦为“过于普遍以至于无需竞争”的范畴。但这是否会发生,还有待观察。

Open source software is very difficult to suppress

开源软件极难被压制

In reality, the argument about suppressing open source models is mostly beside the point. History tells us that suppression of open source software is extremely difficult, and attempting to do so only serves to weaken companies against international competitors. A brief history of encryption is illustrative: Today, PGP is a commonplace tool anyone can use, and most devs are at least familiar with. But when Phil Zimmermann invented it in 1991, the U.S. government considered encryption to be military technology. A criminal investigation was opened against Zimmermann. When Netscape created SSL, the U.S. government allowed it to only release a weakened version of it internationally. These controls backfired: it was much easier to acquire the weakened, “international” version, so even many Americans used it. Export controls did not succeed in limiting encryption as the government wished. SSL, PGP, and similar tools were readily available throughout the world, and the controls disadvantaged Americans. Eventually, courts ruled that releasing encryption source code is protected speech, and the U.S. government relaxed encryption export controls. 事实上,关于压制开源模型的争论大多偏离了重点。历史告诉我们,压制开源软件极其困难,而试图这样做只会削弱本国公司在国际竞争中的地位。加密技术的简史就是一个很好的例子:今天,PGP 是任何人都可以使用的通用工具,大多数开发者至少都很熟悉它。但当 Phil Zimmermann 在 1991 年发明它时,美国政府将其视为军事技术,并对 Zimmermann 展开了刑事调查。当网景(Netscape)创建 SSL 时,美国政府只允许其在国际上发布弱化版本。这些控制措施适得其反:获取弱化的“国际版”要容易得多,以至于许多美国人也在使用它。出口管制并没有像政府希望的那样限制住加密技术。SSL、PGP 和类似的工具在世界各地随处可见,而这些控制措施反而让美国人处于劣势。最终,法院裁定发布加密源代码属于受保护的言论,美国政府才放宽了加密技术的出口管制。

Narrowing suppression to “Chinese” models won’t make things easier. What, exactly, makes an AI model Chinese? Is it Chinese if, as frontier models allege, it was distilled from American models? What about if an American fine-tunes a Chinese model? At best, regulating AI in this way will (temporarily) encumber Americans with red tape and diminished AI access relative to the rest of the world. 将压制范围缩小到“中国”模型并不会让事情变得简单。究竟是什么让一个人工智能模型成为“中国”模型?如果像前沿实验室所声称的那样,它是从美国模型中蒸馏出来的,那它算中国模型吗?如果一个美国人微调了一个中国模型又该怎么算?往好了说,这种监管方式只会(暂时)让美国人陷入繁文缛节,并使他们相对于世界其他地区在获取人工智能方面处于劣势。

Open source AI is not just a Chinese phenomenon

开源人工智能不仅仅是中国现象

There’s an assumption baked into the open source AI debate that open source models are something that only the Chinese government has an incentive to develop. In reality there are many commercial actors with ample incentive to develop open source AI: 在开源人工智能的辩论中,存在一种根深蒂固的假设,即开源模型只有中国政府才有动力去开发。但实际上,许多商业参与者都有充分的动力去开发开源人工智能:

  • Chip makers: Nvidia CEO Jensen Huang has described what Nvidia is building as “token factories”. Nvidia doesn’t care if its chips are used to run frontier models or cheap open source models - it just wants to produce and generate demand for as many tokens as possible. And indeed Nvidia has itself released a suite of open source models. 芯片制造商: 英伟达 CEO 黄仁勋将英伟达正在构建的东西描述为“Token 工厂”。英伟达并不关心它的芯片是被用于运行前沿模型还是廉价的开源模型——它只想尽可能多地生产芯片并创造对 Token 的需求。事实上,英伟达自己也发布了一系列开源模型。
  • American Startups: Thinking Machines Lab recently released a powerful open source model. They and others are betting that models will be commoditized, and a defensible moat can be built around auxiliary services that complement or customize models. 美国初创公司: Thinking Machines Lab 最近发布了一个强大的开源模型。他们和其他公司都在押注模型将实现商品化,并认为可以通过围绕模型进行补充或定制的辅助服务来建立防御性护城河。
  • Enterprise AI users: Frontier model customers aren’t currently all that active in open source AI development, but they will be. They will want lower-cost models for low-complexity tasks, and more fine-grained control over customer-facing features. 企业人工智能用户: 前沿模型的客户目前在开源人工智能开发方面并不活跃,但未来会。他们将需要成本更低的模型来处理低复杂度任务,并希望对面向客户的功能拥有更精细的控制权。
  • BigCos: You can be sure that Google and Meta are watching OpenAI’s new ad product closely. Should frontier model ad products gain traction, it would be well worth it for these behemoths to commoditize ad-free, open source models to squash ad competition. 大型科技公司: 你可以肯定,谷歌和 Meta 正在密切关注 OpenAI 的新广告产品。如果前沿模型的广告产品获得成功,这些巨头完全有理由将无广告的开源模型商品化,以打压广告竞争。

The “AI race” is… what, exactly?

所谓的“人工智能竞赛”究竟是什么?

Much of the angst around China’s models centers on “losing the AI race”. But what’s the goal of this race? Is it to develop the best model? To sell the most tokens? To destroy humanity first? Talking about an “AI Race” doesn’t make more sense than talking about an “Internet Race”. We’re not competing to be the first to send a rocket to the moon, we’re reacting to a new, transformational technology. To the extent there’s a race between nations, it’s to absorb this transition and grow economies. In this framing, free AI models are a boon, not a threat. 围绕中国模型的许多焦虑都集中在“输掉人工智能竞赛”上。但这场竞赛的目标是什么?是开发出最好的模型?卖出最多的 Token?还是率先毁灭人类?谈论“人工智能竞赛”并不比谈论“互联网竞赛”更有意义。我们不是在竞争谁能第一个把火箭送上月球,我们是在应对一种全新的、变革性的技术。如果说国家之间存在竞赛,那也是在比拼谁能更好地吸收这种转型并促进经济增长。在这种框架下,免费的人工智能模型是福音,而不是威胁。

Bad arguments to fear Chinese AI models

对中国人工智能模型感到恐惧的错误论点

China is “AI dumping!” 中国在进行“人工智能倾销”!

Scott Galloway has argued that free Chinese AI is an attempt to eliminate competitors in the long run: This is what China did to solar panels, steel, EVs, and batteries. First, they match Western quality, or they don’t even match it. 89%. Close. Actually, match it with cars, they’ve matched it, but go ahead. Then they cut the price by two thirds, then they own the market. But apart from chips, AI isn’t a physical good. Solar panels and steel require physical supply chain. Scott Galloway 曾指出,免费的中国人工智能是试图在长期内消灭竞争对手:这就是中国在太阳能电池板、钢铁、电动汽车和电池领域所做的事情。首先,他们达到西方产品的质量,或者甚至还没达到,比如 89%,差不多了。实际上,在汽车方面他们已经达到了,但继续说。然后他们将价格削减三分之二,接着就占领了市场。但除了芯片之外,人工智能并不是一种实物商品。太阳能电池板和钢铁需要物理供应链。