The AI Race Just Got Awkward

The AI Race Just Got Awkward

AI 竞赛变得有些尴尬了

If you read the news headlines these days, you would be forgiven for thinking that the Western labs are getting spawn-camped by Chinese labs en masse. 如果你最近阅读新闻头条,你会觉得西方实验室正被中国实验室大规模地“围堵”在出生点,这并不奇怪。

The Distillation Drama: Not a week goes by when Anthropic doesn’t release another article on how the Chinese are distilling their models, becoming a danger to humanity itself, etc. It’s beneficial for them to say that because it sets the ground for these models to be restrained legally and regulatorily later on. But it’s clear that the days of mindless distilling are over. Not just over. The new game in town is adopting Chinese labs’ advances. Note how I call this adoption instead of the more vitriol-infused “stealing” that Anthropic tends to use. That’s because, unlike the Western companies, the Chinese are pretty much giving away their recipes. 蒸馏风波:Anthropic 每周都会发布文章,声称中国实验室如何通过蒸馏模型,从而对人类构成威胁等。这种说法对他们有利,因为这为日后在法律和监管层面限制这些模型奠定了基础。但显而易见,盲目蒸馏的时代已经结束了。不仅如此,现在的游戏规则变成了“采用中国实验室的先进技术”。请注意,我称之为“采用”,而不是 Anthropic 倾向使用的、带有强烈敌意的“窃取”。这是因为,与西方公司不同,中国实验室几乎是在无偿分享他们的技术配方。

A Different Game: The latest one shamelessly copied without acknowledgement is the breakthrough in KV cache optimizations that DeepSeek has generously shared with the world. It is a mind-blowing optimization that basically dropped the KV cache footprint for certain use cases that use a long session context, like coding, by a factor of roughly 437x compared with DeepSeek-V1. They were the first ones to release the MLA architecture, which compressed the cache by roughly 15x, and then followed it up with ‘Compressed Sparse Attention’ and ‘Heavily Compressed Attention.’ The latest DeepSeek-V4.1-Flash pushes it even further with CSA2, cross-layer cache reuse, a causal encoder-decoder architecture, and FP4 caching, bringing the global KV cache down to 890 bytes per token. 不同的游戏:最近被毫无顾忌地“借鉴”且未予鸣谢的,是 DeepSeek 向全球慷慨分享的 KV 缓存优化突破。这是一项令人惊叹的优化,在处理长会话上下文(如编程)的特定用例中,它将 KV 缓存占用量相比 DeepSeek-V1 降低了约 437 倍。他们率先发布了 MLA 架构,将缓存压缩了约 15 倍,随后又推出了“压缩稀疏注意力”(Compressed Sparse Attention)和“重度压缩注意力”(Heavily Compressed Attention)。最新的 DeepSeek-V4.1-Flash 通过 CSA2、跨层缓存重用、因果编码器-解码器架构以及 FP4 缓存技术,将全局 KV 缓存进一步降低至每个 Token 仅 890 字节。

Why do these things matter? Because for serving long-context models, one of the largest costs is the VRAM needed to hold this cache in GPU memory. Below is the graph showing just how crazy this whole thing is: 为什么这些很重要?因为在服务长上下文模型时,最大的成本之一就是 GPU 内存中存储缓存所需的显存(VRAM)。下图展示了这一切有多么疯狂:

Follow the Cache Money: Compare that with what the same tier cost roughly two months ago. All prices below are per 1 million tokens: 追踪缓存成本:将其与大约两个月前同等级别的成本进行比较。以下所有价格均为每 100 万个 Token 的费用:

A Very Quiet Thank-You: All this must mean the Western AI companies are now extremely inference-margin positive. The constraints on access to advanced GPUs forced Chinese labs to make performance optimization a number one goal, and it shows in the results. Now I don’t know why they would freely give away such a breakthrough, but they just did, and for once both Anthropic and OpenAI released models that are basically top-tier and are using these optimizations. They do seem to be a little embarrassed by the copying. Hence the silent releases without much pre-announcement for both Claude Opus 5.5 and GPT-6.1 Sol. The user reviews have been stellar w.r.t. usage, and the quality doesn’t seem to be that far off compared to their flagship models (Claude Fable 5.1 and GPT-6 Astra). The cache read costs are the proof of the adoption. They dropped sharply: Opus 5.5 cut cache-read pricing by 60% versus Opus 5, while GPT-6.1 Sol cut it by 80% versus GPT-5.6 Sol’s late-July pricing. So the Chinese labs have thrown a lifeline to the Western loss-making labs, and I just have no clue as to why. 一声静悄悄的感谢:这一切意味着西方 AI 公司现在的推理利润率极高。获取先进 GPU 的限制迫使中国实验室将性能优化作为首要目标,而结果也证明了这一点。我不知道他们为什么要免费分享这样的突破,但他们确实这么做了。这一次,Anthropic 和 OpenAI 都发布了顶尖模型,并且都使用了这些优化技术。他们似乎对这种“借鉴”感到有些尴尬,因此 Claude Opus 5.5 和 GPT-6.1 Sol 都是在没有太多预告的情况下悄然发布的。用户对它们的使用评价极高,且质量与各自的旗舰模型(Claude Fable 5.1 和 GPT-6 Astra)相比似乎相差无几。缓存读取成本就是采用这些技术的证据,价格大幅下降:Opus 5.5 的缓存读取价格比 Opus 5 降低了 60%,而 GPT-6.1 Sol 比 GPT-5.6 Sol 7 月底的价格降低了 80%。所以,中国实验室向亏损的西方实验室抛出了救命稻草,而我完全不知道原因何在。