AMD acquires Taalas to boost inference performance by etching models in silicon

AMD acquires Taalas to boost inference performance by etching models in silicon

AMD 收购 Taalas,通过将模型刻入硅片提升推理性能

In AMD’s latest bid to upset Nvidia’s dominance in AI hardware, the House of Zen has acquired AI chip company Taalas, which bakes model weights directly into silicon in a process that promises to boost inference performance by an order of magnitude or more. 为了挑战英伟达在 AI 硬件领域的统治地位,AMD 近日收购了 AI 芯片公司 Taalas。该公司通过将模型权重直接“烘焙”进硅片,有望将推理性能提升一个数量级甚至更多。

The deal, announced at market close on Thursday, appears to be framed in much the same context as Nvidia’s $20 billion licensing deal with Groq last December: make high-performance “premium” inference services prized for AI agents, like code assistants, faster and cheaper to run. AMD didn’t disclose the terms of the deal, but from what we understand, this is an actual acquisition rather than an acquihire. 这笔交易于周四收盘时宣布,其背景似乎与去年 12 月英伟达与 Groq 达成的 200 亿美元许可协议如出一辙:旨在让代码助手等 AI 代理所依赖的高性能“高级”推理服务运行得更快、成本更低。AMD 未披露交易条款,但据我们了解,这是一次真正的收购,而非单纯的人才收购(acquihire)。

Founded in 2023 and based in Toronto, Taalas’ approach to inference is radically different from conventional GPUs or the dataflow architectures that underpin Groq LPUs or Cerebras’ waferscale accelerators. The startup’s chips don’t rely on HBM to store the model weights but rather etch them directly into the silicon. In a sense, Taalas’ chips are really model-specific integrated circuits or MSICs. Taalas 成立于 2023 年,总部位于多伦多。其推理方法与传统 GPU 或支撑 Groq LPU 及 Cerebras 晶圆级加速器的数据流架构截然不同。该初创公司的芯片不依赖 HBM(高带宽内存)来存储模型权重,而是将其直接刻入硅片。从某种意义上说,Taalas 的芯片实际上是“模型专用集成电路”(MSIC)。

Perhaps more importantly, Taalas’ tech isn’t just conceptual. In February, the startup revealed its first test chip fabbed on TSMC’s 6nm process tech, which it called the HC1. Initial benchmarks saw the chip serve Meta’s Llama 3.1 8B at a blistering 16,960 tokens a second — when announced last February, that was 48x faster than Nvidia’s GPUs and 8.5x faster than Cerebras’ accelerators. 更重要的是,Taalas 的技术并非仅仅停留在概念阶段。今年 2 月,该公司展示了其首款采用台积电 6nm 工艺制造的测试芯片 HC1。初步基准测试显示,该芯片运行 Meta 的 Llama 3.1 8B 模型时,速度高达每秒 16,960 个 token——在去年 2 月发布时,这一速度比英伟达 GPU 快 48 倍,比 Cerebras 加速器快 8.5 倍。

While Llama 3.1 is ancient by today’s standards, having made its debut all the way back in mid 2024, the reticle-sized chip was really intended to prove the concept. Taalas has been incredibly secretive about how its chips actually work, but we know its processors are comprised of two main regions: the mask-ROM recall fabric where model weights are etched, and the SRAM recall fabric where KV caches and fine-tuning adapters are stored. 虽然以今天的标准来看,2024 年中期发布的 Llama 3.1 已属“古董”,但这款光罩尺寸的芯片主要是为了验证概念。Taalas 对其芯片的工作原理一直守口如瓶,但我们已知其处理器由两个主要区域组成:刻录模型权重的掩模 ROM(mask-ROM)存储区,以及存储 KV 缓存和微调适配器的 SRAM 存储区。

For its second-gen HC2 chip due out this summer, Taalas aims to boost parameter count to 20 billion parameters. That might not sound like much, but just like with GPUs for larger models, weights are simply distributed across multiple accelerators using pipeline parallelism. At 20 billion parameters per chip, you’d need just 50 accelerators to support a trillion-parameter model, and AMD just so happens to have a rack-scale compute platform and in-house system design team that can comfortably accommodate that. 对于今年夏天即将推出的第二代 HC2 芯片,Taalas 的目标是将参数量提升至 200 亿。这听起来可能不多,但就像处理大型模型的 GPU 一样,权重可以通过流水线并行分布在多个加速器上。每颗芯片 200 亿参数,只需 50 个加速器即可支持万亿参数模型,而 AMD 正好拥有能够轻松容纳这一规模的机架级计算平台和内部系统设计团队。

That’s quite a bit more space and power efficient than Nvidia’s recently unveiled LPX systems, which would need a few dozen GPUs and at least 2,000 Groq LPUs to serve the same model. From what we understand, AMD intends to pair its Instinct-based Helios racks with chips based on Taalas’ tech, which implies a disaggregated architecture where compute-heavy prompt processing is done on GPUs while token generation is offloaded to Taalas-based accelerators. 这比英伟达最近推出的 LPX 系统在空间和能效上要高效得多,后者需要几十个 GPU 和至少 2,000 个 Groq LPU 才能服务于同一个模型。据我们了解,AMD 打算将其基于 Instinct 的 Helios 机架与基于 Taalas 技术的芯片配对,这意味着一种解耦架构:计算密集型的提示词处理由 GPU 完成,而 token 生成则卸载到基于 Taalas 的加速器上。

It’s also possible that AMD could adopt a sort of tick-tock cadence in which customers initially deploy and validate models on Instinct accelerators and, once they’re satisfied with them, transition to Taalas accelerators. We can only speculate at this point, but here’s what AMD’s SVP of AI, Vamsi Boppana, had to say about it in a canned statement: “AMD is building a full-stack AI platform that gives customers the flexibility to deploy the right compute solutions for every AI workload.” AMD 还有可能采取一种“钟摆式”策略,即客户最初在 Instinct 加速器上部署和验证模型,一旦满意,再迁移到 Taalas 加速器上。目前我们只能推测,但 AMD AI 高级副总裁 Vamsi Boppana 在一份官方声明中表示:“AMD 正在构建一个全栈 AI 平台,为客户提供灵活性,以便为每种 AI 工作负载部署合适的计算解决方案。”

You better really love that model

你最好真的钟情于那个模型

While the tech is blazing fast, if you hadn’t already figured it out, it comes with a pretty substantial downside. Once the chips are deployed you’re stuck with that model. Any change bigger than something like a LoRA adapter is going to require a re-spin of the chips, which is not only expensive but time-consuming. 虽然这项技术速度极快,但如果你还没意识到,它有一个相当大的缺点:一旦芯片部署完毕,你就被锁定在该模型上了。任何超过 LoRA 适配器规模的改动都需要重新流片(re-spin),这不仅昂贵,而且耗时。

Nearly four years into the AI boom, new models are rolling out on a nearly monthly basis. In order to benefit from Taalas’ tech, AMD’s customers are going to have to be really sure about their choice of models, which will be easier for some than others. However, if the startup is to be believed, the situation isn’t quite as bad as it sounds. While new models will require a re-spin, it doesn’t require starting over from scratch. Instead, just two layers of metal need to be changed, which is a lot cheaper and less time-consuming. 在 AI 热潮持续近四年后,新模型几乎每月都在涌现。为了从 Taalas 的技术中获益,AMD 的客户必须对模型选择有十足的把握,这对某些客户来说比其他人更容易。然而,如果这家初创公司所言属实,情况并没有听起来那么糟。虽然新模型需要重新流片,但并不需要从零开始。只需更改两层金属层即可,这要便宜且省时得多。

With that said, we strongly suspect this tech will largely be deployed by AI model devs, their infrastructure providers, and a handful of inference providers. In an interview with our sibling site The Next Platform in February, the company suggested that etching a model’s weights into silicon is 100x less expensive than training a frontier model. 话虽如此,我们强烈怀疑这项技术将主要由 AI 模型开发者、基础设施提供商以及少数推理服务商部署。在 2 月份接受我们姊妹网站 The Next Platform 采访时,该公司表示,将模型权重刻入硅片的成本比训练一个前沿模型要低 100 倍。

AMD is certainly in a position to negotiate those deals. OpenAI, Anthropic, and Meta are all major Instinct customers. Given the close working relationship between the model houses and the chip designer, it wouldn’t be surprising to see a GPT or Claude deployed on a combination of Taalas and instinct accelerators. AMD 当然有能力促成这些交易。OpenAI、Anthropic 和 Meta 都是 Instinct 的主要客户。考虑到模型厂商与芯片设计商之间紧密的合作关系,看到 GPT 或 Claude 运行在 Taalas 和 Instinct 加速器的组合上也就不足为奇了。

The tech also has implications for model development. One of the ways developers have cut down on hallucinations is by trading time for accuracy. The technique, called test-time scaling, is quite simple in practice, and involves allowing a model to “think” for longer before responding. One drawback of test-time scaling is that it consumes substantially more tokens, which makes it expensive, and means users have to wait longer for the chatbot, code assistant, or agent to respond. 这项技术对模型开发也有影响。开发者减少幻觉的方法之一是用时间换取准确性。这种被称为“测试时扩展”(test-time scaling)的技术在实践中非常简单,即允许模型在响应前“思考”更长时间。测试时扩展的一个缺点是它会消耗大量额外的 token,这不仅昂贵,还意味着用户必须等待更长时间才能得到聊天机器人、代码助手或代理的响应。