OpenAI Jalapeño: Better than Nvidia Blackwell

OpenAI Jalapeño: Better than Nvidia Blackwell

OpenAI’s self-designed ASIC compared with Rubin, Jalapeño’s TCO, throughput per MW, and spicy deets. OpenAI 自研 ASIC 对比 Rubin:Jalapeño 的总拥有成本(TCO)、每兆瓦吞吐量及更多内幕细节。

OpenAI has spent the past couple years quietly building “Jalapeño,” an inference chip just announced at Hot Chips. Rumors of a successful tapeout had been swirling for a while. But now we have details. OpenAI invited us to look at their chip, go to their labs to check out how real it is, and benchmark it with our InferenceX suite. 过去几年,OpenAI 一直在悄悄研发名为“Jalapeño”的推理芯片,该芯片刚刚在 Hot Chips 大会上正式发布。关于其流片成功的传闻早已流传多时,现在我们终于获得了详细信息。OpenAI 邀请我们参观了他们的芯片,前往其实验室验证其真实性,并使用我们的 InferenceX 套件对其进行了基准测试。

In June, OpenAI unveiled the chip program in partnership with Broadcom, built from a blank slate exclusively for LLM inference. Design work began in the middle of 2024, going from initial team hiring to manufacturing tape-out in ~16 months, an extremely fast ASIC development cycle. 今年 6 月,OpenAI 公布了与博通(Broadcom)合作的芯片计划,该芯片从零开始设计,专为大语言模型(LLM)推理而打造。设计工作始于 2024 年年中,从最初的团队组建到最终流片仅耗时约 16 个月,这是一个极快的 ASIC 开发周期。

In general, first-generation chips are not competitive, but OpenAI bucks the trend by being industry-leading and beating every Nvidia, AMD, and Google chip we have been able to test on multiple top open-source models. OpenAI does this with extreme hardware-software codesign. Surprisingly, OpenAI is not over-specializing on any specific part of model inference, but instead by focusing on being a general chip that delivers high performance in all scenarios. In this article, we will go into architectural details, software details, and performance results for Jalapeño on InferenceX. 通常情况下,第一代芯片往往缺乏竞争力,但 OpenAI 打破了这一趋势,其性能处于行业领先地位,在我们测试过的多个顶级开源模型中,它击败了所有我们能接触到的 Nvidia、AMD 和 Google 芯片。OpenAI 通过极致的软硬件协同设计实现了这一点。令人惊讶的是,OpenAI 并没有过度针对模型推理的某个特定部分进行优化,而是专注于打造一款在所有场景下都能提供高性能的通用芯片。在本文中,我们将深入探讨 Jalapeño 的架构细节、软件细节以及在 InferenceX 上的性能表现。

A generalized inference chip

通用推理芯片

Everyone says that OpenAI’s chip is specialized for OpenAI models, but that’s wrong; OpenAI made a generalized chip for AI inference. 每个人都说 OpenAI 的芯片是专门为 OpenAI 模型设计的,但这是错误的;OpenAI 制造的是一款用于 AI 推理的通用芯片。

The timelines are insane. It shows that claims that use of AI is being used to accelerate chip design are real. Regardless of the quick timelines, OpenAI spent a bunch of money, made pragmatic design decisions, and their team is cracked, so this comes as no surprise. 时间线非常疯狂。这表明“利用 AI 加速芯片设计”的说法是真实的。尽管时间紧迫,但 OpenAI 投入了大量资金,做出了务实的设计决策,且其团队实力超群,因此取得这样的成果并不令人意外。

Just looking at the specs, it is an immediate contender. And the use of HBM4 makes it stand out as comparable to flagship GPUs from NVIDIA and AMD. 仅从规格来看,它立刻就成为了强有力的竞争者。而 HBM4 的使用使其脱颖而出,足以与 NVIDIA 和 AMD 的旗舰 GPU 相媲美。

A lot of the media coverage of this chip has followed a few throwaway comments from OpenAI that claim the chip will be optimized for their models in a way that other chips are not. This is wrong. Jalapeño is a generalized inference chip capable of running all sorts of models, and all sorts of workloads, including our benchmark InferenceX, where we ran the benchmark with OpenAI engineers in the lab. As a joke, OpenAI even showed us it running Doom, which was ported to their chip with just Codex prompts. 许多媒体对这款芯片的报道都引用了 OpenAI 的几句随口评论,声称该芯片将以其他芯片无法比拟的方式针对其模型进行优化。这是错误的。Jalapeño 是一款通用的推理芯片,能够运行各种模型和各种工作负载,包括我们的 InferenceX 基准测试——我们是在 OpenAI 实验室与他们的工程师一起运行该测试的。作为玩笑,OpenAI 甚至向我们展示了它运行《毁灭战士》(Doom)的过程,该游戏仅通过 Codex 提示词就被移植到了他们的芯片上。

The following is our headline perf/W result, looking at token throughput per All-in utility MW. Jalapeño smokes every other chip. All this is done without Multi Token Prediction (MTP), while the other chips on the chart are the best performing configs of each respective SKU, all with MTP. 以下是我们主要的每瓦性能(perf/W)结果,考察的是每兆瓦(MW)全效用下的 Token 吞吐量。Jalapeño 完胜其他所有芯片。这一切都是在没有使用多 Token 预测(MTP)的情况下完成的,而图表中的其他芯片均为各自 SKU 的最佳性能配置,且全部启用了 MTP。

Jalapeño beats Blackwell on perf/W across almost all scenarios without being tuned for any specific point in the curve. It excels not only in low-latency scenarios but also in high-throughput scenarios. A more apples-to-apples comparison is against Single Token Prediction results; it knocks every competitor out of the water. At low concurrency scenarios, Jalapeño demonstrates remarkable interactivity, hitting over 700 tokens per sec per user at concurrency 1 on the DeepSeek R1 model. Jalapeño 在几乎所有场景下的每瓦性能都击败了 Blackwell,且无需针对曲线上的任何特定点进行调优。它不仅在低延迟场景中表现出色,在高吞吐量场景中也同样优秀。如果进行更公平的对比(即对比单 Token 预测结果),它将所有竞争对手远远甩在身后。在低并发场景下,Jalapeño 展示了卓越的交互性,在 DeepSeek R1 模型上,并发数为 1 时,每用户每秒可处理超过 700 个 Token。

Incredibly, this is all achieved with single-token prediction (STP), no speculative decoding, and no prefill-decode disaggregation. In addition to DeepSeek R1, we also got to see some other models, including Kimi-K2.5 and GPT-OSS which ran at approximately 1,400 tok/sec/user. For all models, we confirmed that Jalapeño’s GSM8k evals attained results on par with Nvidia chips. 令人难以置信的是,这一切都是通过单 Token 预测(STP)实现的,没有使用推测解码,也没有进行预填充与解码的分离。除了 DeepSeek R1,我们还看到了其他一些模型,包括 Kimi-K2.5 和 GPT-OSS,它们的运行速度约为每用户每秒 1,400 个 Token。对于所有模型,我们确认 Jalapeño 的 GSM8k 评估结果与 Nvidia 芯片持平。

Some caveats on this. First, all numbers are provided to us by OpenAI. We verified the InferenceX runs in person in the lab, but we did not run the full suite of InferenceX benchmarks nor have we seen AgentX results. AgentX is our preferred suite for comparing chip performance due to the datasets’ long context and multi-turn characteristics that reflect the cache behavior of realistic production workflows. 关于这一点需要说明几点。首先,所有数据均由 OpenAI 提供。我们在实验室亲自验证了 InferenceX 的运行情况,但我们没有运行完整的 InferenceX 基准测试套件,也没有看到 AgentX 的结果。AgentX 是我们首选的芯片性能对比套件,因为其数据集具有长上下文和多轮对话特征,能够反映真实生产工作流中的缓存行为。

Second, we believe that comparison to Blackwell is somewhat incomplete and unfair. Jalapeño is really competing against chips like Rubin that also use HBM4. Vera Rubin systems are starting to ship to customers right now, while it will still be some time before OpenAI has anything beyond engineering samples of Jalapeño. 其次,我们认为与 Blackwell 的对比在某种程度上是不完整且不公平的。Jalapeño 真正的竞争对手是像 Rubin 这样同样使用 HBM4 的芯片。Vera Rubin 系统目前已开始向客户发货,而 OpenAI 的 Jalapeño 除了工程样品外,距离量产还需要一段时间。

Third, the models being tested are not on the open frontier. NVIDIA and AMD have published results on larger models such as DeepSeek V4 Pro and Kimi K3, using AgentX. The larger the model and the more recent the release, the more complicated it is to bring up on a new chip. With that said, the models OpenAI has working on Jalapeño aren’t exactly small either. 第三,所测试的模型并非处于开源前沿。NVIDIA 和 AMD 已经使用 AgentX 发布了更大模型(如 DeepSeek V4 Pro 和 Kimi K3)的测试结果。模型越大、发布时间越新,在芯片上的适配就越复杂。话虽如此,OpenAI 在 Jalapeño 上运行的模型也绝非小模型。

Performance Analysis

性能分析

OpenAI designs for perf/W. The reason is simple: OpenAI is currently limited by datacenter power, not by budget or floorspace, and thus tokens per MW is paramount. OpenAI 的设计目标是每瓦性能(perf/W)。原因很简单:OpenAI 目前受限于数据中心的电力供应,而非预算或占地面积,因此每兆瓦的 Token 处理量至关重要。