OpenTPU – An open-source AI accelerator, developed by AI

OpenTPU – An open-source AI accelerator, developed by AI

OpenTPU – 一个由 AI 开发的开源 AI 加速器

openTPU An open-source AI accelerator, developed by AI. openTPU brings the lessons of auto-arch-tournament to AI accelerators. It asks two questions: how far can AI agents go at hardware design, and can they build the chip that runs their own inference? otpu-chat running LFM2.5-230M on the FPGA card (left), with otpu-smi showing the card’s utilization and DRAM bandwidth (right).

openTPU 是一个由 AI 开发的开源 AI 加速器。openTPU 将“自动架构竞赛”(auto-arch-tournament)的经验引入了 AI 加速器领域。它提出了两个问题:AI 智能体在硬件设计方面能走多远?它们能否构建出运行自身推理任务的芯片?图示展示了在 FPGA 卡上运行 LFM2.5-230M 的 otpu-chat(左),以及显示该卡利用率和 DRAM 带宽的 otpu-smi(右)。

A place to learn openTPU is also a learning project. The whole accelerator lives in one small monorepo that you can read end to end: the hardware design (SystemVerilog), the instruction set, a bit-exact simulator, a kernel language and its compiler, and the host software that drives a real PCIe card. If you want to understand how an AI accelerator works, from a matmul in Python down to the wires, this is a good place to start.

openTPU 不仅是一个项目,也是一个学习平台。整个加速器包含在一个小型代码仓库中,你可以从头到尾阅读:包括硬件设计(SystemVerilog)、指令集、位精确模拟器、内核语言及其编译器,以及驱动真实 PCIe 卡的主机软件。如果你想了解 AI 加速器是如何工作的——从 Python 中的矩阵乘法一直到最底层的电路——这是一个很好的起点。

Results

结果

The design runs ten modern models with their real weights on an Inspur YPCB-00338 card (Xilinx Kintex-7 xc7k480t, two DDR3 channels), and the card produces the same tokens as the simulator, bit for bit.

该设计在浪潮 YPCB-00338 卡(Xilinx Kintex-7 xc7k480t,双 DDR3 通道)上运行了十个现代模型及其真实权重,且该卡生成的 Token 与模拟器完全一致,实现了位级对齐。

(Note: Due to the extensive nature of the performance table provided in the source, the following summary captures the key technical findings.)

(注:由于原文包含大量性能数据表格,以下总结了关键的技术发现。)

Measured on the card: the first three models on 2026-09-29 with the production image deploy_champ_e698dcd7. LFM2-2.6B, SmolLM3-3B and Phi-4-mini on 2026-09-30, and Qwen3.5-2B and 4B and Gemma 4 on 2026-10-01, with build B, deploy_fused133c_79c5707a, production since then. Build B decodes LFM2-2.6B, SmolLM3 and Phi-4-mini 8-9% faster than e698dcd7 (Gemma 4 E2B 10%), at 91-94% of the DRAM peak instead of 82-87%.

在卡上进行的测量显示:前三个模型于 2026 年 9 月 29 日使用生产镜像 deploy_champ_e698dcd7 进行测试。LFM2-2.6B、SmolLM3-3B 和 Phi-4-mini 于 9 月 30 日测试,Qwen3.5-2B/4B 和 Gemma 4 于 10 月 1 日使用 Build B (deploy_fused133c_79c5707a) 进行测试,该版本自此成为生产版本。Build B 在解码 LFM2-2.6B、SmolLM3 和 Phi-4-mini 时比 e698dcd7 快 8-9%(Gemma 4 E2B 快 10%),DRAM 带宽利用率从 82-87% 提升至 91-94%。

The image: main e698dcd at 133.33 MHz, one bitstream for all models. It has LiteDRAM controllers calibrated by a small CPU inside the memory core, a four-column systolic matrix unit and the stream engine. DDR3-1066, with a 17.1 GB/s peak.

镜像说明:主镜像 e698dcd 运行在 133.33 MHz,所有模型共用一个比特流。它配备了由内存核心内的小型 CPU 校准的 LiteDRAM 控制器、四列脉动矩阵单元和流引擎。DDR3-1066 内存,峰值带宽为 17.1 GB/s。

Mixture-of-experts models bigger than the card’s 4 GiB run with their experts streamed from host storage. The card routes each token and computes every expert, and it keeps the experts in per-layer slots in its DRAM. The host only copies missing experts from a pool file into those slots, at the link’s rate.

对于超过卡上 4 GiB 容量的混合专家模型(MoE),其专家参数会从主机存储中流式传输。该卡负责路由每个 Token 并计算每个专家,同时将专家参数保留在 DRAM 的分层槽位中。主机仅需以链路速度将缺失的专家参数从池文件中复制到这些槽位中。