Cerebras CS-4

Cerebras CS-4

The Fastest AI Just Got Faster. Introducing the all new Cerebras CS-4, a revolutionary rack-scale solution that delivers up to 30x faster inference compared to GPUs, enhanced economics, and a simple path to deploy hyperscale capacity. It is the architecture for frontier AI.

最快的人工智能变得更快了。 隆重推出全新的 Cerebras CS-4,这是一款革命性的机架级解决方案,与 GPU 相比,其推理速度最高可提升 30 倍,具备更优的经济性,并为部署超大规模算力提供了简便途径。它是面向前沿人工智能的架构。

Three WSE-3 Turbos per System Each wafer delivers up to 2x the speed of the previous generation. Faster Wafer I/O Scales massive models and enables heterogeneous, disaggregated inference. Nexus Rack-Scale Platform Enables rapid deployment in hyperscale datacenters.

每套系统配备三个 WSE-3 Turbo 每片晶圆的运行速度最高可达上一代的 2 倍。 更快的晶圆 I/O 支持大规模模型扩展,并实现异构、解耦的推理。 Nexus 机架级平台 支持在超大规模数据中心进行快速部署。

Up to 30x faster than GPUs Powered by WSE-3 Turbo, CS-4 delivers up to 30x faster inference compared to GPU systems, setting a new record for the fastest inference available in production.

比 GPU 快 30 倍 得益于 WSE-3 Turbo 的驱动,CS-4 的推理速度比 GPU 系统快 30 倍,创下了生产环境中推理速度的最快纪录。

Higher ultrafast throughput The CS-4 solution shifts the inference Pareto frontier, delivering up to 10x more throughput per watt than CS-3 while generating tokens up to 30x faster than production GPU systems. The result is a system designed to deliver both throughput and interactivity.

更高的超快吞吐量 CS-4 解决方案改变了推理的帕累托前沿(Pareto frontier),其每瓦吞吐量比 CS-3 高出 10 倍,同时生成 Token 的速度比生产级 GPU 系统快 30 倍。其成果是一个旨在同时提供高吞吐量和交互性的系统。

Frontier-ready architecture By reducing wafer-to-wafer interconnect latency to 2 microseconds, CS-4 delivers more than 1,000 tokens per second on models exceeding 10 trillion parameters, preserving interactive decode performance at unprecedented scale.

面向前沿的架构 通过将晶圆间的互连延迟降低至 2 微秒,CS-4 在超过 10 万亿参数的模型上每秒可生成超过 1,000 个 Token,从而在史无前例的规模下保持了交互式解码性能。

BUILT FOR HYPERSCALE CS-4 is the first iteration of the new Cerebras Nexus Platform Architecture. It is built around a modular concept with three foundational elements: Compute, Power, and I/O – each with significant innovation to simplify manufacturing, deployment, maintenance, and upgrades.

为超大规模而生 CS-4 是全新 Cerebras Nexus 平台架构的首次迭代。它围绕模块化概念构建,包含三个基础要素:计算、电源和 I/O。每一项都进行了重大创新,以简化制造、部署、维护和升级流程。

Modular compute backpack design Cerebras has fundamentally re-imagined the server. Each Wafer-Scale Backpack is a self-contained assembly that folds the wafer, power conversion, direct liquid cooling, high-speed I/O, and control electronics into a compact 3D package with 50% fewer components. This design simplifies manufacturing and reduces deployment time from days to hours.

模块化计算背包设计 Cerebras 从根本上重新构想了服务器。每个“晶圆级背包”(Wafer-Scale Backpack)都是一个自包含的组件,将晶圆、电源转换、直接液体冷却、高速 I/O 和控制电子设备集成到一个紧凑的 3D 封装中,组件数量减少了 50%。这种设计简化了制造流程,并将部署时间从几天缩短至几小时。

High-density power delivery With power delivery just 0.5 millimeters away from the processor - roughly 100x closer than the roughly 50mm of conventional GPU boards - CS-4 nearly eliminates board-level power loss. This enables the delivery of twice as much power to the WSE-3T, enabling higher operating frequencies and faster token generation.

高密度供电 由于供电距离处理器仅 0.5 毫米(比传统 GPU 板约 50 毫米的距离近了约 100 倍),CS-4 几乎消除了板级功率损耗。这使得 WSE-3T 能够获得两倍的电力供应,从而实现更高的运行频率和更快的 Token 生成速度。

Next-gen wafer I/O interface CS-4 introduces a new programmable I/O subsystem that doubles I/O bandwidth and reduces latency, benefitting both aggregated and disaggregated solutions. The Wafer I/O Module also enables wafers to be linked within and across racks without a switch, for wafer-to-wafer latency as low as two microseconds that is key to interactivity for models with tens of trillions of parameters.

下一代晶圆 I/O 接口 CS-4 引入了全新的可编程 I/O 子系统,该系统将 I/O 带宽翻倍并降低了延迟,使聚合和解耦解决方案均能受益。晶圆 I/O 模块还支持在机架内及机架间无需交换机即可连接晶圆,实现低至 2 微秒的晶圆间延迟,这对于拥有数十万亿参数模型的交互性至关重要。

Deploy infrastructure then compute CS-4 separates the stable power, cooling, and network layer from its modular wafer-scale compute. The Cerebras PowerRack can be installed and facility-qualified before compute arrives. Compute backpacks then slide into place and connect to power, cooling, and data—reducing deployment from days to hours while simplifying service and future upgrades at hyperscale.

先部署基础设施,再部署计算 CS-4 将稳定的电源、冷却和网络层与其模块化晶圆级计算分离开来。Cerebras PowerRack 可以在计算组件到达之前进行安装和设施验证。随后,计算背包只需滑入到位并连接电源、冷却和数据,从而将部署时间从几天缩短至几小时,同时简化了超大规模环境下的服务和未来升级。

CS-4 by the numbers First CS-4 shipments begin this quarter. Bring the fastest AI to your data center.

CS-4 数据概览 首批 CS-4 将于本季度开始发货。 将最快的人工智能引入您的数据中心。