Accelerating GPT-5.6 Sol Ultrafast
Accelerating GPT-5.6 Sol Ultrafast
Aug 13, 2026 | Joyce Er
Today, Cerebras and OpenAI are sharing an early look at Ultrafast Mode, a new service tier launching first in the OpenAI API and powered by Cerebras. Ultrafast is available initially to a select group of customers, with access expanding over time. Cerebras powers GPT-5.6 Sol on Ultrafast mode, delivering up to 750 output tokens per second and without any quality compromise, allowing Sol Ultrafast to accelerate your most time-sensitive, mission-critical work.
今天,Cerebras 和 OpenAI 共同发布了“Ultrafast Mode”(超快模式)的早期预览。这是一项全新的服务层级,将首先在 OpenAI API 中推出,并由 Cerebras 提供算力支持。Ultrafast 目前仅向部分精选客户开放,并将随着时间推移逐步扩大访问范围。Cerebras 为 Ultrafast 模式下的 GPT-5.6 Sol 提供动力,在不牺牲任何质量的前提下,每秒可输出高达 750 个 token,助力 Sol Ultrafast 加速处理您最紧迫、最关键的任务。
Frontier Intelligence at Unprecedented Speed
以史无前例的速度实现前沿智能
AI builders have always needed to choose between speed and intelligence. As models scale up in size and intelligence, they incur higher computational and data movement costs, slowing down response times. Users often need to wait for high-quality results or accept inferior results within a shorter timeframe.
AI 构建者一直需要在速度和智能之间做出取舍。随着模型规模和智能水平的提升,它们会产生更高的计算和数据传输成本,从而拖慢响应速度。用户往往需要在等待高质量结果和在短时间内接受较差结果之间进行选择。
GPT-5.6 Sol Ultrafast resolves this tradeoff, bringing frontier intelligence to products and workflows where every second matters. Compared with output speeds reported by Artificial Analysis, GPT-5.6 Sol on Ultrafast mode runs 11x faster than Fable 5, and 5x faster than Opus 4.8 on Fast mode.
GPT-5.6 Sol Ultrafast 解决了这一权衡难题,将前沿智能带入到每一秒都至关重要的产品和工作流中。根据 Artificial Analysis 报告的输出速度对比,Ultrafast 模式下的 GPT-5.6 Sol 的运行速度比 Fable 5 快 11 倍,比 Fast 模式下的 Opus 4.8 快 5 倍。
At Cerebras, we put Ultrafast to the test by running it head-to-head with popular models on Humanity’s Last Exam. HLE is a challenging model benchmark that consists of 2,500 questions typically answerable only by those holding PhDs in fields such as chemistry, economics, and literature.
在 Cerebras,我们通过“人类终极考试”(Humanity’s Last Exam, HLE)将 Ultrafast 与主流模型进行了正面较量。HLE 是一项极具挑战性的模型基准测试,包含 2,500 个问题,通常只有在化学、经济学和文学等领域拥有博士学位的人才能回答。
In our evaluations, GPT-5.6 Sol on Ultrafast mode answered all 2,500 HLE questions in 11 hours and 11 minutes. Claude Fable 5 needed 78 hours and 27 minutes, more than three days of continuous compute, to arrive at the same conclusions. In other words, Ultrafast worked through the frontier of human knowledge in a single working day, achieving comparable accuracy nearly 7× faster.
在我们的评估中,Ultrafast 模式下的 GPT-5.6 Sol 在 11 小时 11 分钟内回答了全部 2,500 个 HLE 问题。而 Claude Fable 5 则需要 78 小时 27 分钟(超过三天连续计算)才能得出同样的结论。换句话说,Ultrafast 在一个工作日内就完成了对人类知识前沿的探索,以近 7 倍的速度实现了相当的准确率。
Humanity’s Last Exam Benchmark: Benchmarking was performed by Cerebras using GPT 5.6 Sol Ultrafast with Codex on xhigh reasoning on July 10 and Claude Fable 5 with Claude Code on xhigh reasoning on July 13-15.
“人类终极考试”基准测试:由 Cerebras 于 7 月 10 日使用运行于 xhigh 推理模式下的 GPT 5.6 Sol Ultrafast (Codex) 进行,并于 7 月 13-15 日使用运行于 xhigh 推理模式下的 Claude Fable 5 (Claude Code) 进行对比。
As model capabilities continue to advance, the range of applications for fast inference expands. GPT-5.6 Sol is OpenAI’s best model yet for legal briefs, financial models, and engineering reports. On GDP-Val, a benchmark for economically valuable knowledge work tasks, Ultrafast delivered a 5.6x end-to-end speedup with no quality degradation, showing how faster inference can accelerate economically valuable work.
随着模型能力的不断进步,快速推理的应用范围也在扩大。GPT-5.6 Sol 是 OpenAI 目前处理法律摘要、财务模型和工程报告表现最好的模型。在衡量具有经济价值的知识工作任务的基准测试 GDP-Val 上,Ultrafast 在不降低质量的情况下实现了 5.6 倍的端到端加速,展示了更快的推理速度如何加速具有经济价值的工作。
Benchmarking was performed by Cerebras on July 31, 2026 using GPT 5.6 Sol and GPT 5.6 Sol Ultrafast on medium reasoning within Codex.
基准测试由 Cerebras 于 2026 年 7 月 31 日使用 Codex 中等推理模式下的 GPT 5.6 Sol 和 GPT 5.6 Sol Ultrafast 进行。
High-Speed Intelligence Powers High-Stakes Work
高速智能赋能高风险工作
Faster intelligence changes what’s possible for individuals and organizations. With Ultrafast, you can now put agents on the critical path of problems where every second counts.
更快的智能改变了个人和组织的可能性。有了 Ultrafast,您现在可以将 AI 智能体置于每一秒都至关重要的关键问题路径上。
“With GPT-5.6 Sol Ultrafast, Cerebras enables AI that keeps up with how you think, code, and collaborate. We’re excited to see how workflows and applications are transformed by Ultrafast inference.” — Rohan Varma, Product at OpenAI
“借助 GPT-5.6 Sol Ultrafast,Cerebras 实现了能够跟上您思考、编码和协作节奏的 AI。我们非常期待看到 Ultrafast 推理如何改变工作流和应用程序。”—— Rohan Varma,OpenAI 产品负责人
Ultrafast is a persistent edge for organizations using frontier AI to quickly respond to incoming information. Companies operating web services can leverage Ultrafast to root-cause and address production outages, preserving customer trust, preventing lost revenue, and saving downtime minutes against their SLAs. And in adversarial, high-stakes cyberattacks, Ultrafast is an invaluable tool for security teams who must quickly detect and respond to bad actors to contain catastrophic losses.
对于使用前沿 AI 来快速响应信息输入的组织而言,Ultrafast 是一种持久的优势。运营 Web 服务的公司可以利用 Ultrafast 来根除并解决生产故障,从而维护客户信任、防止收入损失,并节省符合 SLA(服务等级协议)的停机时间。在对抗性、高风险的网络攻击中,Ultrafast 对于安全团队来说也是不可或缺的工具,他们必须快速检测并响应恶意行为者,以遏制灾难性的损失。
More broadly, Ultrafast enables entirely new modes of working with agents; it delivers real-time insights and updates, so you don’t have to context-switch across multiple parallel sessions to get the most out of your agents.
更广泛地说,Ultrafast 实现了与智能体协作的全新模式;它提供实时洞察和更新,因此您无需在多个并行会话之间切换上下文,即可充分利用您的智能体。
“Whereas formerly I might have to wait a couple minutes for a task to finish, it now finishes for me before I even have the opportunity to context-switch. It makes me way more productive.” — Jeffrey Wang, OpenAI Researcher
“以前我可能需要等待几分钟才能完成一项任务,但现在它在我还没来得及切换上下文之前就已经完成了。这让我的工作效率大大提高。”—— Jeffrey Wang,OpenAI 研究员
With Ultrafast, researchers and engineers can reserve their attention for going deep on select problems that matter most, while continuing to use Standard processing for parallelizing commodity tasks. Cerebras is excited to power the next wave of AI innovation, raising the ceiling for what individuals and organizations can accomplish with responsive AI.
有了 Ultrafast,研究人员和工程师可以将精力集中在最重要的问题上进行深入研究,同时继续使用标准处理来并行化常规任务。Cerebras 很高兴能为下一波 AI 创新提供动力,提高个人和组织利用响应式 AI 所能达到的上限。
Breakneck Speed is Enabled by Breakthrough Innovation
突破性创新实现极速体验
GPT-5.6 Sol on Ultrafast mode is powered by Cerebras’ revolutionary Wafer-Scale Engine architecture, purpose-built for frontier AI workloads. Fast frontier inference is a data movement problem: on GPUs, inference on large models is bottlenecked by memory bandwidth, as model weights must be repeatedly transferred between on-chip memory and off-chip storage to generate successive tokens within a model response.
Ultrafast 模式下的 GPT-5.6 Sol 由 Cerebras 革命性的晶圆级引擎(Wafer-Scale Engine)架构驱动,该架构专为前沿 AI 工作负载而构建。快速前沿推理本质上是一个数据移动问题:在 GPU 上,大型模型的推理受到内存带宽的限制,因为模型权重必须在片上内存和片外存储之间反复传输,才能在模型响应中生成连续的 token。
Cerebras takes a contrarian approach to eliminating this inefficient data movement: we pack 44 GB of SRAM on each wafer-sized chip. Weights stay on-chip, and tokens flow uninterrupted through model layers pipelined across wafers. This technical approach scales smoothly with model size, paving the way for a continued speed advantage on future frontier models.
Cerebras 采取了一种反直觉的方法来消除这种低效的数据移动:我们在每块晶圆大小的芯片上集成了 44 GB 的 SRAM。权重保留在芯片上,token 在跨晶圆流水线化的模型层中不间断地流动。这种技术方法可以随着模型规模的扩大而平滑扩展,为未来前沿模型持续保持速度优势铺平了道路。
Ultrafast: Now in Limited Preview
Ultrafast:现已开启有限预览
GPT-5.6 Sol on Ultrafast mode is available in a limited preview today to a select group of customers. Access will expand as capacity grows. Sign up for updates.
Ultrafast 模式下的 GPT-5.6 Sol 即日起向部分精选客户开放有限预览。随着容量的增加,访问权限将逐步扩大。欢迎注册以获取最新动态。