Grok 4.6
Grok 4.6
Introducing Grok 4.6 隆重推出 Grok 4.6
Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. Grok 4.6 在 Grok 4.5 的基础上进行了升级,特别侧重于长周期运行的智能体(agents)以及更具挑战性的交互式和视觉化任务。
Today we are releasing Grok 4.6. Grok 4.6 builds on Grok 4.5 with a particular focus on long-running agents and more ambitious interactive and visual work. It stays with complex tasks across many steps, whether researching a topic, analyzing information, working across a codebase, or turning an idea into a polished application or work artifact. 今天,我们正式发布 Grok 4.6。Grok 4.6 在 Grok 4.5 的基础上进行了升级,特别侧重于长周期运行的智能体以及更具挑战性的交互式和视觉化任务。它能够持续处理跨越多个步骤的复杂任务,无论是研究课题、分析信息、处理代码库,还是将一个想法转化为完善的应用程序或工作成果。
Grok 4.6 achieves frontier intelligence across several agentic coding and knowledge work benchmarks. It matches GPT-5.6 Sol on the Artificial Analysis Intelligence Index, which is a composite score of nine benchmarks. Grok 4.6 在多项智能体编码和知识工作基准测试中达到了前沿水平。在由九项基准测试组成的“人工智能分析指数”(Artificial Analysis Intelligence Index)中,其表现与 GPT-5.6 Sol 持平。
Grok 4.6 is available today in Cursor and Grok Build. We’re offering 2x included usage inside Grok Build and Cursor for the first week so you can start trying 4.6 immediately. Grok 4.6 即日起在 Cursor 和 Grok Build 中可用。在发布的第一周,我们为 Grok Build 和 Cursor 用户提供双倍的使用额度,以便您可以立即开始体验 4.6 版本。
Training Grok 4.6 Grok 4.6 的训练
Grok 4.6 underwent a longer supplemental training run than Grok 4.5, with curated model-generated data for reasoning and advanced technical concepts, high-quality engineering data, and an improved optimizer and training recipe. This produced a stronger foundation for the SFT and RL stages that followed. Grok 4.6 经历了比 Grok 4.5 更长的补充训练过程,使用了针对推理和高级技术概念精心筛选的模型生成数据、高质量工程数据,以及改进后的优化器和训练方案。这为随后的 SFT(监督微调)和 RL(强化学习)阶段奠定了更坚实的基础。
We then used Grok 4.5 to regenerate the SFT trajectories across reasoning efforts, agent harnesses, and domains such as STEM, software engineering, and knowledge work, and filtered out problematic traces with model-based checks. The resulting SFT checkpoint shows strong performance and improved behavior. 随后,我们利用 Grok 4.5 在推理任务、智能体框架以及 STEM、软件工程和知识工作等领域重新生成了 SFT 轨迹,并通过基于模型的检查过滤掉了有问题的记录。最终得到的 SFT 检查点展现出了强大的性能和更优的行为表现。
Grok 4.6 is trained on a wide range of agentic RL tasks, including knowledge work, general coding, and domain-specific environments for kernel optimization, web development, computer-aided design, and more. Grok 4.6 在广泛的智能体强化学习任务中进行了训练,涵盖了知识工作、通用编码,以及针对内核优化、Web 开发、计算机辅助设计等领域的特定环境。
Turning ambitious ideas into working projects 将宏伟构想转化为实际项目
We tested Grok 4.6 on projects designed to stretch its range and ability to sustain work over many steps. We found the model is especially strong at turning a broad product idea into a working first version. It can research unfamiliar domains, structure the application, implement the core interactions, and continue refining the result through several rounds of feedback. 我们通过旨在测试其极限和多步骤持续工作能力的各种项目对 Grok 4.6 进行了评估。我们发现,该模型在将宽泛的产品构想转化为可运行的初版产品方面表现尤为出色。它能够研究陌生领域、构建应用程序结构、实现核心交互,并通过多轮反馈不断优化结果。
On longer trajectories, we also started to see more self-testing and verification, with the model checking its own work before moving on. Grok 4.6 produces stronger first passes on visual and interactive projects than we typically saw with Grok 4.5. Given a concrete product idea, it is able to establish structure and visual language for an application in one pass. This has made it especially useful for projects where the fastest route to a good result was to begin with something substantial and then iterate in the loop. 在更长的任务轨迹中,我们还观察到它具备了更强的自我测试和验证能力,模型会在继续下一步之前检查自己的工作。与 Grok 4.5 相比,Grok 4.6 在视觉和交互项目上的初次产出质量更高。给定一个具体的产品构想,它能够一次性建立起应用程序的结构和视觉语言。这使得它在那些“先从实质性内容入手,再进行循环迭代”的快速开发项目中显得尤为高效。
Safety and capabilities 安全与能力
Grok 4.6’s safeguards have been improved and calibrated in line with the model’s capabilities. Our safety stack is designed to maximize utility and security across legitimate use cases, allowing Grok 4.6 to be helpful and safe in domains such as vulnerability patching, accelerating the engineering design cycle, and augmenting AI research. Grok 4.6 的安全防护机制已根据模型的能力进行了改进和校准。我们的安全架构旨在最大限度地提高合法用例的实用性和安全性,使 Grok 4.6 能够在漏洞修补、加速工程设计周期以及增强 AI 研究等领域发挥积极且安全的作用。
Our safeguard evaluation work reflects Grok 4.6’s expanded capabilities, with our widest-ever suite of pre-deployment testing for capabilities and safeguard calibration, as well as extensive post-deployment and third-party testing. 我们的安全评估工作反映了 Grok 4.6 扩展后的能力,我们进行了有史以来规模最大的部署前能力测试和安全校准,并辅以广泛的部署后测试和第三方测试。
Get started with Grok 4.6 开始使用 Grok 4.6
Grok 4.6 is available today in Cursor and Grok Build. It’s also available in the API and other partners like OpenRouter, Vercel, and Cloudflare. Pricing starts at $2 per million input tokens and $6 per million output tokens. Additionally, there is a fast variant which is twice the price. We’re offering 2x included usage inside Grok Build and Cursor for the first week so you can start trying 4.6 immediately. Grok 4.6 即日起在 Cursor 和 Grok Build 中可用。它也已通过 API 以及 OpenRouter、Vercel 和 Cloudflare 等合作伙伴提供。定价为每百万输入 Token 2 美元,每百万输出 Token 6 美元。此外,还有一个价格翻倍的快速版本。在发布的第一周,我们为 Grok Build 和 Cursor 用户提供双倍的使用额度,以便您可以立即开始体验 4.6 版本。