Grok 4.7

Grok 4.7

Introducing Grok 4.7 隆重推出 Grok 4.7

SpaceXAI’s most powerful model for coding and knowledge work. Twice as fast, at half the price of comparable models. 这是 SpaceXAI 在编程和知识工作领域最强大的模型。其速度提升至两倍,价格仅为同类模型的一半。

Grok 4.7 is our most capable model for coding and knowledge work. It works longer on difficult tasks, checks its own work more carefully, and comes with our best-calibrated safeguards to date. Served at the same price and speed as Grok 4.6, it is highly competitive in its class. Grok 4.7 是我们目前在编程和知识工作方面能力最强的模型。它能更持久地处理复杂任务,更仔细地核查自身工作,并配备了我们迄今为止校准效果最好的安全防护机制。在保持与 Grok 4.6 相同价格和速度的同时,它在同类产品中极具竞争力。

On CursorBench 4.0, which stresses longer-running coding tasks, Grok 4.7 is at the frontier in price-performance. 在侧重于长时编程任务的 CursorBench 4.0 测试中,Grok 4.7 在性价比方面处于行业前沿。

Model Improvements 模型改进

Grok 4.7 uses a new, larger base model compared to Grok 4.6. It was trained with a longer reinforcement learning run on a harder mix of tasks, weighted toward problems that take many hours to complete. The model is better at verifying its own work and managing longer context. We also trained Grok 4.7 to natively understand the Grok Bot harness, making it better at conversational tasks and general knowledge work. 与 Grok 4.6 相比,Grok 4.7 使用了全新的、规模更大的基础模型。它通过更长时间的强化学习训练,处理了更具挑战性的任务组合,并重点优化了那些需要耗时数小时才能完成的问题。该模型在自我验证工作和管理长上下文方面表现更佳。我们还对 Grok 4.7 进行了训练,使其能够原生理解 Grok Bot 框架,从而在对话任务和通用知识工作方面表现更出色。

(Table data omitted for brevity, but context provided below) (表格数据略,以下为相关说明)

Token prices and benchmark scores for Grok 4.7, Grok 4.6, GPT-5.6 Sol, and Fable 5.1. Benchmarks are CursorBench 4.0, DeepSWE v1.1, EEBench, AA Briefcase v1.1, Terminal-Bench 4.0, Harvey Legal Agent Benchmark, and HealthBench Professional. An asterisk on Grok 4.7 DeepSWE marks a high-effort score. 以上为 Grok 4.7、Grok 4.6、GPT-5.6 Sol 和 Fable 5.1 的 Token 价格及基准测试分数。基准测试包括 CursorBench 4.0、DeepSWE v1.1、EEBench、AA Briefcase v1.1、Terminal-Bench 4.0、Harvey Legal Agent Benchmark 以及 HealthBench Professional。Grok 4.7 DeepSWE 上的星号表示高强度测试分数。

Grok 4.7 is better at creating documents and presentations. In GDPval and AA Briefcase, AI is asked to work on tasks done by professionals such as lawyers, nurses, and financial analysts. Grok 4.7 improves upon Grok 4.6 on both benchmarks and performs comparably to other frontier models. Grok 4.7 在创建文档和演示文稿方面表现更优。在 GDPval 和 AA Briefcase 测试中,AI 需要完成律师、护士和金融分析师等专业人士的工作任务。Grok 4.7 在这两项基准测试中均优于 Grok 4.6,并与其他前沿模型表现相当。

Safety & Cybersecurity 安全与网络安全

Grok 4.7 was built with an entirely new safeguard stack. It is the strongest model we’ve tested on refusals and jailbreak resistance. In dual-use domains like cybersecurity and biological work, it leads on both utility for benign tasks and safe refusal on dangerous ones, topping LatchBio’s biosafety benchmark at 62.4%. Grok 4.7 构建于一套全新的安全防护体系之上。它是我们测试过的在拒绝响应和防越狱能力方面最强的模型。在网络安全和生物研究等双用途领域,它在良性任务的实用性与危险任务的安全拒绝方面均处于领先地位,并在 LatchBio 的生物安全基准测试中以 62.4% 的成绩位居榜首。

Grok 4.7 balances strong cyber defense capabilities with low refusal rates for legitimate use. It shows the highest safety on HackerBench v0.3, our benchmark for risky and malicious cyber tasks, allowing only 3.3% of risky dual-use prompts through while rarely blocking legitimate security work. We’ve also started giving select cybersecurity partners invite-only access to Grok 4.7’s red-team capabilities for defense research. Grok 4.7 在强大的网络防御能力与合法使用的低拒绝率之间取得了平衡。在针对高风险和恶意网络任务的 HackerBench v0.3 基准测试中,它展现了最高的安全性,仅允许 3.3% 的高风险双用途提示通过,同时极少拦截合法的安全工作。我们还开始向部分网络安全合作伙伴提供 Grok 4.7 红队功能的受邀访问权限,以用于防御研究。

Pricing and availability 定价与可用性

Grok 4.7 is available today in Cursor and Grok Build. It is also available through the Grok API, third-party coding harnesses, and model routers and cloud platforms. The model is priced starting at $2 per million input tokens and $6 per million output tokens. We also serve a fast variant with twice the output speed at twice the price. Grok 4.7 即日起可在 Cursor 和 Grok Build 中使用。它也通过 Grok API、第三方编程框架、模型路由及云平台提供。该模型定价为每百万输入 Token 2 美元,每百万输出 Token 6 美元。我们还提供一个快速版本,其输出速度翻倍,价格也相应翻倍。