Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index

Grok 4.6 在 Artificial Analysis 智能指数中获得 61 分

SpaceXAI’s Grok 4.6 scores 61 on the Artificial Analysis Intelligence Index, joining the frontier in line with GPT-5.6 Sol, with standout agentic performance at lower cost. Grok 4.6 gains 5 points over Grok 4.5 on the Intelligence Index just over one month after its release, or +23 points compared to Grok 4.3. This brings SpaceXAI back to the intelligence frontier alongside OpenAI, behind only Anthropic. SpaceXAI 的 Grok 4.6 在 Artificial Analysis 智能指数中获得 61 分,与 GPT-5.6 Sol 并列进入前沿梯队,并以更低的成本展现出卓越的智能体(Agentic)性能。Grok 4.6 在发布仅一个多月后,其智能指数较 Grok 4.5 提升了 5 分,较 Grok 4.3 提升了 23 分。这使得 SpaceXAI 重回与 OpenAI 并肩的智能前沿,仅次于 Anthropic。

Key takeaways: 核心要点:

➤ Grok 4.6 joins the frontier of the Artificial Analysis Intelligence Index: It scores 61, in line with GPT-5.6 Sol (max), behind Claude Opus 5 (max, 63) and Claude Fable 5 (max with fallback, 62), and just ahead of Kimi K3. ➤ Grok 4.6 进入 Artificial Analysis 智能指数前沿:其得分为 61 分,与 GPT-5.6 Sol (max) 持平,落后于 Claude Opus 5 (max, 63) 和 Claude Fable 5 (max with fallback, 62),略高于 Kimi K3。

➤ Strong agentic performance: Grok 4.6 achieves a GDPval-AA v2 Elo of 1753, behind only Claude Opus 5 and with overlapping confidence intervals with Claude Fable 5 and Qwen3.8 Max. It scores 50.7% on 𝜏³-Banking, among the top two scores alongside Qwen3.8 Max (51.3%), and 88.4% on Terminal-Bench v2.1, in line with the leading models. ➤ 强大的智能体性能:Grok 4.6 在 GDPval-AA v2 评测中获得 1753 的 Elo 分数,仅次于 Claude Opus 5,且与 Claude Fable 5 和 Qwen3.8 Max 的置信区间重叠。它在 𝜏³-Banking 评测中得分 50.7%,与 Qwen3.8 Max (51.3%) 并列前二;在 Terminal-Bench v2.1 上得分 88.4%,与领先模型持平。

➤ Frontier-level intelligence at lower cost: Headline pricing is unchanged from Grok 4.5 at $2/$6 per 1M input/output tokens, 60%+ below Claude Opus 5 ($5/$25) and GPT-5.6 Sol ($5/$30). It cost $0.84 per task, the same as Kimi K3 with slightly higher intelligence, placing it on the Intelligence vs. Cost per Task Pareto frontier. ➤ 更低成本的前沿级智能:标价与 Grok 4.5 保持一致,为每百万输入/输出 token 2 美元/6 美元,比 Claude Opus 5 ($5/$25) 和 GPT-5.6 Sol ($5/$30) 低 60% 以上。其单任务成本为 0.84 美元,与 Kimi K3 相同但智能水平略高,使其处于“智能水平 vs. 单任务成本”的帕累托前沿。

➤ Grok 4.6 sits at Fable 5-tier on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577 - behind the Claude Opus 5 family. It is notably turn-efficient, completing tasks in ~53 turns and ~0.5B input tokens on average vs. ~103 turns and ~2.0B input tokens for Claude Opus 5 (max). ➤ 在我们针对长周期智能体知识工作任务的私有基准测试 AA-Briefcase 中,Grok 4.6 处于 Fable 5 梯队,Elo 得分为 1577,仅次于 Claude Opus 5 系列。它在轮次效率上表现突出,平均以约 53 轮对话和 0.5B 输入 token 完成任务,而 Claude Opus 5 (max) 则需要约 103 轮和 2.0B 输入 token。

Other model details: 其他模型细节:

➤ Context window of 500k tokens (unchanged from Grok 4.5). ➤ 上下文窗口为 500k token(与 Grok 4.5 保持不变)。

➤ Pricing of $2/$6 per 1M tokens of input/output; cache hits discounted to $0.5 per 1M tokens, an increase over Grok 4.5’s $0.3 per 1M tokens for cache hits. ➤ 定价为每百万输入/输出 token 2 美元/6 美元;缓存命中折扣价为每百万 token 0.5 美元,较 Grok 4.5 的 0.3 美元有所上调。

Agentic performance 智能体性能

Grok 4.6’s strongest results are on agentic work rather than static reasoning. On GDPval-AA v2, our leading measure of real-world agentic knowledge work, it scores an Elo of 1753 - behind only Claude Opus 5, and statistically indistinguishable from Claude Fable 5 and Qwen3.8 Max given overlapping confidence intervals. The pattern holds across task types. 𝜏³-Banking (50.7%) tests multi-turn customer service with tool use and places Grok 4.6 in the top two, while Terminal-Bench v2.1 (88.4%) puts it level with the leaders on terminal-based software tasks. Few models are simultaneously competitive across knowledge work, customer service and terminal use; combined with its pricing, this places Grok 4.6 on the cost vs. performance Pareto frontier for every agentic evaluation in the Intelligence Index. Grok 4.6 最强的表现体现在智能体工作而非静态推理上。在衡量真实世界智能体知识工作的领先指标 GDPval-AA v2 中,它获得了 1753 的 Elo 分数,仅次于 Claude Opus 5,且在置信区间重叠的情况下,与 Claude Fable 5 和 Qwen3.8 Max 在统计学上没有显著差异。这种模式在各类任务中均成立。𝜏³-Banking (50.7%) 测试了结合工具使用的多轮客户服务,使 Grok 4.6 位列前二;而 Terminal-Bench v2.1 (88.4%) 则使其在基于终端的软件任务中与领先者持平。很少有模型能同时在知识工作、客户服务和终端使用方面保持竞争力;结合其定价,这使得 Grok 4.6 在智能指数的每一项智能体评估中都处于“成本 vs. 性能”的帕累托前沿。

Cost 成本

Holding headline pricing flat across a generation is unusual at the frontier, where intelligence gains have typically been accompanied by price increases. Grok 4.6 delivers a 5-point Intelligence Index gain at unchanged $2/$6 pricing, and our measured cost per task of $0.84 reflects both that pricing and reasonable token efficiency. The comparison that matters for buyers is against the models scoring within two points of it: Claude Opus 5 at $5/$25 and GPT-5.6 Sol at $5/$30. Grok 4.6 offers effectively the same Intelligence Index score as GPT-5.6 Sol at a fraction of the output token price, which is the dimension that dominates cost in reasoning-heavy workloads. 在模型迭代中保持标价不变在行业前沿并不常见,因为智能水平的提升通常伴随着价格上涨。Grok 4.6 在定价维持 2 美元/6 美元不变的情况下,智能指数提升了 5 分,我们测算的 0.84 美元单任务成本既反映了这一定价,也体现了合理的 token 效率。对于买家而言,最有意义的对比是与得分在 2 分以内的模型进行比较:Claude Opus 5 ($5/$25) 和 GPT-5.6 Sol ($5/$30)。Grok 4.6 以极低的输出 token 价格提供了与 GPT-5.6 Sol 基本相同的智能指数得分,而输出 token 价格正是决定重推理工作负载成本的关键维度。

Long-horizon knowledge work 长周期知识工作

Grok 4.6 debuts on AA-Briefcase, our private benchmark of long-horizon agentic knowledge work tasks, with an Elo of 1577. This places it at Fable 5-tier, behind the Claude Opus 5 family, with consistently strong performance across rubric grading, presentation quality and analytical quality rather than strength in one dimension offsetting weakness in another. The efficiency profile is as notable as the score. Grok 4.6 resolves tasks in ~53 turns and ~0.5B input tokens on average, against ~103 turns and ~2.0B input tokens for Claude Opus 5 (max). Long-horizon agentic work accumulates context rapidly, so a model that reaches a comparable answer in half the turns and a quarter of the input tokens has a cost advantage well beyond its per-token pricing. Grok 4.6 在我们针对长周期智能体知识工作任务的私有基准测试 AA-Briefcase 中首次亮相,Elo 得分为 1577。这使其处于 Fable 5 梯队,仅次于 Claude Opus 5 系列,并在评分标准、演示质量和分析质量方面表现出持续的强劲性能,而非在某一维度上的强项掩盖了另一维度的弱项。其效率表现与得分同样引人注目。Grok 4.6 平均以约 53 轮对话和 0.5B 输入 token 完成任务,而 Claude Opus 5 (max) 则需要约 103 轮和 2.0B 输入 token。长周期智能体工作会迅速积累上下文,因此,一个能以一半轮次和四分之一输入 token 达到同等答案质量的模型,其成本优势远超其每 token 定价本身。