Qwen3.8 Max now ranked as the best overall model by agentic index

Qwen3.8 Max now ranked as the best overall model by agentic index

Artificial Analysis: Independent analysis of AI Understand the AI landscape to choose the best model and provider for your use case. 人工智能分析:AI 独立测评 了解 AI 行业格局,为您的使用场景选择最合适的模型与服务商。


Update: Intelligence Index v4.1.1 Intelligence Index v4.1.1 moves 𝜏³-Banking to v1.0.1 and upgrades the grader for HLE, AA-LCR, and AA-Omniscience to GPT-5.6 Luna (medium). 更新:智能指数 v4.1.1 智能指数 v4.1.1 将 𝜏³-Banking 升级至 v1.0.1 版本,并将 HLE、AA-LCR 和 AA-Omniscience 的评分器升级为 GPT-5.6 Luna (中等版本)。


Artificial Analysis Intelligence Index v4.1.1 Artificial Analysis Intelligence Index v4.1.1 incorporates 9 evaluations: GDPval-AA v2, 𝜏³-Banking, Terminal-Bench v2.1, SciCode, Humanity’s Last Exam, GPQA Diamond, CritPt, AA-Omniscience, and AA-LCR. 人工智能分析智能指数 v4.1.1 人工智能分析智能指数 v4.1.1 纳入了 9 项评估指标:GDPval-AA v2、𝜏³-Banking、Terminal-Bench v2.1、SciCode、Humanity’s Last Exam (HLE)、GPQA Diamond、CritPt、AA-Omniscience 以及 AA-LCR。


Methodology: Cost per Intelligence Index Task Weighted average cost (USD) per Artificial Analysis Intelligence Index task. Each evaluation’s cost is calculated from input, cache hit, cache write, reasoning, and answer token prices, divided by task count, and weighted by its Intelligence Index weight. 方法论:每项智能指数任务的成本 每项人工智能分析智能指数任务的加权平均成本(美元)。每项评估的成本均根据输入、缓存命中、缓存写入、推理和回答 token 的价格计算得出,并除以任务总数,再根据其在智能指数中的权重进行加权。


Intelligence Index vs. Cost per Task The most attractive quadrant represents models that offer high intelligence at a lower cost per task. Current leaders in this space include models from Alibaba (Qwen), DeepSeek, and Anthropic. 智能指数与任务成本对比 最具吸引力的象限代表了那些能以较低任务成本提供高智能水平的模型。目前在该领域处于领先地位的包括来自阿里巴巴(通义千问)、深度求索(DeepSeek)和 Anthropic 的模型。