Best LLM for every budget, updated daily

Best LLM for every budget, updated daily

每日更新:各预算下的最佳大语言模型 (LLM)

Intelligence Coding Math Min score 30 One row per model Maker All Frontier only Label every point Performance vs price Blended cost per 1M tokens (3:1 input:output, log scale) against the Intelligence Index. Hover or tab to a point for details. 智能、编程、数学、最低分数 30、每个模型一行、制造商、全部、仅限前沿模型、标注每个点。性能与价格对比:每 100 万 token 的混合成本(输入与输出比例为 3:1,对数刻度)与智能指数的对比。悬停或点击数据点可查看详情。

On the value frontier Dominated: a cheaper model matches or beats it Value frontier Best model for your budget The frontier as a lookup table. Find the row your budget falls in; the pick is the highest-scoring model you can get at that price, and the runner-up is the next best that also fits. 价值前沿:被支配(更便宜的模型匹配或超越了它)、价值前沿、您预算内的最佳模型。将前沿视为查询表:找到您预算所在的行;“首选”是该价格下得分最高的模型,“次选”是同样符合预算的次优模型。

Budget per 1M tokens Pick Score Price Runner-up Raw capability Ignoring cost entirely. What changed Diff between consecutive daily fetches: new models, removed models, and re-scored or re-priced ones. All figures Click a column header to sort. Names link to the model’s Artificial Analysis page. 每 100 万 token 预算、首选、分数、价格、次选。原始能力:完全忽略成本。变动情况:每日抓取数据之间的差异,包括新模型、已移除模型以及重新评分或重新定价的模型。所有数据:点击列标题可进行排序。名称链接至该模型在 Artificial Analysis 上的页面。

Model Maker Intelligence Coding Math Blended $/1M Input $/1M Output $/1M Tokens/s TTFT s 模型、制造商、智能、编程、数学、混合价格(每 100 万 token)、输入价格(每 100 万 token)、输出价格(每 100 万 token)、每秒 Token 数、首字延迟 (TTFT)。

How to read this Value frontier. Sort by price ascending and keep every model that scores higher than everything cheaper. Ties on price go to the higher score; ties on score go to the cheaper model. Blended price is Artificial Analysis’s 3:1 input:output blend per 1M tokens. Cached-input discounts, batch pricing and fast modes are not included. 如何阅读此表:价值前沿。按价格升序排列,保留所有得分高于更便宜模型的模型。价格相同时取分数较高者;分数相同时取价格较低者。混合价格是 Artificial Analysis 按照 100 万 token 中 3:1 的输入/输出比例计算的。不包含缓存输入折扣、批量定价和快速模式。

“One row per model” keeps the highest-scoring effort or reasoning variant of each model name (ties go to the cheaper one). Untick it to see every variant AA benchmarks separately, such as low, medium, high, xhigh and max effort. “Min score” hides models below that index from both the chart and the frontier calculation, so an old, tiny model at a rock-bottom price does not anchor the line. “每个模型一行”保留每个模型名称中得分最高的努力程度或推理变体(平局时取价格较低者)。取消勾选可查看 AA 单独基准测试的所有变体,例如低、中、高、超高和最大努力程度。“最低分数”会从图表和前沿计算中隐藏低于该指数的模型,以防止价格极低的旧版小型模型影响基准线。

Scores move. AA re-bases the Index between versions, so compare against this page only, not against an older snapshot. Source: Artificial Analysis free data API, fetched daily by a GitHub Actions cron. Code and data: github.com/terryds/bestvaluemodel. 分数会变动。AA 会在不同版本间重置指数,因此请仅与此页面进行比较,不要与旧的快照对比。来源:Artificial Analysis 免费数据 API,通过 GitHub Actions 定时任务每日抓取。代码与数据:github.com/terryds/bestvaluemodel。