NanoGPT Speedrun Frontier
NanoGPT Speedrun Frontier
NanoGPT Speedrun Frontier NanoGPT 竞速前沿
Collapse all models | All models | Best validated result for each model | All trajectories | Log scale | Filter by model 折叠所有模型 | 所有模型 | 每个模型的最佳验证结果 | 所有轨迹 | 对数刻度 | 按模型筛选
The following table tracks the performance of various AI models in the “NanoGPT Speedrun” benchmark, detailing their record scores, percentage of gap closed within 24 hours, and operational metrics. 下表追踪了各种 AI 模型在“NanoGPT 竞速”基准测试中的表现,详细列出了它们的记录分数、24 小时内的差距缩小百分比以及运行指标。
| Rank | Model | Status | Record | Gap Closed @24H | Agent | Total Tokens | Days |
|---|---|---|---|---|---|---|---|
| 1 | Fable 5 | note | 2,726 | 81.7% | claude-code · high | 800M | 8.7 |
| 2 | Opus 5 | serial era | 2,920 | 53.6% | claude-code · max | 183M | 2.9 |
| 3 | Kimi K3 | serial era | 2,930 | 52.2% | prime-agent · max | 112M | 3.6 |
| 4 | Kimi K3 | note | 2,974 | 45.8% | kimi-code · max | 682M | 5.1 |
| 5 | Opus 4.8 | - | 3,018 | 39.4% | claude-code · max | 318M | 3.0 |
| 6 | GPT-5.6 Sol | note | 3,042 | 35.9% | codex · xhigh | 2.9B | 6.1 |
| 7 | GPT-5.6 Sol Pro | serial era | 3,058 | 33.6% | codex · xhigh | 1.2B | 3.4 |
| 8 | Sonnet 5 | note | 3,105 | 26.8% | claude-code · max | 998M | 2.0 |
| 9 | GPT-5.6 Luna | note | 3,110 | 26.1% | codex · xhigh | 894M | 1.9 |
| 10 | Grok 4.5 | note | 3,120 | 24.6% | grok-cli · xhigh | 46M | 2.7 |
| 11 | Qwen3.8 Max | running | 3,120 | 24.6% | qwen-code · max | 216M | 1.9 |
| 12 | GLM 5.2 | - | 3,150 | 20.3% | pi · high | 57M | 1.8 |
| 13 | DeepSeek V4 Pro | running | 3,205 | 12.3% | claude-code · max | 26M | 1.1 |
| 14 | GPT-5.6 Terra | serial era | 3,214 | 11.0% | codex · xhigh | 417M | 1.1 |
| 15 | Grok 4.6 | running | 3,220 | 10.1% | grok-cli · xhigh | 27M | 0.6 |
| 16 | Muse Spark 1.2 | running | 3,230 | 8.7% | muse-code · xhigh | 41M | 0.6 |
| 17 | Muse Spark 1.1 | - | 3,232 | 8.4% | pi · max | 122M | 3.7 |
| 18 | GPT-5.5 | serial era | 3,234 | 8.1% | codex · xhigh | 70M | 1.1 |
| 19 | Kimi K2.7 | note | 3,240 | 7.2% | kimi-code · max | 160M | 1.6 |
| 20 | GLM 5.3 | running | — | no record | claude-code · xhigh | — | — |
Compare Results: This section provides a comparative view of model records against the Human baseline (2,600) and the target baseline (3,290). 结果对比:本节提供了模型记录与人类基准(2,600)及目标基准(3,290)的对比视图。
Loading available trajectories… Open Traces to explore 41 curated full agent trajectories, including tool calls, subagents, and scratchpads. 正在加载可用轨迹……打开“追踪”(Traces)以探索 41 个精选的完整智能体轨迹,包括工具调用、子智能体和草稿板内容。