Brood War Bench

Brood War Bench (星际争霸:母巢之战基准测试)

Key takeaways (核心要点)

  • None of the models played beyond a beginner level. 没有模型能达到初学者以上的水平。
  • Codex Astra is the clear leader beating all other models consistently. Codex Astra 是明显的领先者,持续击败所有其他模型。
  • Grok models are not smart enough to play Brood War yet. Grok 模型目前还不够聪明,无法游玩《母巢之战》。
  • Older models tended to play the RTS as a turn-based game, leading them to get destroyed while they were thinking. Newer models sometimes fell into the same trap, which may explain why some lower-effort settings performed better, but overall were much more cognizant of the cost of thinking. 旧模型倾向于将即时战略游戏(RTS)当作回合制游戏来玩,导致它们在思考时被摧毁。新模型有时也会陷入同样的陷阱,这或许能解释为什么一些低算力设置表现更好,但总体而言,它们对“思考成本”的认知要深刻得多。

Leaderboard (排行榜)

(Table omitted for brevity, showing the top rank) 🥇 Codex Astra / xhigh: 18 Wins, 0 Losses, 100.0% Win rate. 🥇 Codex Astra / xhigh: 18 胜,0 负,胜率 100.0%。


Brood War Bench started after I built a version of Brood War that you could only play through agents as an experiment to play with friends. I played it with a couple friends who did surprisingly well for people who have only played a couple Starcraft games in their lives. When I asked them why, they said they hadn’t done much, they asked their agent to attack and it had built a small army and done the full attack for them. This lead me to wonder how far they can go on their own; this is my answer.

“星际争霸:母巢之战基准测试”的诞生,源于我为了和朋友娱乐而开发的一个只能通过智能体(Agent)操作的《母巢之战》版本。我和几位朋友玩了几局,对于那些一生中只玩过几次《星际争霸》的人来说,他们的表现出奇地好。当我问他们原因时,他们说自己其实没做什么,只是让智能体去进攻,而智能体就自动造了一支小部队并完成了全线进攻。这让我好奇它们独立作战能达到什么程度;这就是我的答案。


What I observed (我的观察)

01 Codex found cheese before it found macro (Codex 先学会了“猥琐流”,后学会了运营)

Codex’s strongest recurring idea was disruption. In Protoss games it often sent a Probe across the map to attack workers or buildings. This worked shockingly well as the opposing agents often spent dozens of seconds thinking about what to do about a probe instead of doing anything else. The same systems were much weaker at sustained production. They delayed tech, trickled one or two basic units into defended bases, and threw workers into last stands.

Codex 最强且反复出现的策略是干扰。在神族对局中,它经常派出一个探测机(Probe)跨越地图去攻击敌方的农民或建筑。这招效果惊人,因为对手智能体往往会花几十秒时间思考如何应对探测机,而不是做其他事。同样的系统在持续生产方面则弱得多。它们会推迟科技升级,零星地将一两个基础单位送入敌方防守严密的基地,并让农民进行最后的抵抗。

I also noticed Codex often created separate subagents to manage the economy, army production, and army control. They didn’t communicate much with one another, so the army agent often sent each new unit straight into an attack, unaware of the larger army the other agents were planning to build. This is a common beginner mistake: sending units in one at a time instead of waiting for a critical mass and a planned attack timing.

我还注意到,Codex 经常创建独立的子智能体来分别管理经济、部队生产和部队控制。它们之间沟通很少,因此部队控制智能体经常将每个新单位直接送去进攻,而不知道其他智能体正在计划建立更大的部队。这是初学者常犯的错误:一次只送一个单位,而不是等待达到临界数量并规划进攻时机。

02 Grok spent the game between actions (Grok 把时间都花在了行动间隙)

Grok 4.6 frequently produced long stretches of reasoning and very few command batches. In G043, the xhigh run logged 11,138 reasoning tokens but issued only six command batches across 43 minutes and never fielded a combat unit. The actions it did take rarely developed into a working control loop.

Grok 4.6 经常产生长时间的推理,却极少下达指令。在 G043 局中,xhigh 设置记录了 11,138 个推理 Token,但在 43 分钟内仅下达了 6 批指令,且从未部署过战斗单位。它所采取的行动也很少能形成有效的控制循环。

03 Fable earnestly tried to play the game (Fable 认真地尝试玩游戏)

I found myself rooting for Claude Fable in more than a few games. Fable usually tried to build an economy and climb the tech tree instead of stopping at the first unit available. It seemed more interested in actually playing the game than any of the other models.

在不少对局中,我发现自己都在为 Claude Fable 加油。Fable 通常会尝试建立经济并攀升科技树,而不是停留在最初可用的单位上。相比其他模型,它似乎对“玩游戏”本身更感兴趣。


No agent here played beyond beginner level (没有智能体达到初学者以上的水平)

Even Astra and Fable were unable to build complex army’s, defend simple attacks or play concrete strategies. A beginner playing photon rush would win every single one of these games. That said, watching the agents play made me more excited than I have been in a while. This benchmark is nowhere near exhausted. There is much more for the agents to learn, and much more for the benchmark to ask them to do. I look forward to watching them get there.

即使是 Astra 和 Fable,也无法组建复杂的部队、防御简单的进攻或执行具体的战术。一个会玩“光子炮台快攻(Photon Rush)”的初学者能赢下这些对局中的每一场。话虽如此,观看这些智能体对战让我感到久违的兴奋。这个基准测试远未达到极限。智能体还有很多东西要学,基准测试也有更多可以要求它们完成的任务。我期待着看到它们不断进步。