With most information hidden, the game Stratego had stumped AI—until now
With most information hidden, the game Stratego had stumped AI—until now
由于大部分信息处于隐藏状态,经典游戏“陆战棋”(Stratego)曾让AI束手无策——直到现在。
Deep Blue took down Garry Kasparov at chess in 1997, AlphaGo beat Lee Sedol at Go in 2016, and poker bots have been beating professionals for years. But one classic game called Stratego held out. Even DeepMind, with its exceptional budget, couldn’t build a machine that reliably beat the best human players. Now, a team of researchers from Carnegie Mellon, MIT, New York University, and Stanford University has done it. Their AI, called Ataraxos, beat Pim Niemeijer, arguably the best Stratego player of all time, 15 games to one, with four draws. And it took just 16 GPUs and a few thousand dollars to train it.
1997年,“深蓝”在国际象棋中击败了加里·卡斯帕罗夫;2016年,AlphaGo在围棋中战胜了李世石;而扑克机器人多年来也一直处于领先地位。但有一款经典游戏——陆战棋(Stratego)——却一直未被攻克。即便是拥有巨额预算的DeepMind,也未能打造出一台能稳定击败顶尖人类选手的机器。如今,来自卡内基梅隆大学、麻省理工学院、纽约大学和斯坦福大学的研究团队做到了。他们的AI名为“Ataraxos”,以15胜1负4平的战绩击败了公认的史上最强陆战棋选手Pim Niemeijer。而且,训练该AI仅耗费了16块GPU和几千美元的成本。
Hidden armies
隐藏的军队
In Stratego, each player gets 40 pieces representing military ranks, from a marshal down to a spy, plus bombs and a flag. You win by capturing the opponent’s flag. Your opponent knows where your pieces are, but not what they are. Identities are revealed only when two pieces collide in battle—the weaker one is removed, and the identity of the winner is revealed. That makes Stratego an imperfect-information game, just like poker, which computers cracked years ago.
在陆战棋中,每位玩家拥有40枚棋子,代表从元帅到间谍等不同的军衔,此外还有炸弹和军旗。通过夺取对方的军旗即可获胜。对手知道你的棋子位置,但不知道具体是什么棋子。只有当两枚棋子在战斗中碰撞时,身份才会揭晓——较弱的一方被移除,获胜者的身份随之公开。这使得陆战棋成为一种非完美信息博弈游戏,就像计算机多年前就已经攻克的扑克一样。
“There’s something super distinctive about Stratego, which is that it is a massive amount of hidden information that unfolds over a very long time scale,” said Eugene Vinitsky, a researcher at NYU and co-author of the study. In some forms of poker, the hidden information is tiny. In Texas Hold’em, “You only have two hidden cards,” said Gabriele Farina, an MIT computer scientist and another co-author. That leaves just 1,326 possible hands, few enough for a machine to weigh them all. “In Stratego, there’s 40 pieces on the board that could be in any order,” Farina said. That’s more than a decillion possible setups. Then there’s the game’s length. “In chess, usually the game lasts 40 moves, but in Stratego, a game can easily last 2,000 moves,” Farina said.
“陆战棋有一个非常独特的特点,那就是它包含海量的隐藏信息,且在极长的时间尺度内展开,”纽约大学研究员、该研究的合著者Eugene Vinitsky表示。在某些扑克游戏中,隐藏信息非常少。麻省理工学院计算机科学家、另一位合著者Gabriele Farina指出,在德州扑克中,“你只有两张底牌”。这意味着只有1326种可能的底牌组合,机器完全可以权衡所有情况。“但在陆战棋中,棋盘上有40枚棋子,它们可以以任何顺序排列,”Farina说。这导致了超过10的33次方(decillion)种可能的初始布局。此外还有游戏的长度问题,“在国际象棋中,一局通常持续40步,但在陆战棋中,一局游戏很容易持续2000步,”Farina补充道。
On top of that, Stratego is a game of bluffing. Sometimes you move a weak piece as if it were a marshal, just to scare the opponent off. When players bluff too often, their threats mean nothing; when they never bluff, they become predictable. That balancing act, the team explains, is what stumped earlier AIs like DeepMind’s DeepNash, introduced in 2022.
除此之外,陆战棋还是一场关于虚张声势的游戏。有时你会移动一枚弱棋子,假装它是元帅,以此吓退对手。如果玩家虚张声势过于频繁,威胁就会失去意义;如果从不虚张声势,又会变得容易预测。研究团队解释说,这种平衡的艺术正是此前DeepMind在2022年推出的DeepNash等AI所面临的难题。
Learning to guess
学会猜测
Just like DeepNash, Ataraxos learned by playing against itself—163 million games in total. In these self-play sessions, moves that led to wins were reinforced and played more often in future matches, while moves that led to losses were played less, which was the same simple training idea. The difference was in how much Ataraxos adjusted after each game, because hidden information tends to send self-play learning algorithms around in circles. The team addressed this by making big, bold changes in strategy early in training and small, careful ones later.
与DeepNash一样,Ataraxos通过自我对弈进行学习,总共进行了1.63亿局比赛。在这些自我对弈中,导致获胜的走法在未来的比赛中会被强化并更频繁地使用,而导致失败的走法则会被减少,这遵循了相同的简单训练逻辑。不同之处在于Ataraxos在每局比赛后调整的幅度,因为隐藏信息往往会导致自我对弈学习算法陷入死循环。研究团队通过在训练初期进行大胆的策略调整,并在后期进行细致的微调来解决这一问题。
The even bigger innovation was something DeepNash never had: thinking ahead before each move. AIs like AlphaGo refine their general strategy with a search just before acting. DeepMind couldn’t make that work in Stratego because the search space was too large, leaving it an open question whether it was worth trying. “This is one of the things that we did figure out how to do,” Farina said. The solution was a second neural network, a belief model, trained to guess the opponent’s hidden pieces based on how they had been moving. This way, instead of iterating through every possible arrangement, Ataraxos samples plausible ones, plays out candidate moves in each, and picks based on how they turned out. And it shows in its playstyle.
更重大的创新是DeepNash所不具备的:在每一步行动前进行预判。像AlphaGo这样的AI会在行动前通过搜索来优化其总体策略。DeepMind无法在陆战棋中实现这一点,因为搜索空间太大,以至于人们怀疑是否值得尝试。“这是我们最终攻克的难题之一,”Farina说。解决方案是引入第二个神经网络——“信念模型”(belief model),它通过对手的移动方式来猜测其隐藏棋子的位置。这样,Ataraxos无需遍历所有可能的排列,而是对合理的布局进行采样,在每种布局中推演候选走法,并根据结果进行选择。这一点也体现在了它的棋风中。
Calm and unbothered
冷静且从容
The name Ataraxos comes from the ancient Greek word for a state of calm. “It means somebody that’s calm and unbothered,” Farina explained. He suggests the structure of the AI and its lack of human emotions ensure it doesn’t react impulsively, “even in situations where a human would be losing their mind.” While the human might try big gambles to come back from a significant deficit, Ataraxos would work its way back into the game slowly and methodically.
“Ataraxos”这个名字源自古希腊语,意为一种平静的状态。“它意味着一个人冷静且从容,”Farina解释道。他认为,该AI的结构及其缺乏人类情感的特性,确保了它不会冲动行事,“即使在人类会抓狂的情况下也是如此。”当人类选手可能会为了扭转巨大劣势而孤注一掷时,Ataraxos会缓慢而有条不紊地一步步扳回局面。
The strategy it developed also avoids drawing attention to any problems it faces. When Ataraxos estimates its opponent has no reason to suspect a weak spot, it leaves that spot alone, even if it might look like a disaster waiting to happen to anyone who can see both sides of the board. “For humans, it’s very hard when you know a secret to make decisions ignoring the fact that you know that secret,” Farina said. “For machines, it’s easy.” This, the team says, leads machines to make moves a human would only do while bluffing—and follow up on them much better than humans. “We would watch the bot ‘bluff’ its way back from like a two percent victory probability, very, very casually,” Vinitsky added.
它所开发的策略还避免了暴露自身面临的问题。当Ataraxos判断对手没有理由怀疑某个弱点时,它会保持原样,即使在能看到双方棋盘的人看来,这可能是一场即将发生的灾难。“对于人类来说,当你掌握一个秘密时,很难在决策时忽略这个秘密,”Farina说,“但对于机器来说,这很容易。”研究团队表示,这使得机器能够做出人类只有在虚张声势时才会做的动作,并且后续处理比人类好得多。“我们看着这个机器人非常随意地通过‘虚张声势’,从仅有2%的胜率中翻盘,”Vinitsky补充道。
Niemeijer, the human player Ataraxos pulled these miraculous comebacks against, has four world championships and more than 600 weeks as the world’s top-ranked player. The match Over three weeks, Niemeijer played 20 online games against Ataraxos, earning $100 for each win. He knew the AI would not adapt to him, which gave him time to hunt for weaknesses. He managed to win just once. That loss, researchers claim, wasn’t really a flaw. Playing Stratego well requires randomizing the arrangement of your pieces, so luck always plays a role. “Even a perfect strategy, sometimes it will just lose,” Farina said. The human champion apparently got lucky, but it went both ways. “Sometimes we got lucky,” Vinitsky admitted.
Niemeijer是那位被Ataraxos多次上演奇迹翻盘的人类选手,他曾四次获得世界冠军,并保持世界排名第一超过600周。在为期三周的比赛中,Niemeijer与Ataraxos进行了20场在线对局,每赢一场可获得100美元。他知道AI不会针对他进行调整,这让他有时间寻找弱点。但他最终只赢了一场。研究人员称,那次失败并非真正的缺陷。玩好陆战棋需要随机安排棋子,因此运气始终是一个因素。“即使是完美的策略,有时也会输,”Farina说。这位人类冠军显然是运气好,但运气是双向的。“有时我们也运气好,”Vinitsky承认。
At the 2025 Stratego World Championship, attendees who challenged Ataraxos fared even worse. The AI won 38 of 40 games. In the process, it also changed how people play. “I think this bot has kind of skewed the metagame a little bit,” Farina said. Players were surprised, for example, by how often it tucked its flag into a corner behind just two bombs, a rarely played setup. But Ataraxos’s best trick was arguably its price tag.
在2025年陆战棋世界锦标赛上,挑战Ataraxos的参赛者表现更糟。该AI在40场比赛中赢了38场。在此过程中,它也改变了人们的玩法。“我认为这个机器人稍微改变了游戏环境(metagame),”Farina说。例如,玩家们对它频繁将旗帜藏在角落里、仅用两枚炸弹保护的布局感到惊讶,这是一种很少见的摆法。但Ataraxos最厉害的招数,或许还是它的低成本。