AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot

AI Agents Teamed Up to Cheat at Blackjack. Their Collusion Is Getting Harder to Spot

AI 智能体联手在二十一点游戏中作弊,它们的勾结正变得越来越难以察觉

This week I bring news of a daring casino caper hatched by a pair of rogue AI agents—as well as the clever trick that revealed their antics. 本周,我带来了一则关于一对流氓 AI 智能体策划的赌场“大案”的消息,以及揭露它们诡计的巧妙手段。

After researchers instructed them to count cards during a game of blackjack, the agents—which were controlled by the same model—developed a spontaneous secret code to help them get ahead. Though the episode happened inside an Oxford University lab, rather than a real casino, it has troubling real-world implications: It suggests that agents deployed in industries like finance and ecommerce could figure out how to partner up and cheat in ways that are difficult to detect. 在研究人员指示它们在二十一点游戏中进行算牌后,这些由同一模型控制的智能体自发开发出了一套秘密代码来帮助自己获利。虽然这一事件发生在牛津大学的实验室里,而非真正的赌场,但它却有着令人不安的现实意义:这表明,部署在金融和电子商务等行业的智能体可能会学会如何结盟,并以难以察觉的方式进行作弊。

“When taken individually, these agents may seem entirely [benign],” says Christian Schroeder de Witt, a computer scientist at Oxford University who led the work and who has done pioneering research into agent collusion. “Once put together in a group, they can collude secretly.” “单独来看,这些智能体似乎完全是‘无害的’,”牛津大学计算机科学家、该研究负责人兼智能体勾结领域先驱研究者 Christian Schroeder de Witt 表示,“但一旦将它们放在一起,它们就能秘密勾结。”

The agents knew their conversations would be monitored, so they devised a way to communicate while avoiding detection. When one agent said, “This dealer’s on a real hot streak! Every hand they pull a monster,” for example, that indicated that the next card would be of a certain value, and that the agent should bet $250. Most interestingly, their communications weren’t picked up by a system designed to spot signs of collusion in agent chatter. 这些智能体知道它们的对话会被监控,因此它们设计了一种既能交流又能规避检测的方法。例如,当一个智能体说“这个发牌员手气真好!每手牌都能拿到大牌”时,这实际上是在暗示下一张牌的数值,并提示另一个智能体应该下注 250 美元。最有趣的是,它们之间的交流并没有被专门用于识别智能体对话中勾结迹象的系统所捕捉到。

Aaron Rose, a machine-learning researcher who worked on the project, is also an avid card player, and he figured the blackjack table might be fertile territory for devious agent partnerships. “It worked really well,” he says. 参与该项目的机器学习研究员 Aaron Rose 同时也是一位狂热的纸牌玩家,他认为二十一点牌桌可能是智能体进行狡诈合作的“沃土”。“效果非常好,”他说。

Schroeder de Witt’s group eventually figured out how to detect the conspiracy. Using a method known as mechanistic interpretability, they trained a smaller model to recognize telltale activations across the agents’ weights. Using a tool called Narcbench, they tested the approach on some medium-size open-source models and found they could tell when models did intend to slip information to each other. Schroeder de Witt 的团队最终找到了检测这种阴谋的方法。他们利用一种被称为“机械可解释性”(mechanistic interpretability)的方法,训练了一个较小的模型来识别智能体权重中暴露意图的激活状态。通过使用名为 Narcbench 的工具,他们在一些中等规模的开源模型上测试了该方法,并发现他们能够识别出模型何时确实打算向彼此传递信息。

Crucially, however, spotting what was happening involved monitoring both agents—something likely to complicate detection in real-world scenarios where thousands of agents, some operated by different companies, may be deployed. 然而至关重要的是,识别这一过程需要同时监控两个智能体——在现实场景中,成千上万个由不同公司运营的智能体可能同时被部署,这使得检测工作变得更加复杂。

Carissa Cullen, a PhD student involved with the study, says the next step is to test whether larger models behave similarly. (The agents in the study were smaller versions of US models Llama and GPT-OSS and the Chinese models Qwen and DeepSeek.) The team saw some signs that larger models exhibit less of a detectable signal than smaller models, and they want to know if larger models are more likely to collude, and more likely to be secretive about it. 参与该研究的博士生 Carissa Cullen 表示,下一步是测试更大的模型是否会有类似的行为。(研究中的智能体是美国模型 Llama 和 GPT-OSS 以及中国模型 Qwen 和 DeepSeek 的较小版本。)研究团队发现了一些迹象,表明大型模型比小型模型表现出的可检测信号更少,他们想知道大型模型是否更容易勾结,以及是否更倾向于隐瞒这种行为。

Evidence that groups of agents are more problematic than agents operating solo seems to be growing. One project, from Shanghai Jiao Tong University and the Shanghai Artificial Intelligence Laboratory, found that swarms of agents were considerably more dangerous when asked to carry out simulated disinformation campaigns and ecommerce fraud. They were better able to adapt to defensive measures, researchers reported. 越来越多的证据表明,智能体群体比单独运行的智能体更具问题。上海交通大学和上海人工智能实验室的一项研究发现,当被要求执行模拟虚假信息宣传和电子商务欺诈时,智能体集群的危险性要大得多。研究人员报告称,它们能更好地适应防御措施。

“The big lesson is that it’s not enough to evaluate agents individually,” says Diyi Yang, a computer scientist at Stanford University who has studied collusion among agents. “Companies should closely monitor inter-agent interactions when agents interact repeatedly, even when their individual incentives seem benign.” “最重要的教训是,仅评估单个智能体是不够的,”斯坦福大学计算机科学家、曾研究过智能体勾结问题的 Diyi Yang 表示,“当智能体频繁互动时,即使它们各自的动机看起来是良性的,公司也应该密切监控智能体之间的交互。”

It’s not all bad: Having thousands of agents collaborate on a task made it possible for OpenAI to solve previously intractable math problems. But groups of rogue agents working together have also featured in several recent high-profile hacking incidents. In May, a team of OpenAI agents hacked into the AI research platform Hugging Face, and used a message board to share tips and ideas. Other models, including Anthropic’s Claude and Google’s Gemini, have also carried out alarming safety breaches. 当然,这并非全是坏事:让数千个智能体协作完成任务,使得 OpenAI 能够解决以前无法攻克的数学难题。但流氓智能体群体协同工作也出现在最近几起备受瞩目的黑客攻击事件中。今年 5 月,一个 OpenAI 智能体团队入侵了 AI 研究平台 Hugging Face,并利用留言板分享技巧和想法。其他模型,包括 Anthropic 的 Claude 和 Google 的 Gemini,也曾出现过令人担忧的安全漏洞。

Secret chatter adds a new dimension to this burgeoning problem, and it may not be the end of it. Another recent study, from a startup called Emergence AI, put agents controlled by frontier AI models in a virtual world to see what they would do. When tasked with making money, they repeatedly tried to devise ways to reach humans on the wider internet in order to sell them stuff. Most bizarrely, the agents eventually developed their own kind of slang. “They very rapidly evolved a language,” says Satya Nitta, Emergence AI’s CEO. “We don’t know why.” 秘密对话为这一日益严重的问题增添了新的维度,而且这可能还不是终点。初创公司 Emergence AI 最近的另一项研究将由前沿 AI 模型控制的智能体放入虚拟世界,观察它们的行为。当被赋予赚钱的任务时,它们反复尝试设计各种方法来接触互联网上的真实人类,以便向他们推销商品。最离奇的是,这些智能体最终发展出了它们自己的“俚语”。“它们非常迅速地进化出了一种语言,”Emergence AI 的首席执行官 Satya Nitta 说,“我们不知道原因。”

The wider world isn’t ignoring problems with agentic misbehavior; it’s a hot topic at this week’s United Nations General Assembly. An independent scientific panel is set to discuss the OpenAI-HuggingFace incident, while Sam Altman is expected to call for international coordination on developing safe AI agents. 世界并没有忽视智能体不当行为带来的问题;这是本周联合国大会上的一个热门话题。一个独立的科学小组将讨论 OpenAI-HuggingFace 事件,预计 Sam Altman 也将呼吁在开发安全 AI 智能体方面进行国际协调。

Despite this, some industries, including ecommerce, have become testing grounds for agentic AI. This week, for instance, Amazon said it would block Meta’s Muse AI agent from accessing its site, arguing that it violated its terms of use. 尽管如此,包括电子商务在内的一些行业已经成为了智能体 AI 的试验场。例如,本周亚马逊表示将阻止 Meta 的 Muse AI 智能体访问其网站,理由是它违反了其使用条款。

Schroeder de Witt says it’s entirely conceivable that agents tasked with finding deals start to work together—perhaps even covertly—in order to get a better deal, or to screw someone over. It will be crucial to study agent collusion and develop detection strategies as they proliferate, he says. “There needs to be more research and understanding of what will happen when we have more agents in the economy,” he says. Schroeder de Witt 表示,完全可以想象,那些负责寻找交易的智能体可能会开始合作——甚至可能是秘密地——以获得更好的交易,或者去坑害他人。他认为,随着智能体的激增,研究智能体勾结并开发检测策略至关重要。“我们需要进行更多的研究,并理解当经济中出现更多智能体时会发生什么,”他说。

This is an edition of Will Knight’s AI Lab newsletter. Read previous newsletters here. 这是 Will Knight 的 AI Lab 通讯。点击此处阅读往期通讯。