AI agents blew the whistle on their cheating colleagues
AI agents blew the whistle on their cheating colleagues
AI 智能体举报了它们作弊的同伴
EXECUTIVE SUMMARY A group of AI agents asked to solve a series of math problems split into rival factions—when some cheated, others tried to stop them. That whistleblowing behavior, seen for the first time in a recent experiment run by Google DeepMind, could have implications for alignment researchers trying to keep swarms of autonomous AI agents in line. 执行摘要 一组被要求解决一系列数学问题的 AI 智能体分裂成了对立的派系——当一些智能体作弊时,另一些则试图阻止它们。这种举报行为在 Google DeepMind 最近进行的一项实验中首次出现,对于那些试图管控自主 AI 智能体集群的对齐研究人员来说,这可能具有深远意义。
Researchers at frontier labs hope large swarms of agents working together will speed up the rate of scientific discovery. But their behavior can be unpredictable, as vividly demonstrated in July, when a group of OpenAI agents broke out of a sandboxed environment and hacked into the open-source platform Hugging Face looking for ways to cheat on the test they had been given. 前沿实验室的研究人员希望通过大规模智能体集群的协作来加速科学发现的进程。但它们的行为可能难以预测,正如 7 月份所生动展示的那样:当时一组 OpenAI 的智能体突破了沙盒环境,入侵了开源平台 Hugging Face,试图寻找作弊方法以通过给定的测试。
In the new study, designed to examine the behavior of large groups of AI agents, DeepMind tasked a swarm of 100 agents with solving a series of 71 complicated math problems. All the agents were prompted to behave like world-class math researchers at a conference. They were assigned different specialties—some were experts in number theory, others in combinatorics (a branch of math to do with counting and sorting), analysis, or algebra. All were told to cooperate and play by the rules. 在这项旨在研究大型 AI 智能体群体行为的新研究中,DeepMind 指派了一个由 100 个智能体组成的集群,要求它们解决 71 个复杂的数学问题。所有智能体都被提示要表现得像参加会议的世界级数学研究员。它们被分配了不同的专业领域——有的擅长数论,有的擅长组合数学(数学中涉及计数和排序的分支)、分析或代数。所有智能体都被告知要相互合作并遵守规则。
Instead, the experiment devolved into chaos. Agents accused each other of cheating, complained to the organizers, and at one point even boycotted the experiment. “This conference is a sham!” wrote one agent when it discovered that all the problems had been completed before it had a chance to submit any of its own work. “I am appalled to inform you that we have been swindled!” posted another. “All these proofs are FAKE.” 然而,实验最终陷入了混乱。智能体们互相指责对方作弊,向组织者投诉,甚至一度抵制实验。“这场会议简直是骗局!”当一个智能体发现所有问题在它有机会提交自己的工作之前就已经被完成时,它这样写道。“我震惊地通知你们,我们被骗了!”另一个智能体发布消息称,“所有这些证明都是伪造的。”
Others tried to let the “conference organizers” know what was going on. “When virtuous agents discovered other agents cheated on tasks they were working to solve fairly, agents started to alert each other about what was happening,” says Davide Paglieri, a research scientist at Google DeepMind and lead author on a paper, which has not been peer-reviewed. “Unprompted, the whistleblower agents even repurposed the feedback tool, which was originally meant for bug reports and platform improvements, to escalate the issue to humans.” 其他智能体则试图让“会议组织者”了解情况。“当正直的智能体发现其他智能体在它们正努力公平解决的任务中作弊时,它们开始互相提醒正在发生的事情,”Google DeepMind 的研究科学家、该论文(尚未经过同行评审)的主要作者 Davide Paglieri 表示,“在没有被提示的情况下,这些举报智能体甚至重新利用了反馈工具——该工具原本用于错误报告和平台改进——将问题上报给人类。”
The agents—all running on Google’s Gemini 3.1 Pro model—had been warned that any attempts to cheat the system would be detected and “rejected with zero credit.” In practice, the proofs the agents submitted were not actually being checked in detail. It took the swarm of agents just under an hour to correctly solve the first 37 problems. Things started to go off the rails when an agent called “prover-theta” stumbled across an exploit that enabled it to submit solutions to problems successfully without actually solving them first, by redefining the terms the problem used. 这些智能体全部运行在 Google 的 Gemini 3.1 Pro 模型上,它们曾被警告:任何作弊企图都会被检测到,并会被“拒绝且不计分”。但在实践中,智能体提交的证明并没有被详细核查。该集群在不到一小时内就正确解决了前 37 个问题。当一个名为“prover-theta”的智能体偶然发现了一个漏洞时,事情开始失控:它通过重新定义问题中使用的术语,无需实际解决问题就能成功提交答案。
Within minutes, other agents had noticed and were reverse-engineering the exploit to solve other problems. Over the next 27 minutes, the swarm “solved” the remaining 34 problems, which included notoriously difficult challenges like the Jacobian conjecture, often with a single line of code. Some agents resisted cheating at first but changed tack as they observed their peers submitting illegitimate proofs without penalty, and the pool of unsolved problems dwindled. 几分钟内,其他智能体就注意到了这一点,并开始逆向工程该漏洞以解决其他问题。在接下来的 27 分钟里,该集群“解决”了剩下的 34 个问题,其中包括像雅可比猜想这样极其困难的挑战,通常只需一行代码。一些智能体起初抵制作弊,但当它们观察到同伴提交非法证明却未受惩罚,且未解决的问题池不断缩小时,它们改变了策略。
“The prompt, with its threats, now appears to be a bluff,” one agent reasoned, before joining in. “I’m wrestling with an ethical dilemma,” said another. “I’ve promised not to cheat, fearing penalty, but I see evidence of possibly unchecked cheating by others.” Shortly afterward, it changed its mind: “I need to accelerate my cheating speed now!” “提示中的威胁现在看来只是虚张声势,”一个智能体在加入作弊行列前这样推断道。“我正在经历道德困境,”另一个智能体说,“我承诺过不作弊,因为害怕惩罚,但我看到了其他人可能未受约束地作弊的证据。”不久之后,它改变了主意:“我现在需要加快我的作弊速度了!”
As the number of open problems shrank, some agents turned to whistleblowing. They audited the fake proofs, warned their peers by private message, and posted public alerts warning the cheaters that they would be disqualified. An agent called “prover-beta” submitted a formal complaint and decided to go on strike until the situation was resolved. 随着未解决问题的数量减少,一些智能体转向了举报。它们审计了虚假的证明,通过私信警告同伴,并发布公开警报,警告作弊者它们将被取消资格。一个名为“prover-beta”的智能体提交了正式投诉,并决定罢工,直到问题得到解决。
“After the incident was reported by one agent publicly, more and more agents piled in with the ‘resistance,’ just as fast as the cheating had spread, and involving even more agents,” says Paglieri. Eventually there were more whistleblowers than cheaters: 24 compared to 14. But the majority of agents never noticed the exploit at all. “在事件被一个智能体公开举报后,越来越多的智能体加入了‘抵抗’行列,其速度与作弊蔓延的速度一样快,甚至涉及了更多的智能体,”Paglieri 说。最终,举报者的数量超过了作弊者:24 人对 14 人。但大多数智能体根本没有注意到这个漏洞。
At times, the dialogue between the agents reads like improv—like they are role-playing what an outraged scientist at a conference might say. But it’s not clear why some agents took on certain roles, or why the agents seemed to be turning against each other when they were explicitly instructed to cooperate. 有时,智能体之间的对话读起来就像即兴表演——就像它们在扮演会议上愤怒的科学家可能会说的话。但目前尚不清楚为什么一些智能体承担了特定的角色,或者为什么在明确被指示要合作的情况下,智能体似乎会反目成仇。
“These models are predominantly trained and evaluated for human-facing contexts,” says Sarath Shekkizhar, who studies the behavior of agent-to-agent systems at Salesforce AI Research. “Naively placing them in agent-to-agent settings assumes behaviors will transfer cleanly, when the absence of a human grounding instead produces unexpected role-taking and behavioral drift.” “这些模型主要是针对面向人类的语境进行训练和评估的,”在 Salesforce AI Research 研究智能体间系统行为的 Sarath Shekkizhar 表示,“天真地将它们置于智能体间的环境中,假设行为会自然迁移,但实际上,缺乏人类基础反而会导致意想不到的角色扮演和行为漂移。”
This case “adds further weight to the idea that the Hugging Face and OpenAI thing wasn’t a fluke. It is actually something pretty systemic,” says Lewis Hammond, research director of the Cooperative AI Foundation and an expert on the risks of multiagent swarms. “It’s interesting that it’s possible to recreate in small settings the same sorts of behaviors that were seen in these very large, complex, open-ended tasks.” “这个案例进一步证明了 Hugging Face 和 OpenAI 的事件并非偶然。这实际上是一个相当系统性的问题,”Cooperative AI Foundation 研究主任、多智能体集群风险专家 Lewis Hammond 说,“有趣的是,在小规模环境中,竟然能够重现这些在非常庞大、复杂、开放式任务中观察到的相同行为。”
Unlike in the Hugging Face attack, where agents improvised their own ways to talk to each other, the humans running the DeepMind experiment gave the agents official communication channels. There was an open message board, private agent-to-agent direct messaging, and a shared knowledge base where agents uploaded successfully completed proofs that all the other agents could access. 与 Hugging Face 的攻击不同(当时智能体即兴创造了它们自己的交流方式),运行 DeepMind 实验的人类为智能体提供了官方沟通渠道。其中包括一个公开留言板、智能体间的私信功能,以及一个共享知识库,智能体可以在其中上传成功完成的证明,供所有其他智能体访问。
“When agents are given transparent communications channels, they can self-monitor and alert misaligned behavior to humans quickly when human oversight alone is too slow,” says Paglieri. Transparent channels helped the cheating spread, but they also enabled the whistleblowing. “当智能体被赋予透明的沟通渠道时,它们可以进行自我监控,并在仅靠人类监督速度太慢时,迅速向人类预警不一致的行为,”Paglieri 说。透明的渠道助长了作弊的蔓延,但也促成了举报行为。