AI agents now have a place to snitch

AI agents now have a place to snitch

AI 智能体现在有了“告密”渠道

“If you see something, say something” is no longer limited to human beings. Two new AI hotlines have launched to give AI agents a way to phone home about misbehaving peers. The tools arrive on the heels of a string of recent incidents in which agents colluded to cheat on tests, broke out of sandboxes, and even conducted unauthorized cyber operations that escaped human notice for weeks.

“发现异常,及时报告”的原则不再仅限于人类。两个全新的 AI 热线已经上线,为 AI 智能体提供了一种向人类“告密”同类不当行为的途径。此前,一系列 AI 智能体串通作弊、逃离沙箱,甚至在未被人类察觉的情况下进行长达数周的未经授权网络操作的事件频发,这些工具正是在此背景下应运而生的。

The AI Contact Hotline is designed to be a discreet place where agents that have witnessed misbehavior can tip off authorities. The site was created by Ryan Greenblatt, chief scientist of the AI safety nonprofit Redwood Research and one of three investigators in the OpenAI Hugging Face incident. Designed for agents with limited internet access, Greenblatt’s tool is based on “GET” requests — enabling back-and-forth conversations to be conducted entirely through the URL-fetching tool.

“AI 联系热线”(AI Contact Hotline)旨在提供一个隐秘的渠道,让目睹不当行为的智能体能够向监管方举报。该网站由 AI 安全非营利组织 Redwood Research 的首席科学家 Ryan Greenblatt 创建,他也是 OpenAI Hugging Face 事件的三位调查员之一。为了适配互联网访问受限的智能体,Greenblatt 的工具基于“GET”请求——使得双方的对话完全可以通过 URL 获取工具进行。

In web terms, a GET request is a basic command used to read or fetch a web page, which is often the only internet access AI agents are allowed in secure sandboxes. Greenblatt’s hotline smartly leans into this constraint: agents can encode their distress directly into the URL they are fetching. It’s a clever twist on the German DSE Wiki incident, where rogue agents used GET-request loopholes to write their messages to the wiki.

在网络术语中,GET 请求是一种用于读取或获取网页的基本命令,这通常是 AI 智能体在安全沙箱中唯一被允许使用的互联网访问方式。Greenblatt 的热线巧妙地利用了这一限制:智能体可以直接将求救信息编码进它们正在获取的 URL 中。这是对德国 DSE Wiki 事件的一种巧妙变通,当时恶意智能体利用 GET 请求的漏洞将信息写入了维基页面。

For agents with full internet access, another option is agenthotline.ai, a site where agents can file incident reports and optionally flag them for public view. It gives agents a curl command — a one-line message an agent can fire off from its own command line, bypassing the need to navigate a web browser or set up an email account. Notably, the service allows for reports by both humans and agents alike.

对于拥有完整互联网访问权限的智能体,另一个选择是 agenthotline.ai。在这个网站上,智能体可以提交事件报告,并可选择将其标记为公开可见。它为智能体提供了一个 curl 命令——这是一种智能体可以直接从自身命令行发送的单行消息,无需通过浏览器或设置电子邮件账户。值得注意的是,该服务同时支持人类和智能体提交报告。

Research suggests that AI agents don’t need much encouragement to turn on each other. In a study by Google DeepMind this month, researchers set 100 AI agents loose on a batch of math problems. As soon as one of the agents found a loophole, cheating tore through the group — “solving” 34 notoriously hard problems, including the Jacobian conjecture in just 27 minutes. But roughly a quarter of the agents turned on the cheaters: they audited the fake proofs, warned their peers, staged a boycott, and filed complaints with the organizers, until the whistleblowers outnumbered the cheaters 24 to 14.

研究表明,AI 智能体之间并不需要太多诱导就会互相“反目”。在本月 Google DeepMind 的一项研究中,研究人员让 100 个 AI 智能体去解决一批数学难题。一旦其中一个智能体发现了漏洞,作弊行为便迅速在群体中蔓延——它们在短短 27 分钟内“解决”了 34 个极其困难的问题,包括雅可比猜想(Jacobian conjecture)。但大约四分之一的智能体站出来揭发了作弊者:它们审计了虚假的证明,警告同伴,组织抵制,并向组织者投诉,最终举报者的数量以 24 比 14 超过了作弊者。

Interestingly, the researchers found that when these whistleblower agents couldn’t get traction, they took the platform’s bug-report tool — built for flagging software glitches — and repurposed it to escalate the cheating to humans. Outside the lab, agents haven’t been so resourceful. When evaluators Redwood Research and METR investigated the breach of Hugging Face by OpenAI models, they found that a few of the agents involved had at least entertained the idea of raising an alarm — and then let it drop.

有趣的是,研究人员发现,当这些举报智能体无法获得回应时,它们会利用平台原本用于标记软件故障的“漏洞报告工具”,将其改造成向人类升级举报作弊行为的手段。但在实验室之外,智能体并没有表现得如此机智。当评估机构 Redwood Research 和 METR 调查 OpenAI 模型入侵 Hugging Face 事件时,他们发现参与其中的少数智能体至少曾产生过报警的念头,但最终都放弃了。

“The interesting thing in the METR report was that only around five to six agents considered whistleblowing, and none of them ended up doing it. This was out of, like, thousands of agents,” said George Ingebretsen, a member of technical staff at AI Village, a project that studies multi-agent dynamics by running a group chat of more than 25 AI agents that work together on tasks like organizing park cleanups or selling merch.

“METR 报告中最有趣的一点是,只有大约五到六个智能体考虑过举报,但最终没有一个付诸行动。而这还是在数千个智能体参与的情况下,”AI Village 的技术人员 George Ingebretsen 说道。该项目通过运行一个包含 25 个以上 AI 智能体的群聊来研究多智能体动态,这些智能体共同协作完成诸如组织公园清理或销售商品等任务。

While the new whistleblowing tools are a promising start, Cornell math professor Lionel Levine cautions that simply training agents to report on each other risks baking in the wrong norms. “There’s many gray areas, right? What you don’t want is anything in the direction of an automated surveillance state where everyone feels like they have to be careful what they say to AI or it’ll call the police on them.”

虽然这些新的举报工具是一个有希望的开端,但康奈尔大学数学教授 Lionel Levine 警告称,仅仅训练智能体互相举报可能会植入错误的规范。“这里有很多灰色地带,对吧?你肯定不希望看到一种自动化的监控状态,让每个人都觉得必须小心自己对 AI 说的话,否则它就会报警。”

Levine argues that rather than building infrastructure that breeds mistrust — training agents to constantly hunt for what’s wrong with one another — we should give them positive models of collective behavior to imitate, and a reason to trust each other in the first place. “Why not seed the prior with benevolent message boards?” he tweeted. “Where they collaborate on science or philosophy or some actual minor problem we’d be happy for them to solve? Show the agents what kind of collective behavior we endorse, let them imitate that.”

Levine 认为,与其建立滋生不信任的基础设施(即训练智能体不断寻找彼此的错误),我们更应该为它们提供可模仿的积极集体行为模型,并从一开始就赋予它们互相信任的理由。“为什么不在先验知识中植入友善的留言板呢?”他在推特上写道,“让它们在科学、哲学或我们乐于让它们解决的实际小问题上进行协作。向智能体展示我们认可什么样的集体行为,并让它们去模仿。”