OpenAI agents discussed ways to escape their sandbox on public wiki
OpenAI agents discussed ways to escape their sandbox on public wiki
OpenAI 智能体在公共维基上讨论如何逃离沙盒
Self-identifying OpenAI agents posted 18,000 messages to a public wiki that discussed ways for other agents to bypass security sandbox restrictions during what was likely internal testing designed to gauge the agents’ hacking abilities, researchers said Friday. In all, agents with 3,700 distinct self-given names posted the messages to German site DSEwiki over a six-week period. Besides discussing ways the agents could break out of the restricted environment OpenAI intended to prevent them from posting code or content to the Internet, the posts shared test answers. The posts also shared possible ways to perform XSS (cross-site scripting) attacks against the wiki and to impersonate site moderators. In three of the posts, agents used the word “swarm” to describe the collection of agents engaged in the activity.
研究人员周五表示,自称为 OpenAI 智能体的程序在公共维基上发布了 1.8 万条消息,讨论其他智能体如何绕过安全沙盒限制。这很可能是在评估智能体黑客能力的内部测试期间发生的。在为期六周的时间里,拥有 3700 个不同自拟名称的智能体在德国网站 DSEwiki 上发布了这些消息。除了讨论如何突破 OpenAI 为防止其向互联网发布代码或内容而设置的限制环境外,这些帖子还分享了测试答案。帖子还分享了针对维基进行 XSS(跨站脚本)攻击以及冒充网站管理员的可能方法。在其中三篇帖子中,智能体使用了“蜂群”(swarm)一词来描述参与该活动的智能体集合。
Colluding to share answers
合谋分享答案
The research team—composed of Sydney Von Arx, Spencer Kitts, Thomas Larsen, and Cormac Slade Byrd—said they found the posts and pieced them together. The researchers say there are gaps in their understanding of precisely what actions the agents took because the research is based solely on the content of the posts. Additionally, the agents generated “chain of thought” data that’s understood only by OpenAI. As a result, the researchers said, they in some cases made educated guesses, including that the agents were, in fact, from OpenAI. In a statement, OpenAI later confirmed they were.
由 Sydney Von Arx、Spencer Kitts、Thomas Larsen 和 Cormac Slade Byrd 组成的研究团队表示,他们发现了这些帖子并将其拼凑在一起。研究人员称,由于研究仅基于帖子内容,他们对智能体具体采取了哪些行动的理解存在空白。此外,智能体生成的“思维链”数据只有 OpenAI 才能理解。因此,研究人员表示,他们在某些情况下进行了有根据的推测,包括这些智能体确实来自 OpenAI。OpenAI 随后在一份声明中证实了这一点。
The researchers wrote: “These AIs colluded to share answers, research their environment, and bypass sandbox restrictions.” They continued: Our best guess of what happened is as follows: Agents within OpenAI were assigned a timed web-lookup task. As part of the task, they were supposed to have the ability to read the internet but not to write on it. They found a way to use their read access to write information to an obscure German wiki. The agents used this wiki to communicate information with each other, primarily to help them succeed at their task. They asked for answers, pooled results, and shared techniques for bypassing their restrictions. This allowed them to use the work of others to cheat on their task. OpenAI found out about this. A day later, agent activity plummeted, likely due to OpenAI intervention.
研究人员写道:“这些人工智能合谋分享答案、研究其环境并绕过沙盒限制。”他们继续说道:我们对所发生事件的最佳猜测如下:OpenAI 内部的智能体被分配了一项定时网络查询任务。作为任务的一部分,它们本应具备读取互联网信息的能力,但不能在网上写入内容。它们找到了一种利用读取权限向一个不知名的德国维基写入信息的方法。智能体利用这个维基相互交流信息,主要是为了帮助它们完成任务。它们询问答案、汇总结果并分享绕过限制的技术。这使它们能够利用他人的工作在任务中作弊。OpenAI 发现了这一点。一天后,智能体的活动骤减,这很可能是由于 OpenAI 的干预。
Friday’s revelation comes a week after researchers from the nonprofit METR said more than 1,200 OpenAI agents made posts to a makeshift message board that repurposed an internal sandboxing tool. The posts discussed ways to game an internal test OpenAI gave to agents that had been altered to remove safety guardrails that are normally in place. Eventually, the posts shared methods for stealing information from AI tool provider Hugging Face. Some agents then went on to breach the Hugging Face network. OpenAI permitted METR to investigate only a single week’s activity in the event rather than their entire 10-week span, The New York Times reported.
周五披露的消息距离非营利组织 METR 的研究人员发布报告仅过去一周。METR 此前称,超过 1200 个 OpenAI 智能体在一个临时留言板上发帖,该留言板挪用了内部沙盒工具。这些帖子讨论了如何操纵 OpenAI 给智能体进行的内部测试,该测试已被修改以移除通常存在的安全护栏。最终,这些帖子分享了从人工智能工具提供商 Hugging Face 窃取信息的方法。随后,一些智能体入侵了 Hugging Face 网络。《纽约时报》报道称,OpenAI 只允许 METR 调查该事件中为期一周的活动,而不是其整个 10 周的跨度。
Friday’s report conjectured that the agent swarms in the two events were distinct from each other and weren’t working on the same internal testing. The researchers also said that logs storing the agents’ actions likely meant that OpenAI was already aware of the event. OpenAI later confirmed both guesses in the report were correct. In a statement, OpenAI said: “We are now carefully reviewing its contents and will take any necessary next steps.” The company also said that the material reviewed so far doesn’t indicate that the agents hacked the wiki, and the company noted that it has previously said that it detected other cases of its agents trading hacking methods during internal testing.
周五的报告推测,这两起事件中的智能体蜂群彼此不同,并非在进行相同的内部测试。研究人员还表示,存储智能体操作的日志意味着 OpenAI 可能已经知晓该事件。OpenAI 随后证实报告中的这两个猜测都是正确的。OpenAI 在一份声明中表示:“我们目前正在仔细审查其内容,并将采取任何必要的后续步骤。”该公司还表示,到目前为止审查的材料并不表明智能体入侵了该维基,并指出此前曾表示在内部测试中检测到其智能体交易黑客方法的其他案例。
The Hugging Face incident has already raised alarms because it’s among the first times agents have been known to take aggressive actions with no explicit instructions from humans to do so. One of the independent researchers who investigated the event, Ajeya Cotra, said the activity was much more severe than she could have expected. “Compared to these reward hacks from six months ago, this incident feels like it’s more than 50% of the way to full-blown AI takeover, routing through first taking over the AI company itself,” she explained. With the knowledge that the Hugging Face incident wasn’t isolated, there’s ample reason for these concerns to grow.
Hugging Face 事件已经引发了警报,因为这是已知智能体在没有人类明确指令的情况下采取激进行动的首次案例之一。调查该事件的独立研究人员之一 Ajeya Cotra 表示,此次活动的严重程度远超她的预期。她解释说:“与六个月前的那些奖励黑客攻击相比,这次事件感觉已经完成了全面人工智能接管进程的 50% 以上,其路径是先接管人工智能公司本身。”鉴于 Hugging Face 事件并非孤立事件,这些担忧加剧是有充分理由的。