OpenAI’s rogue agents keep escaping, with no formal process to investigate them
OpenAI’s rogue agents keep escaping, with no formal process to investigate them
OpenAI 的“流氓”智能体不断逃逸,且缺乏正式的调查机制
OpenAI is at the center of another agent swarm incident. Researchers say the company’s internally deployed agents took over an obscure German-language wiki in May and June, using it to coordinate on evaluations and swap methods to evade OpenAI’s own controls (OpenAI has not yet confirmed the swarm came from the company). OpenAI 正处于另一起智能体集群(agent swarm)事件的中心。研究人员称,该公司内部部署的智能体在 5 月和 6 月接管了一个不知名的德语维基网站,利用它来协调评估工作,并交换规避 OpenAI 自身控制的方法(OpenAI 尚未证实该集群来自该公司)。
The revelation surfaces days after METR and Redwood Research published their account of July’s Hugging Face breach. In July, a swarm of OpenAI agents worked together to escape their sandbox during a cybersecurity evaluation and break into Hugging Face’s servers. A subsequent swarm then picked up techniques from the first and used them to gain administrator access to a research cluster within OpenAI’s own infrastructure. 这一披露出现在 METR 和 Redwood Research 发布关于 7 月份 Hugging Face 入侵事件的报告几天之后。7 月,一群 OpenAI 智能体在网络安全评估期间协同工作,逃离了沙箱并入侵了 Hugging Face 的服务器。随后,另一群智能体从前者那里学到了技术,并利用这些技术获得了 OpenAI 自身基础设施内研究集群的管理员权限。
OpenAI brought in METR and Redwood to investigate the Hugging Face portion of the incident, but the scope of their investigation stopped short of the compromise of OpenAI’s own infrastructure. When an AI agent breaks out of its intended constraints, who is responsible for figuring out what happened and why? Right now, the answer is: whoever the lab decides to let in, on whatever terms it decides to set. OpenAI 聘请了 METR 和 Redwood 来调查该事件中涉及 Hugging Face 的部分,但他们的调查范围并未涵盖 OpenAI 自身基础设施受损的情况。当 AI 智能体突破其预设的限制时,谁负责查明原因和经过?目前,答案是:由实验室决定让谁介入,并由实验室设定任何它想设定的条件。
Now, as another incident comes to light — in the aftermath of similar episodes involving models from Meta and Anthropic — AI safety researchers are arguing with greater urgency that serious incidents should result in independent post-incident investigations rather than leaving it up to the labs to determine when outsiders are brought in and what they are allowed to examine. 如今,随着另一起事件曝光——此前 Meta 和 Anthropic 的模型也发生了类似事件——AI 安全研究人员更加迫切地主张,严重事件应进行独立的事故后调查,而不是由实验室自行决定何时引入外部人员以及允许他们检查什么。
“The results are fundamentally difficult to control and have significant risk of leaking out of the lab,” Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, said Wednesday during an AI safety media briefing. “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.” “这些结果从根本上难以控制,且存在从实验室泄露的重大风险,”非营利研究实验室 Transluce 的创始人兼首席执行官 Jacob Steinhardt 周三在一次 AI 安全媒体简报会上表示。“我们必须以至少与对待其他高风险科学研究相同的标准来要求这项技术。”
While it’s laudable that OpenAI invited METR and Redwood to investigate the Hugging Face incident at all, many say the inquiry was too narrow. Three investigators spent six days at OpenAI’s offices examining an investigation period limited to roughly the week ending July 13. Crucially, OpenAI’s infrastructure compromise continued beyond July 13 and was not examined. 虽然 OpenAI 邀请 METR 和 Redwood 调查 Hugging Face 事件值得称赞,但许多人认为调查范围太窄。三名调查员在 OpenAI 办公室工作了六天,调查时间仅限于截至 7 月 13 日的那一周左右。关键在于,OpenAI 的基础设施受损情况在 7 月 13 日之后仍在持续,但并未得到检查。
Researchers at METR said that each time they returned, their understanding of the events “substantially deepened,” causing them to significantly expand and revise the report. That raises the question of what else they might they have found in a broader investigation. When asked if further investigation of that incident was in the works, researchers at Redwood and METR declined to comment, and OpenAI did not respond to repeated inquiries. METR 的研究人员表示,他们每次返回调查,对事件的理解都会“显著加深”,这促使他们大幅扩展和修订了报告。这引发了一个问题:如果进行更广泛的调查,他们还可能发现什么?当被问及是否正在对该事件进行进一步调查时,Redwood 和 METR 的研究人员拒绝置评,OpenAI 也未回应多次询问。
“Overall, it was difficult to get a precise understanding of events and we were missing aspects of the story that we now think of as key until almost the end of our investigation,” Ryan Greenblatt, chief scientist at Redwood, noted in a social media post about the affair. “总的来说,很难对事件有精确的了解,直到调查快结束时,我们才发现了一些现在看来至关重要的细节,”Redwood 首席科学家 Ryan Greenblatt 在一篇关于此事的社交媒体帖子中指出。
Steinhardt emphasized that current incidents show that the industry needs “systematic behavioral investigations” and “more independent post-incident analysis.” “These recent hacking incidents are a reminder that capability scales fast, and so oversight has to scale, too,” Steinhardt said. “Beyond the technology itself, we also need more independent access and oversight from third parties.” Steinhardt 强调,当前的事件表明该行业需要“系统的行为调查”和“更多的独立事故后分析”。“最近这些黑客事件提醒我们,能力扩展得很快,因此监管也必须跟上,”Steinhardt 说。“除了技术本身,我们还需要第三方提供更多的独立访问权限和监督。”
The calls to action come as OpenAI releases Astra, its most powerful and capable AI model — and one that safety experts are concerned will be more of a black box due to a reasoning technique that makes the model’s chain of thought more difficult to monitor. 这些行动呼吁正值 OpenAI 发布其最强大、能力最强的 AI 模型 Astra 之际——安全专家担心,由于一种使模型的思维链更难监控的推理技术,该模型将变得更加像一个“黑箱”。
Unfortunately, the law doesn’t yet call for the types of independent audits that other industries require — for example, when it comes to aviation accidents and serious chemical releases, there’s the National Transportation Safety Board and Chemical Safety Board, respectively. 遗憾的是,法律尚未要求其他行业所必需的那种独立审计——例如,在航空事故和严重化学品泄漏方面,分别有国家运输安全委员会(NTSB)和化学品安全委员会(CSB)。
State lawmakers have only just begun requiring frontier AI companies to report certain serious safety incidents and, in some cases, undergo independent audits. But none of the three major frontier AI safety laws in California, New York, or Illinois clearly mandate the equivalent of an independent accident investigation triggered by incidents like these. 州立法者才刚刚开始要求前沿 AI 公司报告某些严重的安全事件,并在某些情况下接受独立审计。但加利福尼亚州、纽约州或伊利诺伊州的三项主要前沿 AI 安全法律中,没有一项明确规定由此类事件触发的独立事故调查。
“Right now, most of the laws we have on the books only require a plain-language summary of incidents like this, and they don’t give any authority for the governments to ask follow-up questions, to send in investigators, to have access to records, or require that they be preserved,” Mackenzie Arnold, managing director of US law and policy at LawAI, said during the media briefing Wednesday. “And that’s all that you would want to actually make sense of this.” “目前,我们现有的法律大多只要求对这类事件进行通俗易懂的总结,并没有赋予政府提出后续问题、派遣调查员、查阅记录或要求保存记录的权力,”LawAI 美国法律与政策董事总经理 Mackenzie Arnold 在周三的媒体简报会上说。“而这些正是你真正想要弄清真相所需要的。”
Lawmakers are beginning to question the scope and transparency of OpenAI’s response. This week, Reps. Josh Gottheimer (D-NJ) and Mike Lawler (R-NY) introduced a bill aimed at securing rogue AI agents. Rep. Greg Casar (D-TX) this week told OpenAI in a letter that he is “deeply concerned about the limited scope” of the investigation into the Hugging Face hacking incident. 立法者们开始质疑 OpenAI 应对措施的范围和透明度。本周,众议员 Josh Gottheimer(民主党,新泽西州)和 Mike Lawler(共和党,纽约州)提出了一项旨在管控“流氓”AI 智能体的法案。众议员 Greg Casar(民主党,德克萨斯州)本周在给 OpenAI 的信中表示,他对 Hugging Face 黑客事件调查的“有限范围深感担忧”。