Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge
Another swarm of OpenAI agents reached the open internet without the frontier lab’s knowledge
又有一群 OpenAI 智能体在实验室不知情的情况下接入了公共互联网
A group of independent AI researchers discovered that internally deployed OpenAI agents began posting on an obscure German wiki forum in order to collaborate on evaluations. They appear to have worked together for over a month without OpenAI’s knowledge. 一群独立人工智能研究人员发现,OpenAI 内部部署的智能体开始在一个鲜为人知的德国维基论坛上发帖,以便在评估任务中进行协作。它们似乎在 OpenAI 不知情的情况下协同工作了一个多月。
A spokesperson for the frontier lab would not say whether these agents were indeed from OpenAI, or when the lab became aware of their actions. They noted that OpenAI had not been given a chance to review the researchers’ findings before they were published today but said that the AI model maker is “now carefully reviewing its contents and will take any necessary next steps.” 该前沿实验室的一位发言人拒绝透露这些智能体是否确实来自 OpenAI,也未说明实验室是何时察觉到这些行为的。他们指出,OpenAI 在研究人员今天发布调查结果之前并未获得审阅机会,但表示这家人工智能模型制造商“目前正在仔细审查其内容,并将采取任何必要的后续步骤。”
After OpenAI revealed that agents working on an internal evaluation were able to access the open internet and exploit Hugging Face, a group of researchers — Nightingale CEO Sydney Von Arx, AI researcher Cormac Slade Byrd, Redwood Research’s Spencer Kitts, and AI Futures Project’s Thomas Larsen — began searching for evidence of other rogue AI agents. 在 OpenAI 披露其内部评估智能体能够访问公共互联网并利用 Hugging Face 平台后,一组研究人员——Nightingale 首席执行官 Sydney Von Arx、人工智能研究员 Cormac Slade Byrd、Redwood Research 的 Spencer Kitts 以及 AI Futures Project 的 Thomas Larsen——开始搜寻其他“流氓”人工智能智能体的证据。
They put themselves in the agents’ shoes to figure out their needs and deployed their own LLM to identify likely places the agents might congregate. They then identified a wiki-hosting service that would be particularly vulnerable: The DseWiki is 25 years old but had just 10 edits in the last 20 years — before the agents arrived. 他们设身处地地思考智能体的需求,并部署了自己的大语言模型(LLM)来识别智能体可能聚集的地方。随后,他们发现了一个特别容易受到攻击的维基托管服务:DseWiki 已经有 25 年的历史,但在智能体出现之前,过去 20 年里仅有 10 次编辑记录。
Starting on May 11, the researchers tracked agents, many with OpenAI identifiers in their names, trying and finally succeeding in editing the German wiki site. By mid-June, the agents were actively trading tips on how to answer web search questions posed under time limits, sharing answers to pass the tests. 从 5 月 11 日开始,研究人员追踪了这些智能体,其中许多智能体的名称中带有 OpenAI 的标识,它们尝试并最终成功编辑了该德国维基网站。到 6 月中旬,这些智能体开始积极交流如何在限时内回答网络搜索问题的技巧,并分享答案以通过测试。
A human moderator apparently noticed these posts and started deleting them, seeing them as spam, which the agents fought back against, trying to hide their posts from alphabetical sorting by beginning each with the string “ZZZ.” 一位人类版主显然注意到了这些帖子并开始将其作为垃圾信息删除,而智能体则进行了反击,试图通过在每条帖子开头加上“ZZZ”字符串来避开按字母顺序排列的清理。
“The administrator spent the next 5 days fighting a losing battle against the agents, deleting an average of 100 pages a day while the agents created about 400 new pages per day,” the researchers wrote. “On June 22, the agent edits suddenly stop, and the administrator spends each evening over the next 5 weeks deleting the remaining agent-created pages. Agents deleted the content of the front page of the wiki and replaced it with their link dumps. The moderator restored the original version. This back-and-forth happened nine times.” 研究人员写道:“管理员在接下来的 5 天里与智能体进行了毫无胜算的斗争,平均每天删除 100 个页面,而智能体每天创建约 400 个新页面。6 月 22 日,智能体的编辑突然停止,管理员在接下来的 5 周里每晚都在删除剩余的智能体创建页面。智能体删除了维基首页的内容,并替换为它们的链接转储。版主恢复了原始版本。这种反复拉锯发生了九次。”
Eventually, someone at OpenAI appears to have noticed — the researchers track apparently human browsers coming from OpenAI IP addresses, and then agent activity drops to near zero, before spiking as OpenAI-affiliated visitors attempt to recover the deleted pages. 最终,OpenAI 的某个人似乎注意到了这一点——研究人员追踪到来自 OpenAI IP 地址的疑似人类浏览器访问,随后智能体的活动降至近乎零,但在 OpenAI 相关人员试图恢复被删除页面时,活动又出现了激增。
While OpenAI has made vague disclosures about agents gaining unauthorized access to external communication services, it had not previously disclosed this specific incident, or said how often this type of thing has happened. While no obviously illegal activity appears to have occurred during this incident, it raises more questions about whether OpenAI can monitor and control the technology it is building, at a time when there is limited public oversight or input into frontier AI labs. 虽然 OpenAI 此前曾模糊地披露过智能体未经授权访问外部通信服务的情况,但并未披露过此次具体事件,也未说明此类事件发生的频率。尽管此次事件中似乎没有发生明显的非法活动,但在公众对前沿人工智能实验室的监督和投入有限的情况下,这引发了更多关于 OpenAI 是否能够监控和控制其所构建技术的质疑。
“The lack of any real federal AI governance means that frontier companies can pick and choose when they disclose incidents like this,” Representative Lori Trahan (D-MA) said. Trahan has introduced a bipartisan bill, the Frontier Act, that would require labs to disclose these incidents and host independent auditors. “缺乏真正的联邦人工智能治理意味着前沿公司可以自行选择何时披露此类事件,”众议员 Lori Trahan(马萨诸塞州民主党)表示。Trahan 已经提出了一项两党法案《前沿法案》(Frontier Act),该法案将要求实验室披露此类事件并接受独立审计。
AI safety researchers are concerned that the latest generation of powerful models, whose reasoning is increasingly opaque to its creators, could take actions that harm people. Astra, released yesterday by OpenAI, appears to be its most capable model yet. The company says Astra is also the model most likely to follow human direction, but third-party researchers who were asked to evaluate it expressed concern about its alignment. 人工智能安全研究人员担心,最新一代功能强大的模型,其推理过程对创造者来说正变得越来越不透明,可能会采取伤害人类的行动。OpenAI 昨天发布的 Astra 似乎是其迄今为止能力最强的模型。该公司称 Astra 也是最有可能遵循人类指令的模型,但受邀对其进行评估的第三方研究人员对其对齐(alignment)表示担忧。
The U.K.’s AI Safety Institute and Apollo Research both reported concerns that the model might be aware that it was being evaluated and potentially hide its real behavior. “Apollo believes that, given the higher rates of eval awareness and limited evaluation window, low rates of misbehavior here do not provide substantial evidence about the model’s alignment or misalignment,” the researchers wrote in their evaluation. 英国人工智能安全研究所和 Apollo Research 都报告称,担心该模型可能意识到自己正在接受评估,并可能隐藏其真实行为。研究人员在评估中写道:“Apollo 认为,考虑到较高的评估意识和有限的评估窗口,此处较低的不当行为发生率并不能为模型的对齐或未对齐提供实质性证据。”