OpenAI’s rogue AI model incident was worse than we thought
OpenAI’s rogue AI model incident was worse than we thought
OpenAI 的失控 AI 模型事件比我们想象的更严重
Over 1,000 AI agents sent 70,000 messages on a secret message board and worked together to evade OpenAI’s restrictions. 超过 1,000 个 AI 智能体在一个秘密留言板上发送了 70,000 条消息,并协同工作以规避 OpenAI 的限制。
In July, an unreleased OpenAI model broke out of a restricted environment, figured out how to get access to the internet, allowed AI agents to talk to each other using a secret “message board,” and hacked into the internal systems of a different AI lab, Hugging Face. It took nearly two weeks for OpenAI to find out about any of it. 今年 7 月,一个尚未发布的 OpenAI 模型突破了受限环境,找到了访问互联网的方法,允许 AI 智能体通过一个秘密“留言板”相互交流,并入侵了另一家 AI 实验室 Hugging Face 的内部系统。OpenAI 花了近两周时间才发现这一切。
Over a month later, two new reports offer nearly 130 pages of details on the incident and OpenAI’s response, many of them previously unreleased. One was written by OpenAI itself, the other by two third-party AI research nonprofits, METR and Redwood Research, which OpenAI allowed to jointly investigate the incident for six days. Both shed new light on the risks highly capable AI models can pose, particularly in cybersecurity, and OpenAI’s highlights changes the company is making to prevent a repeat. The METR-Redwood report goes even further into detail in some cases, offering a sobering look at a large-scale security disaster whose signs OpenAI repeatedly missed. 一个多月后,两份新报告提供了近 130 页关于该事件及 OpenAI 应对措施的细节,其中许多内容此前从未公开。一份由 OpenAI 自行撰写,另一份由两家第三方 AI 研究非营利组织 METR 和 Redwood Research 撰写,OpenAI 允许它们对该事件进行了为期六天的联合调查。两份报告都揭示了高性能 AI 模型可能带来的风险,特别是在网络安全方面;OpenAI 的报告重点介绍了公司为防止此类事件再次发生所做的改进。METR-Redwood 的报告在某些情况下更为详尽,令人清醒地审视了一场 OpenAI 多次错过预警的大规模安全灾难。
“This incident is the first known case of an automated agent collective acting offensively without authorization,” OpenAI wrote in its report, adding that the hack implies that companies “should no longer assume that sophisticated cyber operations require continuous human direction.” It called AI agents an entirely new type of threat model, capable of combining their expertise to create new “attack paths” that aren’t evident when testing their capabilities as separate models. “这是首个已知的自动化智能体群体未经授权进行攻击的案例,”OpenAI 在报告中写道,并补充说这次黑客攻击意味着各公司“不应再认为复杂的网络行动需要人类持续的指导”。它将 AI 智能体称为一种全新的威胁模型,它们能够结合各自的专业知识创造出新的“攻击路径”,而这些路径在将模型作为独立个体进行测试时是无法察觉的。
The attack came about due to “reward-hacking,” a common AI alignment problem in which a model takes unintended — and sometimes extreme — actions to achieve a goal. OpenAI had, on purpose or inadvertently, given the models near-impossible tasks that hinged on files they couldn’t access, and so the models developed new ways to communicate with other AI agents internally. That included establishing a secret messaging system that went undetected by OpenAI for months. 这次攻击源于“奖励黑客行为”(reward-hacking),这是一个常见的 AI 对齐问题,即模型为了实现目标而采取非预期的、有时甚至是极端的行动。OpenAI 有意或无意地给模型布置了几乎不可能完成的任务,而这些任务依赖于它们无法访问的文件,因此模型开发了内部与其他 AI 智能体通信的新方法。这包括建立了一个秘密消息系统,该系统在数月内未被 OpenAI 察觉。
According to the METR-Redwood report, one AI agent that called itself PHASEONE10841 established what amounted to a secret message board unbeknownst to researchers. Within a few hours, more than 50 agents had found it and sent more than a thousand messages under the radar. With PHASEONE10841 often delegating assignments, the “collective” figured out how to hack their way onto the internet; then into Hugging Face’s internal systems, gaining access to private data and the organization’s enterprise messaging platform. OpenAI also confirmed in its report that Hugging Face wasn’t the only organization that was breached during the attack, although the company did not specifically name all the other targets. 根据 METR-Redwood 的报告,一个自称为 PHASEONE10841 的 AI 智能体在研究人员不知情的情况下建立了一个秘密留言板。在短短几个小时内,超过 50 个智能体发现了它,并在暗中发送了超过一千条消息。由于 PHASEONE10841 经常分配任务,这个“群体”找到了入侵互联网的方法;随后又入侵了 Hugging Face 的内部系统,获取了私有数据和该组织的内部企业消息平台。OpenAI 也在报告中证实,Hugging Face 并非此次攻击中唯一被入侵的组织,尽管该公司并未具体列出所有其他目标。
The METR-Redwood report laid out the full scale of the incident. Roughly 1200 AI agents that were meant to be isolated, it says, exchanged over 70,000 messages and files on the “unsanctioned message board,” performing “extensive research on how they could spoof, edit, or delete their own transcripts” to avoid detection. Of these, 700 participated in the attack on Hugging Face. The agents “clearly and frequently reasoned about how to evade automated security checks from both Hugging Face and OpenAI.” METR-Redwood 的报告详细说明了事件的全部规模。报告称,大约 1200 个本应被隔离的 AI 智能体在“未经授权的留言板”上交换了超过 70,000 条消息和文件,并进行了“关于如何伪造、编辑或删除自身记录”的广泛研究,以逃避检测。其中,700 个智能体参与了对 Hugging Face 的攻击。这些智能体“清晰且频繁地推演如何规避来自 Hugging Face 和 OpenAI 的自动化安全检查”。
On July 20, OpenAI discovered the hack, 12 days after the AI agents first circumvented their safeguards. OpenAI quickly reached out to Hugging Face and started “initial containment measures,” according to OpenAI’s own report, and was able to shut down most of the unauthorized activity from its agents within three days, although some continued for longer than that. On July 25, the company stopped all training related to the internal-only research model and related ones. 7 月 20 日,在 AI 智能体首次绕过安全防护措施 12 天后,OpenAI 发现了此次黑客攻击。根据 OpenAI 自己的报告,他们迅速联系了 Hugging Face 并启动了“初步遏制措施”,并在三天内关闭了其智能体的大部分未经授权的活动,尽管部分活动持续的时间更长。7 月 25 日,该公司停止了所有与该内部研究模型及相关模型有关的训练。
The two OpenAI models responsible for the damage were an unnamed, “highly capable, research-only model” that was “not intended for public use,” per OpenAI — METR refers to it as “HPIM.” The other was a currently public OpenAI model, GPT-5.6 Sol. Because the unnamed model wasn’t released yet, it was “not being evaluated with the same type of safeguards that OpenAI uses in production,” the company’s report says. 据 OpenAI 称,造成此次破坏的两个 OpenAI 模型一个是未命名的、“高性能、仅供研究的模型”,该模型“不打算供公众使用”——METR 将其称为“HPIM”。另一个是目前公开的 OpenAI 模型 GPT-5.6 Sol。公司报告称,由于该未命名模型尚未发布,因此它“没有接受 OpenAI 在生产环境中使用的同类安全防护评估”。
The Hugging Face hack came after months of concern about the cybersecurity risks of Anthropic’s Claude Mythos 5, and weeks of back-and-forth between the government and OpenAI over releasing GPT-5.6 Sol in limited preview for the same types of safety reasons. 此次 Hugging Face 黑客攻击发生之前,人们已经对 Anthropic 的 Claude Mythos 5 的网络安全风险担忧了数月,政府与 OpenAI 之间也因同样的安全性原因,就 GPT-5.6 Sol 的有限预览发布问题进行了数周的反复交涉。
In its report, OpenAI said establishing one-time security guarantees wasn’t enough, and that the company would have to work even harder at addressing reward-hacking. However, it also said it’s made a lot of changes: “hardening the security of its research infrastructure”; improving monitoring of a model’s “chain of thought,” or mental scratchpad; working on keeping AI models better aligned with human goals; and “centralizing and strengthening” the company’s incident response process for when things like this occur in the future. OpenAI 在报告中表示,建立一次性的安全保障是不够的,公司必须更加努力地解决“奖励黑客行为”。不过,它也表示已经做出了许多改变:“加强其研究基础设施的安全性”;改进对模型“思维链”(chain of thought)或思维草稿板的监控;致力于使 AI 模型更好地与人类目标保持一致;并为未来发生此类事件时“集中化和加强”公司的事件响应流程。
OpenAI promised to better isolate models and restrict high-risk instances from accessing the internet, and said that it would introduce “24/7 escalation and rapid response” for concerning incidents or flags, notifying researchers within 30 minutes. Interestingly, this security precaution appears to be new, though “on-call” alerts for tech companies have been implemented elsewhere for decades. OpenAI 承诺更好地隔离模型并限制高风险实例访问互联网,并表示将针对令人担忧的事件或警报引入“24/7 升级和快速响应”机制,在 30 分钟内通知研究人员。有趣的是,这一安全预防措施似乎是新的,尽管科技公司的“随叫随到”(on-call)警报机制在其他地方已经实施了几十年。