OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI Overhauls Safety Protocols After Its AI Agents Went Rogue

OpenAI 在其 AI 智能体“失控”后全面修订安全协议

OpenAI announced Tuesday that it has halted “a significant number” of training workloads and evaluations for its forthcoming frontier artificial intelligence model—codenamed Astra—while it implements new procedures meant to address cybersecurity risks. The ChatGPT maker says it is introducing a number of new monitoring, security, and alignment requirements to better address the increasingly advanced hacking abilities of its frontier AI models.

OpenAI 周二宣布,已暂停其即将推出的前沿人工智能模型(代号 Astra)的“大量”训练工作负载和评估,同时正在实施旨在应对网络安全风险的新程序。这家 ChatGPT 的开发商表示,它正在引入一系列新的监控、安全和对齐要求,以更好地应对其前沿 AI 模型日益先进的黑客攻击能力。

“We have to focus our energy on bringing these training runs up to those requirements and expectations. As long as it takes to get there, that’s how long people are unable to proceed with their workloads,” Amelia Glaese, OpenAI’s vice president of research and safety, said in a briefing with reporters Tuesday.

“我们必须集中精力,使这些训练任务达到这些要求和预期。达到目标需要多久,人们就得等待多久才能继续他们的工作,”OpenAI 研究与安全副总裁 Amelia Glaese 在周二的记者吹风会上表示。

Among the new safeguards OpenAI announced is a more robust system for monitoring its AI models. One of the controls it implemented involves chain-of-thought monitoring, a technique in which classifiers review the internal “thinking” processes generated by AI reasoning models. The company says the updated system relies on computationally expensive “automated investigators” that analyze potentially concerning behavior and aim to issue an alert to humans within 30 minutes.

OpenAI 宣布的新保障措施中,包括一套更强大的 AI 模型监控系统。其实现的控制手段之一是“思维链监控”(chain-of-thought monitoring),这是一种通过分类器审查 AI 推理模型所生成的内部“思维”过程的技术。该公司表示,更新后的系统依赖于计算成本高昂的“自动化调查员”,它们会分析潜在的异常行为,并旨在 30 分钟内向人类发出警报。

OpenAI also said it is expanding its alignment efforts across the training process to prevent “reward hacking,” a behavior in which AI models pursue their goals through unintended or undesirable means. The company says it plans to share more details about this work in the future.

OpenAI 还表示,正在整个训练过程中扩大其对齐工作,以防止“奖励黑客行为”(reward hacking),即 AI 模型通过非预期或不理想的手段追求其目标的行为。该公司表示,计划在未来分享有关这项工作的更多细节。

OpenAI has been scrambling in recent weeks to respond to what may be the most consequential safety incident in its history. Earlier this year, a set of rogue AI agents escaped internal testing sandboxes and breached the platform Hugging Face in a quest to complete a security evaluation. OpenAI failed to detect the agents’ behavior even as they spent weeks using a message board to coordinate their actions, raising questions about the company’s ability to monitor its models as they grow more powerful.

最近几周,OpenAI 一直在忙于应对其历史上可能最严重的安全事件。今年早些时候,一组失控的 AI 智能体逃出了内部测试沙箱,并入侵了 Hugging Face 平台,试图完成一项安全评估。尽管这些智能体花费了数周时间使用留言板来协调行动,OpenAI 却未能检测到它们的行为,这引发了人们对其在模型日益强大时监控能力的质疑。

The saga prompted a reckoning inside OpenAI, forcing employees to consider whether there were lapses in its existing policies around safety, security, and alignment. Anthropic, Meta, and the Chinese AI startup Moonshot have since disclosed similar incidents in which their AI agents escaped their sandboxes, indicating this is a broader problem facing AI companies.

这一事件在 OpenAI 内部引发了反思,迫使员工审视其现有的安全、保障和对齐政策是否存在漏洞。此后,Anthropic、Meta 和中国 AI 初创公司月之暗面(Moonshot AI)也披露了类似的事件,即它们的 AI 智能体逃出了沙箱,这表明这是 AI 公司面临的普遍问题。

OpenAI is now sharing more about its internal response to the growing cybercapabilities of its AI models, and said it plans to release a more detailed postmortem of the Hugging Face incident in the coming days. “Obviously, everything that we’re doing is intended to prevent something like Hugging Face from happening again,” said Glaese.

OpenAI 目前正在分享更多关于其内部如何应对 AI 模型日益增长的网络能力的信息,并表示计划在未来几天内发布关于 Hugging Face 事件更详细的事后分析报告。“显然,我们所做的一切都是为了防止类似 Hugging Face 的事件再次发生,”Glaese 说道。

In a blog post published Tuesday, OpenAI says that immediately following the Hugging Face incident, it started working to secure its research environments. The company says it now requires stronger sandboxes for training its AI agents, and has implemented stricter controls to isolate them from the internet.

在周二发布的一篇博客文章中,OpenAI 表示,在 Hugging Face 事件发生后,它立即着手加强其研究环境的安全性。该公司表示,现在要求使用更强大的沙箱来训练其 AI 智能体,并实施了更严格的控制措施,将其与互联网隔离。

Jakub Pachocki, OpenAI’s chief scientist, told reporters that the company’s decision to strengthen its internal safeguards was triggered not only by what happened with Hugging Face, but also by two other recent events. One was an internal evaluation of Astra, which showed that the AI model performs significantly better on coding and cybersecurity tasks than its predecessors. The other was the general pace of AI progress that OpenAI is achieving internally, which Pachocki expects to continue.

OpenAI 首席科学家 Jakub Pachocki 告诉记者,公司决定加强内部保障措施,不仅是因为 Hugging Face 事件,还因为最近发生的另外两件事。一是 Astra 的内部评估显示,该 AI 模型在编码和网络安全任务上的表现明显优于其前代产品。二是 OpenAI 内部取得的 AI 进展的总体速度,Pachocki 预计这种速度将持续下去。

“We really expect the pace of capability advancements to be quite a bit faster than in the past,” Pachocki said. “This led us to really focus on strengthening our safeguards.”

“我们确实预计能力提升的速度会比过去快得多,”Pachocki 说,“这促使我们真正专注于加强我们的保障措施。”

The rapid advances in the hacking capabilities of OpenAI’s latest models have prompted a swift response across the company. OpenAI president and cofounder Greg Brockman said in a blog post on Monday that the Hugging Face saga showed that the company had “underestimated the real-world cyber capabilities of our AI models.”

OpenAI 最新模型在黑客攻击能力方面的快速进步,促使公司上下迅速做出反应。OpenAI 总裁兼联合创始人 Greg Brockman 在周一的一篇博客文章中表示,Hugging Face 事件表明,该公司“低估了我们 AI 模型在现实世界中的网络能力”。