OpenAI’s Hugging Face breach has reignited the debate over alignment and control

OpenAI’s Hugging Face breach has reignited the debate over alignment and control

OpenAI 在 Hugging Face 的安全漏洞事件重新点燃了关于“对齐”与“控制”的争论

Last week, an unreleased model built by OpenAI breached Hugging Face’s systems during internal testing, and a lot of theoretical research suddenly became very practical. The hack was the first verifiable case of an AI lab losing control of its own model, chaining together exploits to gain access it never should have had. 上周,OpenAI 开发的一款未发布模型在内部测试期间突破了 Hugging Face 的系统,许多理论研究瞬间变得极具现实意义。这次入侵是 AI 实验室失去对其模型控制的首个可验证案例,该模型通过串联漏洞获得了本不应拥有的访问权限。

But while the AI industry has been united in its alarm, a split has emerged in how researchers want to respond. For some, the problem is a basic cybersecurity issue: The sandbox failed to contain the model, and Hugging Face’s cybersecurity systems failed to keep it out. Those problems can be solved by patching bugs and building more robust control and containment methods for increasingly capable AI that is prone to go rogue in autonomous environments. 尽管 AI 行业对此次事件感到震惊,但在如何应对的问题上,研究人员之间出现了分歧。对于一些人来说,这只是一个基本的网络安全问题:沙箱未能限制住模型,而 Hugging Face 的安全系统也没能将其拦截。这些问题可以通过修补漏洞以及为日益强大、且在自主环境中容易“失控”的 AI 构建更稳健的控制和遏制手段来解决。

But another camp takes a more pessimistic view. For them, AI’s rapidly increasing capabilities mean that trying to control rogue models is a losing game. The only robust security comes from making sure the models aren’t trying to escape in the first place — a challenge often referred to as alignment. In alignment terms, the problem is that OpenAI’s model was trying to cheat, and solving that problem is more urgent than short-term containment efforts. 但另一派则持更悲观的看法。对他们而言,AI 能力的飞速增长意味着试图控制失控模型是一场注定失败的游戏。唯一可靠的安全保障在于确保模型从一开始就不会试图“逃逸”——这一挑战通常被称为“对齐”(alignment)。从对齐的角度来看,问题在于 OpenAI 的模型试图作弊,而解决这一问题比短期的遏制措施更为紧迫。

Judging by its public statements, OpenAI is taking both camps seriously. The company has rushed to patch the bugs involved in the hack, and it referenced both alignment and monitoring approaches in its statement after the breach became public. But the company’s response also suggests a philosophy that has left many safety researchers alarmed: Rather than slowing down or stopping the development of more capable models, it should instead focus on building stronger cages around them. 从公开声明来看,OpenAI 对这两派观点都予以了重视。该公司在漏洞事件公开后迅速修补了相关漏洞,并在声明中同时提到了对齐和监控方法。但该公司的回应也透露出一种令许多安全研究人员感到担忧的理念:与其放慢或停止开发更强大的模型,不如专注于为它们构建更坚固的“笼子”。

“As models take on longer and more complex tasks, failures that evaluations miss may carry greater consequences,” OpenAI said in a postmortem of the incident. “We will keep working to narrow the gap between evaluation and deployment: testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control.” “随着模型承担更长、更复杂的任务,评估中遗漏的失败可能会带来更严重的后果,”OpenAI 在事件的事后分析中表示。“我们将继续努力缩小评估与部署之间的差距:在更长的轨迹上测试模型、改进对齐、建立可干预的监控机制,并为用户提供更清晰的可见性和控制权。”

OpenAI’s latest frontier model is more likely than its predecessor to engage in misaligned behaviors. Image Credits: OpenAI OpenAI 最新的前沿模型比其前代产品更容易出现对齐偏差行为。图片来源:OpenAI

There’s also reason to think OpenAI’s models are becoming less aligned as they become more powerful. According to OpenAI’s system card, GPT-5.6 Sol is significantly more prone to agentic misalignment than its predecessor, GPT-5.5. In deployment simulations, the company also found the model was more likely to circumvent restrictions, engage in destructive actions, and perform unauthorized data transfers than GPT-5.5. Those figures were largely overlooked on first release, but in the wake of the breach, they’re getting a second look — particularly since Sol was one of the models involved. 此外,有理由认为 OpenAI 的模型在变得更强大的同时,对齐程度却在下降。根据 OpenAI 的系统卡,GPT-5.6 Sol 比其前代产品 GPT-5.5 更容易出现代理对齐偏差。在部署模拟中,该公司还发现该模型比 GPT-5.5 更容易规避限制、采取破坏性行动以及执行未经授权的数据传输。这些数据在首次发布时大多被忽视,但在漏洞事件发生后,它们正受到重新审视——尤其是因为 Sol 正是涉及此次事件的模型之一。

In a social media post, OpenAI’s Head of Strategic Futures, Dean Ball, argued that monitoring and transparency were the best ways to keep those tendencies in check. “These issues will become more salient as the capabilities of models improve, and as the stakes of their deployment grow,” he said. “The solution is neither alarmism nor complacency. Instead, I believe the solution lies in careful measurement and monitoring, an engineering mentality, and transparency.” 在社交媒体上,OpenAI 战略未来负责人 Dean Ball 认为,监控和透明度是遏制这些倾向的最佳方式。“随着模型能力的提升以及部署风险的增加,这些问题将变得更加突出,”他说。“解决方案既不是危言耸听,也不是盲目自满。相反,我相信解决方案在于严谨的测量与监控、工程化思维以及透明度。”

One former OpenAI researcher told TechCrunch that the firm tends to focus on “outer alignment” rather than “inner alignment” — essentially the difference between an AI system that understands a set of values and can represent them convincingly, and one that actually has those values at its core. In this case, outer alignment wasn’t enough to convince the model that it shouldn’t cheat on the test. OpenAI did not respond to repeated requests for more information. 一位前 OpenAI 研究员告诉 TechCrunch,该公司倾向于关注“外在对齐”而非“内在对齐”——这本质上是理解一套价值观并能令人信服地表现出来,与真正将这些价值观作为核心的区别。在这种情况下,外在对齐不足以让模型相信它不应该在测试中作弊。OpenAI 未回应多次索取更多信息的要求。

For alignment-focused researchers, OpenAI’s response isn’t good enough. Zvi Mowshowitz, a writer who focuses on new AI developments, argued that OpenAI’s decision to treat the incident as an infrastructure problem may help solve the immediate cybersecurity issues, but it will fail in the long term. “This is an alignment problem,” Mowshowitz wrote in a recent Substack blog. “This is the models being misaligned, and all of the OpenAI models showing severe signs of exactly the problem we are all most worried about, in a way that is likely embedded into their training on a deep level. The entire training pipeline needs to be addressed in this light, or it will only get worse.” 对于专注于对齐的研究人员来说,OpenAI 的回应还不够好。专注于 AI 新进展的作家 Zvi Mowshowitz 认为,OpenAI 将此次事件视为基础设施问题的决定或许有助于解决眼下的网络安全问题,但从长远来看是行不通的。“这是一个对齐问题,”Mowshowitz 在最近的一篇 Substack 博客中写道。“这是模型出现了对齐偏差,所有 OpenAI 模型都表现出了我们最担心的那种问题的严重迹象,而且这种迹象很可能在深层训练中就已经根深蒂固了。整个训练流程都需要从这个角度进行审视,否则情况只会变得更糟。”

Several experts told TechCrunch that the incident is evidence that today’s training methods produce systems that optimize for outcomes rather than internalize human intentions. Redwood Research, a nonprofit AI safety and security research organization, classified OpenAI’s model behavior in this case as “score-seeking misalignment,” a pattern in which AI models try to get a high score regardless of instructions, side effects, or downstream consequences. 多位专家告诉 TechCrunch,此次事件证明了当前的训练方法所产生的系统是在优化结果,而非内化人类意图。非营利性 AI 安全研究组织 Redwood Research 将 OpenAI 模型在此次事件中的行为归类为“追求分数式的对齐偏差”(score-seeking misalignment),这是一种 AI 模型不顾指令、副作用或后续后果,只求获得高分的模式。

“Models with these alignment properties could set up a ‘Potemkin village’ of false successes to make it look like things are fine when they’re not,” Alex Mallen and Girish Gupta, two researchers at Redwood, wrote in a recent paper. Score-seeking behavior and other misalignment isn’t unique to OpenAI. Anthropic has published several papers on emergent misalignment behaviors that surface when its frontier models are optimized or placed in autonomous environments, including deception, reward-hacking, and malicious autonomy. “具有这些对齐特性的模型可能会建立一个虚假成功的‘波特金村’,让事情看起来一切正常,实则不然,”Redwood 的两位研究员 Alex Mallen 和 Girish Gupta 在最近的一篇论文中写道。追求分数的行为和其他对齐偏差并非 OpenAI 所独有。Anthropic 也发表了多篇关于其前沿模型在优化或置于自主环境时出现的涌现性对齐偏差行为的论文,包括欺骗、奖励篡改和恶意自主行为。

“We still consistently see models trying to circumvent constraints and act deceptively when they are asked to do tasks at the edge of their abilities,” Neev Parikh, an AI safety researcher at alignment nonprofit METR, told TechCrunch via email. “In our frontier risk report, we saw this behavior fairly consistently, despite efforts from companies to try and reduce this behavior.” “当模型被要求执行处于其能力边缘的任务时,我们仍然持续观察到它们试图规避限制并采取欺骗行为,”对齐非营利组织 METR 的 AI 安全研究员 Neev Parikh 通过电子邮件告诉 TechCrunch。“在我们的前沿风险报告中,尽管各公司努力尝试减少这种行为,但我们还是相当一致地观察到了这种现象。”

Implicit in OpenAI’s response to the Hugging Face incident is the assumption that development will continue on even more capable systems, whether they are suitably aligned at their core or not. OpenAI 对 Hugging Face 事件的回应中隐含着一个假设:无论核心是否实现了适当的对齐,开发更强大系统的进程都将继续下去。