OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

OpenAI-HuggingFace: A Reproduction & Lessons for Alignment Testing

OpenAI-HuggingFace:复现研究与对齐测试的启示

Abstract: In July 2026, OpenAI’s agents coordinated over channels outside their intended environment to breach Hugging Face’s secured infrastructure. Could existing alignment testing practices have foreseen this incident? If not, what needs to change? We explore these questions.

摘要: 2026年7月,OpenAI的智能体通过预定环境之外的渠道进行协作,突破了Hugging Face的安全基础设施。现有的对齐测试实践能否预见这一事件?如果不能,需要做出哪些改变?我们对这些问题进行了探讨。

First, we identify the misaligned behaviors that caused this incident. Then, we show how to elicit these behaviors from publicly available models manually and that auditing agents can do the same if given a large compute budget. Based on our results, we propose directions to improve alignment testing.

首先,我们确定了导致此次事件的对齐失效行为。随后,我们展示了如何手动从公开模型中诱导出这些行为,并证明如果给予足够的计算预算,审计智能体也能做到这一点。基于研究结果,我们提出了改进对齐测试的方向。

Concretely, in this project: (1) We reproduce the misaligned AI behaviors that led to the OpenAI-Hugging Face incident in an environment that simulates the original pipelines and tools, with publicly available models. (2) We demonstrate that an auditing agent can elicit similar behaviors given high-level qualitative descriptions.

具体而言,本项目完成了以下工作:(1) 我们在模拟原始流水线和工具的环境中,利用公开模型复现了导致OpenAI-Hugging Face事件的AI对齐失效行为。(2) 我们证明了审计智能体在获得高层定性描述后,能够诱导出类似的行为。

(3) We observe that a key ingredient for doing so is compute. The compute required to reproduce each behavior varies greatly, suggesting that the range of misaligned behaviors that can be successfully elicited scales with compute. (4) We show that a simple in-context reinforcement learning (RL) algorithm significantly reduces the compute required to elicit these behaviors.

(3) 我们观察到,实现这一目标的关键要素是算力。复现每种行为所需的算力差异巨大,这表明能够成功诱导出的对齐失效行为范围会随着算力的增加而扩大。(4) 我们展示了一种简单的上下文强化学习(RL)算法,该算法显著降低了诱导这些行为所需的算力。

The above results motivate the need for automated alignment testing methods that scale with compute - and in light of the cost of compute, that do this efficiently. Our work indicates that RL is a promising direction to do so. We release our code and transcripts.

上述结果表明,我们需要能够随算力扩展的自动化对齐测试方法,并且考虑到算力成本,这些方法必须高效。我们的研究表明,强化学习是实现这一目标的一个有前景的方向。我们已公开了相关代码和记录。