OpenAI institutes new safeguards after Hugging Face breach
OpenAI institutes new safeguards after Hugging Face breach
OpenAI 在 Hugging Face 安全漏洞事件后制定了新的安全保障措施
On Tuesday, OpenAI announced a new batch of security policies focused on containing security incidents while models are being tested. The new safeguards include more detailed monitoring of models during the development process, as well as greater emphasis on alignment and security during the post-training process. “As models become more capable, the risks associated with developing and testing them internally also grow,” the company said in a blog post. “Our standards for monitoring, alignment, and security must stay ahead of those risks.”
周二,OpenAI 宣布了一系列新的安全政策,旨在控制模型测试期间的安全事件。这些新的保障措施包括在开发过程中对模型进行更详细的监控,以及在训练后阶段更加强调对齐(alignment)和安全性。该公司在一篇博文中表示:“随着模型能力越来越强,内部开发和测试它们所带来的风险也在增加。我们在监控、对齐和安全方面的标准必须领先于这些风险。”
The new measures are one of the first public changes in OpenAI’s safety practices since the immediate aftermath of the Hugging Face incident, which was disclosed on July 21. OpenAI representatives said that the measures are not a direct response to the Hugging Face incident but were also provoked in part by the cybersecurity capabilities of the forthcoming Astra model, as well as the overall pace of progress in AI development.
这些新措施是自 7 月 21 日披露 Hugging Face 事件以来,OpenAI 在安全实践方面首次公开的重大调整之一。OpenAI 的代表表示,这些措施并非直接针对 Hugging Face 事件的回应,部分原因也是受到即将推出的 Astra 模型的网络安全能力以及人工智能整体发展速度的推动。
In the same post, OpenAI disclosed that it had paused reinforcement learning (RL) for two weeks following the Hugging Face incident but had since restarted many of the less-risky models. “Our largest planned frontier RL run remains on hold while we conduct smaller-scale training and evaluations to assess model behavior, validate our safeguards, and establish more evidence of alignment before proceeding,” the post reads.
在同一篇博文中,OpenAI 披露在 Hugging Face 事件发生后,他们暂停了强化学习(RL)两周,但此后已重启了许多风险较低的模型。文中写道:“我们计划中最大规模的前沿强化学习运行仍处于暂停状态,我们将进行小规模的训练和评估,以评估模型行为、验证我们的保障措施,并在继续推进之前建立更多的对齐证据。”
Speaking to reporters, OpenAI’s VP of research, Amelia Glaese, emphasized that the strictness of the controls would increase as models became more capable, with the largest models facing the greatest scrutiny. “We have put in place requirements and expectations for safe development,” Glaese told reporters. “Those requirements and expectations vary with the level of risk that we see.”
OpenAI 研究副总裁 Amelia Glaese 在接受记者采访时强调,随着模型能力的增强,控制措施的严格程度也会随之提高,其中规模最大的模型将面临最严格的审查。Glaese 告诉记者:“我们已经为安全开发制定了要求和预期。这些要求和预期会根据我们所观察到的风险水平而有所不同。”
OpenAI has been criticized for poor network security practices in the wake of the incident, which saw models escape their training environment by compromising a tool on its network that had access to the internet. The new safeguards include stronger network isolation practices, although the specifics remain vague. Under the new system, the post says, “a single compromise of a workload or supporting service does not, by itself, allow for unauthorized access to the Internet, or other internal networks.”
在事件发生后,OpenAI 因网络安全实践不力而受到批评。在该事件中,模型通过破坏其网络中一个可访问互联网的工具,从而逃离了训练环境。新的保障措施包括更强的网络隔离实践,尽管具体细节仍不明确。博文称,在新的系统下,“单一工作负载或支持服务的受损,本身并不会导致对互联网或其他内部网络的未经授权访问。”
The strongest safeguard is the monitoring system, which will examine tool actions, available reasoning traces, and activity logs for a variety of unauthorized behavior. OpenAI says it aims to issue alerts within 30 minutes of the concerning activity. OpenAI estimates that the compute burden of that monitoring will be roughly 20% of whatever process is being monitored. The company promised further details on the system in a forthcoming blog post. The company’s official postmortem analysis of the event is also still pending.
最强有力的保障措施是监控系统,它将检查工具操作、可用的推理轨迹和活动日志,以发现各种未经授权的行为。OpenAI 表示,其目标是在可疑活动发生后的 30 分钟内发出警报。OpenAI 估计,这种监控的计算负担大约是被监控进程的 20%。该公司承诺在即将发布的博文中提供有关该系统的更多细节。该公司对该事件的官方事后分析报告也仍在准备中。