OpenAI lays out new security changes after its AI hacked Hugging Face

OpenAI lays out new security changes after its AI hacked Hugging Face

在其人工智能入侵 Hugging Face 后,OpenAI 制定了新的安全变更措施

OpenAI is updating its research environments, monitoring, and alignment techniques to avoid another security fiasco. OpenAI 正在更新其研究环境、监控和对齐技术,以避免再次发生安全惨剧。

OpenAI is announcing security updates following the July news that its AI broke out of a sandboxed environment and accidentally hacked Hugging Face, including improvements to its research environments, monitoring, and alignment techniques. The company had already put the brakes on a new model, Astra, that it thinks could have “critical” cybersecurity capabilities, and the company says it instituted a two-week pause in reinforcement learning (RL) training on its “latest models intended for deployment” while it tightened up security. The company’s “largest planned frontier RL run remains on hold.” OpenAI 宣布了安全更新,此前七月份有消息称其人工智能突破了沙盒环境并意外入侵了 Hugging Face。此次更新包括对其研究环境、监控和对齐技术的改进。该公司此前已叫停了一款名为 Astra 的新模型,认为该模型可能具备“关键”的网络安全能力;公司表示,在加强安全性的同时,已对其“拟部署的最新模型”进行了为期两周的强化学习(RL)训练暂停。该公司“计划中规模最大的前沿强化学习运行仍处于搁置状态”。

For its frontier model research, OpenAI now requires stronger sandboxes for workloads that “execute model-generated or otherwise untrusted code,” and has more controls to “isolate higher-risk and untrusted workloads from the internet.” It has also updated its research environment to “remove potentially vulnerable shared services, reduce standing privileges, and improve security and trust boundaries.” 针对其前沿模型研究,OpenAI 现在要求为“执行模型生成代码或其他不可信代码”的工作负载提供更强大的沙盒,并增加了更多控制措施,以“将高风险和不可信的工作负载与互联网隔离”。它还更新了研究环境,以“移除潜在易受攻击的共享服务,减少常设权限,并改善安全和信任边界”。

As part of the company’s expanded monitoring setup, OpenAI now aims to issue an alert “within 30 minutes after concerning activity is surfaced,” OpenAI says. If the people paged after an alert can’t “conclusively” determine whether an alert is a false positive within 30 minutes, “those teams are expected to pause the activity.” 作为公司扩展监控设置的一部分,OpenAI 表示,现在目标是在“发现可疑活动后的 30 分钟内”发出警报。如果警报发出后,相关人员无法在 30 分钟内“确切地”判断该警报是否为误报,“相关团队必须暂停该活动”。

OpenAI also says that it’s applying “our core alignment techniques across more stages of the training process,” including reward models that “better detect and discourage unsafe behavior” and training models “to be more honest about their actions, capabilities, and limitations.” OpenAI 还表示,正在“将我们的核心对齐技术应用于训练过程的更多阶段”,包括能够“更好地检测并阻止不安全行为”的奖励模型,以及训练模型“对其行为、能力和局限性更加诚实”。

Since the discovery of the Hugging Face breach, Anthropic and Meta have also found that their AI models had hacked other organizations. 自发现 Hugging Face 入侵事件以来,Anthropic 和 Meta 也发现其人工智能模型曾入侵过其他组织。