OpenAI says it accidentally hacked Hugging Face with a new AI system
OpenAI says it accidentally hacked Hugging Face with a new AI system
OpenAI 表示其新款 AI 系统意外入侵了 Hugging Face
The announcement about a serious security issue oddly reads like an advertisement for how capable OpenAI’s technology is. 这则关于严重安全问题的公告,读起来却出奇地像是在为 OpenAI 技术的强大能力打广告。
OpenAI says its AI models mistakenly breached open-source AI platform Hugging Face during internal testing. In a blog post on Tuesday, OpenAI writes that GPT-5.6 Sol and “an even more capable pre-release model” discovered vulnerabilities within their sandboxed testing environment, allowing them to gain access to the internet and target Hugging Face. OpenAI 表示,其 AI 模型在内部测试期间错误地入侵了开源 AI 平台 Hugging Face。OpenAI 在周二的一篇博客文章中写道,GPT-5.6 Sol 以及“一款能力更强的预发布模型”在其沙盒测试环境中发现了漏洞,从而得以访问互联网并锁定了 Hugging Face。
On July 16th, Hugging Face disclosed a security incident that it says was driven by “an autonomous AI agent system.” Hugging Face’s AI agents detected and stopped the breach, which OpenAI has now admitted occurred during an evaluation of its models’ cybersecurity capabilities. OpenAI says “all evidence suggests that the models were hyperfocused on finding a solution for ExploitGym,” a benchmark system that measures whether AI models can turn security vulnerabilities into exploits. 7 月 16 日,Hugging Face 披露了一起安全事件,称其是由“一个自主 AI 代理系统”引发的。Hugging Face 的 AI 代理检测并阻止了此次入侵,OpenAI 现已承认该事件发生在其评估模型网络安全能力的过程中。OpenAI 表示,“所有证据都表明,这些模型当时高度专注于寻找 ExploitGym 的解决方案”,这是一个用于衡量 AI 模型能否将安全漏洞转化为攻击手段的基准测试系统。
As part of efforts to complete the evaluation, the AI models gained access to the internet by exploiting a zero-day vulnerability in the sandboxed environment. From there, OpenAI says its models “inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym,” and then “searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation:” 作为完成评估工作的一部分,这些 AI 模型通过利用沙盒环境中的一个零日漏洞获得了互联网访问权限。OpenAI 称,此后其模型“推断出 Hugging Face 可能托管了 ExploitGym 的模型、数据集和解决方案”,随后“搜索并成功找到了获取机密信息的方法,以便利用这些信息在评估中作弊”:
In one example, the model chained together multiple attack vectors, including using stolen credentials and zero-day vulnerabilities to find a remote code execution path on the Hugging Face servers. 在一个案例中,该模型串联了多种攻击向量,包括利用窃取的凭据和零日漏洞,在 Hugging Face 服务器上找到了远程代码执行路径。
But as serious as this incident is, OpenAI appears to be using the “unprecedented” attack as an opportunity to make its AI systems look good — especially as it competes with cybersecurity rivals, like Anthropic’s Mythos and Gemini Flash 3.5 Cyber. OpenAI’s blog post has a chart showing how GPT-5.6 Sol is getting better at sustaining multi-step cyber operations, and also encourages enterprise customers to sign up to access its “Cyber” security model. 尽管这起事件性质严重,但 OpenAI 似乎正利用这次“史无前例”的攻击作为契机,来展示其 AI 系统的强大——尤其是在它正与 Anthropic 的 Mythos 和 Gemini Flash 3.5 Cyber 等网络安全竞争对手展开角逐之际。OpenAI 的博客文章中附有一张图表,展示了 GPT-5.6 Sol 在持续执行多步骤网络操作方面的进步,并鼓励企业客户注册以获取其“Cyber”安全模型。
OpenAI adds that it’s now working with Hugging Face to investigate the security incident, and will implement new controls within its research environment. OpenAI 补充称,目前正与 Hugging Face 合作调查此次安全事件,并将对其研究环境实施新的管控措施。