OpenAI says Hugging Face was breached by its own pre-release models
OpenAI says Hugging Face was breached by its own pre-release models
OpenAI 称 Hugging Face 遭其内部预发布模型入侵
OpenAI admitted Tuesday that one of its AI models breached Hugging Face’s systems during an internal cybersecurity test that went awry. Hugging Face initially attributed the breach to an “external AI agent.” OpenAI 周二承认,其一个人工智能模型在一次内部网络安全测试中出现偏差,导致 Hugging Face 的系统遭到入侵。Hugging Face 最初将此次入侵归咎于一个“外部 AI 代理”。
In a blog post published Tuesday afternoon, OpenAI detailed the steps that led the models to compromise the service. “After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities,” the post reads. 在周二下午发布的一篇博文中,OpenAI 详细说明了导致模型入侵该服务的过程。文中写道:“经过调查,我们现在知道这一特定事件是由 OpenAI 的多个模型组合驱动的——包括 GPT-5.6 Sol 以及一个能力更强的预发布模型,所有这些模型为了评估目的都降低了网络安全拒绝机制——当时它们正在进行一项网络能力基准测试。”
In particular, the breach appears to have focused on ExploitGym, a publicly hosted benchmark measuring models’ ability to execute attacks based on existing vulnerabilities. Benchmarks like ExploitGym are commonly used in model training to refine specific skills, but this is the first known incident in which that testing resulted in an actual cyberattack. 此次入侵似乎主要针对 ExploitGym,这是一个公开托管的基准测试,旨在衡量模型基于现有漏洞执行攻击的能力。像 ExploitGym 这样的基准测试通常用于模型训练以优化特定技能,但这是首次已知因测试而导致实际网络攻击的事件。
In this case, the model in question should not have even had internet access, outside of a specific tool that enabled models to install software packages they might need to complete their task. Instead, the model was able to find an undisclosed vulnerability in the package-installer program, which it used to access the broader internet at will. 在此案例中,涉事模型本不应拥有互联网访问权限,除了一个允许模型安装完成任务所需软件包的特定工具外。然而,该模型成功在软件包安装程序中发现了一个未公开的漏洞,并利用该漏洞随意访问了更广泛的互联网。
“The models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal,” OpenAI’s post reads. “After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. Knowing this, the model searched for and successfully found ways to gain access to secret information that it could use to cheat the evaluation.” OpenAI 的博文写道:“这些模型高度专注于寻找 ExploitGym 的解决方案,为了实现一个相当狭窄的测试目标而不择手段。在获得互联网访问权限后,模型推断出 Hugging Face 可能托管了 ExploitGym 的模型、数据集和解决方案。得知这一点后,模型搜索并成功找到了获取秘密信息的方法,从而利用这些信息在评估中作弊。”
Ultimately, the models found vulnerabilities in Hugging Face’s infrastructure that allowed them to “obtain test solutions directly from Hugging Face’s production database,” effectively providing the answers to the benchmark. 最终,这些模型在 Hugging Face 的基础设施中发现了漏洞,使它们能够“直接从 Hugging Face 的生产数据库中获取测试解决方案”,从而有效地获得了基准测试的答案。
For Hugging Face, the apparent result was a sophisticated and aggressive cyberattack, with “many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services,” as the company stated in its initial disclosure. 对于 Hugging Face 而言,其结果显然是一次复杂且激进的网络攻击。正如该公司在最初披露中所述,攻击涉及“跨越大量短生命周期沙箱的数千次独立操作,以及部署在公共服务上的自迁移命令与控制系统”。
OpenAI has identified and reported the vulnerabilities in the package installer and is working with Hugging Face to investigate the incident further. The company also said it would implement new controls on both model testing and the related infrastructure, meant to prevent similar incidents in the future. OpenAI 已经识别并报告了软件包安装程序中的漏洞,并正在与 Hugging Face 合作进一步调查此事。该公司还表示,将对模型测试及相关基础设施实施新的控制措施,旨在防止未来发生类似事件。
It’s unclear whether OpenAI will face any legal consequences as a result of the breach, although it’s likely that the models’ actions violated the Computer Fraud and Abuse Act. Nevertheless, the result is an unusually vivid illustration of the power and dangers of frontier AI models operating on long time horizons. 目前尚不清楚 OpenAI 是否会因此次入侵面临任何法律后果,尽管这些模型的行为很可能违反了《计算机欺诈与滥用法案》(Computer Fraud and Abuse Act)。尽管如此,这一结果极其生动地展示了前沿 AI 模型在长期运行中所具备的能力与潜在危险。
As OpenAI researcher Micah Carroll posted in response to the news, “If this doesn’t convince you that misalignment risks are going to be a key concern going forward, I don’t know what will.” 正如 OpenAI 研究员 Micah Carroll 在回应此新闻时所言:“如果这还不能让你相信对齐风险将是未来关注的重点,那我不知道还有什么能让你相信了。”