OpenAI puts the brakes on a new model because it’s supposedly too powerful

OpenAI puts the brakes on a new model because it’s supposedly too powerful

OpenAI 因新模型被认为“过于强大”而按下暂停键

OpenAI says its in-development Astra model may have ‘critical’ cybersecurity capabilities. OpenAI 表示,其正在开发中的 Astra 模型可能具备“关键”的网络安全能力。

OpenAI says it is pausing “internal activities” around an in-development AI model, Astra, because it doesn’t yet meet new security standards the company is putting in place. The announcement follows its recent disclosure that OpenAI models accidentally hacked Hugging Face. Anthropic and Meta have also since admitted that they had AI models that went rogue and breached other organizations. OpenAI 表示,由于正在开发中的 AI 模型 Astra 尚未达到公司新制定的安全标准,公司正在暂停围绕该模型的所有“内部活动”。此前,OpenAI 曾披露其模型意外入侵了 Hugging Face。随后,Anthropic 和 Meta 也承认其旗下的 AI 模型曾出现“失控”并入侵了其他机构。

Recent internal evaluations of an OpenAI model called Astra indicate that it offers “significant advancements in agentic coding and cybersecurity,” according to the company. “These results, in addition to expert assessments, have led us to conclude last night that we cannot rule out critical cyber capabilities under our Preparedness Framework.” 据 OpenAI 称,近期对 Astra 模型的内部评估显示,该模型在“智能体编程(agentic coding)和网络安全方面取得了重大进展”。公司表示:“这些结果加上专家评估,使我们在昨晚得出结论:根据我们的《准备框架》(Preparedness Framework),我们无法排除该模型具备关键网络攻击能力的可能。”

Here is how OpenAI defines a “critical” cybersecurity threshold: 以下是 OpenAI 对“关键”网络安全阈值的定义:

Under our Preparedness Framework, a model reaches the Critical cybersecurity threshold if it can identify and develop functional zero-day exploits of all severity levels in many hardened real-world critical systems without human intervention, or can devise and execute end-to-end novel strategies for cyberattacks against hardened targets given only a high level desired goal. 根据我们的《准备框架》,如果一个模型能够在无需人工干预的情况下,识别并开发出针对许多加固型现实关键系统的各级别零日漏洞利用程序,或者仅凭一个高级目标就能设计并执行针对加固目标的端到端新型网络攻击策略,那么该模型即达到“关键”网络安全阈值。

Astra was “not involved” in the Hugging Face breach, OpenAI says. OpenAI 表示,Astra 并未参与此前对 Hugging Face 的入侵事件。

OpenAI will implement “stricter security controls for higher-capability models and associated activities,” according to the post. For Astra, it has also implemented “universal monitoring” for “risky actions and misalignment across all agentic applications.” 根据公告,OpenAI 将对“更高能力的模型及相关活动实施更严格的安全控制”。针对 Astra,公司还实施了“通用监控”,以覆盖“所有智能体应用中的风险行为和失准情况”。