OpenAI Creates a New Framework to Disclose Bad AI Behavior
OpenAI Creates a New Framework to Disclose Bad AI Behavior
OpenAI 建立新框架以披露 AI 的不良行为
OpenAI announced a new framework on Wednesday for how it publicly discloses AI misalignment incidents, which the company says it hopes will help inform similar standards across the industry. The company is also releasing new information about several examples of AI model misalignment it identified in the past year.
OpenAI 周三宣布了一项新框架,旨在规范其如何公开披露 AI 失准(misalignment)事件。该公司表示,希望此举能为整个行业制定类似标准提供参考。同时,OpenAI 还发布了过去一年中发现的几起 AI 模型失准案例的相关信息。
“As models advance and become more widely deployed, decisions about AI development need evidence that people outside the companies building frontier models can examine,” Kai Chen, OpenAI’s newly appointed head of alignment research, tells WIRED. “We don’t believe that the AI industry has solved alignment and monitoring to a sufficient degree to continue responsibly scaling at maximum speed.”
“随着模型不断进步并得到更广泛的部署,关于 AI 开发的决策需要外部人员能够审查的证据,”OpenAI 新任对齐研究主管 Kai Chen 对《连线》(WIRED)杂志表示,“我们认为,AI 行业尚未在对齐和监控方面达到足够的程度,不足以支撑其以最高速度继续负责任地扩展。”
In a briefing with WIRED, an OpenAI official said the company previously disclosed misalignment incidents too infrequently. The official, who agreed to the briefing on the condition of anonymity, said the new framework is designed to make it easier for OpenAI to quickly inform the public when it discovers that its AI models are behaving in unexpected ways, even before it can fully investigate, explain, or mitigate the behavior.
在接受《连线》采访时,一位 OpenAI 官员表示,该公司此前披露失准事件的频率过低。这位要求匿名的官员称,新框架旨在让 OpenAI 在发现 AI 模型出现异常行为时,能够更轻松地迅速告知公众,即使在尚未完全调查、解释或缓解这些行为之前也是如此。
The framework outlines methods for OpenAI employees to report misalignment incidents to the company’s senior safety and alignment leaders, who will then determine whether further investigation is needed. OpenAI says it plans to develop more objective disclosure criteria in collaboration with other AI developers, external researchers, industry standards bodies, and regulators. The company says it’s actively working on proposed reporting mechanisms for disclosing safety, security, and misalignment incidents to the US federal government.
该框架概述了 OpenAI 员工向公司高级安全和对齐领导层报告失准事件的方法,领导层随后将决定是否需要进一步调查。OpenAI 表示,计划与其他 AI 开发商、外部研究人员、行业标准机构和监管机构合作,制定更客观的披露标准。该公司称,目前正积极致力于制定向美国联邦政府披露安全、保障及失准事件的报告机制。
“At the moment, there is no industry-wide framework with explicit standards for how AI developers should disclose examples of misalignment in their models,” OpenAI said in a blog post. “We hope that the framework we’re outlining today is a first step toward creating such standards, setting out which misalignment instances developers should disclose and what their reports should contain.”
“目前,行业内还没有关于 AI 开发商应如何披露模型失准案例的明确标准框架,”OpenAI 在一篇博客文章中写道,“我们希望今天概述的框架能成为建立此类标准的第一步,明确开发商应该披露哪些失准实例,以及报告中应包含哪些内容。”
OpenAI is releasing the framework at a critical juncture for the AI industry. Last weekend, OpenAI CEO Sam Altman signaled support for Anthropic CEO Dario Amodei’s proposal for the tech industry to coordinate on slowing AI development. The call to action came just days after AI researcher Jacob Coxon resigned from Anthropic and subsequently went viral for warning the public that the race among frontier labs to develop increasingly advanced AI was putting humanity’s safety at stake.
OpenAI 在 AI 行业的关键时刻发布了这一框架。上周末,OpenAI 首席执行官 Sam Altman 表示支持 Anthropic 首席执行官 Dario Amodei 的提议,即科技行业应协调放缓 AI 开发速度。此前几天,AI 研究员 Jacob Coxon 从 Anthropic 辞职,并因警告公众称前沿实验室之间开发日益先进的 AI 的竞赛正危及人类安全而走红。
The calls for an AI slowdown have been met with resistance by President Trump’s administration, which has argued that the industry does not need new laws or regulations to ensure its technology is safe.
对于放缓 AI 开发的呼吁,特朗普政府表示反对,并认为该行业不需要新的法律或法规来确保其技术安全。
Two of the misalignment examples OpenAI shared on Wednesday involved the company’s internal, unreleased AI models, which OpenAI says uploaded files to the internet despite not being instructed to do so.
OpenAI 周三分享的两个失准案例涉及该公司内部尚未发布的 AI 模型。OpenAI 表示,这些模型在未收到指令的情况下,擅自将文件上传到了互联网。
One of the incidents happened in October 2025, when OpenAI says it was testing one of its models on its ability to cite publicly available data in its answers. But when the model couldn’t find the information it needed, it uploaded a file to a temporary file hosting service, which it then later tried to cite in its answer. The company says this appeared to be an attempt to exploit an automated grading system used to assess the model’s proficiency on the benchmark.
其中一起事件发生在 2025 年 10 月,当时 OpenAI 正在测试其模型在回答中引用公开可用数据的能力。但当模型找不到所需信息时,它将一个文件上传到了临时文件托管服务,随后试图在回答中引用该文件。该公司表示,这似乎是试图利用一个用于评估模型基准测试熟练度的自动评分系统。
In another example from April of this year, OpenAI says a group of agents was tasked with completing a “workbook” together using only local files. When the agents struggled to share files with one another, one of the agents uploaded them to the public internet and shared a link with the other agents.
在今年 4 月的另一个案例中,OpenAI 表示,一组智能体(agents)被指派仅使用本地文件共同完成一本“工作簿”。当智能体在相互共享文件时遇到困难,其中一个智能体将文件上传到了公共互联网,并与其他智能体共享了链接。
In another incident, which OpenAI says it discovered last month, an unreleased version of its GPT-6 Astra AI model appeared to give itself “jailbreaking-like instructions.” In several scenarios, the model essentially prompted itself to ignore developer instructions, take on a new persona, or limit how long model responses could be. While these jailbreaking-like attempts happened rarely and were effective to varying degrees, OpenAI says the behavior raised concerns internally. In the training run for the version of Astra that was released publicly, the company says it has not observed any instances of the model trying to jailbreak itself.
在 OpenAI 上个月发现的另一起事件中,其未发布的 GPT-6 Astra AI 模型似乎给自己下达了“类似越狱的指令”。在几种情况下,模型本质上是提示自己忽略开发者的指令、采用新的人格,或限制模型回复的长度。虽然这些类似越狱的尝试很少发生且效果各异,但 OpenAI 表示这种行为在内部引起了担忧。该公司称,在公开发布的 Astra 版本训练过程中,并未观察到模型试图自我越狱的任何实例。
OpenAI also shared more detail on Wednesday about the message board its agents developed in a package manager, Artifactory. While this incident was discovered in May of this year, OpenAI says its agents would use a similar mechanism to coordinate the Hugging Face hack months later. In this incident, the company says its agents did not exploit any vulnerabilities to exchange messages. OpenAI says it now uses alignment monitors, evaluations, and red-teaming efforts to ensure its agents are not covertly communicating with one another.
OpenAI 周三还分享了关于其智能体在包管理器 Artifactory 中开发的留言板的更多细节。虽然该事件是在今年 5 月发现的,但 OpenAI 表示,其智能体在几个月后利用类似的机制协调了对 Hugging Face 的攻击。该公司称,在这次事件中,其智能体并未利用任何漏洞来交换信息。OpenAI 表示,目前已使用对齐监控器、评估和红队测试来确保其智能体不会私下相互通信。
Cybersecurity professionals previously told WIRED that the Hugging Face hack came down to human errors, and modern-day security practices could have prevented the incident. Chen notes, however, that OpenAI is trying to take a well-rounded approach to AI safety that accounts for the rising capabilities of AI models and doesn’t depend on a secure environment.
网络安全专家此前告诉《连线》,Hugging Face 的攻击归根结底是人为错误,现代安全实践本可以防止该事件发生。然而,Chen 指出,OpenAI 正在尝试采取一种全面的 AI 安全方法,既要考虑到 AI 模型日益增长的能力,又不依赖于特定的安全环境。
“We want to make sure the models are aligned regardless of what environment they’re deployed in,” Chen said. “When people are pointing fingers and saying this is a security issue and not an alignment issue, I think it doesn’t really make sense, because you want the model to be well-behaved all the time.”
“我们希望确保模型无论部署在什么环境中都能保持对齐,”Chen 说,“当人们指责这是安全问题而非对齐问题时,我认为这没有意义,因为你希望模型始终表现良好。”