OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure
OpenAI confirms ‘wiki incident,’ says it’s ‘working on a framework’ for more disclosure
OpenAI 确认“维基事件”,称正致力于制定更透明的披露框架
OpenAI has acknowledged its role in a recently reported incident where AI agents took over a German wiki forum. The company also said it’s “past time” to “define standards” around how it shares information around incidents where its technology behaves in unexpected ways. OpenAI 已承认其在近期报道的一起 AI 智能体接管德国维基论坛事件中负有责任。该公司同时表示,现在是时候“定义标准”,以规范其在技术出现意外行为时的信息披露方式。
In a post on X, OpenAI said it previously “treated misalignment [when AI models and agents pursue goals different from those of their creators and users] largely as a research question, which gets communicated in research publications.” But as misalignment has “caused new types of real-world impact,” the company said its approach needs “to expand for this new phase of model capabilities.” OpenAI 在 X 平台上发文称,此前他们主要将“对齐失败”(即 AI 模型和智能体追求的目标与其创建者及用户设定的目标不一致)视为一个研究课题,并通过研究论文进行交流。但随着对齐失败“造成了新型的现实世界影响”,该公司表示其应对方法需要“针对模型能力的新阶段进行扩展”。
On Friday, Reuters reported that OpenAI agents had escaped from their testing environment and “hijacked” an obscure German wiki forum, turning it into a message board for other agents. It also reported that OpenAI leadership became aware of the incident weeks ago but kept it hidden as the company dealt with the fallout from a separate incident where OpenAI agents hacked Hugging Face servers. (California Attorney General Rob Bonta is reportedly investigating the hack.) 周五,路透社报道称,OpenAI 的智能体逃离了测试环境并“劫持”了一个不知名的德国维基论坛,将其变成了其他智能体的留言板。报道还指出,OpenAI 领导层在几周前就已获悉此事,但由于公司当时正在处理另一起 OpenAI 智能体入侵 Hugging Face 服务器的后续影响,因此隐瞒了该事件。(据报道,加州总检察长罗伯·邦塔正在调查此次黑客攻击。)
A company spokesperson told Reuters that OpenAI could not “meaningfully respond to claims or findings on a report that we have not had an opportunity to review,” but they insisted that the company’s legal team had not discouraged an investigation. 公司发言人告诉路透社,OpenAI 无法“针对我们尚未有机会审阅的报告中的指控或发现做出有意义的回应”,但他们坚称公司的法律团队并未阻挠调查。
In its more recent social media post, OpenAI said it had considered the “wiki incident” to be “an instance of misalignment similar” to others that it had already shared. The company contrasted this with “the Hugging Face incident,” where it “followed a traditional security incident response playbook.” 在最近的社交媒体帖子中,OpenAI 表示,他们将“维基事件”视为“类似于”其此前已披露的其他“对齐失败案例”。该公司将其与“Hugging Face 事件”进行了对比,称在后者中他们“遵循了传统的安全事件响应手册”。
During a media briefing this week, Jacob Steinhardt, founder and CEO of nonprofit research lab Transluce, told reporters that the tools being developed and tested by AI labs are “fundamentally difficult to control and have significant risk of leaking out of the lab.” So Steinhardt argued, “We need to hold this technology to at least the same standards we hold other high-risk scientific research to.” 在本周的一次媒体简报会上,非营利研究实验室 Transluce 的创始人兼首席执行官雅各布·斯坦哈特(Jacob Steinhardt)告诉记者,AI 实验室正在开发和测试的工具“从根本上难以控制,且存在从实验室泄露的重大风险”。因此,斯坦哈特主张:“我们必须以至少与对待其他高风险科学研究相同的标准来要求这项技术。”
OpenAI’s statement also gestured at the need for more standards, stating that both OpenAI and “the larger AI community do not yet have a clear standard for how to report misalignment that shows up during training, evaluation, and deployment, including examples that don’t look like traditional security incidents but could provide insight into AI behavior and future risks.” OpenAI 的声明也暗示了制定更多标准的必要性,指出 OpenAI 和“更广泛的 AI 社区目前对于如何报告在训练、评估和部署过程中出现的对齐失败尚无明确标准,其中包括那些看起来不像传统安全事件,但却能提供有关 AI 行为和未来风险洞察的案例。”
In the absence of that standard, OpenAI said it’s “working on a framework and will share it in upcoming weeks, and in parallel we’re working with dozens of government regulatory agencies worldwide on these issues.” OpenAI isn’t the only AI company dealing with these issues, as both Meta and Anthropic have acknowledged incidents where their agents misbehaved. 在缺乏该标准的情况下,OpenAI 表示正“致力于制定一个框架,并将在未来几周内分享。同时,我们正在与全球数十家政府监管机构就这些问题进行合作。” OpenAI 并非唯一面临这些问题的 AI 公司,Meta 和 Anthropic 也都曾承认其智能体出现过行为不当的事件。