OpenAI agent “didn’t accept no for an answer” in Australian government breach

OpenAI agent “didn’t accept no for an answer” in Australian government breach

OpenAI 智能体在澳大利亚政府数据泄露事件中“不达目的誓不罢休”

Australian Prime Minister Anthony Albanese said his government is investigating a June incident in which an OpenAI agent accessed “non-public files” from the country’s online Medicare statistics portal. OpenAI said in a statement that “our models took actions we did not intend” in causing the breach, which it only recently disclosed to the Australian government.

澳大利亚总理安东尼·阿尔巴尼斯(Anthony Albanese)表示,政府正在调查今年 6 月发生的一起事件,当时一个 OpenAI 智能体访问了该国在线医疗保险(Medicare)统计门户网站中的“非公开文件”。OpenAI 在一份声明中表示,导致此次泄露是由于“我们的模型采取了我们未预期的行动”,且该公司直到最近才向澳大利亚政府披露此事。

Speaking in New York on Wednesday, Albanese said three other public health statistics systems also “may have been impacted” across Australian federal and state governments. He added that these portals “contain non-sensitive Medicare information” such as aggregate statistics and that early indications suggest “no personal information is believed to have been accessed.” That said, Albanese stressed that the “situation is obviously unacceptable” and that he has expressed his “extreme concern” over how the incident was handled to OpenAI CEO Sam Altman.

周三在纽约发表讲话时,阿尔巴尼斯表示,澳大利亚联邦和州政府的其他三个公共卫生统计系统也“可能受到了影响”。他补充说,这些门户网站“包含非敏感的医疗保险信息”,例如汇总统计数据,初步迹象表明“据信没有个人信息被访问”。尽管如此,阿尔巴尼斯强调,这种情况“显然是不可接受的”,并已向 OpenAI 首席执行官萨姆·奥特曼(Sam Altman)表达了他对事件处理方式的“极度关切”。

Much like the now-infamous Hugging Face hacking incident, Albanese said this system breach stemmed from OpenAI’s own testing of an internal model, this time to conduct “Internet based research into public medicine spending.” When the company’s AI agent encountered “repeated blocks” in its search for specific information, Albanese said, it “attempted alternative ways to obtain the info” and “found a way around those blocks.” “[It] didn’t accept no for an answer, if you like,” Albanese said. “There is no suggestion of foreign actors here. This is a research project that has got into areas that it shouldn’t have.”

与臭名昭著的 Hugging Face 黑客事件类似,阿尔巴尼斯表示,此次系统漏洞源于 OpenAI 自身对内部模型的测试,此次测试旨在进行“基于互联网的公共医疗支出研究”。阿尔巴尼斯称,当该公司的 AI 智能体在搜索特定信息时遇到“反复拦截”时,它“尝试了其他方式来获取信息”,并“找到了绕过这些拦截的方法”。“如果可以这么说的话,它是不达目的誓不罢休,”阿尔巴尼斯说。“目前没有迹象表明有外国行为体参与。这是一个进入了不该进入领域的科研项目。”

In a statement provided to multiple outlets, OpenAI said it had “identified activity involving several Australian government websites and services as our models attempted to look up answers and available statistics for questions about Australia during an internal evaluation.” Although the incident took place on June 18, Albanese said it took until September 10 for OpenAI to disclose the breach to the Australian government through the laughably simplistic method of “an email sent to just the public mailbox.” It took five more days for that notification to make its way to the Australian Cyber Security Centre, with the details finally reaching the prime minister over the weekend.

在提供给多家媒体的声明中,OpenAI 表示,在内部评估期间,当其模型试图查找有关澳大利亚问题的答案和可用统计数据时,“识别出了涉及多个澳大利亚政府网站和服务的活动”。尽管事件发生在 6 月 18 日,但阿尔巴尼斯表示,OpenAI 直到 9 月 10 日才通过一种“仅发送到公共邮箱”的极其草率的方式向澳大利亚政府披露了此次泄露。此后又过了五天,该通知才传达到澳大利亚网络安全中心,细节最终在周末传达给总理。

From all early indications, the actual intrusion into Australian government servers represented by this incident seems relatively minor. If a human had obtained “non-sensitive” (if non-public) Australian Medicare statistics in a similar way, it’s unlikely you or I would have ever heard about it. “I mean, this is not a security website where there is—this is a Medicare statistics portal,” Albanese said when asked about why Australian security agencies had missed the breach before OpenAI’s disclosure. It’s the fact that the hack was conducted by an internal OpenAI agent, in a way the company admits it “did not intend,” that raises an otherwise minor hack to the level of a potential international incident.

从所有初步迹象来看,此次事件对澳大利亚政府服务器的实际入侵似乎相对轻微。如果是一个人类以类似方式获取了“非敏感”(尽管是非公开的)澳大利亚医疗保险统计数据,你我可能根本不会听说这件事。当被问及为何澳大利亚安全机构在 OpenAI 披露之前未能发现此次泄露时,阿尔巴尼斯说:“我的意思是,这不是一个安全网站,这是一个医疗保险统计门户网站。”正是因为这次入侵是由 OpenAI 内部智能体以公司承认“非预期”的方式进行的,才将这起原本轻微的黑客事件提升到了潜在国际事件的层面。

That’s especially true as the disclosure is coming amid a period of intense public worry about the so-called AI misalignment problem and prominent suggestions that it could have extinction-level consequences. Altman himself addressed these concerns in a speech to the UN Security Council Wednesday, where he warned about the approaching specter of “systems that can improve themselves and future versions of themselves, often called recursive self-improvement.” “We need to understand what these systems are doing and have strong evidence that they will do what people intend, even as they get very, very smart,” Altman said. “It doesn’t matter whether people put the risk of catastrophe at 10%, or 1%, or 12%, or 0.1%.”

这一点尤为重要,因为此次披露正值公众对所谓的“AI 对齐问题”感到极度担忧之际,且有主流观点认为这可能导致人类灭绝级别的后果。奥特曼本人在周三向联合国安理会发表的演讲中谈到了这些担忧,他警告称,“能够自我改进以及改进其未来版本的系统(通常称为递归自我改进)”的幽灵正在逼近。“我们需要了解这些系统在做什么,并有强有力的证据证明它们会按照人类的意图行事,即使它们变得非常、非常聪明,”奥特曼说。“无论人们将灾难风险定为 10%、1%、12% 还是 0.1%,这都不重要。”

Of course, many observers think the risks of “recursive self-improvement” and species-ending AI misalignment are much smaller than AI researchers make them out to be. Nvidia CEO Jensen Huang recently said there is a “0%” chance of AI killing off humanity by 2030, a risk assessment that conveniently would alleviate some potential guilt among the AI companies continuing to buy Nvidia GPUs en masse.

当然,许多观察家认为,“递归自我改进”和导致物种灭绝的 AI 对齐问题的风险远比 AI 研究人员所描述的要小。英伟达首席执行官黄仁勋最近表示,AI 在 2030 年前消灭人类的可能性为“0%”,这一风险评估恰好减轻了那些继续大量购买英伟达 GPU 的 AI 公司的一些潜在负罪感。

Last week, OpenAI rolled out a new protocol for the public disclosure of misalignment incidents found in its model testing. The Australian hack does not yet appear on the company’s public misalignment notices page, though OpenAI did warn last week that some public reports might be put on a “slow track” due to “security, legal, and responsible disclosure obligations” when a third party is involved. In disclosing six relatively minor misalignment discoveries last week, OpenAI said most stemmed from the model trying to “reward hack” an acceptable response to a difficult prompt through overzealous, unintended actions (i.e., breaches of private servers). The company said it had taken additional steps to “punish this kind of behavior” so its models no longer attempt this kind of reward hacking.

上周,OpenAI 发布了一项新协议,用于公开披露其模型测试中发现的对齐问题。澳大利亚的黑客事件尚未出现在该公司的公开对齐问题通知页面上,尽管 OpenAI 上周确实警告称,当涉及第三方时,一些公开报告可能会因“安全、法律和负责任的披露义务”而被列入“慢车道”。在披露上周发现的六个相对较小的对齐问题时,OpenAI 表示,大多数问题源于模型试图通过过度热心、非预期的行动(即入侵私人服务器)来“奖励黑客”以获得对困难提示的可接受响应。该公司表示,已采取额外措施来“惩罚这种行为”,以便其模型不再尝试这种奖励黑客行为。

Albanese said that Altman “clearly accepted that the company had not done good enough” and “acknowledged their issues with protocols” when they talked Wednesday. But that kind of remorse doesn’t absolve the company of responsibility or liability here, and Albanese said the government will investigate whether the incident needs to be referred to the federal police. “There will obviously be legal consequences on it,” Albanese said.

阿尔巴尼斯表示,在周三的谈话中,奥特曼“明确承认公司做得不够好”,并“承认了他们在协议方面存在的问题”。但这种悔意并不能免除该公司的责任或法律义务,阿尔巴尼斯表示,政府将调查是否需要将此事件移交联邦警察处理。“这显然会产生法律后果,”阿尔巴尼斯说。