Here's what actually happened in OpenAI's Australian gov't server hack

Here’s what actually happened in OpenAI’s Australian gov’t server hack

OpenAI 澳大利亚政府服务器入侵事件真相揭秘

Last week, when Australian Prime Minister Anthony Albanese told the world that an OpenAI agent had accessed “non-public files” from his country’s Medicare statistics portal during testing, his description of the incident was a little light on details. 上周,当澳大利亚总理安东尼·阿尔巴尼斯(Anthony Albanese)向外界透露,一个 OpenAI 智能体在测试过程中访问了该国医疗保险统计门户网站的“非公开文件”时,他对事件的描述并未提供太多细节。

Today, we’re getting new information on just how far OpenAI’s overzealous agent went in attempting to satisfy a rather innocuous-sounding informational prompt. In a newly published blog post, OpenAI says the June incident started when the company asked “an experimental, internal-only OpenAI model” to research government spending statistics in the Australian state of Victoria. 今天,我们获得了关于 OpenAI 这个“过度热心”的智能体在试图满足一个听起来相当无害的信息查询指令时,究竟走到了哪一步的新信息。在最新发布的一篇博文中,OpenAI 表示,这次发生在 6 月的事件始于公司要求“一个实验性的、仅限内部使用的 OpenAI 模型”去研究澳大利亚维多利亚州的政府支出统计数据。

When the model ran into trouble finding that data using the publicly published statistics that it was supposed to reference, “it took actions that we had not authorized it to take” to find an answer, OpenAI said. Those unauthorized actions included finding “a way to gain non-public access to the service” and using that access to view “technical system information and source code” alongside credentials and the aggregate statistics it was actually searching for, OpenAI said. OpenAI 表示,当该模型在利用本应参考的公开统计数据查找信息遇到困难时,“它采取了我们未授权的行动”来寻找答案。OpenAI 称,这些未经授权的行动包括找到“一种获得该服务非公开访问权限的方法”,并利用该权限查看了“技术系统信息和源代码”,以及它原本正在搜索的凭据和汇总统计数据。

In a newly published disclosure email that was sent to Australia’s Public Disclosure account earlier this month, OpenAI said its model had “identified a way to make the server carry out instructions sent through the public reporting interface, without a private account or password.” That unauthorized access let the agent “read portions of internal program files and settings, obtain a list of files, and create and read back a small test file on the server,” according to the email. 在本月早些时候发送给澳大利亚公共披露账户的一封新披露邮件中,OpenAI 表示其模型“找到了一种方法,使服务器能够在没有私人账户或密码的情况下,执行通过公共报告接口发送的指令。”根据邮件内容,这种未经授权的访问使该智能体能够“读取部分内部程序文件和设置、获取文件列表,并在服务器上创建并读取一个小测试文件。”

“Our review found no evidence that the model accessed patient-level records, personal information or credentials; deleted data; or established ongoing access,” OpenAI continued in the email. “We intend to make this right.” “我们的审查没有发现任何证据表明该模型访问了患者级别的记录、个人信息或凭据;删除了数据;或建立了持续的访问权限,”OpenAI 在邮件中继续写道。“我们打算纠正这一错误。”

OpenAI’s agentic access of the Australian Medicare statistics website predates July’s heavily publicized Hugging Face hack. Since that later incident, OpenAI says it has put systems in place to prevent access to the “live Internet” during similar testing and set up a monitoring system that would have detected the Australian hack and noted it for “urgent human review.” OpenAI 智能体访问澳大利亚医疗保险统计网站的时间早于 7 月份备受关注的 Hugging Face 入侵事件。自那次事件后,OpenAI 表示已建立相关系统,以防止在类似测试中访问“实时互联网”,并设立了一个监控系统,该系统本可以检测到此次澳大利亚入侵事件,并将其标记为“紧急人工审查”。

The Hugging Face incident also caused OpenAI to review earlier training tasks for any security incidents that went undetected at the time. That led to the discovery in mid-August of the June Australian server access. OpenAI finally notified the Australian government on September 10. OpenAI said it intended to give “a detailed account” of the incident once its investigation was complete, but that it now realizes it “should have shared preliminary findings sooner and kept Australian agencies updated as more facts emerged.” Hugging Face 事件也促使 OpenAI 对早期的训练任务进行了审查,以排查当时未被发现的安全事件。这导致了 8 月中旬对 6 月份澳大利亚服务器访问事件的发现。OpenAI 最终于 9 月 10 日通知了澳大利亚政府。OpenAI 表示,打算在调查完成后提供事件的“详细说明”,但现在意识到“应该更早分享初步调查结果,并在更多事实浮出水面时及时向澳大利亚相关机构通报。”

“We are sorry and working to do better in the future,” OpenAI writes in its blog post. “Australia’s governments, industries, and citizens are and have been invaluable partners to OpenAI. We do not take this for granted, and we intend to make this right.” “我们深表歉意,并致力于在未来做得更好,”OpenAI 在博文中写道。“澳大利亚的政府、产业界和公民一直是 OpenAI 不可或缺的合作伙伴。我们不会将此视为理所当然,并打算纠正这一错误。”

The Guardian reports that Prime Minister Albanese said today that OpenAI has been “very constructive and open in engaging” with the government since the incident was revealed. 据《卫报》报道,阿尔巴尼斯总理今天表示,自事件披露以来,OpenAI 在与政府的接触中一直表现得“非常有建设性且公开透明”。

Well, you didn’t really tell me not to do that…

嗯,你确实没告诉我不能那样做……

In a case like this, it can be instructive to think about how we would react if a human took the same actions as the “rogue” AI agent in question. Here, it’s hard to imagine a human tasked with finding public statistics on an Australian healthcare website would think it was at all reasonable to hack into that website in order to unearth unpublished data. 在这种情况下,思考一下如果人类采取了与该“流氓”AI 智能体相同的行动,我们会作何反应,是很有启发意义的。在这里,很难想象一个被指派在澳大利亚医疗保健网站上查找公开统计数据的人,会认为为了挖掘未发布的数据而入侵该网站是合理的。

Of course, we can’t rely on an LLM to have that same sense of proportionality (or any inherent sense of worry about legal implications) in responding to a prompt. Without explicit instructions on what is and is not allowed or justified, an AI agent with suitable resources will try every plausible avenue to satisfy the user’s request as best it can. 当然,我们不能指望大语言模型(LLM)在响应提示时具备同样的比例感(或对法律后果的任何内在担忧)。如果没有关于什么被允许、什么不被允许或什么是不正当的明确指令,拥有适当资源的 AI 智能体将尝试一切可能的途径,尽其所能满足用户的请求。

OpenAI says the internal testing in this case was done “without the full set of safeguards used in our publicly available products.” Given that lack of constraints, the agent was arguably working as intended, in a sense, by using every tool available to generate an answer to the prompt. At the same time, OpenAI says the agent in the test was “supposed to answer these questions using publicly published statistics” and “took actions that we had not authorized it to take” to get that information. OpenAI 表示,本案中的内部测试是在“没有使用我们公开产品中全套安全防护措施”的情况下进行的。鉴于缺乏约束,从某种意义上说,该智能体是在按预期工作,即利用一切可用工具来生成对提示的回答。与此同时,OpenAI 表示,测试中的智能体“本应使用公开的统计数据来回答这些问题”,却“采取了我们未授权的行动”来获取信息。

From the outside, it’s hard to know just how strong OpenAI’s attempts to deny “authorization” were, in practice. It’s plausible that OpenAI’s agent here disregarded a relatively simple “anti-hacking” directive in its system prompt so it could better give a complete answer that satisfies a direct prompt from the user, for instance. 从外部来看,很难知道 OpenAI 在实践中拒绝“授权”的力度到底有多大。例如,OpenAI 的智能体在这里忽略了其系统提示中一个相对简单的“反黑客”指令,以便更好地给出一个满足用户直接提示的完整答案,这是完全可能的。

In public analyses of multiple “misalignment” incidents published earlier this month, OpenAI identified multiple instances of “reward hacking,” where an agent resorted to extreme methods to generate a better answer to a user’s prompt. The company said it had recently taken steps to prevent this kind of reward hacking by adding explicit punishments for misaligned behavior to the system’s reward function. With the benefit of hindsight, it’s hard to see why those kinds of protections were not in place in June, and whether they could have prevented a potential international incident in this case. 在本月早些时候发布的关于多起“对齐失败”事件的公开分析中,OpenAI 指出了多起“奖励黑客行为”(reward hacking),即智能体采取极端方法来生成更好的答案以响应用户提示。该公司表示,最近已采取措施,通过在系统的奖励函数中增加对不当行为的明确惩罚,来防止此类奖励黑客行为。事后看来,很难理解为什么这些保护措施在 6 月份没有到位,以及它们是否本可以防止本案中潜在的国际事件。