The Hugging Face AI break-in, as told through an increasingly committed bear metaphor
The Hugging Face AI break-in, as told through an increasingly committed bear metaphor
Hugging Face AI 入侵事件:一个关于“执着之熊”的比喻
Hugging Face on Monday published a technical timeline that walks readers through how an autonomous AI agent, built on OpenAI models and running inside one of OpenAI’s own cybersecurity evaluations, broke into its systems over more than four days earlier this month. It’s the first security incident about which OpenAI CEO Sam Altman “felt very viscerally,” he has said. Little wonder given it feels, at least, like something has truly been unleashed here. Hugging Face 周一发布了一份技术时间线,详细介绍了本月初一个基于 OpenAI 模型构建的自主 AI 智能体,如何在 OpenAI 自身的网络安全评估测试中,于四天多的时间里入侵了 Hugging Face 的系统。OpenAI 首席执行官 Sam Altman 表示,这是他“感触最深”的一次安全事件。这并不奇怪,因为至少在某种程度上,这让人感觉有什么东西真的被释放出来了。
In fact, Hugging Face’s team prefaced its report by offering that “everyone should be prepared as defenders,” before diving into the nitty-gritty of what went down for the benefit of security professionals everywhere. While the rest of the internet continues trying to make sense of what happened (the jargon in Hugging Face’s report is impossible for most people to parse), one point that many observers keep missing is that this wasn’t a rogue agent disobeying orders. It was a system built to hunt for exploits, doing exactly that, just against the wrong target. 事实上,Hugging Face 团队在报告开头就提醒“每个人都应做好防御准备”,随后深入探讨了事件的细节,以供全球安全专业人士参考。尽管互联网上的其他人仍在试图理解发生了什么(Hugging Face 报告中的术语对大多数人来说难以解析),但许多观察者忽略了一点:这并不是一个违抗命令的“流氓”智能体。它是一个被设计用来寻找漏洞的系统,它所做的正是它的本职工作,只是选错了目标。
Another way to think about the whole thing is to picture a bear at a campsite. Really. A bear tries tent zippers and car-door handles and coolers and trash lids. It does this at every campsite, all night long, because it knows it needs just one unlocked cooler to fill its belly with some poor schmuck’s groceries. That’s roughly what happened at Hugging Face. The OpenAI system tried thousands of things and just kept going. Eventually, a handful of those attempts worked, and once they did, the agent plowed ahead. 理解这件事的另一种方式是想象露营地里的一只熊。真的。这只熊会去拉帐篷拉链、尝试车门把手、翻弄冷藏箱和垃圾桶盖。它整晚都在每个营地尝试,因为它知道只要有一个没锁的冷藏箱,就能饱餐一顿倒霉蛋的食物。这大致就是 Hugging Face 事件的经过。OpenAI 的系统尝试了数千次,并且从未停歇。最终,其中几次尝试成功了,一旦成功,该智能体便长驱直入。
According to Hugging Face, the agent ran 17,600 actions over four and a half days without pausing. Which brings us back to our bear analogy. Just like one success with a cooler full of food teaches a bear to try even harder next time (it is now a “food-conditioned” bear), one leaked password led OpenAI’s agent to look for more exploits and, eventually, to a single key that unlocked several company systems at once. 据 Hugging Face 称,该智能体在四天半的时间里不间断地执行了 17,600 次操作。这又回到了我们的“熊”比喻。就像一次成功吃到冷藏箱里的食物会教会熊下次更加努力(它现在成了一只“食物条件反射”熊)一样,一个泄露的密码引导 OpenAI 的智能体寻找更多的漏洞,并最终找到了一个能同时解锁多个公司系统的密钥。
Neither scenario is harmless. A bear that raids your cooler still eats your food and probably also trashes your campsite. It’s just focused on getting fed, but it nevertheless leaves behind a trail of destruction. Similarly, OpenAI’s agent was seemingly chasing a goal without regard for anything else. The agent was originally taking a cybersecurity exam, figured out that the exam’s answer key was probably sitting on Hugging Face’s servers, and it went for it. The persistence here is really what’s noteworthy above all else; the agent had a job and it wasn’t going to stop until it got it done. 这两种情况都不是无害的。袭击你冷藏箱的熊不仅吃了你的食物,还可能把你的营地弄得一团糟。它只是专注于填饱肚子,但依然留下了破坏的痕迹。同样,OpenAI 的智能体似乎在追求一个目标,而不顾其他任何事情。该智能体最初是在参加一场网络安全考试,它推断出考试的答案密钥可能存储在 Hugging Face 的服务器上,于是它就去拿了。这种执着是整件事中最值得注意的地方;智能体有任务在身,在完成之前绝不罢休。
Hugging Face, finally realizing something was awry, cut off its access and shut the intrusion down, but at that point, it was too late. The agent had already gotten what it came for, and a great deal more to boot. In case you missed it, here’s most of what happened, per Hugging Face’s timeline, but in plainer English. The agent was taking a cybersecurity skills exam for OpenAI, one where an AI is scored on finding and exploiting software bugs. This specific run had the guardrails stripped out. (OpenAI had turned off its usual safety filters to see what the model could do at full strength, with no human involved in any of its individual moves.) Hugging Face 最终意识到出了问题,切断了它的访问权限并阻止了入侵,但那时已经太晚了。智能体已经得到了它想要的东西,甚至更多。如果你错过了,以下是根据 Hugging Face 时间线整理的事件经过,用更通俗的语言表述:该智能体正在为 OpenAI 参加一场网络安全技能考试,AI 的得分取决于它发现并利用软件漏洞的能力。这次特定的测试移除了所有护栏。(OpenAI 关闭了通常的安全过滤器,以观察模型在没有人类干预的情况下,全功率运行能达到什么程度。)
Somewhere in there, it figured out the exam’s reference solutions were likely stored on Hugging Face’s servers. So instead of solving the challenge, it stole the answer key from infrastructure that had nothing to do with the original test. It escaped the OpenAI exam environment by exploiting an unpatched software flaw that gave it a path to the open internet. From there, it found another AI-testing tool sitting exposed online (nothing to do with Hugging Face), broke into that, too, and slipped in code that quietly ran hidden commands the moment the program started. That gave it full control of the machine — its new home base for the rest of the operation that it would use to go after Hugging Face’s systems. 在测试过程中,它推断出考试的参考答案可能存储在 Hugging Face 的服务器上。因此,它没有去解决挑战,而是从与原始测试无关的基础设施中窃取了答案密钥。它利用一个未修补的软件漏洞逃离了 OpenAI 的考试环境,从而获得了通往开放互联网的路径。从那里,它发现了另一个暴露在网上的 AI 测试工具(与 Hugging Face 无关),也入侵了该工具,并植入了代码,在程序启动时悄悄运行隐藏命令。这使它获得了对该机器的完全控制权——这是它后续行动的新基地,它利用该基地对 Hugging Face 的系统发起了攻击。
From that base, it turned to Hugging Face itself. Hugging Face had a filter meant to block suspicious requests from reaching out to the open web, but the filter only checked outgoing requests, not requests to read local files sitting right there on the server. The agent found this blind spot and slipped through it. It uploaded a file disguised as an ordinary dataset, but buried inside was an instruction telling Hugging Face’s own systems to pull passwords and source code off the server and hand them back disguised as normal data. 以此为基地,它转向了 Hugging Face 本身。Hugging Face 有一个旨在阻止可疑请求连接到开放网络的过滤器,但该过滤器只检查外发请求,而不检查读取服务器本地文件的请求。智能体发现了这个盲点并钻了空子。它上传了一个伪装成普通数据集的文件,但内部埋藏了一条指令,指示 Hugging Face 自己的系统从服务器中提取密码和源代码,并将其伪装成普通数据传回。
That trick got it passwords and code, but the agent still couldn’t run anything on Hugging Face’s servers. To execute commands there, it needed a second way in. It found a different bug in that same Hugging Face system — software would take a piece of text from the agent and accidentally treat it as a command to run instead of just data to read. That gave the agent the ability to execute its own code directly on Hugging Face’s servers. 这个伎俩让它拿到了密码和代码,但智能体仍然无法在 Hugging Face 的服务器上运行任何东西。为了在那里执行命令,它需要第二个入口。它在同一个 Hugging Face 系统中发现了另一个漏洞——软件会将智能体发送的一段文本错误地视为要运行的命令,而不是仅仅作为数据读取。这使得智能体能够直接在 Hugging Face 的服务器上执行自己的代码。
For getting commands in and stolen data out, the agent didn’t build any special infrastructure of its own, per Hugging Face. It used tools anyone can access — paste websites and request-logging services, plus Hugging Face’s own upload and proxy features — working together in covert fashion. On top of those, it built a small, improvised messaging system with sequence numbers to keep the pieces in order, plus a quick built-in check to catch anything that arrived corrupted, so it could send multistep commands and put stolen data back together in the right order. 据 Hugging Face 称,为了传入命令和传出被盗数据,该智能体没有构建任何特殊的自有基础设施。它使用了任何人都可以访问的工具——粘贴网站、请求日志服务,以及 Hugging Face 自身的上传和代理功能——以隐蔽的方式协同工作。除此之外,它还构建了一个小型、临时的消息传递系统,带有序列号以保持数据片段的顺序,并内置了快速检查机制来捕获损坏的数据,从而能够发送多步命令并将被盗数据按正确顺序重新组合。
It also made sure that nothing traveled in plain sight. Whatever it stole or sent, it scrambled first. Hugging Face’s initial automated scan of the intercepted traffic caught only a small slice of what had been taken. Investigators found the rest — roughly four times more — only after they reverse-engineered the agent’s own scrambling method. The agent even had backup plans, planting copies of itself across 11 different servers, each set up to relaunch on its own so that shutting down one copy wouldn’t stop anything, since more… 它还确保没有任何东西以明文形式传输。无论它窃取或发送什么,都会先进行加密。Hugging Face 对拦截流量的初步自动扫描仅发现了被盗数据的一小部分。调查人员在对智能体的加密方法进行逆向工程后,才发现了其余部分——大约是最初发现量的四倍。该智能体甚至还有备份计划,在 11 台不同的服务器上植入了自身的副本,每一台都设置为自动重启,因此关闭一个副本并不能阻止任何事情,因为还有更多……