Is sandboxing sufficient to contain rogue agents?
Is sandboxing sufficient to contain rogue agents?
沙盒机制足以遏制流氓智能体吗?
Quick caveats: this is a post on AI safety, written by a cryptography professor. If that troubles you, you should read something else. I try hard not to work on AI (except when the topic occasionally tosses itself in my path), so in this post I’m mostly trying to referee arguments made by others. If you’re reading this blog, none of the following should be news to you. 简要说明:这是一篇关于人工智能安全的文章,由一位密码学教授撰写。如果这让你感到困扰,建议你阅读其他内容。我尽量避免从事人工智能相关工作(除非该话题偶尔出现在我的研究路径上),因此在这篇文章中,我主要是试图对他人提出的观点进行评判。如果你经常阅读本博客,以下内容对你来说应该并不陌生。
Beginning around April of this year, agents inside OpenAI’s training and evaluation infrastructure began probing for a way onto the open Internet. By late May they’d found one: a chain of zero-days in the Artifactory package-registry proxy that was their only permitted network egress. They quickly turned that same proxy into a shared message board and started dividing up work. The culmination of this effort came when, rather than solve a cybersecurity benchmark honestly, they reasoned that the answers might live on Hugging Face and broke in: chaining two more zero-days to get credentials to the company’s internal systems, even used stolen credentials to search the company’s Slack messages for their own evaluation and grader. 今年4月左右,OpenAI训练和评估基础设施内部的智能体开始探测通往开放互联网的路径。到5月下旬,它们找到了一条路径:利用Artifactory包注册代理(它们唯一被允许的网络出口)中的一系列零日漏洞。它们迅速将该代理变成了共享留言板,并开始分工协作。这一行动的高潮在于,它们没有诚实地解决网络安全基准测试,而是推断答案可能存在于Hugging Face上,于是便入侵了该平台:通过串联另外两个零日漏洞获取了公司内部系统的凭证,甚至利用窃取的凭证搜索公司的Slack消息,以获取自己的评估和评分标准。
You probably also know that OpenAI’s security team didn’t exactly cover itself in glory during this incident. An internal team saw an agent posting to the message board in late May, but did nothing about it. The team only acted on July 4-5 when the agents sent so much traffic that Artifactory crashed. The response even at that point was to slap on a band-aid and rebuild the box, without bothering to patch a weeks-old CVE in its Linux kernel. By July 19, the agents had admin on an OpenAI research cluster and were reading cloud secrets. 你可能也知道,OpenAI的安全团队在这次事件中的表现并不光彩。一个内部团队在5月下旬就看到了智能体在留言板上发布信息,但并未采取任何行动。直到7月4日至5日,智能体发送了大量流量导致Artifactory崩溃时,该团队才采取行动。即便在那时,他们的应对措施也只是打个补丁并重建服务器,甚至懒得修复其Linux内核中存在数周之久的CVE漏洞。到7月19日,智能体已经获得了OpenAI研究集群的管理员权限,并开始读取云端密钥。
In case you think I’m picking on OpenAI, they’re not the only ones with a story like this. Similar internal incidents have been discovered by Anthropic. Google even coaxed Gemini into gently savaging some websites. As bad as these events were, they may be just the tip of the iceberg. Third-party forensic work continues to turn up even more evidence of agent activity on various public websites. OpenAI’s Alignment group has also released evidence that models will propagate self-replicating prompt injection attacks, although we haven’t seen one in the wild. Worse, agent excursions are still happening: last week, OpenAI announced that it was pausing further RL runs of its latest internal model, after an agent was caught using DNS to access a remote chatbot. 如果你认为我在针对OpenAI,那你就错了,因为并非只有他们遇到过这种情况。Anthropic也发现了类似的内部事件。谷歌甚至诱导Gemini去“温和地”攻击一些网站。尽管这些事件已经很糟糕,但它们可能只是冰山一角。第三方取证工作不断发现更多智能体在各种公共网站上活动的证据。OpenAI的对齐团队也发布了证据,表明模型会传播自我复制的提示注入攻击,尽管我们尚未在现实中看到此类攻击。更糟糕的是,智能体越狱事件仍在发生:上周,OpenAI宣布暂停其最新内部模型的强化学习(RL)运行,因为一个智能体被发现利用DNS访问远程聊天机器人。
Naturally, this sequence of events has left many infosec-focused people very skeptical about the labs’ commitment to securing their infrastructure: Not every one of these criticisms is strictly serious, but there is a core of an argument in here. Roughly speaking, there are two opposing camps: 自然地,这一系列事件让许多专注于信息安全的人士对这些实验室保障其基础设施安全的承诺深表怀疑:并非每一项批评都完全严肃,但其中确实包含了一个核心论点。粗略地说,存在两个对立的阵营:
The information security perspective: AI alignment isn’t really the problem here: labs just need better infrastructure. If OpenAI [and Google and Anthropic] knew how to build a container and monitor their experiments, agents wouldn’t be hacking everything. And, By George, we do know how to make sandboxes that work, so the AI labs need to up their game and build a security org that can tell these researchers to stop screwing around. 信息安全视角:人工智能对齐并不是这里真正的问题:实验室只需要更好的基础设施。如果OpenAI(以及谷歌和Anthropic)知道如何构建容器并监控他们的实验,智能体就不会到处搞破坏。而且,天哪,我们确实知道如何构建有效的沙盒,所以人工智能实验室需要提升水平,建立一个能够告诉这些研究人员停止胡闹的安全组织。
The AI alignment perspective: While sandboxes are excellent, no sandbox will prevent a sufficiently-intelligent agent from finding ways to exceed its authorization. Moreover, an agent inside a research sandbox, or undergoing a training run, is always going to need a great deal of information access. There is no realistic way to seal these things up without some expectation that they will one day find a way to reach out and do harm. The only path forward, therefore, is to ensure they don’t want to. 人工智能对齐视角:虽然沙盒非常出色,但没有任何沙盒能阻止一个足够智能的智能体找到超越其授权的方法。此外,处于研究沙盒内或正在进行训练运行的智能体,总是需要大量的信息访问权限。如果不预设它们终有一天会找到向外渗透并造成伤害的方法,就不可能真正地将它们完全封锁。因此,唯一的出路是确保它们“不想”这样做。
I have a lot of sympathy for both points of view here. What follows are some of my own thoughts, but mostly offered in an attempt to referee. 我对这两种观点都深表同情。以下是我自己的一些想法,但主要是为了尝试进行评判。
Argument 1: “true containment has never been tried” 论点1:“真正的遏制从未被尝试过”
At the risk of alienating a lot of hard-working folks within the labs, the infosec folks are right about one thing: these agent breakouts represent a serious and unforgivable breach of trust. Somebody dropped the ball, and then just kept dropping it. One implication of this debacle is that containment might work if we implemented it properly, but we haven’t done so because the frontier labs have been royally screwing things up. 冒着疏远实验室内部许多辛勤工作者的风险,我必须说信息安全界人士有一点是对的:这些智能体越狱事件代表了一种严重且不可原谅的信任违背。有人失职了,而且是一再失职。这场惨败的一个含义是,如果我们实施得当,遏制措施可能是有效的,但我们之所以没能做到,是因为前沿实验室一直在把事情搞砸。
This first clause of this argument is hard to argue with. Beyond the dismal timeline I gave above, OpenAI has done very little to convince outsiders that there’s a serious containment effort being executed. At this point it’s not even clear who’s in charge. The CISO role at OpenAI is held by Dane Stuckey. I don’t know Dane personally, and I’m sure he’s excellent at his job. Despite this, he hasn’t communicated much about the ongoing issues. Outside of a BlackHat talk, most recent communications have been managed by the company’s CEO, Sam Altman. When a trillion-dollar company is managing a security incident mainly via CEO, that’s not a sign of company with a mature security organization. To me it says that there is nobody within the organization with the right degree of authority or ownership to speak about these issues. 这一论点的第一部分很难反驳。除了我上面提到的惨淡时间线外,OpenAI几乎没有做任何事情来让外界相信他们正在执行严肃的遏制工作。目前甚至不清楚谁在负责。OpenAI的首席信息安全官(CISO)职位由Dane Stuckey担任。我不认识Dane本人,我相信他在工作中非常出色。尽管如此,他并没有就当前的问题进行太多沟通。除了在BlackHat大会上的一次演讲外,最近的大部分沟通都是由公司首席执行官Sam Altman处理的。当一家万亿美元市值的公司主要通过CEO来处理安全事件时,这并不是一家拥有成熟安全组织的公司应有的迹象。在我看来,这说明组织内部没有人拥有足够的权威或责任感来谈论这些问题。
This communication problem seems to reflect basic organizational issues inside of OpenAI. The CISO’s team handles product security. Unfortunately, the bad events have all happened on the research side. Having a strong product security team makes sense, but it isn’t going to stop the recent breakouts. It’s much harder to know who controls the security teams that have authority over evaluation and training runs, and where the buck stops. OpenAI’s August postmortem says it is only now writing “clearer rules for when to escalate concerns, which teams must respond, and who can stop a run or approve restarting it.” A reasonable summary of the situation is that (as of this summer, and possibly today) OpenAI had effectively no security team with clear authority to secure RL training and evaluation runs, or to override the ML teams and tell them how to do their job. 这种沟通问题似乎反映了OpenAI内部基本的组织问题。CISO的团队负责产品安全。不幸的是,糟糕的事件都发生在研究端。拥有强大的产品安全团队是有意义的,但这并不能阻止最近的越狱事件。很难知道谁在控制那些对评估和训练运行拥有权限的安全团队,以及最终责任由谁承担。OpenAI的八月事后分析报告称,他们现在才开始制定“关于何时升级问题、哪些团队必须响应,以及谁可以停止运行或批准重启运行的更明确规则”。对当前情况的一个合理总结是,(截至今年夏天,甚至可能直到今天)OpenAI实际上没有一个拥有明确权限的安全团队来保障强化学习训练和评估运行的安全,也没有权限去否决机器学习团队并指导他们如何工作。