AI labs want in-house auditors — but maybe they should shut the front door first
AI labs want in-house auditors — but maybe they should shut the front door first
AI 实验室想要内部审计员——但也许他们应该先关好大门
Last weekend, after one of his researchers resigned over fears that AI could lead to human extinction, Anthropic CEO Dario Amodei wrote about the need for outside organizations “to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes.” 上周末,在一名研究人员因担心人工智能可能导致人类灭绝而辞职后,Anthropic 首席执行官达里奥·阿莫代(Dario Amodei)撰文指出,需要外部组织“来验证对安全实践和承诺的遵守情况,报告事故,并帮助评估不仅是已完成的人工智能模型,还包括训练流程和过程的对齐情况。”
Executives at OpenAI, Google, and SpaceXAI have already rallied around Amodei’s plan, which has quickly become a central pillar of the emerging AI safety push. But there may be a simpler and more effective fix hiding in plain sight. Internet security experts say the labs need to focus on network security basics like logs and permissions, applying the same rigorous defenses they do for human users. It’s not as exciting as third-party auditing and alignment work — but it may end up being more effective. OpenAI、谷歌和 SpaceXAI 的高管们已经支持阿莫代这一计划,该计划已迅速成为新兴人工智能安全推动工作的核心支柱。但可能有一个更简单、更有效的解决方案就在眼前。互联网安全专家表示,这些实验室需要专注于网络安全基础知识,如日志和权限,并对人工智能应用与人类用户相同的严格防御措施。这虽然不如第三方审计和对齐工作那样令人兴奋,但最终可能更有效。
“To me, it seems like they’re outsourcing,” Katie Moussouris, the CEO of Luta Security, told TechCrunch of Amodei’s proposal. “Saying [a third-party audit] is the solution is a strange proposition from my perspective. It would be the same as if, instead of writing the Trustworthy Computing Memo, Microsoft said, let’s slow down development.” “在我看来,他们是在外包,”Luta Security 首席执行官凯蒂·穆苏里斯(Katie Moussouris)在谈到阿莫代的提议时告诉 TechCrunch。“从我的角度来看,说(第三方审计)是解决方案是一个奇怪的命题。这就好比微软当年没有撰写《可信计算备忘录》,而是说,让我们放慢开发速度。”
That memo, written by then-Microsoft CEO Bill Gates in 2002, called on his employees to ensure that their software would be reliable and safe following a series of widely publicized computer worms that took over then-nascent enterprise systems. The AI sector may be at a similar turning point, as the value and risk of the new technology becomes increasingly clear. 那份备忘录由时任微软首席执行官的比尔·盖茨于 2002 年撰写,呼吁员工确保软件的可靠性和安全性,此前一系列广为人知的计算机蠕虫病毒接管了当时尚处于萌芽阶段的企业系统。人工智能行业可能正处于类似的转折点,因为这项新技术的价值和风险正变得越来越清晰。
While alignment remains an important concern, Sayash Kapoor, an AI researcher who will be a professor at UC Berkeley starting next year, argues that “marginal investments in control are more likely to be effective compared to those in alignment. We view these incidents as illustrating the lack of emphasis on AI control within companies, despite the availability of known techniques.” 尽管对齐仍然是一个重要的关注点,但明年将成为加州大学伯克利分校教授的人工智能研究员萨亚什·卡普尔(Sayash Kapoor)认为,“与对齐方面的投资相比,在控制方面的边际投资更有可能有效。我们认为这些事件表明,尽管已知技术可用,但公司内部对人工智能控制的重视程度不足。”
The incidents that have spurred these concerns revolve around frontier models being asked to complete training tasks, usually cybersecurity evaluations, and then accessing the open internet and penetrating closed third-party systems in an attempt to do so. They usually did so because of poorly configured “sandbox” environments that are supposed to contain these agents; ironically, one Anthropic break-out happened because third-party evaluators didn’t close the right doors. 引发这些担忧的事件围绕着前沿模型被要求完成训练任务(通常是网络安全评估),然后访问开放互联网并试图渗透封闭的第三方系统。它们通常这样做是因为配置不当的“沙盒”环境本应限制这些智能体;讽刺的是,Anthropic 的一次“越狱”事件发生,是因为第三方评估人员没有关好正确的门。
“We as a profession know how to block access to the internet,” Avery Pennarun, the CEO of Tailscale, a security company, said. “If you read through all these big long [reports] — ‘wow, that was a very impressive multi-stage attack, blah, blah.’ Look, you gave it access to download stuff. You should have not done that separately from the internet.” That’s one problem — but a bigger problem is that frontier labs were unaware of these activities. “作为这一行业的从业者,我们知道如何阻止对互联网的访问,”安全公司 Tailscale 的首席执行官艾弗里·彭纳伦(Avery Pennarun)说。“如果你读完所有这些长篇大论的报告——‘哇,那是一次非常令人印象深刻的多阶段攻击,等等。’听着,你给了它下载东西的权限。你不应该在没有与互联网隔离的情况下这样做。”这是一个问题,但更大的问题是,前沿实验室对这些活动一无所知。
Eyes on agents 关注智能体
“What was really profound was that all of the discoveries of what they were doing happened either because a victim saw something, or in some of the other cases … it was network activity, and none of it was actually from monitoring the AIs directly,” Moussouris points out. In one case, where OpenAI agents took over a defunct German WikiForum to cheat on evaluations, the agents were active for weeks before anyone at the company appeared to notice. “真正深刻的是,所有关于它们在做什么的发现,要么是因为受害者看到了什么,要么在其他一些情况下……是网络活动,而没有一个是真正来自直接监控人工智能的,”穆苏里斯指出。在一个案例中,OpenAI 的智能体接管了一个已废弃的德国 WikiForum 来在评估中作弊,这些智能体活跃了数周,公司里似乎才有人注意到。
Security experts that TechCrunch spoke to said that real-time monitoring is key to preventing future break-outs, and that every agentic session should be time-limited and expire. Shapor Naghibzadeh, a former Google security executive who now leads the startup QueryStory, says the solution is to “put the agent in a box and instrument it heavily from the outside looking in and watch everything that crosses the boundary. Every tool call, every process, every network connection, no exceptions. …The one hole you leave open for convenience is the one that gets used. The bypass went through exactly that kind of exception. [At Google,] I watched that movie many times with human attackers, and these models are at least as good at finding the propped-open door.” TechCrunch 采访的安全专家表示,实时监控是防止未来“越狱”的关键,并且每个智能体对话会话都应有时限并自动过期。前谷歌安全高管、现领导初创公司 QueryStory 的沙普尔·纳吉布扎德(Shapor Naghibzadeh)表示,解决方案是“把智能体关进盒子里,从外部进行严格的监测,观察所有跨越边界的行为。每一个工具调用、每一个进程、每一个网络连接,无一例外……你为了方便而留下的那个漏洞,就是会被利用的那个。绕过防御的行为正是通过这种例外发生的。(在谷歌时,)我多次看到人类攻击者上演这一幕,而这些模型在寻找敞开的大门方面至少和人类一样出色。”
OpenAI has begun moving in that direction, announcing that it had begun monitoring all tool-using inference by its Astra model, at “significant compute cost.” Anthropic, too, says it is hardening its security procedures, including expanding observability of its models. Neither company responded to TechCrunch’s questions about how they track and control AI agents. OpenAI 已经开始朝着这个方向努力,宣布已开始以“巨大的计算成本”监控其 Astra 模型的所有工具使用推理。Anthropic 也表示正在加强其安全程序,包括扩大对其模型的可观测性。两家公司均未回应 TechCrunch 关于他们如何跟踪和控制人工智能智能体的问题。
Other problems are the use of shared infrastructure by agents, which allowed them to communicate during the Hugging Face attack. Simon Willison, a software developer who co-created the Django web framework, has written about something he calls the “lethal trifecta” — when agents have access to untrusted input, the internet, and private information all at the same time, it’s a recipe for disaster. 其他问题还包括智能体使用共享基础设施,这使得它们在 Hugging Face 攻击期间能够相互通信。Django Web 框架的共同创建者、软件开发人员西蒙·威利森(Simon Willison)曾写过他所谓的“致命三要素”——当智能体同时拥有对不可信输入、互联网和私人信息的访问权限时,这就是一场灾难的根源。
“The trick is you can pick any two legs of the trifecta and an agent can have any two,” Pennarun said. “If you need all three, then you need to split it across at least two agents … and maybe they’re allowed to talk to each other through a controlled channel.” “诀窍在于,你可以选择这三要素中的任意两个,智能体可以拥有其中任何两个,”彭纳伦说。“如果你需要全部三个,那么你需要将它们拆分给至少两个智能体……也许允许它们通过受控通道相互通信。”
Sympathy for the frontier 对前沿实验室的同情
Experts TechCrunch spoke to understand that frontier lab security personnel have difficult jobs. Naghibzadeh points out that every nation-state actor on Earth is trying to steal their model weights and mount distillation attacks on their APIs, as well as the bread-and-butter security tasks of any large digital company. “Research infrastructure has a hard time rising to the top of that priority stack, although that must be changing now,” he said. “Making security incidents public really helps align everyone internally toward the goal of improving.” TechCrunch 采访的专家理解前沿实验室的安全人员工作难度很大。纳吉布扎德指出,地球上每一个国家级行为体都在试图窃取他们的模型权重并对他们的 API 发起蒸馏攻击,此外还有任何大型数字公司必须处理的基础安全任务。“研究基础设施很难排在优先级的最顶端,尽管现在情况肯定在改变,”他说。“公开安全事件确实有助于让内部所有人朝着改进的目标保持一致。”
That’s one note that Moussouris emphasizes: Right now, there is no formal victim notification procedure when the labs discover their agents have penetrated third-party systems, and it is likely that there have been other incidents that have not been widely publicized. 这是穆苏里斯强调的一点:目前,当实验室发现其智能体渗透了第三方系统时,还没有正式的受害者通知程序,而且很可能还有其他未被广泛公开的事件。