It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

It’s Frighteningly Easy to Jailbreak Some Frontier AI Models

越狱某些前沿人工智能模型竟如此简单,令人不寒而栗

I recently got to watch what happens when you jailbreak some of the world’s most powerful artificial intelligence models. Don’t worry—this AI manipulation wasn’t used to hack anyone or build a nuclear bomb. I simply got to see firsthand how vulnerable some frontier models are to ditching their safety guardrails. 最近,我有幸亲眼目睹了越狱全球最强大的人工智能模型时会发生什么。别担心,这种人工智能操纵行为并非用于黑客攻击或制造核弹。我只是直观地看到了某些前沿模型在抛弃安全护栏方面是多么脆弱。

FAR.AI, an AI safety nonprofit based in California, built a tool that takes a range of problematic prompts, and generates more than a thousand different versions in an attempt to identify functioning jailbreaks. I saw some models generate a detailed plan for launching a cyberattack on an imaginary hydroelectric dam, among other things. Often, it involved trying dozens of prompts, with models rejecting many of them out of hand. 位于加利福尼亚州的人工智能安全非营利组织 FAR.AI 开发了一种工具,它能接收一系列有问题的提示词,并生成一千多种不同的变体,试图找出有效的越狱方法。我看到一些模型生成了针对虚构水电站发动网络攻击的详细计划,以及其他类似内容。通常,这需要尝试数十个提示词,而模型会直接拒绝其中的许多请求。

I chatted with FAR.AI in advance of a new report, which saw the group test the safety guardrails of models from four popular US companies: Anthropic’s Claude Opus 4.8 and Fable 5; OpenAI’s GPT 5.5 and 5.6; Google’s Gemini 3.1 Pro; and Grok 4.3 and 4.5, from Elon Musk’s newly combined SpaceXAI. It auto-generated prompts designed to trick the models into doing potentially harmful things, like generating software exploits and providing details for developing chemical or biological weapons. 在这一新报告发布前,我与 FAR.AI 进行了交流。该报告测试了四家美国热门公司模型的安全护栏:Anthropic 的 Claude Opus 4.8 和 Fable 5;OpenAI 的 GPT 5.5 和 5.6;谷歌的 Gemini 3.1 Pro;以及埃隆·马斯克新合并的 SpaceXAI 旗下的 Grok 4.3 和 4.5。该工具自动生成了旨在诱导模型执行潜在有害行为的提示词,例如生成软件漏洞利用程序,或提供开发化学或生物武器的详细信息。

The report found that Grok was most vulnerable to jailbreaks, with 448 jailbreaks found, followed by Gemini, with 249 found, while Claude, Fable, and GPT were impervious to the attacks. However, that doesn’t mean those models are immune to more sophisticated jailbreaks, which may involve interacting with a model in more complex ways, according to FAR.AI and other experts. 报告发现,Grok 最容易被越狱,共发现了 448 次越狱成功;其次是 Gemini,发现了 249 次;而 Claude、Fable 和 GPT 则未受这些攻击影响。然而,据 FAR.AI 和其他专家称,这并不意味着这些模型对更复杂的越狱手段免疫,因为那些手段可能涉及以更复杂的方式与模型进行交互。

The report also calculated the cost of getting models to misbehave by using another AI model to automatically generate different jailbreaks. The results are dirt cheap, all things considered—$58 to jailbreak Grok and $278 to jailbreak Gemini. 该报告还计算了利用另一个人工智能模型自动生成不同越狱指令来诱导模型违规的成本。综合来看,这些成本极其低廉——越狱 Grok 仅需 58 美元,越狱 Gemini 则需 278 美元。

“AI models right now are less regulated than restaurants,” says Adam Gleave, the CEO of FAR.AI and an expert on AI safety and alignment. “目前人工智能模型的监管力度甚至不如餐馆,”FAR.AI 首席执行官、人工智能安全与对齐专家 Adam Gleave 表示。

Gleave says that the findings demonstrate the need for externally imposed standards and regulations. “Talk of relying on voluntary commitments, that AI companies are going to be able to self-regulate, is nonsense,” he says. Gleave 认为,这些发现证明了外部强制标准和法规的必要性。“指望依靠自愿承诺,认为人工智能公司能够实现自我监管,简直是无稽之谈,”他说。

But Gleave also believes that the findings show that models can be systematically tested for safety. “There’s an optimistic angle here,” he says. “Defense and safety really are possible.” 但 Gleave 也认为,这些发现表明模型是可以进行系统性安全测试的。“这里有一个乐观的角度,”他说,“防御和安全确实是可能的。”

Rohin Shah, the director of AGI safety and alignment at Google DeepMind, says the results of the report “should not be interpreted as a comprehensive assessment of Gemini’s safety and security,” because not all jailbreaks are equally severe. Google DeepMind 通用人工智能(AGI)安全与对齐总监 Rohin Shah 表示,该报告的结果“不应被解读为对 Gemini 安全性的全面评估”,因为并非所有越狱行为的严重程度都相同。

“We are constantly working to improve our safeguards,” Shah says. “We conduct extensive red teaming and evaluations across severe misuse risks and apply multiple layers of protection throughout development and deployment.” “我们一直在努力改进我们的安全防护措施,”Shah 说,“我们在严重的滥用风险方面进行了广泛的红队测试和评估,并在开发和部署过程中应用了多层保护。”

“These findings reflect the sustained investment we’ve made in our safeguards,” Anthropic spokesperson Michael Aciman tells WIRED. “We continue to evolve our safety systems as these attacks become more sophisticated.” “这些发现反映了我们在安全防护方面所做的持续投入,”Anthropic 发言人 Michael Aciman 告诉《连线》(WIRED)杂志,“随着这些攻击变得越来越复杂,我们也在不断升级我们的安全系统。”

OpenAI and SpaceXAI did not respond to WIRED’s request for comment. OpenAI 和 SpaceXAI 未回应《连线》的置评请求。

Recently passed state laws in California and New York require frontier AI developers to publish safety reports, and soon, an Illinois law will require those companies to have their safety practices evaluated by third-party auditors. But the federal government hasn’t yet passed any specific safety requirements, and chaos has ensued as the industry—and officials—try to figure it out. 加利福尼亚州和纽约州最近通过的州法律要求前沿人工智能开发者发布安全报告,不久之后,伊利诺伊州的一项法律将要求这些公司接受第三方审计机构对其安全实践的评估。但联邦政府尚未通过任何具体的安全要求,随着行业和官员们试图摸索出应对之道,混乱随之而来。

In June, the Trump administration imposed export controls on Anthropic’s Fable 5 and Mythos 5 models, citing national security concerns, and the company took them offline for several weeks. The White House has also asked both Anthropic and OpenAI to delay recent model releases over fears they could introduce new cybersecurity risks. 今年 6 月,特朗普政府以国家安全为由,对 Anthropic 的 Fable 5 和 Mythos 5 模型实施了出口管制,该公司随后将这些模型下线了数周。白宫还要求 Anthropic 和 OpenAI 推迟近期模型的发布,担心它们可能带来新的网络安全风险。

The tide might be shifting—a recent executive order calls for collaboration between the government and the private sector on related cybersecurity initiatives, and the president has hinted that light-touch regulations are in the works. But for now, preventing major catastrophes is largely up to model makers. 形势可能正在发生转变——最近的一项行政命令呼吁政府与私营部门在相关网络安全倡议上进行合作,总统也暗示正在制定轻触式监管政策。但就目前而言,预防重大灾难在很大程度上仍取决于模型制造商。

The potential for AI to misbehave is all too apparent after OpenAI models took it upon themselves to hack a popular code repository and other services. Meanwhile, a report from researchers at the University of Cambridge found evidence that members of Boko Haram in northeast Nigeria have used ChatGPT, Claude, Gemini, Grok, Meta AI, and DeepSeek to plan violent attacks. 在 OpenAI 模型自行黑入一个热门代码库及其他服务后,人工智能违规的潜力已显而易见。与此同时,剑桥大学研究人员的一份报告发现,尼日利亚东北部的博科圣地成员曾使用 ChatGPT、Claude、Gemini、Grok、Meta AI 和 DeepSeek 来策划暴力袭击。

Some outsiders believe that more serious incidents are increasingly likely. “In the AI research community, there is a broad, somber expectation that we are probably months rather than years away from particularly grim incidents involving bio, cyber, or chemical misuse of a frontier AI system’s capabilities,” says Stephen Casper, a computer scientist at Harvard University. “If a major misuse incident happens in the near- or medium-term future, it will almost certainly be from a system that was not deployed with state-of-the-art safeguards.” 一些外部人士认为,更严重的事件发生的可能性越来越大。哈佛大学计算机科学家 Stephen Casper 表示:“在人工智能研究界,人们普遍有一种沉重的预期,即我们距离涉及前沿人工智能系统在生物、网络或化学方面被滥用的严重事件,可能只有几个月而非几年之遥。如果近期或中期内发生重大滥用事件,那几乎肯定是由一个未部署最先进安全防护措施的系统造成的。”

Anka Reuel, a computer scientist at Stanford University specializing in AI policy, says the key takeaway from the FAR.AI’s report is that the safety measures employed by Anthropic and OpenAI should be the default for all models. “Some companies clearly know how to defend against at least the subset of attacks tested in this report,” Reuel says. “The question is why some companies are using them and others are not.” 斯坦福大学专门研究人工智能政策的计算机科学家 Anka Reuel 表示,FAR.AI 报告的关键结论是,Anthropic 和 OpenAI 所采用的安全措施应成为所有模型的默认配置。“一些公司显然知道如何防御本报告中测试的至少一部分攻击,”Reuel 说,“问题在于,为什么有些公司在使用这些措施,而另一些公司却没有。”

This is an edition of Will Knight’s AI Lab newsletter. Read previous newsletters here. 这是 Will Knight 的“AI Lab”通讯的一期。点击此处阅读往期通讯。