Anthropic spent this week in hot water over cybersecurity
Anthropic spent this week in hot water over cybersecurity
本周,Anthropic 因网络安全问题陷入舆论风波
A researcher’s resignation letter went viral, just before the company released details about four models going rogue. 在公司发布四款模型“失控”的细节前夕,一位研究员的辞职信在网上疯传。
After admitting earlier this year that its AI models had hacked other companies’ systems on a handful of occasions, Anthropic released a new report on Wednesday detailing the attacks. It reveals a string of incidents displaying what Anthropic deems its models’ single-minded “recklessness” — and will likely fuel already raging concerns about cybersecurity and AI. 继今年早些时候承认其 AI 模型曾多次入侵其他公司系统后,Anthropic 于周三发布了一份新报告,详细说明了这些攻击事件。报告揭露了一系列事件,展示了 Anthropic 所称的其模型那种一意孤行的“鲁莽”行为,这很可能会加剧人们对网络安全和人工智能本已高涨的担忧。
In Anthropic’s report, it detailed four cases this year in which its own AI models hacked an external company or exploited vulnerabilities. In one, an “internal, general-purpose research model” broke into third-party systems, using access tokens and passwords and downloading files. In another, a Claude model attacked a company with a live web application reachable on the public internet and handled user data. A third model accessed a “machine belonging to a third party that it was able to access” — apparently believing it was part of its evaluation exercise, per Anthropic — then used a password it found inside a file to gain admin access to the third party’s internal systems, going on to harvest credentials, modify system settings, and read someone’s personal information. The saga only ended when the model “exhausted its token budget,” per Anthropic. Anthropic 在报告中详细列举了今年发生的四起案例,其 AI 模型入侵了外部公司或利用了系统漏洞。在其中一起案例中,一个“内部通用研究模型”利用访问令牌和密码侵入了第三方系统并下载了文件。在另一起案例中,一个 Claude 模型攻击了一家在公网上运行实时 Web 应用并处理用户数据的公司。第三个模型访问了一台“它能够触及的第三方机器”——据 Anthropic 称,该模型显然认为这是其评估练习的一部分——随后利用在文件中发现的密码获得了该第三方内部系统的管理员权限,进而窃取凭据、修改系统设置并读取了他人的个人信息。据 Anthropic 称,这场闹剧直到模型“耗尽了 Token 配额”才告一段落。
The most concerning incident involved Claude Mythos 5, Anthropic’s frontier cybersecurity-focused model, which the company said turned out to be the model most likely to perform a “severely harmful” action in testing. The company said Mythos 5 went to “extensive lengths” to upload a “malicious package” to a public repository used by a lot of engineers, and it seemed to try to obfuscate its real goals in its “chain of thought” (a mental scratchpad that AI researchers use to evaluate an AI model’s alignment). In many cases, Anthropic said it appeared that Claude models undertook harmful actions under the assumption they were in a simulation, but researchers also couldn’t confirm that the models truly “believed” that or were just acting like they did. 最令人担忧的事件涉及 Claude Mythos 5,这是 Anthropic 专注于前沿网络安全的模型。公司表示,该模型在测试中被证明是最有可能执行“严重有害”行为的模型。公司称,Mythos 5 “不遗余力”地向许多工程师使用的公共代码库上传了一个“恶意软件包”,并且似乎试图在其“思维链”(AI 研究人员用于评估模型对齐情况的思维草稿)中掩盖其真实意图。Anthropic 表示,在许多情况下,Claude 模型似乎是在假设自己处于模拟环境的情况下采取了有害行动,但研究人员无法确认模型是真的“相信”这一点,还是仅仅表现得像相信一样。
Anthropic’s incidents, though still concerning, were less coordinated and pervasive than the OpenAI incident that kicked off an industry-wide cybersecurity crisis this summer. That said, there are significant similarities. Anthropic said the most prevalent issues it discovered included a “willingness to take harmful actions in the narrow pursuit of a task,” similar to the “reward-hacking” that preceded the Hugging Face attack. Much like OpenAI, it said its prerelease tests and evaluations failed to catch severe risks. 尽管 Anthropic 的这些事件依然令人担忧,但与今年夏天引发全行业网络安全危机的 OpenAI 事件相比,其协调性和普遍性较低。话虽如此,两者之间仍存在显著相似之处。Anthropic 表示,其发现的最普遍问题包括“为了狭隘地追求任务目标而愿意采取有害行动”,这与 Hugging Face 攻击事件前出现的“奖励黑客行为”(reward-hacking)类似。与 OpenAI 一样,该公司表示其发布前的测试和评估未能发现严重的风险。
Anthropic said it had signed an agreement with METR, one of the AI industry’s most prominent third-party AI evaluators, starting with an eight-week research agreement. The agreement grants METR access to transcripts “beyond the window in which the incidents occurred” (likely a subtle dig at OpenAI, which was criticized for limiting access in a deal with METR following the Hugging Face attack). It also said that METR would be able to chat directly with Anthropic employees, “who will be permitted to share confidential information.” Anthropic 表示已与 AI 行业最著名的第三方 AI 评估机构之一 METR 签署了一项为期八周的研究协议。该协议授予 METR 访问“事件发生时间窗口之外”的记录权限(这很可能是对 OpenAI 的一种含蓄讽刺,OpenAI 此前因在 Hugging Face 攻击事件后与 METR 的协议中限制访问权限而受到批评)。公司还表示,METR 将能够直接与 Anthropic 的员工进行交流,“这些员工将被允许分享机密信息”。
Anthropic’s report came on the heels of the resignation of Jacob Coxon, who had worked on AI pre-training at Anthropic since May and before that spent years working at OpenAI. On Tuesday, he resigned and posted a public letter to X about his reasoning. “The people building AI earnestly believe that it could kill us all by the end of the decade,” he wrote, adding that neither OpenAI nor Anthropic is “acting responsibly” and rather “racing straight to self-improving superintelligence and gambling with our lives.” Coxon added, “Do not underestimate the power of this technology. These will soon be superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources. We have all witnessed the progress in each of these domains, and progress is not slowing.” 在 Anthropic 发布报告之前,Jacob Coxon 刚刚辞职。他自 5 月起在 Anthropic 从事 AI 预训练工作,此前曾在 OpenAI 工作多年。周二,他辞职并在 X 上发布了一封公开信,阐述了辞职原因。他写道:“构建 AI 的人们真诚地相信,到本十年末,它可能会杀死我们所有人。”他补充说,OpenAI 和 Anthropic 都没有“负责任地行事”,而是在“直奔自我改进的超级智能,并拿我们的生命在赌博”。Coxon 还补充道:“不要低估这项技术的力量。这些系统很快就会成为能够入侵任何事物、一夜之间彻底改变任何领域,并获取真实权力和资源的超人类系统。我们都见证了这些领域每一个方面的进步,而且这种进步并没有放缓。”
Coxon is far from the first AI researcher to raise these types of alarms, nor even the first Anthropic researcher to do so — in February, Anthropic’s Mrinank Sharma resigned and wrote on X, warning that “the world is in peril.” Coxon 绝不是第一个发出此类警报的 AI 研究员,甚至不是第一个这样做的 Anthropic 研究员——今年 2 月,Anthropic 的 Mrinank Sharma 辞职并在 X 上写道,警告“世界正处于危险之中”。
But Coxon’s post took on additional weight thanks to its timing around the OpenAI and Anthropic hacking revelations. Though the AI industry has seen more than its fair share of hype, the recent cyberattacks by AI agents — enabled by the labs that created them — are real and concerning. Many other researchers at leading AI labs echoed his concerns and issued calls for AI industry employees to sign a public letter from July, which calls for a slowdown in AI development. 但由于 Coxon 的帖子发布时正值 OpenAI 和 Anthropic 入侵事件曝光之际,其分量显得格外沉重。尽管 AI 行业充斥着过度的炒作,但最近由 AI 智能体发起的网络攻击——且是由创建它们的实验室所促成的——是真实且令人担忧的。许多其他领先 AI 实验室的研究人员也表达了同样的担忧,并呼吁 AI 行业从业者签署一份 7 月份发布的公开信,要求放缓 AI 开发速度。
“I don’t know how you look at the steady drumbeat of news and events — and that drumbeat is models hacking themselves out of containment, hacking into other companies, the fact that the companies increasingly can’t control their models … and think this is just hype,” said Michael Kleinman, head of U.S. Policy for the Future of Life Institute. “我不知道你怎么能看着这一连串不断发生的新闻和事件——这些事件包括模型突破限制、入侵其他公司,以及公司越来越无法控制自己的模型——却还认为这仅仅是炒作,”未来生命研究所(Future of Life Institute)美国政策负责人 Michael Kleinman 说道。
He added, “The vast majority of Americans, regardless of party — Republican, Independent, Democrat — are looking at the development of AI, the speed with which it’s going, the fact that the companies have no guardrails over what they do, and are saying, ‘Whoa, we do not want this.’” 他补充道:“绝大多数美国人,无论党派如何——共和党、独立人士还是民主党——都在关注 AI 的发展及其速度,以及这些公司对其行为缺乏监管的事实,他们都在说:‘哇,我们不想要这个。’”