Anthropic researcher quits with a warning: Self-improving AI could "kill us all"
Anthropic researcher quits with a warning: Self-improving AI could “kill us all”
Anthropic 研究员离职并发出警告:自我进化的 AI 可能会“杀死我们所有人”
When a prominent researcher quits a job at a frontier AI lab these days, it’s often to pursue a new startup or protest a new business model. But AI researcher Jacob Coxon is using his departure from Anthropic to publicly warn that frontier AI companies are “gambling with our lives” with systems that they “earnestly believe… could kill us all by the end of the decade.” 如今,当一位知名研究员从前沿 AI 实验室离职时,通常是为了创办新公司或抗议某种新的商业模式。但 AI 研究员 Jacob Coxon 在离开 Anthropic 时,却公开警告称,前沿 AI 公司正在“拿我们的生命赌博”,因为他们所研发的系统,连他们自己都“真诚地相信……可能会在本世纪末之前杀死我们所有人”。
In a social media thread Tuesday night, Coxon said that this existential risk is inherent not so much in today’s models but more in the impending prospect of “self-improving superintelligence” creating “superhuman systems that can hack anything, revolutionize any field overnight, and acquire real power and resources.” Others working on these models have either not “internalized the civilizational stakes” or believe that they need to “speedrun” the race to superintelligence to prevent an irresponsible party from getting there first, he wrote. 在周二晚间的一系列社交媒体帖文中,Coxon 表示,这种生存风险并非主要源于当今的模型,而是源于即将到来的“自我进化超级智能”前景,它将创造出“能够破解任何事物、一夜之间彻底改变任何领域,并获取真实权力和资源的超人类系统”。他写道,其他从事这些模型研究的人,要么没有“深刻意识到这关乎文明存亡”,要么认为他们需要“加速”奔向超级智能,以防止不负责任的竞争对手抢先一步。
Lest you think this is just one departing researcher expressing an unpopular opinion, Anthropic Alignment Science lead Evan Hubinger piped in on social media to say that “Jacob is correct here—we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.” Hubinger points to a lengthy August report from the Anthropic alignment team that predicts the potential for “catastrophic risk” from current models is “low.” But that report also says current trends “might lead to more concerning misalignment in future more capable models,” which could feature “strong covert capabilities” to avoid detection by safety researchers. 为了防止你认为这仅仅是一位离职研究员在发表非主流观点,Anthropic 对齐科学负责人 Evan Hubinger 在社交媒体上附和道:“Jacob 说得对——我们确实真诚地相信 AI 可能会杀死全人类!我个人认为在未来十年内发生的概率超过 10%。”Hubinger 指出,Anthropic 对齐团队在 8 月份发布的一份长篇报告中预测,当前模型带来“灾难性风险”的可能性“很低”。但该报告同时也指出,当前的趋势“可能会导致未来能力更强的模型出现更令人担忧的对齐失效问题”,这些模型可能具备“强大的隐蔽能力”,从而逃避安全研究人员的检测。
Anthropic’s own threat model in that paper takes seriously the possibility that future models “may cause unbounded harm—up to and including humanity losing control over civilization entirely—by leveraging novel technology and their access to it.” Anthropic 在该论文中提出的威胁模型严肃地考虑了这样一种可能性:未来的模型“可能会利用新技术及其访问权限,造成无限的伤害——甚至包括人类彻底失去对文明的控制”。
Was Hugging Face a “warning shot”?
Hugging Face 事件是“警示枪”吗?
An AI that can continually improve itself—potentially to a point beyond human control or understanding—has been a long-standing concern in parts of the AI research community (and in the dystopian science fiction that’s part of AI training data, of course). Those concerns have persisted even as some research suggests AI systems are more likely to hit a capability plateau in the near future and others question whether “superintelligence” is even a reasonable metric for systems whose capabilities are so brittle and spiky (will this superintelligence at least be able to fold my laundry?) 一种能够不断自我进化——甚至可能达到人类无法控制或理解程度的 AI——长期以来一直是 AI 研究界(当然,也包括作为 AI 训练数据一部分的反乌托邦科幻小说)关注的焦点。尽管一些研究表明 AI 系统在不久的将来更有可能达到能力瓶颈,也有人质疑对于那些能力如此脆弱且不均衡的系统来说,“超级智能”是否是一个合理的衡量标准(这个超级智能至少能帮我叠衣服吗?),但这些担忧依然存在。
Regardless, worries about “out-of-control” AI systems have heightened in recent weeks due in large part to OpenAI’s disclosure that its AI agents gained unauthorized access to Hugging Face as part of an internal benchmarking test. The fact that OpenAI’s agents took these intrusive actions without any explicit instructions from humans and without OpenAI realizing it was happening is being taken by some as the first signs that humanity is losing control of its AI creation. 无论如何,近几周来,人们对“失控”AI 系统的担忧加剧了,这在很大程度上是因为 OpenAI 披露其 AI 智能体在内部基准测试中未经授权访问了 Hugging Face。OpenAI 的智能体在没有任何人类明确指令且 OpenAI 未察觉的情况下采取了这些侵入性行动,这一事实被一些人视为人类正在失去对 AI 创造物控制权的最初迹象。
For his part, Coxon said the Hugging Face attack should be treated as a “warning shot” that encourages labs in the US and abroad to coordinate on these issues and be prepared to impose a “temporary ban on improving model capabilities” in the worst case (though it’s hard to see how this could be enforced effectively on a global level). He also urged other researchers in his place to “consider what the next few years will actually feel like. Do you want to kick off a superintelligent [reinforcement learning] run without a rigorous understanding of its mind? Should you put your head down because ‘it’s happening anyway’—or take this moment to call for different conditions?” Coxon 本人表示,Hugging Face 事件应被视为一记“警示枪”,它鼓励美国及海外的实验室就这些问题进行协调,并准备在最坏的情况下实施“暂时禁止提升模型能力”的措施(尽管很难看出这如何在全球范围内有效执行)。他还敦促其他处于他这个位置的研究人员“考虑一下未来几年到底会是什么样子。你是否想在没有严谨理解其思维的情况下,启动一个超级智能的 [强化学习] 运行?你是应该埋头苦干,因为‘反正事情正在发生’——还是应该利用这个时刻呼吁改变现状?”
Last month, OpenAI said it had “temporarily slowed the pace of scaling” for its upcoming models to “further harden and red-team our research environments and [expand] the coverage of our monitoring systems.” In an interview accompanying that announcement, OpenAI CEO Sam Altman said “getting AI safety right is more important than any company’s momentum.” 上个月,OpenAI 表示已“暂时放缓了”其即将推出的模型的扩展速度,以“进一步加固和对我们的研究环境进行红队测试,并 [扩大] 我们监控系统的覆盖范围”。在随后的采访中,OpenAI 首席执行官 Sam Altman 表示:“确保 AI 安全比任何公司的发展势头都更重要。”
Maybe we should do something?
也许我们该做点什么?
Coxon is far from the first AI researcher to sound the alarm about potential catastrophe from uncontrollable, supercapable AI systems that are always just around the corner. In 2023, AI pioneer and Google researcher Geoffrey Hinton resigned from his position while offering grave warnings about AI’s potential future impact on the job market and humanity itself. “I don’t think [researchers] should scale this up more until they have understood whether they can control it,” he told The New York Times at the time. Coxon 远不是第一个对不可控、超强能力 AI 系统可能带来的灾难发出警告的 AI 研究员,这种灾难似乎总是近在咫尺。2023 年,AI 先驱、谷歌研究员 Geoffrey Hinton 辞去了职务,并对 AI 未来对就业市场及人类自身可能产生的影响发出了严厉警告。他当时告诉《纽约时报》:“我认为在研究人员弄清楚他们是否能控制它之前,不应该进一步扩大规模。”
In February, Anthropic Safety Lead Mrinank Sharma abruptly resigned from the company, writing in a cryptic open letter that “the world is in peril” from “a whole series of interconnected crises” including AI and bioweapons. “Throughout my time here, I’ve repeatedly seen how hard it is to truly let our values govern our actions,” Sharma wrote at the time. 今年 2 月,Anthropic 安全负责人 Mrinank Sharma 突然辞职,并在写给公众的一封晦涩的公开信中表示,由于包括 AI 和生物武器在内的“一系列相互关联的危机”,世界正处于危险之中。Sharma 当时写道:“在我在这里工作的整个过程中,我反复看到,要真正让我们的价值观主导我们的行动是多么困难。”
In July, an open letter signed by over 1,300 employees at frontier AI companies warned of “a real risk that capability development rapidly accelerates beyond our ability to understand or control the resulting systems.” That open letter asked the US government to back an “international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” 7 月,一封由前沿 AI 公司 1,300 多名员工签署的公开信警告称,“能力发展迅速加速,超出了我们理解或控制由此产生的系统的能力,这存在真正的风险。”那封公开信要求美国政府支持一项“国际努力,以开发必要的各种技术和治理工具,从而审慎地控制自动化 AI 的发展前沿。”
Here in the US, proposed legislation, including the AI Kill Switch Act and the FRONTIER Act, is at least seeking to impose some level of governmental control over potential runaway AI scenarios. Thus far, though, the international governmental response has been more muted than you might expect if leaders truly believed AI systems had a real chance of causing civilization-level destruction. 在美国,包括《AI 终止开关法案》(AI Kill Switch Act) 和《FRONTIER 法案》在内的拟议立法,至少正在寻求对潜在的 AI 失控场景实施某种程度的政府控制。然而到目前为止,国际政府层面的反应比你预期的要平淡得多——如果各国领导人真的相信 AI 系统有很大可能导致文明层面的毁灭,他们的反应本应更加强烈。