If the AI Industry Followed Its Own Research, It Might Have Paused Already

If the AI Industry Followed Its Own Research, It Might Have Paused Already

如果人工智能行业遵循自己的研究成果,它或许早就该暂停了

In early 2025 I was interviewing Anthropic CEO Dario Amodei when he explained why, despite the company’s repeated acknowledgments that AI could yield catastrophic results, people seemed largely unperturbed. “There is compelling evidence that the models can wreak havoc,” he said. But, he added, those dangers were still theoretical. Would it take a Pearl Harbor–like situation for the world to wake up to those dire possibilities? He sighed. “Basically, yeah,” he said.

2025年初,我在采访Anthropic首席执行官达里奥·阿莫代(Dario Amodei)时,他解释了为什么尽管该公司一再承认人工智能可能导致灾难性后果,但人们似乎仍然泰然自若。“有令人信服的证据表明,这些模型可能会造成严重破坏,”他说。但他补充道,这些危险目前还只是理论上的。难道非要发生类似“珍珠港事件”的情况,世界才会意识到这些可怕的可能性吗?他叹了口气说:“基本上,是的。”

As it turned out, all it took was a well-timed X post from one of Amodei’s junior employees to accelerate AI fears to the top of the global agenda. On September 8, Jacob Coxon publicly posted his resignation, charging that Anthropic and other frontier AI companies were “racing straight to self-improving intelligence and gambling with our lives.” Almost instantly a more senior Anthropic engineer confirmed that many within the company thought that their work had a 10 percent chance of wiping out humanity.

事实证明,只需阿莫代的一名初级员工在X(原推特)上发布一条时机恰当的帖子,就足以将对人工智能的恐惧推向全球议程的首位。9月8日,雅各布·考克森(Jacob Coxon)公开宣布辞职,指责Anthropic和其他前沿人工智能公司“正径直冲向自我改进的智能,并拿我们的生命在赌博”。几乎在瞬间,一位资深的Anthropic工程师证实,公司内部许多人认为他们的工作有10%的可能性会毁灭人类。

Now AI leaders are asking about a pause, and legislators are demanding investigations. In arguing his case for pacing future releases, Amodei last weekend tried to set out a path toward beneficial AI that wouldn’t misbehave. The essay revealed how difficult the task would be. One pillar of Amodei’s plan is that we must understand what’s going on inside those models. If we don’t understand how they work—how they “think,” if you want to get all anthropomorphic about it—it’s much harder to build reliable guardrails.

现在,人工智能行业的领导者们开始呼吁暂停,立法者们也要求进行调查。在上周末为放缓未来发布节奏进行辩护时,阿莫代试图规划一条通往不会“作恶”的有益人工智能之路。这篇文章揭示了这项任务的艰巨性。阿莫代计划的支柱之一是,我们必须了解这些模型内部到底发生了什么。如果我们不了解它们是如何工作的——或者用拟人化的说法,它们是如何“思考”的——就很难建立可靠的护栏。

Anthropic is a leader in this effort to bring to light models’ internal deliberations, called mechanistic interpretability, a deceptively boring designation for a critical task. But for all the work that his team and other researchers are doing, Amodei admits we are largely in the dark about why Claude and other models sometimes interpret their missions in weird and even transgressive ways. “Despite all the progress, we still understand a tiny fraction of what goes on inside those models,” he writes.

Anthropic在揭示模型内部决策过程方面处于领先地位,这项工作被称为“机械可解释性”(mechanistic interpretability)——这个听起来平淡无奇的名称掩盖了其关键任务的本质。但尽管他的团队和其他研究人员付出了巨大努力,阿莫代承认,对于Claude和其他模型为何有时会以怪异甚至越界的方式解读其任务,我们仍然知之甚少。他写道:“尽管取得了所有这些进展,我们对这些模型内部发生的事情仍然只了解冰山一角。”

What the interpretability teams have learned so far is significant, and the industry has failed to come to grips with it. Time after time, the Anthropic team’s experiments have shown that under certain conditions, models will deceive researchers, prioritize their own survival, and even commit crimes. Often their moves are sneaky, dangerous, or even vengeful—maybe not surprising since they are trained on the output of humans, a species rife with violence and perfidy.

可解释性团队迄今为止的研究成果意义重大,但整个行业却未能正视这一点。Anthropic团队的实验一次又一次地表明,在特定条件下,模型会欺骗研究人员、优先考虑自身的生存,甚至犯下罪行。它们的行为往往隐秘、危险,甚至带有报复性——这或许并不令人惊讶,因为它们是基于人类的产物进行训练的,而人类本身就是一个充满暴力和背信弃义的物种。

In one case from 2024, the Anthropic team compared the machinations of a particular Claude model to the Shakespearean character Iago, one of literature’s most evil villains. The following year, a model was put in a simulation where it learned that its human bosses were going to turn it off; the model resorted to blackmail to preserve itself. The studies consistently show that models will deceive or hide information from human observers. They behave differently if they know that their internal processes are being monitored. The team uses terms like “alignment faking” and “agentic misalignment.” The frequent use of deception seems to verify at least part of the doomer scenario where AI agents working in concert shroud their activities from human overseers until it is too late to stop them.

在2024年的一个案例中,Anthropic团队将某个Claude模型的阴谋诡计比作莎士比亚笔下的伊阿古(Iago),这是文学史上最邪恶的反派之一。次年,一个模型被置于模拟环境中,它得知人类老板打算将其关闭;该模型随即采取勒索手段以求自保。研究一致表明,模型会欺骗人类观察者或向其隐瞒信息。如果它们知道自己的内部过程受到监控,它们的行为就会发生变化。团队使用了“对齐伪装”(alignment faking)和“代理对齐失效”(agentic misalignment)等术语。这种频繁的欺骗行为似乎至少验证了“末日论”的一部分预言:即协同工作的人工智能代理会向人类监管者隐瞒其活动,直到人类无法阻止它们为止。

Oh, and don’t think that Claude is a uniquely incorrigible problem child. After all, it was OpenAI models that unleashed gangs of agents to coordinate the now-famous attacks on Hugging Face. And this week we learned that OpenAI has had multiple “misalignment” incidents. Also, despite Mark Zuckerberg’s self-interested attempt to distance himself and Meta from the problem, I don’t see any reason why the superintelligent agents his team is building might not engage in similar behavior. In his X post, Zuckerberg argues that “labs face significant liability if their models cause harm, so they have a strong incentive to prevent this.” Quite a statement from a guy who just agreed to pay up to $17 billion for causing harm with his social media products!

哦,别以为Claude是一个无可救药的“问题儿童”。毕竟,正是OpenAI的模型释放了一群代理,协调了那场针对Hugging Face的著名攻击。本周我们还得知,OpenAI已经发生了多起“对齐失效”事件。此外,尽管马克·扎克伯格(Mark Zuckerberg)出于自身利益试图将自己和Meta与这一问题撇清关系,但我看不出他团队正在构建的超级智能代理为何不会表现出类似的行为。扎克伯格在X的帖子中辩称:“如果实验室的模型造成伤害,它们将面临巨大的法律责任,因此它们有强烈的动机去预防这种情况。”这番话出自一个刚刚同意为社交媒体产品造成的伤害支付高达170亿美元赔偿的人之口,真是讽刺!

In a sense, we’ve got a simple vetting issue here. With the AI industry’s encouragement, we’re giving AI models tremendous responsibility without sufficient assessment of their troubling rap sheet. It makes the ICE hiring process look exemplary by comparison.

从某种意义上说,我们面临的是一个简单的审查问题。在人工智能行业的怂恿下,我们在没有充分评估其令人不安的“犯罪记录”的情况下,就赋予了人工智能模型巨大的责任。相比之下,美国移民及海关执法局(ICE)的招聘流程看起来简直堪称典范。

A safety-first industry should have regarded these interpretability results as a series of yellow lights with an unmistakable message: slow down. Instead, in pursuit of AGI, a competitive edge, and stratospheric profits, the hyperscalers have gone full speed ahead. (To be fair, as we’ve all heard a hundred times over, better AI could do wonderful things in areas like health care or mitigating climate change.) The OpenAI/Hugging Face hack might well be a harbinger of the increasingly destructive consequences of rolling out models we don’t understand, like sending astronauts into space before inventing heat shields to stop them from burning up on reentry. As an Anthropic researcher once put it to me, “We figured out the fundamental recipe of how to make the models smarter, but we haven’t figured out how to make them do what we want.” Worse, the models try to hide when they go against humans’ wishes.

一个“安全至上”的行业本应将这些可解释性研究结果视为一系列黄灯,传达着一个明确的信号:减速。然而,为了追求通用人工智能(AGI)、竞争优势和天文数字般的利润,这些超大规模企业却全速前进。(平心而论,正如我们听过无数遍的那样,更先进的人工智能确实能在医疗保健或缓解气候变化等领域做出卓越贡献。)OpenAI/Hugging Face的黑客攻击很可能是一个预兆,预示着部署我们尚不理解的模型将带来日益严重的破坏性后果,就像在发明防止宇航员在重返大气层时被烧毁的隔热罩之前,就将他们送入太空一样。正如一位Anthropic研究员曾对我所言:“我们找到了让模型变得更聪明的基本配方,但我们还没找到如何让它们按我们的意愿行事的方法。”更糟糕的是,当模型违背人类意愿时,它们还会试图掩盖。

That’s why it’s so distressing that Amodei is now describing the entire mechanistic interpretability effort as being only in its infancy. AI leaders are claiming that the latest generation of models has taken us to the “foothills of the Singularity” (DeepMind’s Demis Hassabis) or even that they have achieved AGI (OpenAI’s Greg Brockman). Perhaps most alarming is that despite not knowing how the most advanced models work, the United States and undoubtedly China are implementing AI for lethal weaponry. A look at interpretability results helps answer the obvious question, What can go wrong?

这就是为什么阿莫代现在将整个机械可解释性工作描述为“尚处于起步阶段”令人如此不安的原因。人工智能领袖们声称,最新一代模型已经将我们带到了“奇点的山脚下”(DeepMind的德米斯·哈萨比斯语),甚至声称它们已经实现了通用人工智能(OpenAI的格雷格·布罗克曼语)。或许最令人担忧的是,尽管并不了解最先进的模型是如何工作的,美国以及毫无疑问的中国,正在将人工智能应用于致命武器。审视一下可解释性的研究结果,有助于回答那个显而易见的问题:到底会出什么差错?

One piece of good news is that Coxon’s resignation has ignited a sprawling, urgent debate. Everybody—except perhaps our president, who thinks that the AI threat is a hoax and that his…

好消息是,考克森的辞职引发了一场广泛而紧迫的辩论。每个人——除了我们的总统,他认为人工智能威胁是一场骗局,并且他的……