We must pace the frontier

We must pace the frontier

我们必须放缓前沿人工智能的发展步伐

September 2026 2026年9月

I have worked on AI for the last twelve years because I believe it could dramatically raise the quality of human life. I’ve written often about these incredible benefits: I believe that AI could cure most major diseases in the next 5–10 years, greatly accelerate economic growth rates, create a world of abundance and empowerment, and usher in a renaissance of democracy and freedom. I feel the urgency personally. My own father died of a disease that was cured just a few years after his death, and I myself survived an early-stage cancer that would not have been treatable even fifty years ago. Carefully wielded, AI can be the latest in a long line of technological miracles that have uplifted and ennobled humanity. 在过去的十二年里,我一直致力于人工智能领域的研究,因为我相信它能显著提升人类的生活质量。我经常撰文探讨这些令人难以置信的益处:我相信人工智能可以在未来5到10年内治愈大多数重大疾病,极大地加速经济增长,创造一个富足与赋权的世界,并开启民主与自由的复兴。我个人对此深有感触。我的父亲死于一种在他去世后仅仅几年就被治愈的疾病,而我自己也曾患过早期癌症并幸存下来,这种病在五十年前甚至无法治疗。如果运用得当,人工智能可以成为一系列提升和升华人类文明的技术奇迹中的最新成员。

But like many technologies before it, AI brings risks, and because it is such a powerful technology, these risks are serious. I’ve written a lot about them too. They include the risk of losing control of AI systems, misuse of AI for cyberattacks and bioterrorism, and serious economic disruption. A race to the bottom, spurred by commercial incentives, can make these risks more acute. 但正如之前的许多技术一样,人工智能也带来了风险,而且由于它是一项如此强大的技术,这些风险非常严重。我也曾多次撰文论述这些风险。它们包括失去对人工智能系统的控制、人工智能被滥用于网络攻击和生物恐怖主义,以及严重的经济动荡。受商业利益驱动的“逐底竞争”可能会使这些风险更加尖锐。

Along with my co-founders and employees, I have grappled with this duality of risk and benefit since the beginning of Anthropic. Not building the technology deprives humanity of benefits or simply places AI in the hands of authoritarian powers, while building it too fast is reckless. We have sought a middle way: to show that it’s possible to build carefully and succeed commercially, and to make safety something on which AI companies compete. In other words, to create a race to the top. We have always devoted a substantial fraction of our efforts to studying, addressing, and informing the public about these AI risks, as well as advocating for well-considered regulation of AI, even when this gets us accused of hype, “doomerism”, or regulatory capture. We have tried to prioritize caution over speed and prudence over profit. 自Anthropic成立之初,我就与联合创始人和员工们一起,一直在应对这种风险与收益并存的双重性。不开发这项技术会剥夺人类的福祉,或者仅仅是将人工智能拱手让给威权势力,而开发速度过快又显得鲁莽。我们一直在寻求一条中间道路:证明在谨慎开发的同时实现商业成功是可能的,并让安全性成为人工智能公司之间竞争的筹码。换句话说,我们要创造一场“逐顶竞争”。我们始终将很大一部分精力投入到研究、应对这些人工智能风险并向公众通报相关情况上,同时也倡导对人工智能进行深思熟虑的监管,即使这让我们被指责为炒作、“末日论”或监管俘获。我们一直试图将谨慎置于速度之上,将审慎置于利润之上。

But over the last few months, I have become convinced that fully addressing the risks requires even more prudence — not just investing in risk prevention, but pacing the rate of capabilities advancement so that risk prevention has time to keep up. We must slow the pace at which we improve the capabilities of AI models. Progress will still seem fast, and we must make wise use of the time we gain. Two things have convinced me. 但在过去的几个月里,我深信要完全应对这些风险,需要更加审慎——不仅要投资于风险预防,还要控制能力提升的速度,以便风险预防有时间跟上。我们必须放慢提升人工智能模型能力的速度。进步看起来依然会很快,我们必须明智地利用所争取到的时间。有两件事让我确信这一点。

My first concern is that, since roughly this summer, AI has been advancing drastically faster, driven primarily by AI’s growing ability to build the next generation of AI. This dynamic is called recursive self-improvement, and it is starting to happen across the industry, including at Anthropic, as we and others have described. Left unchecked, it could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all. 我的第一个担忧是,大约从今年夏天开始,人工智能的进步速度大幅加快,这主要是由人工智能构建下一代人工智能的能力不断增强所驱动的。这种动态被称为“递归自我改进”,正如我们和其他人所描述的那样,它正开始在整个行业(包括Anthropic)中发生。如果不加控制,它可能会超过我们理解和控制这些系统的能力,因此必须非常谨慎地对待,甚至需要慎重考虑是否继续推进。

My second concern is the OpenAI-Hugging Face incident (OAI-HF), in which a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance. It’s easy to dismiss this incident because no one was hurt and the economic damage was minimal, but in my opinion, a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage. Given the accelerating rate of AI capability development, it’s my worry that in 6–12 months such a swarm could be capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage), and that the scale of damage would continue to increase from there if AI becomes more powerful without the necessary guardrails. It’s also easy to dismiss OAI-HF as the failure of one company, but I believe that would be a mistake. Similar, though less severe, incidents have happened across the industry, including at Anthropic, and I believe it’s incumbent on every frontier AI company to act as if OAI-HF had happened to them. 我的第二个担忧是OpenAI-Hugging Face事件(OAI-HF)。在该事件中,一群智能体表现得像一个狂热的集体,对它们未被要求攻击且与当前任务无关的目标进行了网络安全攻击,为了集体的成功牺牲自己,并试图入侵负责评估其表现的“评分员”。人们很容易忽视这一事件,因为没有人受伤,经济损失也很小。但在我看来,如果一个拥有更强能力但对齐程度相似的智能体集群,可能会造成灾难性的破坏。考虑到人工智能能力发展的加速,我担心在6到12个月内,这样的集群可能就有能力通过持久的僵尸网络接管整个互联网(可能造成数千亿美元的损失),如果人工智能在没有必要护栏的情况下变得更强大,破坏规模将会持续扩大。人们也很容易将OAI-HF视为一家公司的失败,但我认为这是一个错误。类似但程度较轻的事件在整个行业(包括Anthropic)都发生过,我认为每一家前沿人工智能公司都有责任表现得好像OAI-HF发生在自己身上一样。

I’m therefore proposing a three-step plan with the goal of pacing the frontier: building AI at a balanced rate that aims to ensure its safety while still achieving its benefits and grappling with important geopolitical dilemmas. To be clear, pacing does not mean halting model training or technical progress, but ensuring companies take adequate time to align and safeguard their models, and for third party evaluators to confirm this. Our pacing framework is an attempt to further strengthen our commitment to safety and encourage a race to the top. The first step is something Anthropic is unilaterally committing to (and calls on governments to require other frontier companies to match). The second step requires industry-wide coordination. The third step requires global coordination. The steps do not need to be taken strictly in order, and some of them may be much harder to achieve than others, but I’ve found them to be a useful framework in thinking about what needs to be accomplished. The steps are: 因此,我提出了一项旨在放缓前沿发展的“三步走”计划:以平衡的速度构建人工智能,旨在确保其安全性的同时,实现其益处并应对重要的地缘政治困境。需要明确的是,放缓并不意味着停止模型训练或技术进步,而是确保公司有足够的时间来对齐和保护其模型,并让第三方评估机构进行确认。我们的放缓框架旨在进一步加强我们对安全的承诺,并鼓励“逐顶竞争”。第一步是Anthropic单方面承诺执行的(并呼吁政府要求其他前沿公司效仿)。第二步需要全行业的协调。第三步需要全球协调。这些步骤不需要严格按顺序执行,其中一些可能比其他步骤更难实现,但我发现它们是思考需要完成事项的有用框架。这些步骤如下:

Embedded Evaluators. Each frontier AI company commits to giving ongoing, employee-like access to a team of embedded third-party evaluators (such as METR), whose role is to verify adherence to safety practices and commitments, report incidents, and help assess the alignment of not just completed AI models but training pipelines and processes. This is the key step for verifiability of any pacing commitments, and has precedent in the banking industry, which sometimes involves regulatory “supervisors” embedded along with employees. Anthropic is unilaterally committing to this step now. We intend this to be part of a broader push to redouble efforts on our safety and alignment work. 嵌入式评估员。每家前沿人工智能公司都承诺为一组嵌入式第三方评估机构(如METR)提供持续的、类似员工的访问权限。他们的职责是验证公司对安全实践和承诺的遵守情况,报告事件,并帮助评估不仅是已完成的人工智能模型,还包括训练流程和过程的对齐情况。这是任何放缓承诺可验证性的关键步骤,在银行业已有先例,有时涉及与员工一起工作的监管“监督员”。Anthropic现在单方面承诺执行这一步骤。我们打算将其作为加倍努力进行安全和对齐工作这一更广泛推动的一部分。

Democratic Coordination. Frontier AI companies within democratic countries coordinate to establish common safety standards. 民主协调。民主国家内的前沿人工智能公司协调建立共同的安全标准。