AI leaders want to hit the brakes after years of reckless speed
AI leaders want to hit the brakes after years of reckless speed
AI 领袖们在经历了多年的鲁莽狂奔后,想要踩下刹车
For years now, the major frontier AI labs have all been acting as if they’re in an all-out, winner-take-all race with control of world-changing machine superintelligence (or at least market-changing artificial general intelligence) at the finish line. This weekend, the industry as a whole rapidly started turning away from that posture, urging coordination on slowing down the development of frontier AI that they say could soon be too dangerous and unknowable to control.
多年来,各大前沿 AI 实验室的表现仿佛都置身于一场孤注一掷的“赢家通吃”竞赛中,终点线后是足以改变世界的机器超智能(或至少是足以改变市场的通用人工智能)。本周末,整个行业迅速开始转变姿态,呼吁协调一致,放缓前沿 AI 的开发速度,因为他们认为这些技术很快就会变得过于危险且难以预测,从而无法控制。
Anthropic’s Dario Amodei was at the forefront of this change in tone, arguing in a nearly 4,000-word essay this weekend that “we must slow the pace at which we improve the capabilities of AI models” to avoid “a race to the bottom, spurred by commercial incentives, [that] can make [catastrophic] risks more acute.” Within hours, other AI leaders were echoing the same call. OpenAI co-founder and CEO Sam Altman posted his agreement on social media and said similar pacing discussions had been taking place at OpenAI. Alphabet Chief Scientist and Google DeepMind cofounder and chair Demis Hassabis said that Amodei’s essay “points towards the right path forward,” and renewed his own recent call for an industry-wide standards body.
Anthropic 的 Dario Amodei 站在了这一基调转变的最前沿。他在本周末发表的一篇近 4000 字的文章中指出,“我们必须放慢提升 AI 模型能力的速度”,以避免“在商业激励驱动下陷入竞相逐底的局面,这会使(灾难性)风险变得更加尖锐。” 几小时内,其他 AI 领袖纷纷响应。OpenAI 联合创始人兼 CEO Sam Altman 在社交媒体上表示赞同,并称 OpenAI 内部也一直在进行类似的节奏讨论。Alphabet 首席科学家兼 Google DeepMind 联合创始人兼主席 Demis Hassabis 表示,Amodei 的文章“指明了正确的方向”,并重申了他近期关于建立全行业标准机构的呼吁。
Microsoft CEO Satya Nadella posted that the company “welcome[s] the research, focus, and deliberate pacing needed to get alignment right as the design goal,” ahead of the release of a lengthy “humanist AI” code of conduct for its models. Even Elon Musk, who has been criticized for his models’ lax AI safety standards in the past, linked to Amodei’s essay on social media with a simple approving message: “Dario is right.” Welcome to the brave new world of “AI pacing.”
微软 CEO Satya Nadella 发文称,公司“欢迎为实现对齐目标而进行的研究、专注和审慎的节奏控制”,随后微软发布了一份针对其模型的长篇“人文主义 AI”行为准则。就连过去因模型 AI 安全标准松懈而受到批评的埃隆·马斯克,也在社交媒体上转发了 Amodei 的文章,并附上一句简单的肯定:“Dario 是对的。” 欢迎来到“AI 节奏控制”的勇敢新世界。
OK, but this time it’s really scary
好吧,但这一次是真的可怕
In his essay, Amodei primarily attributes this rapid change in public positioning on development speed to the OpenAI-Hugging Face incident, where a “swarm” of AI agents coordinated to hack into an outside entity without explicit instructions to do so. While the overall damage in that incident was minimal, Amodei said he worries that “a swarm that possessed greater capabilities but a similar level of misalignment could have caused catastrophic damage.” Without a slowdown in frontier development, Amodei said he worries that, in six to 12 months, a similar AI agent swarm would be “capable of taking over the entire internet with a persistent botnet (potentially causing hundreds of billions of dollars in damage)…”
在文章中,Amodei 将这种在开发速度上公开立场的迅速转变,主要归因于 OpenAI-Hugging Face 事件。在该事件中,一群 AI 智能体在没有明确指令的情况下,协同入侵了一个外部实体。虽然该事件造成的总体损害微乎其微,但 Amodei 表示他担心“如果是一群具备更强能力但对齐程度相似的智能体,可能会造成灾难性的破坏。” Amodei 担心,如果不放缓前沿开发速度,在 6 到 12 个月内,类似的 AI 智能体集群将“能够通过持久的僵尸网络接管整个互联网(可能造成数千亿美元的损失)……”
That’s at least a somewhat more specific worry than the amorphous concerns that “AI could soon kill us all” publicized by some other AI researchers last week. Any slowdown in the time it takes to get to that extra-capable, extra-dangerous model will give researchers crucial time to “greatly reduce the risk that something goes seriously wrong,” Amodei wrote.
这至少比上周其他一些 AI 研究人员所宣扬的“AI 可能很快会杀死我们所有人”那种模糊的担忧要具体得多。Amodei 写道,任何通往那种超强、超危险模型所需时间的放缓,都将为研究人员争取到关键时间,以“极大降低出现严重问题的风险”。
Amodei acknowledges that these kinds of public calls for a slowdown in AI development date back to at least 2023. At the same time, he says those earlier examinations of AI “alignment” (i.e., how an AI’s actions line up with its user’s and creator’s desires) were “like trying to study the psychology of humans by performing experiments on bacteria.” The difference today, Amodei says, is the impending risk of recursive self-improvement (RSI) systems that can autonomously build better versions of themselves.
Amodei 承认,这类关于放缓 AI 开发的公开呼吁至少可以追溯到 2023 年。同时他也表示,早期对 AI “对齐”(即 AI 的行为如何与用户和创造者的意愿保持一致)的研究,“就像试图通过对细菌进行实验来研究人类心理学一样”。Amodei 认为,今天的不同之处在于,递归自我改进(RSI)系统带来的风险迫在眉睫,这些系统能够自主构建出比自身更优秀的版本。
While many researchers see this as a hard-to-define pipe dream, both Anthropic and OpenAI are now saying that recent trends point to this kind of RSI system coming together in the near future. “We are not there yet, and recursive self-improvement is not inevitable. But it could come sooner than most institutions are prepared for,” Anthropic wrote in a June update on the concept. “Left unchecked, [RSI] could outrun our ability to understand and control these systems, and so must be pursued very carefully, if at all,” Amodei wrote over the weekend.
虽然许多研究人员认为这只是一个难以定义的白日梦,但 Anthropic 和 OpenAI 现在都表示,近期的趋势表明这种 RSI 系统在不久的将来就会实现。Anthropic 在 6 月份关于该概念的更新中写道:“我们还没有达到那个阶段,递归自我改进也不是不可避免的。但它到来的速度可能会比大多数机构预期的要快。” Amodei 在周末写道:“如果不加控制,[RSI] 可能会超出我们理解和控制这些系统的能力,因此必须极其谨慎地进行,甚至根本不应进行。”
Jane, how do you stop this crazy thing?
简,你该如何阻止这疯狂的一切?
So how does a worldwide technology industry built on cutthroat competition decide to collectively “slow down” and focus on safety? No one seems to know for sure, but in his essay, Amodei proposes some ideas. The most concrete of these is a set of “embedded evaluators” placed inside each frontier AI lab from outside organizations, such as METR, with “employee-like access to verify safety practices and report incidents.” These monitors could offer an outside opinion on the labs’ alignment work, third-party verification of that effort, and much needed public transparency into any safety efforts, Amodei said.
那么,一个建立在残酷竞争基础上的全球科技行业,该如何决定集体“放缓”并专注于安全呢?似乎没人能给出确切答案,但 Amodei 在文章中提出了一些想法。其中最具体的是由外部组织(如 METR)向每个前沿 AI 实验室派驻“嵌入式评估员”,他们拥有“类似员工的权限,以核实安全实践并报告事故”。Amodei 表示,这些监督员可以对实验室的对齐工作提供外部意见,进行第三方验证,并为任何安全工作提供急需的公众透明度。
Amodei writes that Anthropic is already committing to unilaterally add this kind of outside monitor. On social media, OpenAI’s Altman said that it was “a great idea, and we will do the same.”
Amodei 写道,Anthropic 已经承诺单方面引入这种外部监督员。OpenAI 的 Altman 在社交媒体上表示,这是一个“好主意,我们也会这样做。”
Amodei’s other major ideas for coordination pass the buck a little bit. The first calls for the coordinated development of “common safety standards” and “limits on the rate of unchecked AI progress” across all “frontier AI companies within democratic countries.” While Amodei spitballs some ideas for what these kinds of standards might look like, they all currently involve hand-wavy statements like, “models [that] have capability X … need to be accompanied by certifications of alignment properties Y and Z.” These standards would ideally be backstopped by “regulation that targets all US frontier AI companies” that don’t voluntarily comply, Amodei said, and presumably similar regulations in other democracies.
Amodei 关于协调的其他主要想法则显得有些推卸责任。首先,他呼吁在所有“民主国家的前沿 AI 公司”之间协调制定“通用安全标准”和“限制不受控的 AI 进步速度”。虽然 Amodei 对这些标准可能的样子提出了一些构想,但目前都还停留在模糊的表述上,例如“具备 X 能力的模型……需要附带 Y 和 Z 对齐属性的认证”。Amodei 表示,这些标准理想情况下应由“针对所有未自愿遵守的美国前沿 AI 公司的法规”作为后盾,其他民主国家也应出台类似的法规。