Here’s How an AI Slowdown Could Actually Be Enforced
Here’s How an AI Slowdown Could Actually Be Enforced
人工智能发展放缓究竟该如何落实?
Many AI researchers seem to firmly believe that the technology they are developing could someday prove very dangerous. What’s less clear—even among AI’s technical elite—is precisely how to keep these mercurial algorithms in check. 许多人工智能研究人员似乎坚信,他们正在开发的这项技术有朝一日可能会变得非常危险。然而,即便是在人工智能领域的技术精英中,对于如何确切地控制这些变幻莫测的算法,依然缺乏清晰的共识。
In recent years researchers have thrown around all sorts of ideas for preventing AI from turning nasty. They include less controversial plans such as tighter government regulations, new ways of measuring progress, and probing the inner workings of models, as well as more outlandish proposals like placing tracking devices inside GPUs, and even ceremonially destroying large numbers of AI chips. 近年来,研究人员提出了各种各样的想法来防止人工智能“变坏”。这些想法包括一些争议较小的方案,如加强政府监管、建立衡量进展的新方法以及探测模型的内部运作机制;也包括一些更为离奇的提议,例如在 GPU 中安装追踪设备,甚至以仪式性的方式销毁大量人工智能芯片。
With political and public pressure now growing for a more measured approach to building AI, however, the answer to keeping AI safe is still unclear. 然而,随着政治和公众对人工智能开发采取更审慎态度的呼声日益高涨,如何确保人工智能安全的答案依然不明朗。
“We need to start treating this as a research problem,” says Raymond Douglas, an AI researcher at the University of Toronto and coauthor of a new report titled Pacing the Frontier, A Research Agenda, which warns that slowing down AI development remains an unsolved puzzle. “We don’t really understand what our options even are or what they will do.” “我们需要开始将此视为一个研究课题,”多伦多大学人工智能研究员、新报告《放缓前沿步伐:研究议程》(Pacing the Frontier, A Research Agenda)的合著者雷蒙德·道格拉斯(Raymond Douglas)表示。该报告警告称,放缓人工智能发展仍是一个未解之谜。“我们甚至不真正了解我们有哪些选择,也不清楚这些选择会带来什么后果。”
Talk of AI doom has reached a fever pitch in recent weeks after an Anthropic researcher left the company and warned that within a couple of years, AI might be on course to wipe out humanity. The head of Anthropic’s AI safety lab swiftly echoed his concerns. 近几周,关于人工智能末日的讨论达到了白热化程度。此前,一名 Anthropic 的研究人员离职并警告称,在几年内,人工智能可能会走上毁灭人类的道路。Anthropic 人工智能安全实验室的负责人迅速响应了他的担忧。
The leaders of America’s big AI companies—Dario Amodei of Anthropic, Sam Altman of OpenAI, Elon Musk of SpaceXAI, and Demis Hassabis of Google DeepMind—have all now chimed in to offer support for some sort of AI slowdown or pause. 美国各大人工智能公司的领导者——Anthropic 的达里奥·阿莫代(Dario Amodei)、OpenAI 的萨姆·奥特曼(Sam Altman)、SpaceXAI 的埃隆·马斯克(Elon Musk)以及 Google DeepMind 的德米斯·哈萨比斯(Demis Hassabis)——现在都已表态,支持以某种形式放缓或暂停人工智能的发展。
The issue seems especially pressing because AI companies are now using AI itself to build ever-more powerful models. This has sparked fears of an accelerating recursive self-improvement (RSI) loop that would see AI outstrip humans’ ability to comprehend what it is up to within a few years. 这个问题显得尤为紧迫,因为人工智能公司现在正利用人工智能本身来构建更强大的模型。这引发了人们对“递归自我改进”(RSI)循环加速的担忧,即人工智能可能在几年内超越人类理解其行为的能力。
The AI labs are already touting new approaches of their own. This week Anthropic announced several new ways to track how rapidly—and perhaps dangerously—artificial intelligence is advancing. The techniques show, for example, that Claude now does 26 percent of Anthropic’s AI research, compared to zero at the beginning of 2026. They also reveal that Anthropic spent 6 percent of its compute budget on figuring out how to make its AI safer. 各大人工智能实验室已经在推行他们自己的新方法。本周,Anthropic 宣布了几种追踪人工智能发展速度(以及潜在危险性)的新方法。例如,这些技术显示,Claude 目前承担了 Anthropic 26% 的人工智能研究工作,而 2026 年初这一比例为零。他们还透露,Anthropic 将其计算预算的 6% 用于研究如何提高人工智能的安全性。
But Douglas and other experts say controlling AI development effectively and reliably will require funding and expertise from outside the AI labs themselves. Some of the proposed solutions—both from this latest report and beyond—seem more within reach than others. 但道格拉斯和其他专家表示,要有效且可靠地控制人工智能的发展,需要来自人工智能实验室外部的资金和专业知识。其中一些提议的解决方案——无论是来自这份最新报告还是其他渠道——似乎比其他方案更具可行性。
‘Independent’ Evaluators / “独立”评估者
One idea often floated by AI companies is giving third-party evaluators greater access to their models. These evaluators test models to assess their capabilities and “red team” them by trying to elicit misbehavior within trusted environments. 人工智能公司经常提出的一个想法是,给予第三方评估者对其模型更大的访问权限。这些评估者通过在受信任的环境中测试模型以评估其能力,并进行“红队测试”,试图诱导模型出现错误行为。
Geoffrey Irving, former chief scientist at the UK AI Security Institute, and before that a researcher at Google DeepMind, believes rigorous inspections could effectively pause the development of frontier AI for now. “In the near term, inspections and audits work, or even just mutual agreements,” Irving says. “I do think the companies are afraid of RSI and misaligned takeoff.” 英国人工智能安全研究所前首席科学家、曾任 Google DeepMind 研究员的杰弗里·欧文(Geoffrey Irving)认为,严格的检查目前可以有效地暂停前沿人工智能的发展。“在短期内,检查和审计是有效的,甚至仅仅是相互协议也行,”欧文说,“我确实认为这些公司害怕 RSI 和失控的风险。”
Some doomsayers argue that such inspections would need to be more independent and scientifically rigorous than they currently are. The fact that some AI agents have recently escaped containment during testing certainly seems to suggest that more rigor may be required. 一些末日论者认为,此类检查需要比目前更加独立且科学严谨。事实上,近期一些人工智能代理在测试中逃脱了限制,这似乎确实表明需要更严格的监管。
Connor Leahy, head of Control AI, a nonprofit that advocates for AI controls, says inspections should involve the FBI or the NSA. “When [big AI companies] say ‘independent evaluators,’ they mean ‘I want to pay my friends who live in my group houses to look at my prompts.’” 倡导人工智能管控的非营利组织 Control AI 的负责人康纳·莱希(Connor Leahy)表示,检查工作应涉及联邦调查局(FBI)或国家安全局(NSA)。“当(大型人工智能公司)说‘独立评估者’时,他们的意思是‘我想付钱给我住在一起的朋友,让他们来看看我的提示词。’”
Douglas says new research could also improve model evaluations. He points to recent work showing how outsiders can examine usage of models without disclosing any confidential information. Other techniques that may prove helpful include new ways of peering inside AI models to get a better sense of what they are doing. 道格拉斯表示,新的研究也可以改进模型评估。他指出,最近的研究表明,外部人员如何在不泄露任何机密信息的情况下检查模型的使用情况。其他可能有所帮助的技术还包括深入观察人工智能模型内部的新方法,以更好地了解它们在做什么。
Leahy agrees there is a need for more research on model evaluation as well as what it actually means to “align” a model, or make it reflect human values, in the first place. “There has been a very deliberate marketing campaign from these companies to try to present evaluations as scientific,” he says. “But we don’t actually understand how AI works.” 莱希同意,需要对模型评估进行更多研究,并明确“对齐”模型(即使其反映人类价值观)的真正含义。“这些公司进行了一场非常刻意的营销活动,试图将评估包装成科学,”他说,“但我们实际上并不了解人工智能是如何运作的。”
How much the US government is willing to step in to restrict the development of AI is uncertain. President Trump has largely dismissed the need to regulate the industry, but there are signs that bipartisan support is growing for reigning in big AI. 美国政府愿意在多大程度上介入以限制人工智能的发展尚不确定。特朗普总统在很大程度上否认了监管该行业的必要性,但有迹象表明,两党在加强对大型人工智能公司监管方面的支持正在增加。
Trusted Compute / 可信计算
Some experts believe that imposing limits on the development of AI should ultimately involve checks on the raw compute required. The most powerful models are trained using thousands of cutting-edge Nvidia GPUs inside vast data centers. 一些专家认为,对人工智能的发展施加限制,最终应涉及对所需原始计算能力的检查。最强大的模型是在大型数据中心内使用数千个尖端的英伟达(Nvidia)GPU 训练出来的。
The government has dabbled with tracking this already, through a 2023 Biden-era AI executive order that required companies to report training runs above a certain compute threshold. 政府已经尝试通过 2023 年拜登政府时期的一项人工智能行政命令来追踪这一点,该命令要求公司报告超过特定计算阈值的训练任务。
A policy white paper from March 2024 argues that cloud providers could be crucial to future efforts because of their visibility into major AI training runs. The white paper suggests that tracking billing records, GPU utilization, network traffic, and power consumption could provide proxies for AI capabilities. 2024 年 3 月的一份政策白皮书指出,云服务提供商对未来的工作至关重要,因为他们能够洞察主要的人工智能训练任务。该白皮书建议,追踪账单记录、GPU 利用率、网络流量和电力消耗可以作为衡量人工智能能力的代理指标。
Experts have also proposed ways of tracking and controlling efforts to build advanced AI by modifying chips themselves. 专家们还提出了通过修改芯片本身来追踪和控制构建先进人工智能工作的方法。
One idea, put forward by researchers at RAND in 2024, would involve modifying an existing component on GPUs used to measure performance so that it performs a cryptographically secured record of compute runs that can be inspected periodically. This could reveal, for example, that a company has been training AI above a certain threshold. 兰德公司(RAND)的研究人员在 2024 年提出的一个想法是,修改 GPU 上用于衡量性能的现有组件,使其能够对计算任务进行加密安全记录,并可定期检查。例如,这可以揭示某家公司是否一直在超过特定阈值的情况下训练人工智能。
Others have suggested building new kinds of tamper-proof components. 其他人则建议构建新型的防篡改组件。