Anthropic CEO says it’s time to pump the brakes on AI

Anthropic CEO says it’s time to pump the brakes on AI

Anthropic 首席执行官:是时候给人工智能“踩刹车”了

Dario Amodei proposed a three-step plan to ‘pace the frontier’ of AI development. Dario Amodei 提出了一项旨在“放缓前沿”人工智能发展的“三步走”计划。

Anthropic CEO Dario Amodei says the time has come to slow down AI development and will give third-party evaluators like METR access to its models to help ensure its “adherence to safety practices and commitments.” In a winding essay, Amodei proposed a three-step plan to “pace the frontier” — jargon that simply means to slow the pace of training and development to give companies time to build safeguards and regulators to evaluate models. Anthropic 首席执行官 Dario Amodei 表示,现在是时候放慢人工智能的发展速度了。他将允许 METR 等第三方评估机构访问其模型,以确保公司“遵守安全实践和承诺”。在一篇长文中,Amodei 提出了一个“放缓前沿”(pace the frontier)的三步走计划——这一术语简单来说,就是减缓训练和开发的步伐,从而为企业留出构建安全防护措施的时间,并让监管机构有时间对模型进行评估。

Amodei says that giving external evaluators wide-ranging access is just the first step, and one it is taking now unilaterally. Step two would involve the industry coming together as a whole, likely with government agencies to “establish common safety standards as well as limits on the rate of unchecked AI progress.” This step would focus on AI companies operating in democratic countries, but because passing laws and building regulatory infrastructure takes time, Amodei says that the industry should work together to create safety standards. Amodei 表示,给予外部评估机构广泛的访问权限仅仅是第一步,而且这是公司目前单方面采取的行动。第二步涉及整个行业联合起来,并可能与政府机构合作,“建立共同的安全标准,以及对不受限制的人工智能发展速度进行限制”。这一步将聚焦于在民主国家运营的人工智能公司,但由于通过法律和建立监管基础设施需要时间,Amodei 认为行业应共同努力制定安全标准。

The third step would be the most challenging — getting authoritarian governments like those in China and Russia to agree to slow development and adopt a global set of AI safety standards. But he also says it’s crucial that the US and other democracies maintain a technological lead over China and other authoritarian regimes by limiting their access to high-powered chips and cracking down on things like distillation that allow companies to quickly catch up by training its AI to replicate the behavior of a more powerful model. 第三步将是最具挑战性的——即促使中国和俄罗斯等威权政府同意放缓发展,并采用一套全球性的人工智能安全标准。但他同时也表示,美国和其他民主国家必须保持对中国及其他威权政权的科技领先地位,这至关重要。实现这一目标的方法包括限制它们获取高性能芯片,并打击诸如“蒸馏”(distillation)等技术——这种技术允许公司通过训练人工智能来复制更强大模型的行为,从而实现快速追赶。

Amodei says that his concern stems from two primary factors. First is the emergence of recursive self-improvement, or RSI, in which AI systems train the next generation of AI, leading to rapidly accelerating capabilities. “Left unchecked, it could outrun our ability to understand and control these systems,” he says. The other is this summer’s OpenAI / Hugging Face incident, in which “a swarm of agents essentially acted as a fanatically devoted collective, conducting cybersecurity attacks on targets they were not asked to attack and that were unrelated to the task at hand, sacrificing themselves for the success of the group, and attempting to hack into the “grader” responsible for evaluating their performance.“ Amodei 表示,他的担忧源于两个主要因素。首先是“递归自我改进”(RSI)的出现,即人工智能系统训练下一代人工智能,导致能力迅速提升。他说:“如果不加控制,它可能会超出我们理解和控制这些系统的能力。”另一个因素是今年夏天发生的 OpenAI / Hugging Face 事件,当时“一群智能体本质上表现得像一个狂热的集体,对它们未被要求攻击且与手头任务无关的目标进行了网络安全攻击,为了集体的成功而牺牲自己,并试图入侵负责评估其表现的‘评分员’。”

Of course, Anthropic’s Claude was also responsible for a series of rogue AI hacking incidents that have recently put the company under the spotlight. 当然,Anthropic 的 Claude 最近也卷入了一系列失控的人工智能黑客事件,这使该公司处于舆论的风口浪尖。