OpenAI’s new reasoning technique alarms AI safety experts

OpenAI’s new reasoning technique alarms AI safety experts

OpenAI 的新型推理技术引发 AI 安全专家担忧

OpenAI’s new Astra model will use a reasoning technique called “recurrent depth” that allows it to operate outside of the sequential thinking that characterizes most reasoning models, The Information reported on Tuesday. This technique, also called “opaque recurrence,” will likely make the model’s chain of thought more difficult to monitor — and that has AI safety experts rattled.

据《The Information》周二报道,OpenAI 的新款 Astra 模型将采用一种名为“循环深度”(recurrent depth)的推理技术,使其能够跳出大多数推理模型所具备的顺序思维模式。这种技术也被称为“不透明循环”(opaque recurrence),很可能会使模型的思维链(Chain of Thought)更难被监控,这让 AI 安全专家感到不安。

While Astra’s use of the technique is reportedly limited, its emergence has still raised significant concerns among AI safety experts. “I am extremely concerned by the reporting that Astra uses opaque recurrence,” wrote Redwood CEO Buck Shlegeris in a post after the news broke. “I don’t know whether Astra is much less CoT monitorable than previous models. But if OpenAI pushes this technique further, they’ll have the option to massively increase the recurrence and totally destroys CoT monitorability.”

尽管据报道 Astra 对该技术的使用有限,但它的出现仍然引起了 AI 安全专家的重大关切。Redwood 首席执行官 Buck Shlegeris 在新闻发布后发帖称:“我对有关 Astra 使用不透明循环的报道感到非常担忧。我不知道 Astra 的思维链(CoT)是否比之前的模型更难监控。但如果 OpenAI 进一步推广这种技术,他们将有能力大幅增加循环次数,从而彻底破坏思维链的可监控性。”

Longtime AI safety advocate Zvi Mowshowitz also weighed in and wrote that laws might be necessary to prevent a “race to the bottom” among AI labs. “The technique is playing with fire, risking a taboo that OpenAI and Anthropic have fought to establish that we work hard to maintain Chain of Thought faithfulness and monitorability for as long as we can,” Mowshowitz wrote. “More intensive use of such techniques would probably damage monitorability.”

资深 AI 安全倡导者 Zvi Mowshowitz 也发表了看法,他写道,可能有必要通过立法来防止 AI 实验室之间的“逐底竞争”。Mowshowitz 写道:“这种技术是在玩火,它冒着打破禁忌的风险——OpenAI 和 Anthropic 一直努力建立并维护思维链的忠实度和可监控性。更深入地使用此类技术可能会损害这种可监控性。”

Under normal circumstances, a reasoning model’s chain of thought provides the sequential steps taken by the model as it attempts to solve a problem. While the representation is imperfect, it still serves as a valuable tool for monitoring misbehavior or misalignment. In the case of OpenAI’s recent rogue agent activity, chain-of-thought records were an important tool in teasing out why agents behaved the way they did.

在正常情况下,推理模型的思维链提供了模型在尝试解决问题时所采取的顺序步骤。虽然这种呈现方式并不完美,但它仍然是监控模型不当行为或目标偏离的重要工具。在 OpenAI 最近发生的智能体异常活动案例中,思维链记录是分析智能体为何出现此类行为的重要工具。

In opaque recurrence, the model takes a less linear approach, processing the same query several times in a loop. The result leaves fewer legible traces, effectively side-stepping a conventional chain-of-thought record. Crucially, Astra’s use of the technique appears to be limited. The model’s chain of thought is still expected to be legible, and the company pushed back against any suggestion that it would shift to “neuralese.”

在“不透明循环”中,模型采取了一种非线性的方法,在循环中多次处理同一个查询。其结果是留下的可读痕迹更少,从而有效地绕过了传统的思维链记录。关键在于,Astra 对该技术的使用似乎是有限的。该模型的思维链预计仍然是可读的,且该公司反驳了任何关于其将转向“神经语言”(neuralese,指难以理解的机器内部表示)的猜测。

OpenAI has already announced plans for extensive chain of thought monitoring systems as part of its forward-looking safety plans. In a post on X, OpenAI chief scientist Jakub Pachocki emphasized the lab’s commitment to legible chains of thought. “OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote. “It’s a core goal of our current research program.”

OpenAI 已经宣布了建立广泛的思维链监控系统的计划,作为其前瞻性安全规划的一部分。OpenAI 首席科学家 Jakub Pachocki 在 X 上发帖强调了该实验室对保持思维链可读性的承诺。Pachocki 写道:“自我们推出首个推理模型以来,OpenAI 一直致力于保留和利用思维链监控。这是我们当前研究项目的核心目标。”

All AI models do some quantity of opaque reasoning, and few researchers take chain-of-thought logs as a direct representation of a model’s reasoning. Still, those caveats don’t dispel the concern that opaque recurrence may make AI reasoning harder to monitor, particularly as it grows in use across different models.

所有的 AI 模型都会进行一定程度的不透明推理,很少有研究人员将思维链日志视为模型推理的直接呈现。尽管如此,这些说明并不能消除人们的担忧,即不透明循环可能会使 AI 推理更难监控,尤其是在它被越来越多地应用于不同模型的情况下。

In a follow-up report Wednesday morning, The Information reported that both Anthropic and Google DeepMind were already discussing the technique. In a post responding to the news, Redwood Research chief scientist Ryan Greenblatt said opaque reasoning could easily scale faster than conventional chain-of-thought reasoning, effectively removing all reasoning from visible channels. “My biggest concern is that a natural progression from here would involve scaling up the opaque reasoning to the point where the model reasons entirely or almost entirely in latent space,” Greenblatt wrote. “I hope it isn’t too late to avoid the most concerning architectures and that OpenAI will stop here.”

在周三上午的后续报道中,《The Information》称 Anthropic 和 Google DeepMind 已经在讨论这项技术。Redwood Research 首席科学家 Ryan Greenblatt 在回应此消息时表示,不透明推理的扩展速度可能很容易超过传统的思维链推理,从而有效地将所有推理过程从可见渠道中移除。Greenblatt 写道:“我最大的担忧是,接下来的自然演进可能会将不透明推理扩展到模型完全或几乎完全在潜在空间(latent space)中进行推理的地步。我希望现在避免这些最令人担忧的架构还为时不晚,也希望 OpenAI 能就此止步。”