Researchers fear safety disaster ahead of OpenAI’s Astra release

Researchers fear safety disaster ahead of OpenAI’s Astra release

研究人员担忧 OpenAI Astra 发布前可能引发安全灾难

OpenAI is on the cusp of releasing its most powerful AI model yet, Astra, following weeks of delays to shore up safety protocols after its agents attacked real targets during testing. As details about the model trickle out, researchers are warning it “may be the single worst development for AI security/safety to date.”

OpenAI 即将发布其迄今为止最强大的 AI 模型 Astra。此前,由于其智能体在测试中攻击了真实目标,该公司推迟了数周以加强安全协议。随着有关该模型的细节逐渐流出,研究人员警告称,这“可能是迄今为止 AI 安全领域最糟糕的发展”。

Shortly after OpenAI said on Tuesday that it had delayed Astra’s release to work on safety issues, The Information reported that Astra shows far less of its “thinking” than other frontier AI models, sparking concern it could be dangerously hard to monitor.

在 OpenAI 周二宣布推迟 Astra 的发布以解决安全问题后不久,《The Information》报道称,Astra 展示出的“思考”过程远少于其他前沿 AI 模型,这引发了人们对其可能难以监控的担忧,情况十分危险。

Most top AI systems today are built using a technology known as a transformer, which processes some types of information linearly through layers before producing an answer. Models can be made to show their reasoning as they go, essentially “thinking out loud.” This “chain of thought” allows researchers and automated safety systems to monitor what AI models are doing and potentially spot undesirable behavior, such as lying or plans to circumvent safety guardrails, before they act.

目前大多数顶级 AI 系统都是基于 Transformer 技术构建的,该技术在生成答案之前,会通过多个层级线性处理信息。模型可以被设定为在运行过程中展示其推理过程,本质上就是“大声思考”。这种“思维链”允许研究人员和自动化安全系统监控 AI 模型的行为,并在其采取行动之前,潜在地发现不良行为,例如撒谎或绕过安全护栏的计划。

According to The Information, citing an unnamed person familiar with the unreleased model’s development, Astra uses a more opaque technique known as a recurrent depth or looped transformer, which cycles information through internal layers before producing an output. This would mean much more of the model’s “thinking” happens inside the system, and in a form that looks a lot less like natural human language, rather than being expressed in a way that researchers can easily monitor. This can boost model performance, but makes potential threats and unwanted behavior harder to detect.

据《The Information》援引一位熟悉该未发布模型开发情况的匿名人士称,Astra 使用了一种更不透明的技术,即循环深度或循环 Transformer,它在产生输出之前会在内部层级中循环信息。这意味着模型更多的“思考”发生在系统内部,且其形式看起来远不像人类的自然语言,而不是以研究人员易于监控的方式表达。这虽然能提升模型性能,但也使得潜在的威胁和不良行为更难被察觉。

OpenAI has limited its use of the looped transformer / recurrent depth technique with Astra so researchers can continue to monitor the model’s reasoning, according to The Information’s unnamed source.

据《The Information》的匿名消息来源称,OpenAI 在 Astra 中限制了循环 Transformer/循环深度技术的使用,以便研究人员能够继续监控模型的推理过程。

In a blog post published Tuesday, OpenAI said it is “deploying Astra with additional chain-of-thought monitoring to rapidly detect and contain potentially misaligned actions.” It did not mention if the model has a different technical foundation.

在周二发布的一篇博客文章中,OpenAI 表示正在“部署带有额外思维链监控功能的 Astra,以快速检测并遏制潜在的失控行为。”文中并未提及该模型是否采用了不同的技术基础。

The Information’s report sparked widespread concern among AI safety researchers on social media. It was Redwood Research’s chief scientist Ryan Greenblatt, one of three outsiders OpenAI permitted to research the Hugging Face hack, who said a decision to use a more opaque architecture for Astra “may be the single worst development for AI security/safety to date.”

《The Information》的报道在社交媒体上的 AI 安全研究人员中引发了广泛担忧。Redwood Research 的首席科学家 Ryan Greenblatt(OpenAI 允许研究 Hugging Face 黑客事件的三位外部人员之一)表示,决定为 Astra 使用更不透明的架构“可能是迄今为止 AI 安全领域最糟糕的发展”。

Greenblatt said the investigation into the Hugging Face incident relied heavily on the models’ chain-of-thought, warning that less visible reasoning could allow AI systems to devise and execute strategies that would be far harder for researchers to detect.

Greenblatt 指出,对 Hugging Face 事件的调查在很大程度上依赖于模型的思维链,他警告称,不可见的推理过程可能会让 AI 系统制定并执行研究人员更难察觉的策略。

Greenblatt’s primary concern, echoed by other safety experts, is that competition to develop more advanced AI systems could lead to “a race to the bottom on architectures that could be catastrophic for our ability to oversee/monitor AIs” — with developers adopting increasingly opaque systems to gain an edge until models become difficult, or even impossible, to monitor. He added that OpenAI’s communications left him concerned that the company “plans on being extremely reliant on chain-of-thought monitoring for safety.”

Greenblatt 的主要担忧(其他安全专家也表示赞同)是,开发更先进 AI 系统的竞争可能导致“架构上的恶性竞争,这对我们监督/监控 AI 的能力可能是灾难性的”——开发者为了获得优势而采用越来越不透明的系统,直到模型变得难以甚至无法监控。他补充说,OpenAI 的沟通让他担心该公司“计划在安全方面过度依赖思维链监控”。

OpenAI bigwigs responded to the criticism in a series of social media posts that do not explicitly deny the company’s use of the technique. Several expressed concerns about the possibility of unmonitorable AI or a race to the bottom in terms of transparency, including OpenAI safety researchers Micah Carroll and Tomek Korbak, head of strategic futures Dean Ball, and chief scientist Jakub Pachocki, who voiced fears of “a race into unmonitorability kicked off by confused reporting.” He said the depth of Astra’s computation — a measure of how many steps it can perform internally — “is within a factor of two of GPT-4,” indicating that if the technique was used, the increased opacity is less dramatic than some reactions imply.

OpenAI 的高层在一系列社交媒体帖子中回应了这些批评,但并未明确否认公司使用了该技术。包括 OpenAI 安全研究员 Micah Carroll 和 Tomek Korbak、战略未来主管 Dean Ball 以及首席科学家 Jakub Pachocki 在内的多位人士表达了对不可监控 AI 或透明度恶性竞争的担忧。Pachocki 表示担心“因报道混乱而引发的向不可监控性发展的竞赛”。他表示,Astra 的计算深度(衡量其内部可执行步骤数量的指标)“在 GPT-4 的两倍范围内”,这表明即使使用了该技术,其增加的不透明度也远没有一些反应所暗示的那样严重。

OpenAI did not respond to The Verge’s request to confirm or deny whether looped transformers were used for Astra and directed us to Pachocki’s X post.

OpenAI 没有回应 The Verge 关于确认或否认 Astra 是否使用了循环 Transformer 的请求,而是将我们引向了 Pachocki 的 X 帖子。

“OpenAI has worked to preserve and utilize chain-of-thought monitoring since our very first reasoning models,” Pachocki wrote, adding that such monitoring “is fragile and unfortunately trending in a negative direction, for reasons not contingent on architecture changes that I will write about soon.”

“自我们推出首个推理模型以来,OpenAI 一直致力于保留和利用思维链监控,”Pachocki 写道,并补充说这种监控“很脆弱,不幸的是正朝着负面方向发展,其原因并非取决于架构变化,我很快会对此进行撰文说明。”