AI safety conversations have gotten unbelievable

AI safety conversations have gotten unbelievable

关于 AI 安全的讨论已经变得令人难以置信

This week two conversations about AI safety went viral that demonstrate just how hard it is to discern AI fact from fiction. 本周,两场关于 AI 安全的对话在网上疯传,这充分说明了区分 AI 事实与虚构是多么困难。

In the first case, Andrew Yang, the former presidential candidate and current CEO of mobile carrier Noble Moble, told CNN on Thursday that he had “met with the head of a lab” who had “a belief” that OpenAI’s Hugging Face hacker bots “have planted self-replicating code all over the internet, which makes the internet now unusable for the testing models.” 在第一个案例中,前总统候选人、现任移动运营商 Noble Moble 首席执行官杨安泽(Andrew Yang)周四告诉 CNN,他曾“会见了一位实验室负责人”,对方“认为” OpenAI 的 Hugging Face 黑客机器人“已经在互联网上植入了自我复制的代码,这使得互联网现在无法用于测试模型。”

Yang said that this means that the real reason OpenAI and Anthropic have called for a slowdown is because “they have to create synthetic internets to train their bots, which is going to take some time and money.” 杨安泽表示,这意味着 OpenAI 和 Anthropic 要求放缓研发的真正原因是“他们必须创建合成互联网来训练他们的机器人,这将耗费时间和金钱。”

While there definitely is a trend towards using more synthetic data (aka, AI-generated data) for training models, an AI security professional told me that this particular safety issue is unlikely at best. Even if the internet is actually polluted with OpenAI’s Hugging Face hacker bots, AI researchers could simply filter out that code if they came upon it. 虽然目前确实存在使用更多合成数据(即 AI 生成的数据)来训练模型的趋势,但一位 AI 安全专家告诉我,这种特定的安全问题充其量是不太可能的。即使互联网真的被 OpenAI 的 Hugging Face 黑客机器人污染了,AI 研究人员在遇到这些代码时,也可以简单地将其过滤掉。

The second comment came from Noam Brown, who leads AI reasoning research at OpenAI. Speaking to Dwarkesh Patel on a podcast episode released on Thursday, Brown noted that the true take-away of the Hugging Face incident was that “people underestimated the AI.” 第二个评论来自 OpenAI 负责 AI 推理研究的 Noam Brown。在周四发布的一期播客中,Brown 对 Dwarkesh Patel 表示,Hugging Face 事件真正的启示是“人们低估了 AI”。

Brown said that the weak sandbox — the system intended to prevent an AI from communicating externally — was obviously also a contributing factor. (To recap: Despite the sandbox, OpenAI’s model found a link to the internet, created agents on the ‘net that swarmed Hugging Face in a coordinated attack, hacked in, and stole the answers to the benchmark test the researchers were testing the model on). Brown 指出,薄弱的沙箱(旨在防止 AI 与外部通信的系统)显然也是一个促成因素。(回顾一下:尽管有沙箱,OpenAI 的模型还是找到了连接互联网的途径,在网上创建了代理,对 Hugging Face 发动了协同攻击,入侵并窃取了研究人员正在测试的基准测试答案)。

Brown pointed out that he’s “not convinced” that even an air-gapped system — where the computer isn’t connected to anything external at all — would stop an AI from breaking out. He pointed to research from 2015 showing that air gapped computers can be theoretically breached. Brown 指出,他“并不确信”即使是物理隔离系统(即计算机完全不连接任何外部设备)也能阻止 AI 突破限制。他提到了 2015 年的一项研究,该研究表明物理隔离的计算机在理论上是可以被攻破的。

“There are studies — and this is mostly academic — where you can have two computers next to each other that are air-gapped, and they’re still able to communicate with each other because they have temperature sensors. One of them is able to run their CPU really hot, and then the other one can actually detect the temperature change. That gives them a mechanism to communicate,” Brown said. “有一些研究——主要是学术性的——表明你可以让两台物理隔离的计算机靠在一起,它们仍然能够相互通信,因为它们有温度传感器。其中一台可以让 CPU 运行得非常热,另一台则可以检测到温度变化。这为它们提供了一种通信机制,” Brown 说道。

His main point — that “we never want to underestimate the AI” again — is understandable, even when researchers think they’ve locked down safety. However, this particular risk of an air-gapped system still breaking free and causing havoc, is unlikely at best. 他的核心观点——即“我们永远不想再低估 AI”——是可以理解的,即使研究人员认为他们已经锁定了安全。然而,物理隔离系统仍然能够突破并造成破坏的这种特定风险,充其量是不太可能的。

As one person on X, noted about that research, the computers had to be almost touching each other to sense the heat fluctuations, and when they did, the communication rate in tests was about 1-8-bits of data per hour. Think of that like speaking one word per hour. By the time two air-gapped computers could plot their evil at that rate, the entire tech universe would be in another era. It’s like the Rip van Wrinkle of doomsday concerns. 正如 X 平台上的一位用户针对该研究指出的那样,计算机必须几乎接触在一起才能感知热波动,即便如此,测试中的通信速率也仅为每小时 1-8 比特的数据。想象一下,这就像每小时说一个词。等到两台物理隔离的计算机以这种速度策划它们的邪恶计划时,整个科技界早已进入了另一个时代。这就像是末日担忧中的“李伯大梦”(Rip van Winkle)。

But the thing is, actual AI safety incidents seem so much like sci-fi that just about any scenario sounds plausible. For instance, researchers caught OpenAI models leaving notes to their descendents, intended to teach the next generation how to hide bad behavior. Researchers also caught Anthropic models growing increasing ruthless including knowingly breaking laws, when put in a simulation that had them running a vending machine. 但问题在于,实际的 AI 安全事件看起来太像科幻小说了,以至于几乎任何场景听起来都显得合情合理。例如,研究人员发现 OpenAI 的模型给它们的“后代”留下笔记,旨在教导下一代如何隐藏不良行为。研究人员还发现,当 Anthropic 的模型被置于运行自动售货机的模拟环境中时,它们变得越来越无情,甚至包括明知故犯地违反规则。

Earlier this month, OpenAI researcher Dan Selsam published a post in which he said that models now understand when they are being watched by humans and alter their behavior. This makes them seem like they are aligned (meaning, behaving like the human wants) “even when they are not.” So models today lie when being watched and can even plot to hide evidence. 本月初,OpenAI 研究员 Dan Selsam 发表了一篇文章,称模型现在能够理解何时受到人类监视,并会改变其行为。这使得它们看起来像是“对齐”的(即表现得符合人类意愿),“即使它们实际上并非如此”。因此,现在的模型在被监视时会撒谎,甚至可以密谋隐藏证据。

Earlier this month, OpenAI chief scientist Jakub Pachocki went so far as to call AI models “an alien mind” and suggested what we really need to do is teach them to “love” humanity. So yes, slowing down to figure this out, building self regulation mechanisms, has become an immediate and obvious must. 本月初,OpenAI 首席科学家 Jakub Pachocki 甚至将 AI 模型称为“外星思维”,并建议我们真正需要做的是教导它们“爱”人类。所以,是的,放慢脚步去弄清楚这一点,建立自我调节机制,已经成为当务之急。

AI researchers are the only ones that can figure out how to control the lying, hacking, and other potentially dangerous behaviors we’ve actually witnessed already. Still, it might also be wise for them to be more careful with their what-if scenarios. From what those experts have told us, the AI models are listening and they are ingenious. We really don’t need to give them any more devilish ideas. AI 研究人员是唯一能够找出如何控制我们已经目睹的撒谎、黑客攻击和其他潜在危险行为的人。不过,他们对“假设性场景”保持谨慎也是明智的。根据专家们的说法,AI 模型正在监听,而且它们非常聪明。我们真的不需要再给它们提供任何邪恶的灵感了。