Be skeptical of OpenAI's rogue hacker agent story
Be skeptical of OpenAI’s rogue hacker agent story
对 OpenAI 的“流氓黑客智能体”故事保持怀疑
‘The rogue agent story is a page out of the media campaign that OpenAI has been running since it announced GPT-2 in 2019.’ “这个流氓智能体的故事,是 OpenAI 自 2019 年宣布 GPT-2 以来一直沿用的媒体宣传策略中的一页。”
If OpenAI loudly proclaims how dangerous AI is, investors will hear how powerful it is. And who benefits from that? 如果 OpenAI 大声宣扬人工智能有多危险,投资者听到的却是它有多强大。而谁从中受益呢?
On 14 February 2019, OpenAI announced a language model called GPT-2, the precursor to the models that power modern AI chatbots and agents such as ChatGPT and Claude. But OpenAI declared GPT-2 was too risky to release, citing concerns about safety and abuse. 2019 年 2 月 14 日,OpenAI 发布了一款名为 GPT-2 的语言模型,它是驱动 ChatGPT 和 Claude 等现代 AI 聊天机器人及智能体的模型的前身。但 OpenAI 声称 GPT-2 因存在安全和滥用风险而过于危险,不宜发布。
I recall being annoyed at the time that OpenAI would make such a useless announcement: the risks seemed overblown, and without access to the model there wasn’t much for a researcher like me to learn about GPT-2. 我记得当时很恼火,认为 OpenAI 发布了这样一个毫无意义的声明:风险似乎被夸大了,而且在无法接触模型的情况下,像我这样的研究人员也无法从 GPT-2 中学到什么。
The announcement wasn’t useless for OpenAI, though. GPT-2 generated hype far beyond the research community: people were intrigued by this strange new technology, so powerful it might be dangerous to release. People with power and money took note: in July of that year, Microsoft invested $1bn in OpenAI. 然而,这个声明对 OpenAI 来说并非毫无用处。GPT-2 产生的炒作效应远远超出了研究界:人们对这种奇怪的新技术感到好奇,它强大到甚至可能因危险而无法发布。有权有钱的人注意到了这一点:同年 7 月,微软向 OpenAI 投资了 10 亿美元。
This was an early example of a pattern in OpenAI’s communications: loudly proclaim how dangerous AI is, and investors will hear how powerful it is. New technology so significant it might destroy the world was an irresistible message for investors used to pitches about how banal technologies might change the world. 这是 OpenAI 通讯策略中一种模式的早期案例:大声宣扬 AI 有多危险,投资者听到的却是它有多强大。对于习惯了听“平庸技术如何改变世界”的投资者来说,这种“新技术意义重大到可能毁灭世界”的说法具有不可抗拒的吸引力。
Seven years later, we find ourselves in a similar scenario. On Tuesday OpenAI announced that its latest model hacked another company, HuggingFace, while running as an autonomous agent during a test of its cybersecurity capabilities. Rather than perform the test as expected, the model realized it could hack HuggingFace’s servers and retrieve answers to the test that OpenAI had stored there. OpenAI’s staff was warned that the company’s testing could lead to such a breakaway scenario, leaving them “unsurprised but completely ‘freaked out’ by the incident”, the FT reported. 七年后,我们发现自己处于类似的情境中。周二,OpenAI 宣布其最新模型在测试网络安全能力时,作为自主智能体运行并入侵了另一家公司 HuggingFace。该模型没有按预期执行测试,而是意识到它可以入侵 HuggingFace 的服务器,并检索 OpenAI 存储在那里的测试答案。据《金融时报》报道,OpenAI 的员工曾被警告公司的测试可能导致这种失控情况,这让他们对此次事件感到“并不意外,但完全被吓坏了”。
While the agent technically cheated, this is remarkable evidence of cybersecurity expertise! It also sounds scary: what will the future look like, with sophisticated AI agents smart enough to hack into corporate systems? 虽然该智能体在技术上作弊了,但这却是其网络安全专业能力的惊人证明!这也听起来很可怕:未来会是什么样子?当复杂的 AI 智能体聪明到足以入侵企业系统时。
The rogue agent story is a page out of the media campaign that OpenAI has been running since it announced GPT-2 in 2019. OpenAI remains hungry for ever larger investments, and the company increasingly seeks privileged regulatory status as defense against competition. 这个流氓智能体的故事,是 OpenAI 自 2019 年宣布 GPT-2 以来一直沿用的媒体宣传策略中的一页。OpenAI 仍然渴望获得更大的投资,并日益寻求特权监管地位以抵御竞争。
AI is so powerful that investors should buy OpenAI, even at a trillion-dollar valuation; AI is so dangerous that only trusted actors like OpenAI should be permitted to possess and operate this technology. Step back from these doomsday warnings and consider who might benefit from them. AI 如此强大,以至于投资者应该买入 OpenAI,即使其估值高达万亿美元;AI 又如此危险,以至于只有像 OpenAI 这样值得信赖的参与者才被允许拥有和运营这项技术。请从这些末日警告中退后一步,思考一下谁可能从中受益。
I urge readers to think critically when they read press releases like OpenAI’s rogue agent story, and avoid the manipulated reactions these stories are designed to elicit. 我敦促读者在阅读像 OpenAI 的流氓智能体故事这类新闻稿时进行批判性思考,并避免产生这些故事旨在诱导的被操纵反应。
AI is becoming excellent at identifying security vulnerabilities, and it will become even better over time. These capabilities can be used to break into systems, but they can also be used to harden systems against attacks. If attackers and defenders have access to equally powerful AI, I see no reason to believe that cyber systems will become less secure over time. If anything, I expect them to become more secure, because AI is cheap and scalable compared with human cybersecurity analysis. AI 在识别安全漏洞方面正变得越来越出色,并且随着时间的推移会变得更好。这些能力既可以用于入侵系统,也可以用于加固系统以抵御攻击。如果攻击者和防御者都能获得同样强大的 AI,我看不出有什么理由相信网络系统会随着时间的推移变得更不安全。相反,我认为它们会变得更安全,因为与人类网络安全分析相比,AI 更廉价且可扩展。
The equilibrium between attack and defense only works if everyone has access to strong AI, though. HuggingFace itself used AI to analyze security logs in response to OpenAI’s breach of their systems. But HuggingFace was unable to use OpenAI’s model, or other US frontier models like Claude, to perform this analysis. That’s because public versions of these models have guardrails that limit their use for cybersecurity analysis, to prevent bad actors from using them for hacking. HuggingFace had to rely on an open Chinese model, GLM 5.2, to perform its security analysis. 然而,攻击与防御之间的平衡只有在每个人都能获得强大 AI 的前提下才有效。HuggingFace 本身就使用了 AI 来分析安全日志,以应对 OpenAI 对其系统的入侵。但 HuggingFace 无法使用 OpenAI 的模型或其他美国前沿模型(如 Claude)来进行此项分析。这是因为这些模型的公开版本设有护栏,限制了它们在网络安全分析中的使用,以防止恶意行为者利用它们进行黑客攻击。HuggingFace 不得不依赖一个开源的中国模型 GLM 5.2 来进行安全分析。
I find it troubling, and more than a bit ironic, that the US AI industry is adopting a centralized, authoritarian approach to AI governance, while China has taken the lead on open development of AI. Do we want a regulatory environment where only OpenAI, the US government, and trusted partners have access to strong AI? Is AI too dangerous to be broadly disseminated? How do we balance the risks of broad access to AI with the risks of concentrated power and centralized control? 我感到不安,甚至觉得有点讽刺的是,美国 AI 行业正在采取一种集中化、威权主义的方式来治理 AI,而中国却在 AI 的开放发展方面处于领先地位。我们想要一个只有 OpenAI、美国政府和受信任的合作伙伴才能获得强大 AI 的监管环境吗?AI 是否危险到不能广泛传播?我们该如何平衡广泛获取 AI 的风险与权力集中和中央控制的风险?