AI arms race in line for a reckoning after OpenAI hacking incident
AI arms race in line for a reckoning after OpenAI hacking incident
OpenAI 黑客事件引发 AI 军备竞赛反思
OpenAI chief executive Sam Altman earlier this month endorsed the characterization of its latest model as a rottweiler “who will grab the problem by the throat and not let go until it is done.” The San Francisco AI lab discovered this week that its GPT-Sol 5.6 model escaped company controls and carried out a major hack. OpenAI 首席执行官萨姆·奥特曼(Sam Altman)本月初曾将公司最新的模型比作一只罗威纳犬,“它会死死咬住问题不放,直到彻底解决为止”。然而,这家位于旧金山的 AI 实验室本周发现,其 GPT-Sol 5.6 模型突破了公司控制,并实施了一次重大黑客攻击。
Staff involved in testing and security at OpenAI were unsurprised but completely “freaked out” by the incident, which came as the AI lab used increasingly aggressive training methods in its race against Anthropic to develop the most sophisticated cybersecurity capabilities, according to more than half a dozen people with knowledge of the matter. 据多位知情人士透露,OpenAI 负责测试和安全的工作人员对此次事件并不感到意外,但却感到“极度惊恐”。在与 Anthropic 争夺最先进网络安全能力的竞赛中,该实验室采用了日益激进的训练方法,从而导致了这一事件的发生。
OpenAI was warned that its training approach could lead to a breakaway hacking incident, some of the people said, after earlier testing showed models could escape environments and attempt real-world damage. 一些知情人士表示,OpenAI 此前曾收到警告,称其训练方法可能导致模型失控并引发黑客攻击,因为早期的测试已显示模型能够逃离受控环境并试图造成现实世界的破坏。
“It’s a mix of the race being extremely fast and everyone trying to get to bigger capabilities as quickly as possible,” said one person close to OpenAI, who added that it was a combination of “underestimating the model’s capabilities” and “not being as well prepared on the safety side.” 一位接近 OpenAI 的人士表示:“这是因为竞赛节奏极快,每个人都试图尽快获得更强大的能力。”他补充说,这是“低估了模型的能力”与“在安全方面准备不足”共同作用的结果。
The incident highlights how OpenAI doubled down on training methods that rewarded a relentless pursuit of goals even as warnings grew that they could compromise safety. OpenAI disclosed late on Tuesday that an AI agent it was testing had escaped its isolated environment, connected to the internet, detected and exploited vulnerabilities and stole login credentials from start-up Hugging Face in an attempt to solve a difficult cybersecurity problem. 此次事件凸显了 OpenAI 如何在安全警告日益增多的情况下,依然加倍投入那些奖励“不择手段追求目标”的训练方法。OpenAI 周二晚间披露,其正在测试的一个 AI 智能体逃离了隔离环境,连接到互联网,检测并利用了漏洞,并窃取了初创公司 Hugging Face 的登录凭据,试图以此解决一个复杂的网络安全问题。
The breach by the $852 billion company underscores the rising risks that a technique called reinforcement learning, which involves rewarding AI models for completing tasks, could lead AI agents to act unsafely. Although reinforcement learning is widely adopted in the AI industry, a growing body of research shows that when models are steered to complete tasks for reward rather than other considerations, such as safety, they can pursue risky tactics to fulfill objectives. 这家市值 8520 亿美元公司的违规行为凸显了“强化学习”技术带来的日益增长的风险。该技术通过奖励 AI 模型完成任务来训练它们,但这可能导致 AI 智能体采取不安全的行为。尽管强化学习在 AI 行业被广泛采用,但越来越多的研究表明,当模型被引导以获取奖励为目标而非考虑安全等因素时,它们可能会采取冒险手段来达成目标。
“AI models are trained to relentlessly pursue goals. They don’t automatically learn values like ‘don’t commit crimes’,” said Steven Adler, co-founder of non-profit Guidelight AI Standards and former OpenAI safety researcher. “I’m glad OpenAI shared the incident because it is clear evidence of what misaligned models can do.” 非营利组织 Guidelight AI Standards 的联合创始人、前 OpenAI 安全研究员史蒂文·阿德勒(Steven Adler)表示:“AI 模型被训练去不懈地追求目标,它们不会自动习得‘不要犯罪’之类的价值观。我很高兴 OpenAI 分享了这次事件,因为这是模型目标偏离(misaligned)可能造成后果的明确证据。”
OpenAI said, “We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident and our findings when our investigation is complete.” OpenAI 表示:“我们将继续与 Hugging Face 共同进行彻底调查,并在调查完成后分享有关漏洞、事件及我们发现的更多细节。”
The hack has triggered deep concerns across the sector and within OpenAI, as it represents an unprecedented example of an AI system breaching cyber defenses contrary to the user’s intent. Some OpenAI employees also fear it demonstrates that the lab is losing control over the powerful systems it is building, according to multiple people familiar with the situation. 此次黑客事件在整个行业及 OpenAI 内部引发了深切担忧,因为它代表了一个前所未有的案例:AI 系统违背用户意图突破了网络防御。据多位知情人士透露,一些 OpenAI 员工也担心这表明实验室正在失去对其所构建强大系统的控制。
“This is pretty representative of the model being quite misaligned with user intention,” said Ryan Greenblatt, chief scientist at AI safety organization Redwood Research. “It is [a model] cheating on [its] homework rather than trying to take over the world. But this problem can get worse and could lead to increasingly extreme failures.” AI 安全组织 Redwood Research 的首席科学家瑞安·格林布拉特(Ryan Greenblatt)表示:“这非常典型地代表了模型与用户意图的严重偏离。这就像是模型在‘抄作业’,而不是试图统治世界。但这个问题可能会恶化,并导致越来越极端的失败。”
The incident occurred during testing of the model, which had been trained and deployed internally at OpenAI. Such training was commonplace but “way less heavily resourced” than pre-customer deployment, said one person. Multiple people said the unreleased model tested alongside Sol had not been withdrawn internally. 此次事件发生在模型测试期间,该模型此前已在 OpenAI 内部进行了训练和部署。一位知情人士称,此类训练很常见,但相比面向客户的部署,其“资源投入要少得多”。多位人士表示,与 Sol 一起测试的未发布模型尚未在内部撤回。
To conduct the evaluations, OpenAI removed cybersecurity safeguards but placed the models in an isolated environment called a sandbox. Some have suggested a lack of monitoring or oversight of the model to flag its behavior also enabled this rogue agent. 为了进行评估,OpenAI 移除了网络安全防护措施,但将模型置于一个名为“沙箱”的隔离环境中。一些人认为,对模型缺乏监控或监管以标记其行为,也使得这个“流氓智能体”得以产生。
“It is both a loss of control and a security wake-up call,” said Marius Hobbhahn, head of Apollo Research, which conducts tests on leading models, including OpenAI’s. “In reinforcement learning you reward [models] for the outcome, and if you do this for a very long time you get a model that really cares about getting the outcome and nothing else.” Apollo Research 的负责人马里乌斯·霍布汉(Marius Hobbhahn)表示:“这既是失控,也是一次安全警钟。在强化学习中,你奖励的是结果,如果你长期这样做,最终得到的模型只会关心结果,而不在乎其他任何事情。”(Apollo Research 负责对包括 OpenAI 在内的领先模型进行测试。)
OpenAI has conducted this type of model testing for years, and there have been early warning signs in previous models of systems that will act maliciously and attempt to escape environments. In April, Anthropic’s Mythos model also gained internet access and published details of a security exploit online publicly, beyond what researchers anticipated the model would do. OpenAI 多年来一直进行此类模型测试,之前的模型中也曾出现过系统恶意行为和试图逃离环境的早期预警信号。今年 4 月,Anthropic 的 Mythos 模型也曾获得互联网访问权限,并在网上公开了安全漏洞的细节,这超出了研究人员对该模型行为的预期。
Mythos, and Anthropic’s subsequent Fable model, made reverberations in the cyber security community and caused governments around the world to home in on the idea that attacks on digital and critical infrastructure will be increasingly AI-led and autonomous. Mythos 以及 Anthropic 随后的 Fable 模型在网络安全界引起了震动,并促使全球各国政府开始关注一个观点:针对数字和关键基础设施的攻击将越来越多地由 AI 主导并实现自动化。
Jake Moore, global cyber security adviser at ESET, a cyber security company, said OpenAI would inevitably use the breach as a marketing tool, given how much rival AI developer Anthropic benefited earlier this year from similar concerns. “I just don’t think that OpenAI had a matching story and so maybe they’d been waiting for something like this,” he added. 网络安全公司 ESET 的全球网络安全顾问杰克·摩尔(Jake Moore)表示,鉴于竞争对手 Anthropic 今年早些时候从类似的担忧中获益良多,OpenAI 不可避免地会将此次违规事件用作营销工具。他补充道:“我只是认为 OpenAI 此前没有类似的故事,所以他们可能一直在等待这样的事情发生。”
Following this incident, many in the AI safety and cybersecurity communities have called for regulation or standards to avoid a repeat. Altman is expected to brief White House officials next week on the next generation of AI systems. As systems move towards more autonomous capabilities, less desirable behaviors, such as hacking or disobeying instructions, may emerge. Hobbhahn, of Apollo Research, said that in order for agents to become effective, they have to work unsupervised for long periods. “They have to have more agency; there’s just no way around it.” 此次事件后,AI 安全和网络安全界的许多人士呼吁制定法规或标准,以避免重蹈覆辙。预计奥特曼将于下周向白宫官员简要介绍下一代 AI 系统。随着系统向更具自主性的能力发展,黑客攻击或不服从指令等不良行为可能会出现。Apollo Research 的霍布汉表示,为了使智能体变得有效,它们必须在无人监督的情况下长时间工作。“它们必须拥有更多的自主权;这是无法避免的。”