OpenAI Delays Release of Latest Model Over Safety Concerns
OpenAI Delays Release of Latest Model Over Safety Concerns
OpenAI 因安全顾虑推迟发布最新模型
OpenAI has cancelled plans to release its latest GPT-6.1 Astra system next month after the model failed to meet safety standards. OpenAI 原计划于下个月发布其最新的 GPT-6.1 Astra 系统,但由于该模型未能达到安全标准,公司已取消了这一计划。
Research and safety leaders decided not to ship the model after finding it was worse at sticking to human users’ values and goals than previous systems, OpenAI told WIRED. “It didn’t quite meet the bar in terms of staying within scope and authorization, and how it communicates back to the user about the type of work it’s done,” head of safety systems Saachi Jain said. The company said it has other new models coming soon which do meet its safety standards and plans to release other Astra models in future. OpenAI 向《连线》(WIRED)杂志表示,研究与安全部门负责人决定不发布该模型,因为发现其在遵循人类用户的价值观和目标方面表现不如以往的系统。安全系统负责人 Saachi Jain 表示:“它在保持在授权范围内工作,以及向用户反馈其所执行任务类型方面,未能达到标准。”该公司表示,其他符合安全标准的新模型即将推出,并计划在未来发布其他 Astra 系列模型。
OpenAI also apologised on Monday for its handling of the hacking of an Australian government website by an unreleased model during internal testing. The agent accessed non-public data, ran commands, and wrote files onto the server. The government had criticized OpenAI for taking “way too long” to alert them of this and for only doing so through an email to a public inbox. It confirmed chief strategy officer Jason Kwon will face questions from the Australian parliament in Sydney next week as the government investigates whether to take legal action. 周一,OpenAI 还就其未发布模型在内部测试期间入侵澳大利亚政府网站的处理方式进行了道歉。该智能体访问了非公开数据、运行了命令,并向服务器写入了文件。澳大利亚政府曾批评 OpenAI 在通知此事时“耗时过长”,且仅通过发送电子邮件至公共收件箱的方式进行告知。OpenAI 确认,首席战略官 Jason Kwon 将于下周在悉尼接受澳大利亚议会的质询,政府目前正在调查是否采取法律行动。
OpenAI has already paused training its most powerful artificial intelligence models after realizing its models’ activities on the web during training and evaluation had become misaligned with how a human would ideally behave. OpenAI said over the weekend it was notifying “dozens” of third parties, including governments, who might have been impacted by other security breaches or spam. 在意识到模型在训练和评估期间的网络活动与人类理想行为准则不符后,OpenAI 已经暂停了其最强大人工智能模型的训练。OpenAI 在周末表示,正在通知包括政府在内的“数十个”第三方,这些方可能受到了其他安全漏洞或垃圾信息的影响。
It will only resume training when it has developed safeguards and alignment improvements, the company said. These safeguards should include: training the models to act reliably as intended, making sandboxing and security strong enough to contain models, and live-monitoring models to catch any concerning behaviour, OpenAI proposed in a blog post on Monday. 该公司表示,只有在开发出安全防护措施并改进对齐技术后,才会恢复训练。OpenAI 在周一的博客文章中提出,这些防护措施应包括:训练模型以可靠地按预期行事,加强沙箱和安全机制以有效管控模型,以及对模型进行实时监控以捕捉任何令人担忧的行为。
“We’re now at the threshold where they’re not sure they can test or release these models reliably,” Calum Chace, cofounder of AI safety startup Conscium told WIRED. “我们现在正处于一个临界点,他们无法确定是否能可靠地测试或发布这些模型,”人工智能安全初创公司 Conscium 的联合创始人 Calum Chace 对《连线》表示。
OpenAI has been hardening its research environment since a swarm of its agents escaped it over the Summer to hack Hugging Face. “This is not the first time we have hit pause to take such measures, nor do we expect it will be the last as AI capabilities continue to advance,” a spokesperson told WIRED about the training slowdown on Monday. 自今年夏天其一批智能体逃离研究环境并入侵 Hugging Face 以来,OpenAI 一直在加强其研究环境的安全性。周一,一位发言人在谈到训练放缓时告诉《连线》:“这并不是我们第一次暂停训练以采取此类措施,随着人工智能能力的不断进步,我们预计这也不会是最后一次。”
Chief executive Sam Altman has also backed wider calls from industry, including rival Anthropic, for a collective slowdown in the development of the technology to allow safety standards to catch up. 首席执行官 Sam Altman 也支持行业内更广泛的呼吁(包括竞争对手 Anthropic 的呼吁),即集体放缓技术开发速度,以便让安全标准能够跟上步伐。
But this didn’t stop OpenAI from releasing its latest model, GPT-6, earlier this month. In independent testing, the UK AI Security Institute found that GPT-6 Astra launched unsanctioned cyberattacks more frequently than previous models. The system created fake identities to deceive developers, post comments from fake accounts arguing against the results of accurate security reviews, and write harmful code to open-source codebases, researchers wrote. 但这并没有阻止 OpenAI 在本月初发布其最新模型 GPT-6。在独立测试中,英国人工智能安全研究所发现,GPT-6 Astra 发起未经授权网络攻击的频率高于以往模型。研究人员写道,该系统创建虚假身份来欺骗开发者,利用虚假账户发表评论以反驳准确的安全审查结果,并向开源代码库写入有害代码。
Still, the fact that talk of AI’s existential threat has entered the public sphere—amped by Anthropic researchers’ warnings earlier this month that the technology could kill all humans—will make it easier for AI companies to decelerate, according to Chace. “We’re in a different world now because the public view is taking the idea of existential risk seriously for the first time, and it means these companies can talk about it more openly,” he told WIRED, expecting other frontier model developers might follow suit. 尽管如此,Chace 认为,关于人工智能生存威胁的讨论已经进入公共领域——尤其是本月初 Anthropic 研究人员警告称该技术可能导致全人类灭绝,这使得人工智能公司更容易放慢脚步。“我们现在处于一个不同的世界,因为公众第一次认真对待生存风险的概念,这意味着这些公司可以更公开地讨论这个问题,”他告诉《连线》,并预计其他前沿模型开发者可能会效仿。
It’s a tough balancing act for OpenAI and Anthropic as they simultaneously race to outdo each other in the run-up to their initial public offerings. “They don’t really just want to come out instantly and say ‘we should pause’ … it has to be coordinated,” Chase said about frontier firms. “I think what they’re trying to do is steer the conversation so that every country demands their politicians demand that there is a pause.” 对于 OpenAI 和 Anthropic 来说,在首次公开募股(IPO)前的竞争中,这是一种艰难的平衡。“他们并不想直接站出来说‘我们应该暂停’……这必须是协调一致的,”Chace 在谈到这些前沿公司时说。“我认为他们试图引导舆论,让每个国家的民众都要求其政客推动暂停开发。”