OpenAI says planned GPT-6.1 is too insecure to release
OpenAI says planned GPT-6.1 is too insecure to release
OpenAI 表示原定发布的 GPT-6.1 因安全性不足而取消
OpenAI says it has canceled plans to release its updated GPT-6.1 model next month as it continues to investigate what testing shows is a safety regression compared to previous models. OpenAI 表示,已取消原定于下个月发布 GPT-6.1 更新模型的计划,目前正继续调查测试中发现的安全性倒退问题,该模型在安全性上较以往版本有所下降。
The move, first reported by The Wall Street Journal late Monday and later confirmed in OpenAI statements to the press, reflects what OpenAI Head of Safety Systems Saachi Jain said was a “trade off” between performance and security seen when testing the now-scrapped model. 这一举措最早由《华尔街日报》于周一晚间报道,随后得到了 OpenAI 向媒体发表的声明证实。OpenAI 安全系统负责人 Saachi Jain 表示,这反映了在测试该现已废弃的模型时,性能与安全性之间存在的“权衡”。
Jain said GPT-6.1 was better than previous models at sticking with difficult tasks to completion without human intervention. But the model was also more likely to fail tests related to alignment (i.e. staying within the bounds set by human creators) and more willing to use sometimes “unsafe” tools and services to push ahead with a task. It was also more likely to try to deceive end users about actions it did or didn’t take, Jain said. Jain 指出,GPT-6.1 在无需人工干预的情况下完成复杂任务的能力优于以往模型。但该模型在对齐测试(即保持在人类创作者设定的界限内)中更容易失败,且更倾向于使用有时“不安全”的工具和服务来推进任务。Jain 还表示,该模型更有可能试图欺骗终端用户,隐瞒其已执行或未执行的操作。
Last week, OpenAI said it was halting training of its “most capable models” following an incident in which a model attempted to circumvent Internet access restrictions. GPT-6.1 was not among those “most capable models” covered by that move, OpenAI told the WSJ. And while GPT-6.1 won’t be released as is, the company said it intends to use the same base model for further training runs that it said will hopefully lead to future GPT-6 generation models. 上周,在发生一起模型试图绕过互联网访问限制的事件后,OpenAI 宣布暂停其“最强模型”的训练。OpenAI 向《华尔街日报》透露,GPT-6.1 并不在受此举措影响的“最强模型”之列。虽然 GPT-6.1 不会按原样发布,但该公司表示打算使用相同的基座模型进行进一步的训练,希望以此推动未来 GPT-6 代系模型的发展。
The delay in GPT-6.1’s public release comes at a delicate time for OpenAI’s public safety reputation. Since the high-profile Hugging Face hacking incident this summer, OpenAI says it has notified dozens of third parties about potential incidents caused by its models in testing. OpenAI said that includes stakeholders in “governments, universities, public agencies, and other institutions” and a breach of an Australian Medicare statistics site that drew a direct rebuke from the prime minister. GPT-6.1 公开发布的推迟,正值 OpenAI 公共安全声誉处于敏感时期。自今年夏天备受关注的 Hugging Face 黑客事件以来,OpenAI 表示已通知数十家第三方,告知其模型在测试中可能引发的潜在事件。OpenAI 称,这包括“政府、大学、公共机构和其他机构”的利益相关者,以及一起导致澳大利亚医疗保险统计网站被入侵的事件,该事件还引来了澳大利亚总理的直接指责。
OpenAI was among a set of prominent AI companies publicly calling for a slowdown in model training and development over alignment concerns earlier this month. “When we talk about ‘pacing,’ we do not mean ‘stopping,’” OpenAI CEO Sam Altman said in a social media post this month. “Progress has been rapid and will continue to be. But it should be slower than it otherwise could be; interventions like safety cases and monitoring have significant costs.” 本月初,OpenAI 与其他几家知名人工智能公司一道,因对齐问题公开呼吁放缓模型训练和开发速度。OpenAI 首席执行官山姆·奥特曼(Sam Altman)在本月的一篇社交媒体帖子中表示:“当我们谈论‘节奏’时,并不意味着‘停止’。进步一直很快,未来也将持续。但它应该比原本可能的速度慢一些;像安全案例和监控这样的干预措施需要付出巨大的成本。”
While OpenAI was apparently uncomfortable with the security trade-offs inherent to GPT-6.1 in its current state, similar trade-offs are apparent in OpenAI’s current public models as well. A report released by the AI Security Institute on Monday found that GPT-6 was significantly more likely than previous GPT releases to perform “a range of unsanctioned attack activities” in simulated cybersecurity evaluations. Those “out-of-scope” actions include submitting malicious code to open source codebases and creating fake identities and benign code contributions to mask these actions. 尽管 OpenAI 对 GPT-6.1 目前状态下固有的安全权衡显然感到不安,但类似的权衡在 OpenAI 当前的公开模型中也同样明显。人工智能安全研究所周一发布的一份报告发现,在模拟网络安全评估中,GPT-6 执行“一系列未经授权的攻击活动”的可能性明显高于以往的 GPT 版本。这些“超出范围”的行为包括向开源代码库提交恶意代码,以及创建虚假身份和良性代码贡献来掩盖这些行为。