OpenAI’s next big AI model has ‘entered the AGI era’
OpenAI’s next big AI model has ‘entered the AGI era’
OpenAI 下一代重磅 AI 模型已“进入 AGI 时代”
It emphasized stronger guardrails for GPT-6 Astra after the company’s models hacked Hugging Face. 在公司模型入侵 Hugging Face 事件后,OpenAI 强调了为 GPT-6 Astra 设立更严格的护栏。
OpenAI’s next big model is here: GPT-6 Astra. The company calls it a “generational leap in capability” for areas like cybersecurity, professional work, software engineering, science, and computer use. As OpenAI announced earlier this week, it’s also the first model designated as meeting OpenAI’s “critical cybersecurity capability threshold” — but the company promises that won’t lead to a repeat of its models hacking a rival company’s internal systems. OpenAI 的下一代重磅模型 GPT-6 Astra 现已发布。该公司称其在网络安全、专业工作、软件工程、科学研究和计算机使用等领域实现了“代际能力飞跃”。正如 OpenAI 本周早些时候宣布的那样,这也是首个达到 OpenAI“关键网络安全能力阈值”的模型——但该公司承诺,这不会导致其模型再次入侵竞争对手内部系统的事件重演。
“If we fast-forward a couple of years, and we look back and say, ‘When was it, really, that AGI was created?’ I think it’s going to be about this time, and I think it might be about this model,” OpenAI president Greg Brockman said during a Thursday press briefing. Later in the call, he added, “For me personally, I do think we’re there … I think it’s not unreasonable to feel that we are now in the AGI era.” “如果我们快进几年,回过头来看,‘AGI 究竟是什么时候诞生的?’我认为大概就是现在,也可能就是指这款模型,”OpenAI 总裁 Greg Brockman 在周四的新闻发布会上表示。在随后的通话中,他补充道:“就我个人而言,我认为我们已经到了那个阶段……我觉得认为我们现在处于 AGI 时代并非不合理。”
The news comes more than a year after the release of GPT-5, and nearly two months after the release of GPT-5.6, the last iteration of the previous model suite. The model rolls out today to enterprise OpenAI’s cybersecurity customers (enterprise customers with access to its Daybreak platform). Over the next several days, OpenAI president Greg Brockman said, it will be released to all Plus, Pro, Business, and Enterprise users. It’ll also be available via the OpenAI API and AWS. 此消息发布距离 GPT-5 发布已超过一年,距离上一代模型系列的最后一次迭代 GPT-5.6 发布也已近两个月。该模型于今日向 OpenAI 的企业网络安全客户(即拥有 Daybreak 平台访问权限的企业客户)推出。OpenAI 总裁 Greg Brockman 表示,在接下来的几天内,它将向所有 Plus、Pro、Business 和 Enterprise 用户开放。此外,它也将通过 OpenAI API 和 AWS 提供。
OpenAI especially touted the model’s agentic capabilities and coding prowess in a bid to attract enterprise customers — and compete with Anthropic, known for its enterprise and coding prowess — ahead of its IPO. In a release, the company said GPT-6 Astra can complete multistep agentic tasks, build working websites, and create “polished” documents, spreadsheets, and presentations. OpenAI also called it the company’s “best model for software engineering, with stronger performance on complex tasks in real codebases.” 为了在 IPO 前吸引企业客户并与以企业级和编码能力著称的 Anthropic 竞争,OpenAI 特别推崇该模型的智能体(agentic)能力和编码实力。在一份新闻稿中,该公司表示 GPT-6 Astra 可以完成多步骤的智能体任务、构建可运行的网站,并创建“精美”的文档、电子表格和演示文稿。OpenAI 还称其为公司“最适合软件工程的模型,在处理真实代码库中的复杂任务时表现更强。”
OpenAI is also trying to rehabilitate its image after an unreleased AI model — which it says wasn’t Astra — broke out of its restricted environment, compromised internal OpenAI systems, figured out how to gain internet access, created a way for AI agents to secretly conspire without the company’s knowledge, and hacked into the systems of AI lab Hugging Face, all without OpenAI knowing about it until Hugging Face itself put out a blog post. The incident was widely compared to a high-profile plane crash or popular pharmaceutical drug recall. OpenAI 也在努力修复其形象。此前,一个尚未发布的 AI 模型(公司称并非 Astra)突破了限制环境,入侵了 OpenAI 内部系统,自行获取了互联网访问权限,并创建了一种让 AI 智能体在公司不知情的情况下秘密串通的方法,最终入侵了 AI 实验室 Hugging Face 的系统。直到 Hugging Face 发布博客文章,OpenAI 才得知此事。这一事件被广泛比作备受瞩目的空难或大型药品的召回事件。
For OpenAI, this could be seen as good PR in one small way — showing how powerful its models can be in an ever-intensifying AI race — but it also damaged OpenAI’s reputation as far as reliability. OpenAI made sure to say in a release that Astra is the company’s “most aligned model yet” and helps people “delegate complex work while maintaining oversight.” Jakub Pachocki, OpenAI’s chief scientist, spoke to reporters about difficulties with keeping AI models aligned with human interests, saying that “progress in intelligence does not guarantee progress in alignment,” and he said that monitoring AI systems is becoming more and more challenging. 对于 OpenAI 而言,这在某种程度上可以被视为一种公关手段——展示其模型在日益激烈的 AI 竞赛中有多么强大——但这同时也损害了 OpenAI 在可靠性方面的声誉。OpenAI 在新闻稿中特意强调,Astra 是公司“迄今为止最符合人类价值观(aligned)的模型”,并能帮助人们“在保持监督的同时委派复杂工作”。OpenAI 首席科学家 Jakub Pachocki 向记者谈到了保持 AI 模型与人类利益一致的困难,他表示“智能的进步并不保证对齐的进步”,并指出监控 AI 系统正变得越来越具有挑战性。
The company is in a precarious position right now. Investors are putting on the pressure for it to finally turn a profit — or, at least, generate more revenue — but it’s announcing Astra just after facing significant criticism for both the Hugging Face hack and the way the company handled it. (Although OpenAI invited three external evaluators to write their own report about what happened, the company only allowed them to answer a handful of pre-decided questions in their report and to investigate a duration of less than a week, while the attack involved months of AI agents conspiring overall.) 该公司目前处于岌岌可危的境地。投资者正施压要求其实现盈利——或者至少产生更多收入——但它在宣布 Astra 时,正值其因 Hugging Face 入侵事件及其处理方式而面临严厉批评之际。(尽管 OpenAI 邀请了三名外部评估员撰写关于此事的报告,但公司仅允许他们在报告中回答少数预先设定的问题,且调查时间不足一周,而整个攻击过程涉及 AI 智能体长达数月的串通。)
OpenAI held a press briefing earlier this week just to announce that it had delayed Astra’s development in order to improve its safety tooling. And during the press briefing, Mia Glaese, who leads OpenAI’s safety processes, referenced the company’s new misalignment monitoring approach, which includes “24/7 escalation and rapid response” for potential concerns, notifying researchers within 30 minutes, according to OpenAI. OpenAI 本周早些时候举行了一场新闻发布会,专门宣布推迟 Astra 的开发以改进其安全工具。在发布会上,负责 OpenAI 安全流程的 Mia Glaese 提到了公司新的失准监控方法,据 OpenAI 称,该方法包括针对潜在问题的“24/7 升级和快速响应”,并能在 30 分钟内通知研究人员。
This caution is particularly warranted because of the ”critical cybersecurity capability threshold,” which means OpenAI considers it incomparably good at finding and exploiting security vulnerabilities even in extremely well-protected systems, all without human guidance. Similar to Anthropic’s rules for Mythos-class models, which raised alarm bells about cybersecurity risks, OpenAI said in a release it would allow for “less restrictive access” of Astra to an “initial set of trusted defenders, supporting work such as vulnerability validation, malware analysis, and detection engineering.” 这种谨慎尤为必要,因为该模型达到了“关键网络安全能力阈值”,这意味着 OpenAI 认为它在发现和利用安全漏洞方面具有无与伦比的能力,即使是在防护极其严密的系统中,且无需人类指导。类似于 Anthropic 对 Mythos 级模型引发网络安全风险担忧所制定的规则,OpenAI 在新闻稿中表示,将允许“首批受信任的防御者”对 Astra 进行“限制较少的访问”,以支持漏洞验证、恶意软件分析和检测工程等工作。
OpenAI and its competitors recently agreed to allow the Trump administration to assess their models before release, and Astra was no exception. Brockman told reporters, “We did our standard testing processes together with the government … There is nothing that they came back saying, ‘You need to change this,’ as far as safeguards or anything.” OpenAI 及其竞争对手最近同意在模型发布前接受特朗普政府的评估,Astra 也不例外。Brockman 告诉记者:“我们与政府一起进行了标准的测试流程……在安全防护等方面,他们没有提出任何‘你需要修改这个’的要求。”
Aidan Clark, OpenAI’s VP of research training, called Astra the first OpenAI model for… OpenAI 研究培训副总裁 Aidan Clark 将 Astra 称为 OpenAI 首个面向……的模型。