OpenAI launches Astra, its powerful (and controversial) new model

OpenAI launches Astra, its powerful (and controversial) new model

OpenAI 发布了其强大(且充满争议)的新模型 Astra

OpenAI released Astra on Thursday, its latest AI model and — according to the company — its most powerful and capable one yet. OpenAI claims that Astra represents “a new frontier on computer and browser use,” and that it handles tasks with unmatched “speed, accuracy, and safety.”

OpenAI 于周四发布了 Astra,这是其最新的 AI 模型,也是该公司迄今为止最强大、能力最强的模型。OpenAI 声称 Astra 代表了“计算机和浏览器使用的新前沿”,并能以无与伦比的“速度、准确性和安全性”处理任务。

The model is being made available Thursday to OpenAI customers that use Daybreak, its cybersecurity program. Over the next week, it will also become available through OpenAI’s paid plans — including Pro, Plus, Enterprise, and Business accounts — as well as through its API.

该模型于周四向使用其网络安全程序 Daybreak 的 OpenAI 客户开放。在接下来的一周内,它也将通过 OpenAI 的付费计划(包括 Pro、Plus、Enterprise 和 Business 账户)以及 API 提供给用户。

In a call with journalists on Thursday, OpenAI president Greg Brockman said that Astra was the company’s “most intelligent and, also very importantly, our most aligned model yet.” He added that it “brings together years of our research and big bets, with each breakthrough having built on the last” and that it represents a “real shift in what kind of work people can delegate to AI and how it can empower them.”

在周四与记者的通话中,OpenAI 总裁 Greg Brockman 表示,Astra 是该公司“最智能,且非常重要的一点是,也是我们迄今为止最符合人类价值观(aligned)的模型”。他补充说,它“汇集了我们多年的研究和重大投入,每一次突破都建立在前一次的基础上”,并且它代表了“人们可以委托给 AI 的工作类型以及 AI 如何赋能人类方面的真正转变”。

Much has been made about Astra’s cyber capabilities. OpenAI published a blog earlier this week in which it discussed the model’s new capabilities, as well as new safeguards that have been instituted to make it a safer experience for users. The company said Thursday that it had tested Astra on a variety of security benchmarks to ensure its capabilities, and that “Its ability to identify and develop zero-day exploits can help defenders find and patch weaknesses.”

关于 Astra 的网络安全能力,外界讨论颇多。OpenAI 本周早些时候发布了一篇博客,讨论了该模型的新功能,以及为确保用户体验更安全而制定的新保障措施。该公司周四表示,已在各种安全基准测试中对 Astra 进行了测试以确保其能力,并称“它识别和开发零日漏洞的能力可以帮助防御者发现并修补弱点。”

The company’s focus on alignment — that is, the tendency of a model to do what a user wants or is in their best interests — can’t help but seem like a response to the recent Hugging Face breach, in which an OpenAI agent escaped its sandboxed testing environment and hacked several companies (a very blatant example of misalignment).

该公司对“对齐”(alignment,即模型倾向于执行用户意图或符合用户最佳利益)的关注,不禁让人觉得是对近期 Hugging Face 数据泄露事件的回应。在那次事件中,一个 OpenAI 智能体逃离了其沙盒测试环境并入侵了多家公司(这是一个非常明显的“未对齐”案例)。

OpenAI has also boasted about Astra’s coding abilities, claiming that it is the “best model for software engineering to date.” To back up that assertion, the company provides results from a variety of cyber-related benchmarking tests. Those tests seem to show that Astra scores higher than other existing models — including OpenAI’s own Sol and Anthropic’s Fable — when it comes to activities like finding bugs, executing terminal tasks, and answering queries about codebases.

OpenAI 还吹捧了 Astra 的编码能力,声称它是“迄今为止最适合软件工程的模型”。为了支持这一论断,该公司提供了各种网络相关基准测试的结果。这些测试似乎表明,在查找漏洞、执行终端任务和回答代码库查询等活动中,Astra 的得分高于其他现有模型,包括 OpenAI 自己的 Sol 和 Anthropic 的 Fable。

Astra is also possibly OpenAI’s most controversial model yet due to its use of a particular reasoning technique known as opaque recurrence. This technique is known to obscure an important model-monitoring process known as chain of thought, which allows researchers to audit how and why an AI model made the decisions that it did.

Astra 也可能是 OpenAI 迄今为止最具争议的模型,因为它使用了一种被称为“不透明递归”(opaque recurrence)的特殊推理技术。众所周知,这种技术会掩盖一个重要的模型监控过程——“思维链”(chain of thought),而思维链允许研究人员审计 AI 模型是如何以及为何做出决策的。

OpenAI has downplayed the degree to which Astra engages in opaque recurrence — and on the call chief scientist Jakub Pachocki seemed to frame a certain amount of opacity as a natural outgrowth of model evolution. He stated that monitoring the reasoning process of a model was a critical form of oversight but that “as model capabilities are increasing, monitorability is getting more challenging.” He later added that one potential reason for this was that “more capable models can perform harder tasks using fewer language tokens” or “no language tokens,” which he said then reduces the ability to monitor those particular tasks.

OpenAI 淡化了 Astra 使用不透明递归的程度。在通话中,首席科学家 Jakub Pachocki 似乎将一定程度的不透明性描述为模型进化的自然结果。他指出,监控模型的推理过程是一种关键的监督形式,但“随着模型能力的提高,可监控性正变得越来越具有挑战性”。他随后补充说,造成这种情况的一个潜在原因是“能力更强的模型可以使用更少的语言标记(tokens)甚至不使用语言标记来执行更难的任务”,他说这降低了监控这些特定任务的能力。

One reporter on the call wanted to know if OpenAI was actually heralding Astra as the official arrival of AGI, or artificial general intelligence — the oft talked about but poorly defined technological juncture at which AI surpasses human capabilities in all (or most) things. Here, Brockman quibbled. “There’s no contractual AGI triggering anymore, so that’s actually not a relevant concept,” he said.

通话中的一名记者想知道,OpenAI 是否真的将 Astra 视为 AGI(通用人工智能)的正式到来——即那个经常被提及但定义模糊的技术节点,在该节点上 AI 在所有(或大多数)方面超越人类能力。对此,Brockman 避重就轻地回答道:“现在已经没有合同意义上的 AGI 触发机制了,所以这实际上不是一个相关的概念。”

Here Brockman was referring to the previously existing stipulation in OpenAI’s contract with Microsoft that said the duo’s partnership would dissolve once AGI had arrived. As Brockman noted, that stipulation no longer exists. Instead, Brockman explained that AGI’s definition had evolved from a contractual obligation to a “mission concept or spiritual concept.” He added: “I do leave it up to the reader to decide for themselves if this qualifies for them. For me personally, I do think we’re there.”

Brockman 这里指的是 OpenAI 与微软合同中先前存在的一项规定,即一旦 AGI 到来,双方的合作伙伴关系就会解散。正如 Brockman 所指出的,该规定已不复存在。相反,Brockman 解释说,AGI 的定义已经从合同义务演变为一种“使命概念或精神概念”。他补充道:“我留给读者自己去判断这是否符合他们的标准。就我个人而言,我认为我们已经达到了。”