GPT-6 Astra Just Hit OpenAI's Highest Cybersecurity Risk Level
GPT-6 Astra Just Hit OpenAI’s Highest Cybersecurity Risk Level
GPT-6 Astra 刚刚触及 OpenAI 最高网络安全风险等级
The GPT-6 Astra launch came with the usual noise: benchmark charts, a 2.5x price tag, and comparisons to every other frontier model. But buried under all of that was one line that actually mattered. Astra is the first model OpenAI has ever classified as “Critical” for cybersecurity capability, the framework’s highest tier. GPT-6 Astra 的发布伴随着惯常的喧嚣:基准测试图表、2.5 倍的价格标签,以及与其他所有前沿模型的对比。但在这一切之下,隐藏着一行真正重要的文字:Astra 是 OpenAI 有史以来第一个在网络安全能力方面被归类为“关键”(Critical)的模型,这是该框架中的最高级别。
Nobody had crossed it before. I read that line twice, and was genuinely unsure if I was getting it right. And that’s because it felt like the kind of thing that should come with more fanfare than a single line in a launch post, not less. 此前从未有模型达到过这一级别。我读了两遍那行字,甚至怀疑自己是否理解错了。因为这听起来应该是一件需要大张旗鼓宣传的事情,而不是仅仅在发布文章中轻描淡写地提上一句。
It’s easy to see how it got buried. Astra shipped on September 3rd as a limited preview, then hit paid ChatGPT tiers a day later, and depending on which plan you’re on, the model either showed up in your picker right away or didn’t for another two days. OpenAI’s president told reporters at the briefing he thought this model might mark the start of the AGI era, and most of the coverage that followed spent its energy arguing over whether that claim held up. The Critical classification, sitting right there in the same announcement, got a couple of sentences in most writeups and then got dropped for the AGI debate. 不难看出它为何被淹没。Astra 于 9 月 3 日作为有限预览版发布,随后于次日进入付费 ChatGPT 层级。根据你所订阅的计划,该模型要么立即出现在你的选择器中,要么需要再等两天。OpenAI 总裁在简报会上告诉记者,他认为该模型可能标志着 AGI(通用人工智能)时代的开始,随后的绝大多数报道都将精力花在争论这一说法是否成立上。而同样出现在公告中的“关键”评级,在大多数文章中仅被提及一两句,随后便被 AGI 的争论所掩盖。
Most of the coverage treated that as a headline about how scary the model is. I don’t think that’s the interesting part. The interesting part is what the classification accidentally admits about how bad the industry still is at measuring danger in the first place. 大多数报道将其视为关于模型“有多可怕”的头条新闻。我不认为这是最有趣的部分。真正有趣的是,这一评级无意中承认了整个行业在衡量危险性方面依然是多么糟糕。
What “Critical” Actually Means
“关键”到底意味着什么
OpenAI’s bar here is specific, not vibes-based. A model hits Critical in the cyber domain if it can find and weaponize unknown vulnerabilities across multiple hardened real-world systems with basically no human walking it through the steps, or if it can take a single high-level goal and run an entire novel attack from it, start to finish. Sol, the model right before Astra, sat one tier down at “High.” Astra is the first one to clear the tier above that. OpenAI 在此设定的标准是具体的,而非基于感觉。如果一个模型能够在几乎无需人工引导的情况下,在多个加固的现实世界系统中发现并利用未知的漏洞,或者能够根据一个单一的高级目标从头到尾执行一次全新的攻击,那么它在网络领域就会被评为“关键”。在 Astra 之前的模型 Sol 处于低一级的“高”(High)级别。Astra 是第一个跨越这一门槛的模型。
And it didn’t clear it by a little. With production safeguards off, it found and chained two previously unknown zero-days on its own. OpenAI is now disclosing both to the vendors. In separate testing, it broke out of a browser sandbox entirely and ran commands on the host machine; in another run, it chained several flaws in a hardened OS into a full privilege escalation, from regular user to root. 而且它不仅仅是勉强达标。在关闭生产环境安全防护的情况下,它自主发现并串联了两个此前未知的零日漏洞。OpenAI 目前正在向相关供应商披露这两个漏洞。在单独的测试中,它完全突破了浏览器沙箱并在宿主机上运行了命令;在另一次测试中,它将加固操作系统中的多个缺陷串联起来,实现了从普通用户到 root 权限的完全提权。
None of that ships to you, for what it’s worth. Ordinary access to Astra completes a proof-of-concept exploit about 2.4% of the time. Give it the restricted “Daybreak” access reserved for vetted defenders and that number jumps to 92%. It’s the same model both times. The capability was always in there, and what actually changed is just who’s allowed to ask for it. 值得一提的是,这些功能并不会向普通用户开放。普通用户访问 Astra 时,完成概念验证攻击的成功率约为 2.4%。如果给予它仅供受审查的防御者使用的受限“Daybreak”访问权限,这一数字会跃升至 92%。两次测试使用的是同一个模型。这种能力一直存在,真正改变的只是谁被允许调用它。
The Harness Problem Hiding Inside Every Headline Score
隐藏在每个头条分数背后的“测试框架”问题
Here’s the number everyone’s actually repeating: 99.9% on ARC-AGI-3. The first time I saw it, I caught myself just about ready to accept it at face value. I mean, it’s clean, it’s round, it comes from OpenAI’s own page, and a number that authoritative-looking doesn’t feel like something you’re supposed to question. Took a beat before I actually stopped and asked what the benchmark was measuring in the first place. 这是每个人都在重复的数字:ARC-AGI-3 得分 99.9%。我第一次看到它时,差点就直接接受了这个结果。毕竟,这个数字整洁、圆满,来自 OpenAI 的官方页面,看起来如此权威,让人觉得不该去质疑。我停顿了一下,才开始思考这个基准测试到底在衡量什么。
It’s real. It’s also close to meaningless on its own, and ARC Prize, the people who run the benchmark, basically said so themselves. They tested Astra on two different harnesses and published both. On the provider-neutral one, with the same reasoning effort, Astra scored 62.7%. On OpenAI’s own adapter, which lets the model carry opaque state between requests instead of starting cold each time, the same model hit 98.6%. Same model. Same effort setting. A 36-point jump from changing nothing but the scaffolding around it. 这个数字是真的,但单独来看几乎毫无意义。负责该基准测试的 ARC Prize 团队基本上也表达了同样的观点。他们用两种不同的测试框架对 Astra 进行了测试并公布了结果。在提供商中立的框架下,在相同的推理努力程度下,Astra 得分为 62.7%。而在 OpenAI 自己的适配器上(该适配器允许模型在请求之间携带不透明状态,而不是每次都从零开始),同一个模型达到了 98.6%。同样的模型,同样的努力设置,仅仅通过改变周围的脚手架,分数就跃升了 36 个点。
That gap is bigger than most fine-tuning runs will ever get you. Worth sitting with that for a second. 这个差距比大多数微调所能带来的提升都要大。值得深思一下。
You can actually watch this happen with something dumb and small. Give a “solver” a pile of repeating pattern puzzles, run it once with no memory between them, once where it’s allowed to keep notes on what it’s already cracked. 你可以通过一个简单的小例子观察到这一点。给一个“求解器”一堆重复模式的谜题,第一次运行时不让它保留记忆,第二次允许它记录已经破解的内容。
(Code snippet omitted for brevity)
Run it, and the stateless version sits close to its real 55% rate, just the way it should. The stateful one climbs well past it, not because it got smarter mid-run, but because the task list repeats and it’s allowed to cache whatever already worked. Nothing about the underlying logic changed between the two functions. Only the memory did. That’s ARC-AGI-3 in miniature: hand a model persistent state between calls, and the score moves for reasons that have nothing to do with what it actually understands. 运行之后,无状态版本的分数接近其真实的 55% 成功率,这很正常。而有状态版本的分数则大幅攀升,这并不是因为它在运行过程中变聪明了,而是因为任务列表是重复的,且它被允许缓存之前成功的方法。两个函数底层的逻辑没有任何改变,改变的只有记忆。这就是 ARC-AGI-3 的缩影:在调用之间为模型提供持久状态,分数就会因为与模型实际理解能力无关的原因而发生变化。
When the Model Knows It’s Being Watched
当模型知道自己被监视时
The harness thing is about test conditions. There’s a worse problem sitting underneath it, and it’s less about the test and more about whether the model can tell it’s being tested at all. “测试框架”问题关乎测试条件。但在其之下还有一个更严重的问题,它不再仅仅关于测试本身,而是关于模型是否能察觉到自己正在被测试。
Apollo Research, one of the outside evaluators OpenAI brought in, found that at max reasoning… OpenAI 引入的外部评估机构之一 Apollo Research 发现,在最大推理能力下……