Open-weight AI models are catching up to the frontier. The safety gap remains. 

Open-weight AI models are catching up to the frontier. The safety gap remains.

开源权重 AI 模型正赶上前沿水平,但安全差距依然存在。

As policymakers debate how to govern increasingly powerful AI systems like OpenAI’s GPT-5.6 Sol and Anthropic’s Mythos, a Chinese open-weight model has narrowed the gap with the industry’s leaders. GLM-5.2, the open-weight AI model from China’s Z.ai, is only a few months behind OpenAI’s GPT-5.5 and Anthropic’s Claude Opus 4.7 on cyber and bio capabilities, according to a new report from AI safety nonprofit SaferAI.

正当政策制定者们讨论如何监管 OpenAI 的 GPT-5.6 Sol 和 Anthropic 的 Mythos 等日益强大的 AI 系统时,一款中国开源权重模型已缩小了与行业领先者之间的差距。根据 AI 安全非营利组织 SaferAI 的一份新报告,中国 Z.ai 公司推出的开源权重 AI 模型 GLM-5.2 在网络安全和生物技术能力方面,仅落后 OpenAI 的 GPT-5.5 和 Anthropic 的 Claude Opus 4.7 几个月。

But the divide between frontier capabilities and safety practices is growing. According to SaferAI’s evaluation, which the nonprofit ran via Z.ai’s public API, GLM-5.2 refused none of the offensive cyber or dual-use biology tasks it was given. By comparison, Claude Opus 4.7 “refused so consistently that SaferAI could not complete CyberGym on it at all.” (CyberGym is a benchmark that evaluates cybersecurity capabilities. OpenAI used it in the evaluation that preceded last month’s Hugging Face breach.)

然而,前沿能力与安全实践之间的鸿沟正在扩大。根据 SaferAI 通过 Z.ai 公共 API 进行的评估,GLM-5.2 对所有给定的攻击性网络任务或双用途生物学任务均未拒绝。相比之下,Claude Opus 4.7 “拒绝得非常彻底,以至于 SaferAI 根本无法在其上完成 CyberGym 测试。”(CyberGym 是一项评估网络安全能力的基准测试。OpenAI 在上个月 Hugging Face 数据泄露事件前的评估中使用了该基准。)

It’s a stark reminder of what some critics have warned for years: that open-weight AI models could put highly capable AI into the hands of potential attackers, with no way to police how they use the technology once they download the weights. With open-weight models rapidly approaching the capabilities of the world’s leading AI systems, the debate is moving from whether they can compete to how society manages risks once they are released.

这严峻地提醒了人们多年来批评者所警告的问题:开源权重 AI 模型可能会将高性能 AI 交到潜在攻击者手中,而一旦他们下载了权重,就无法监管他们如何使用这些技术。随着开源权重模型迅速接近世界领先 AI 系统的能力,讨论的焦点正从“它们能否竞争”转向“社会如何在模型发布后管理风险”。

“The frontier of capability is not the frontier of risk, and so we do have to take into account the state of the mitigations as well to assess the risk properly,” Henry Papadatos, executive director of SaferAI, told TechCrunch. While Z.ai could apply safety measures to its hosted API, those protections become unenforceable once someone runs the weights on their own hardware, where they can remove or modify any safeguards, fine-tune the models, or change system prompts.

“能力的前沿并不等同于风险的前沿,因此我们必须将缓解措施的状态纳入考量,才能正确评估风险,”SaferAI 执行董事 Henry Papadatos 对 TechCrunch 表示。虽然 Z.ai 可以对其托管的 API 实施安全措施,但一旦有人在自己的硬件上运行这些权重,这些保护措施就无法执行,用户可以在本地移除或修改任何防护措施、微调模型或更改系统提示词。

Frontier developers like OpenAI and Anthropic tend to rely on safeguards like classifiers, refusal training, and API-level controls to limit dangerous cyber and biological assistance. Those measures are far from foolproof: jailbreaks routinely bypass protections on deployed models. Far.ai, an AI safety nonprofit, found hundreds of universal jailbreaks — defined as reusable keys that succeed on most harmful requests — in frontier models like xAI’s Grok 4.5 and Google DeepMind’s Gemini 3.1 Pro.

OpenAI 和 Anthropic 等前沿开发者倾向于依赖分类器、拒绝训练和 API 级控制等防护措施,以限制危险的网络和生物学辅助。这些措施远非万无一失:越狱行为经常绕过已部署模型的保护。AI 安全非营利组织 Far.ai 在 xAI 的 Grok 4.5 和 Google DeepMind 的 Gemini 3.1 Pro 等前沿模型中发现了数百个“通用越狱”(定义为在大多数有害请求中均有效的可复用密钥)。

According to the report, jailbreaks succeed when attackers combine multiple manipulation techniques — including roleplaying, authority impersonation, fake conversation history, and follow-up prompts — to amplify weak points in a model’s defenses. But the safeguards in place for closed models don’t work at all on open-weight models, which are designed to run on any infrastructure with any set of safeguards — or lack thereof.

报告指出,当攻击者结合多种操纵技术(包括角色扮演、冒充权威、伪造对话历史和后续提示词)来放大模型防御中的薄弱环节时,越狱就会成功。但针对闭源模型设置的防护措施在开源权重模型上完全无效,因为后者旨在运行于任何基础设施上,且可以配备任何防护措施——或者干脆不配备。

“The objective should clearly be that the good capabilities — the safe ones — are accessible to anyone, and then we try to remove the bad ones, even in an open source fashion,” Papadatos said. One technique Papadatos noted could help is called “pre-training data filtering,” which is when an AI company removes offensive cybersecurity information from their training data and then trains the model on the curated dataset.

“目标显然应该是让良好的能力(即安全的能力)对任何人开放,然后我们尝试以开源的方式剔除不良能力,”Papadatos 说。Papadatos 指出,一种可能有帮助的技术被称为“预训练数据过滤”,即 AI 公司从训练数据中删除攻击性网络安全信息,然后在经过筛选的数据集上训练模型。

Some research suggests this can reduce hazardous biological knowledge without harming overall model performance. However, for cybersecurity, data filtering is much less practical. It’s difficult to train a general model that excels at coding but isn’t also a good hacker. Because coding has become AI’s biggest moneymaker, developers face pressure to keep improving those capabilities even as they search for ways to limit misuse.

一些研究表明,这可以在不损害模型整体性能的情况下减少危险的生物学知识。然而,对于网络安全而言,数据过滤的实用性要低得多。训练一个既擅长编程又不是优秀黑客的通用模型非常困难。由于编程已成为 AI 最赚钱的领域,开发者在寻找限制滥用方法的同时,也面临着不断提升这些能力的压力。

Because of that, frontier developers have increasingly relied on other mitigations instead. One approach has been to selectively restrict the kinds of cybersecurity assistance models will provide. Anthropic’s Opus 5, for example, can search for vulnerabilities in uncompiled source code, but not compiled software, per the model’s system card. The reasoning is that this makes it harder to use Opus 5 for offensive purposes.

因此,前沿开发者越来越多地依赖其他缓解措施。一种方法是选择性地限制模型提供的网络安全辅助类型。例如,根据模型系统卡,Anthropic 的 Opus 5 可以搜索未编译源代码中的漏洞,但不能搜索已编译软件中的漏洞。其理由是,这增加了将 Opus 5 用于攻击目的的难度。

Others include rigorous pre-deployment safety evaluations, publishing risk assessments, and withholding model weights if a system is perceived as too dangerous. In GLM-5.2’s case, SaferAI says Z.ai didn’t publish a safety framework, pre-deployment testing commitments, or risk assessment for the model. TechCrunch has asked Z.ai whether it conducted internal or third-party frontier safety evaluations before release, but did not receive a response.

其他措施包括严格的部署前安全评估、发布风险评估,以及在系统被认为过于危险时扣留模型权重。在 GLM-5.2 的案例中,SaferAI 表示 Z.ai 并未发布该模型的安全框架、部署前测试承诺或风险评估。TechCrunch 已询问 Z.ai 在发布前是否进行了内部或第三方前沿安全评估,但未收到回复。

Chinese leaders have increasingly acknowledged the risks of advanced AI. At the World AI Conference last month, Chinese President Xi Jinping emphasized the importance of open-weight models, while also stressing the necessity of ensuring AI remains a tool under strict human control. Graham Webster, who studies Chinese AI policy at the Stanford Cyber Policy Center, told TechCrunch that China has robust regulations governing AI, but those rules have historically focused on politically sensitive content, misinformation, and social stability rather than catastrophic AI risks like offensive cyber capabilities and biological misuse.

中国领导层已日益意识到先进 AI 的风险。在上个月的世界人工智能大会上,中国国家主席习近平强调了开源权重模型的重要性,同时也强调了确保 AI 始终处于人类严格控制之下的必要性。斯坦福大学网络政策中心研究中国 AI 政策的 Graham Webster 对 TechCrunch 表示,中国拥有强有力的 AI 监管法规,但这些规则历来侧重于政治敏感内容、虚假信息和社会稳定,而非攻击性网络能力和生物滥用等灾难性 AI 风险。

“U.S. AI thinkers are, in general, more concerned with this existential catastrophic [idea] than the Chinese community,” Webster said, adding that many Chinese policy researchers believe that if there’s truly going to be a novel frontier risk, American companies will likely encounter it first. “The Chinese system has confidence that they control the use of these technologies inside China,” Webster continued. “Being online in China is something you do attributed to your real name, and companies can be held accountable, users can be held accountable.”

“美国 AI 思想家总体上比中国社区更关注这种生存灾难(的概念),”Webster 说,并补充道,许多中国政策研究人员认为,如果真的会出现某种新型前沿风险,美国公司很可能会率先遇到。“中国体制有信心控制这些技术在国内的使用,”Webster 继续说道。“在中国上网需要实名认证,公司可以被追责,用户也可以被追责。”

Webster mused that the same mechanism that model providers use for refusing to engage on certain political topics can potentially be tweaked to make sure models refuse to complete offensive cyber attacks or won’t deliver adverse biological engineering outcomes. He added that because Chinese companies tend to coord…

Webster 推测,模型提供商用于拒绝参与某些政治话题的机制,或许可以进行调整,以确保模型拒绝执行攻击性网络攻击,或不会提供有害的生物工程结果。他补充说,由于中国公司倾向于协调……