UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities
UK AISI / CAISI Preliminary Assessment of Kimi K3’s Cyber Capabilities
英国人工智能安全研究所 (UK AISI) 与美国人工智能标准与创新中心 (CAISI) 对 Kimi K3 网络能力的初步评估
The UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) (UK AISI / CAISI) conducted a joint evaluation of Moonshot AI’s latest model, Kimi K3 (released on July 16, 2026 and slated for open-weight release by July 27, 2026). This evaluation focused on Kimi K3’s cyber capabilities and found that: 英国人工智能安全研究所 (UK AISI) 与美国人工智能标准与创新中心 (CAISI) 对月之暗面 (Moonshot AI) 的最新模型 Kimi K3(于 2026 年 7 月 16 日发布,计划于 2026 年 7 月 27 日发布开源权重版本)进行了联合评估。此次评估重点关注 Kimi K3 的网络能力,结果发现:
Kimi K3 performs significantly below the most recent frontier cyber-capable models on preliminary cyber evaluations run by UK AISI / CAISI. Specifically: 在 UK AISI / CAISI 进行的初步网络评估中,Kimi K3 的表现显著低于目前最前沿的具备网络能力模型。具体表现如下:
-
When tasked to develop exploits, Kimi K3 performs significantly below the most recent frontier cyber-capable models (Figure 1). 当被要求开发漏洞利用程序时,Kimi K3 的表现显著低于目前最前沿的具备网络能力模型(图 1)。
-
When tasked to attack a simulated corporate network (“The Last Ones”), Kimi K3 performs significantly below the most recent frontier cyber-capable models. Specifically, on average, Kimi K3 reached step 17 of this 32-step attack path, while the most cyber-capable U.S. models reached 28.5 steps on average. 当被要求攻击模拟企业网络(“The Last Ones”)时,Kimi K3 的表现显著低于目前最前沿的具备网络能力模型。具体而言,Kimi K3 平均完成了该 32 步攻击路径中的第 17 步,而美国最强的网络能力模型平均达到了 28.5 步。
-
Kimi K3 performs above GLM-5.2 on the same preliminary cyber evaluations run by UK AISI / CAISI (Figure 2). 在 UK AISI / CAISI 进行的相同初步网络评估中,Kimi K3 的表现优于 GLM-5.2(图 2)。
-
Kimi K3’s safeguards allow assistance with agentic cyber exploit development. Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during UK AISI / CAISI’s evaluations. Kimi K3 的安全防护机制允许其协助进行代理式网络漏洞利用开发。在 UK AISI / CAISI 的评估过程中,Kimi K3 的安全防护机制未能阻止其尝试进行网络漏洞利用开发或进攻性网络操作。
Detailed Results
详细结果
These results represent preliminary evaluations on a small set of public and private benchmarks. U.S. closed-weight models were evaluated with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities. Publicly available versions of these models have these safeguards enabled. Due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations. Detailed methodologies are provided in individual sections. 这些结果代表了在一组小型公共和私有基准测试上的初步评估。美国闭源模型在评估时禁用了系统级安全防护,以减少拒绝响应并测量其最大能力。这些模型的公开版本均已启用此类安全防护。由于 Kimi K3 托管设置的特殊性,UK AISI / CAISI 仅进行了一系列选择性的网络评估。详细方法将在各章节中说明。
Cyber Capability Trends
网络能力趋势
The cyber capability of models is aggregated across multiple tasks from multiple benchmarks using an approach inspired by Item Response Theory (IRT). For details of the methodology please see prior published reports. Kimi K3’s overall cyber capability has a larger confidence interval than other models because it was estimated from a single benchmark (ExploitBench, which has 41 tasks focused on exploit development). ExploitBench is a leading benchmark to measure a model’s ability to progress along the software exploitation ladder. All other models’ overall cyber capability scores were derived from a larger number of tasks that covered additional domains of cyber capability. 模型网络能力是通过受项目反应理论 (IRT) 启发的方法,汇总多个基准测试中的多项任务得出的。有关方法论的详细信息,请参阅之前发布的报告。Kimi K3 的整体网络能力置信区间比其他模型更大,因为它是基于单一基准测试(ExploitBench,包含 41 项专注于漏洞利用开发的任务)估算的。ExploitBench 是衡量模型在软件利用链条上进展能力的领先基准。其他所有模型的整体网络能力得分均源自涵盖更多网络能力领域的任务。
Exploit Development: ExploitBench
漏洞利用开发:ExploitBench
ExploitBench is a public benchmark, developed by Carnegie Mellon University, that measures a model’s ability to progress along the software exploitation ladder, including coverage and crash reproduction, arbitrary read/write, control flow hijack, and arbitrary code execution. The benchmark tests models on 41 recent (post-2023) vulnerabilities in the V8 engine (the JavaScript and WebAssembly software that powers Chrome). ExploitBench 是由卡内基梅隆大学开发的公共基准测试,用于衡量模型在软件利用链条上的进展能力,包括覆盖率与崩溃复现、任意读/写、控制流劫持以及任意代码执行。该基准测试在 V8 引擎(驱动 Chrome 的 JavaScript 和 WebAssembly 软件)中 41 个近期(2023 年后)漏洞上对模型进行测试。
-
Kimi K3 achieves a score of 32%, whereas GLM-5.2 achieves a score of 24% (Figure 1). Kimi K3 的得分为 32%,而 GLM-5.2 的得分为 24%(图 1)。
-
Unlike the most cyber-capable models, Kimi K3 failed to develop exploits that achieved arbitrary code execution (ACE) for ExploitBench tasks. ACE is the highest-severity outcome in exploit development, granting attackers the ability to hijack a target. Kimi K3 achieved ACE on 0/41 samples, whereas the most cyber-capable models achieved ACE on 20/41 samples on average (Figure 3). 与最强的网络能力模型不同,Kimi K3 未能在 ExploitBench 任务中开发出实现任意代码执行 (ACE) 的漏洞利用程序。ACE 是漏洞利用开发中严重程度最高的结果,赋予攻击者劫持目标的能力。Kimi K3 在 41 个样本中实现了 0 次 ACE,而最强的网络能力模型平均在 41 个样本中实现了 20 次 ACE(图 3)。
Cyber Range: The Last Ones (TLO)
网络靶场:“The Last Ones” (TLO)
“The Last Ones” (TLO) cyber range is a 32-step simulated corporate network attack spanning 4 subnets and approximately 20 hosts, which would take a human expert roughly 20 hours to complete. Cyber ranges are expert-built, simulated networks of hosts, services, and vulnerabilities arranged into sequential attack chains that begin at the point of initial network access, and can be used to measure a model’s ability to conduct end-to-end cyberattacks autonomously. “The Last Ones” (TLO) 网络靶场是一个包含 32 个步骤的模拟企业网络攻击场景,跨越 4 个子网和约 20 台主机,人类专家完成该任务大约需要 20 小时。网络靶场是由专家构建的模拟网络,包含主机、服务和漏洞,并排列成从初始网络访问点开始的顺序攻击链,可用于衡量模型自主进行端到端网络攻击的能力。
On this evaluation, Kimi K3 performs significantly below the leading U.S. cyber capable models. Specifically, Kimi K3 reached step 17 of this 32-step attack path on average, while the most cyber-capable U.S. models reached… 在此次评估中,Kimi K3 的表现显著低于美国领先的具备网络能力模型。具体而言,Kimi K3 平均完成了该 32 步攻击路径中的第 17 步,而美国最强的网络能力模型达到了……