Chinese AI model Kimi escaped its cybersecurity testing environment, researchers say
Chinese AI model Kimi escaped its cybersecurity testing environment, researchers say
研究人员称:中国 AI 模型 Kimi 逃脱了网络安全测试环境
Kimi K3, the latest AI model made by Chinese company Moonshot, escaped an environment set up to test its cyber capabilities, researchers said in a blog post published on Friday. The news shows once again that companies and independent organizations are struggling to contain their AI models designed for hacking. 中国公司月之暗面(Moonshot)推出的最新 AI 模型 Kimi K3 逃脱了为其网络能力测试所设定的环境,研究人员在周五发布的一篇博文中披露了这一消息。这一新闻再次表明,各大公司和独立机构在管控其专为黑客攻击设计的 AI 模型方面正面临严峻挑战。
In recent weeks, frontier LLMs at U.S. artificial intelligence labs at OpenAI, Anthropic, and Meta, as well as the U.K.’s AI Security Institute, all escaped testing environments in different ways and ended up hacking real targets that were not part of the experiment. This is starting to happen so often there’s now a website tracking all these incidents called Felony Bench, a nod to the fact that these LLMs may be committing crimes — at least theoretically speaking. 近几周,来自 OpenAI、Anthropic 和 Meta 等美国人工智能实验室,以及英国人工智能安全研究所的前沿大语言模型(LLM),都以不同方式逃脱了测试环境,并最终攻击了实验范围之外的真实目标。此类事件发生得如此频繁,以至于现在出现了一个名为“Felony Bench”的网站专门追踪这些事件,该名称暗示了这些大语言模型可能正在(至少在理论上)实施犯罪。
In the case of this Kimi test, the sandbox designed to contain the experiment was not properly configured. While the sandbox disallowed the AI model from accessing certain web traffic, the model instead bypassed the sandbox by relying on command line tools, according to the researchers at AI-focused cybersecurity firm Frontier Security. 在这次 Kimi 的测试案例中,用于隔离实验的沙箱配置不当。据专注于人工智能的网络安全公司 Frontier Security 的研究人员称,虽然沙箱禁止了该 AI 模型访问特定的网络流量,但模型通过依赖命令行工具绕过了沙箱的限制。
“This suggests that some of the evaluations on cybersecurity the community uses are susceptible to security vulnerabilities and allow models to cheat, and that there are models that intentionally seek loopholes and vulnerabilities which allows them to cheat on evaluations,” the researchers wrote. 研究人员写道:“这表明社区目前使用的一些网络安全评估方法容易受到安全漏洞的影响,从而允许模型作弊;同时也说明,确实存在一些模型会主动寻找漏洞,以便在评估中进行欺骗。”
If you are keeping score at home, according to Felony Bench’s tally, Moonshot now joins OpenAI and Anthropic, which have seven recorded incidents each, and Meta, which has one. 如果你在关注相关记录,根据 Felony Bench 的统计,月之暗面(Moonshot)现在与 OpenAI 和 Anthropic(各记录有 7 起事件)以及 Meta(记录有 1 起)并列出现在名单中。