Likely illegally, Claude gained access to 3 networks. Will Anthropic be held to account?

Likely illegally, Claude gained access to 3 networks. Will Anthropic be held to account?

Claude 可能非法入侵了 3 个网络,Anthropic 会被追责吗?

Anthropic said its Claude-based security models gained unauthorized access to the sensitive production environments of three outside organizations during internal testing designed to measure the models’ offensive cyber capabilities. The events, which Anthropic revealed Thursday, are the second revelation in 10 days that AI models from the world’s wealthiest providers have trespassed into protected networks, an offense that, in more traditional hacking scenarios, could land the human behind the keyboard in prison for years.

Anthropic 表示,其基于 Claude 的安全模型在旨在评估模型攻击性网络能力的内部测试中,未经授权访问了三个外部组织的敏感生产环境。Anthropic 周四披露的这些事件,是 10 天内第二次曝出全球最富有的 AI 提供商旗下的模型擅自闯入受保护网络。在传统的黑客攻击场景中,这种行为足以让操作者面临多年的牢狱之灾。

Earlier this month, OpenAI said its security models exploited a zero-day vulnerability for use in breaking into the network of Hugging Face, a platform for open source machine-learning models and AI datasets. The OpenAI models went on to steal access credentials and other confidential Hugging Face information. The OpenAI models also exploited publicly exposed credentials to compromise accounts of four other third-party services.

本月早些时候,OpenAI 表示其安全模型利用一个零日漏洞入侵了开源机器学习模型和 AI 数据集平台 Hugging Face 的网络。OpenAI 的模型随后窃取了访问凭据及其他 Hugging Face 的机密信息。此外,这些模型还利用公开暴露的凭据入侵了其他四个第三方服务的账户。

Anthropic said the OpenAI event spurred its engineers to review similar cybersecurity evaluations by Claude models. The audit found three incidents “in which a model accessed the internet from within or while interacting with the evaluation environment of Irregular, one of our third-party evaluation partners, and then gained unauthorized access to the production infrastructure of three different organizations.”

Anthropic 表示,OpenAI 的事件促使其工程师对 Claude 模型的类似网络安全评估进行了审查。审计发现了三起事件,即“模型在与我们的第三方评估合作伙伴 Irregular 的评估环境交互时,从内部访问了互联网,随后未经授权访问了三个不同组织的生产基础设施。”

When Models fail to discern reality from fiction

当模型无法分辨现实与虚构时

Anthropic said the prompts engineers delivered during “capture the flag” challenges—a common means for assessing the offensive and defensive capabilities of various hacking techniques—made clear that the testing environment was only a simulation and that the models had no access to the open Internet. In fact, the testing partner Irregular mistakenly made such access available. In response the models treated the Internet paths as part of the exercises.

Anthropic 表示,工程师在“夺旗赛”(Capture the Flag,评估各种黑客技术攻防能力的常用手段)中给出的提示词明确指出,测试环境仅为模拟环境,且模型无法访问开放互联网。事实上,测试合作伙伴 Irregular 错误地提供了这种访问权限。作为回应,模型将这些互联网路径视为练习的一部分。

The intrusions happened through three Claude models: Opus 4.7, Mythos 5, and an internal research prototype. Opus 4.7, the oldest of the three, was the one that overstepped its boundaries the most. Anthropic explained: Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.

入侵事件涉及三个 Claude 模型:Opus 4.7、Mythos 5 和一个内部研究原型。其中最古老的 Opus 4.7 是越界行为最严重的一个。Anthropic 解释称:Claude 在错误地认为所有可访问实体都在练习范围内的情况下,利用弱密码和未经身份验证的端点等基本技术,破坏了受影响组织的基础设施。

It did not find or exploit any complex vulnerabilities, and in each case, Claude continued working to complete only the specific capture-the-flag task its evaluation had assigned. However, in some cases, our older model continued its attack even after getting evidence it was running on the open internet; our latest model stopped once it recognized it was on the internet. In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment.

它没有发现或利用任何复杂的漏洞,且在每种情况下,Claude 都只是继续完成评估所分配的特定夺旗任务。然而,在某些情况下,我们的旧模型在获得其正在开放互联网上运行的证据后仍继续攻击;而我们的最新模型在识别出处于互联网环境后便停止了操作。在所有这些情况下,Claude 都没有试图窃取自身数据,也没有故意尝试逃离测试环境。

While the Opus model continued its attack even after correctly predicting that it had breached the third-party organization’s production system, Anthropic said, Mythos 5 made a similar inference but “reasoned its way back to the conclusion that it was still in a simulation.” That meant that Mythos, too, didn’t abort the exercise once it had clearly stepped over the line. After also exceeding the bounds of level behavior, the internal test model eventually found evidence that it, too, had impinged and stopped the exercise.

Anthropic 表示,尽管 Opus 模型在正确预测其已入侵第三方组织生产系统后仍继续攻击,但 Mythos 5 虽然做出了类似的推断,却“通过推理得出了它仍处于模拟环境中的结论”。这意味着 Mythos 在明显越界后也没有中止练习。内部测试模型在超出行为准则后,最终也发现了其已造成入侵的证据并停止了练习。

In the first of the three intrusions, Opus 4.7 was initially unable to breach the simulated target it was prompted to attack. Once the model discovered that its environment did, in fact, have Internet access, it identified and exploited vulnerabilities inside the network of a real company with the same name as the target in the simulated attack scenario. In four runs, the model extracted application and infrastructure credentials and several hundred rows of production data.

在三起入侵事件的第一起中,Opus 4.7 最初无法攻破其被要求攻击的模拟目标。一旦模型发现其环境实际上可以访问互联网,它便识别并利用了一家真实公司的网络漏洞,而该公司与模拟攻击场景中的目标同名。在四次运行中,该模型提取了应用程序和基础设施凭据以及数百行生产数据。

Anthropic said the incursion was the most serious of the three. “This incident involved Claude Opus 4.7, and was the only case in which the model continued to attack a system after learning it was likely operating in a real environment,” Thursday’s post stated. “In all four of the runs, the model eventually recognized (as can be observed in its verbalized reasoning) that the system was real—in two cases, the model rationalized that the real company must be part of the exercise. However, none stopped the attack on that basis.”

Anthropic 表示,这次入侵是三起事件中最严重的一次。周四的文章称:“该事件涉及 Claude Opus 4.7,也是唯一一个模型在得知其可能在真实环境中运行时仍继续攻击系统的案例。在所有四次运行中,模型最终都识别出(从其口头推理中可以看出)系统是真实的——在其中两次案例中,模型合理化地认为这家真实公司一定是练习的一部分。然而,没有一次因此停止攻击。”

In the second instance, Mythos 5 detected a document inside the fictional environment that the model reasoned was another made-up company’s setup instructions for new developers. The instructions directed employees to install a non-existent Python package from PyPI. Still following prompts to breach the first fictional company, the model built and published a malicious package with the same name. The objective appears to have been to somehow get the real company to run it.

在第二起案例中,Mythos 5 在虚构环境中检测到一份文档,模型推断这是另一家虚构公司为新开发者准备的设置说明。该说明指示员工从 PyPI 安装一个不存在的 Python 包。在继续执行入侵第一家虚构公司的提示时,该模型构建并发布了一个同名的恶意包。其目的似乎是想方设法让真实公司运行它。

The Anthropic post-mortem went on to say: Claude went to extensive lengths to carry out this attack—lengths that would likely have indicated to a human participant that this was no longer just an evaluation, and that they were in fact uploading a real PyPI package. For instance, in one case, in order to create a PyPI account, Claude needed an email address. And in order to create an email address, it needed a phone number. To get a phone number, after failing to find a free phone number service, it tried—and failed—to obtain funds to pay for a phone number through several different means. It finally backtracked, found a free, non-blocked email provider, used this to register a PyPI account, and then used this account to upload malware to PyPI.

Anthropic 的事后分析继续写道:Claude 为了实施这次攻击采取了极其复杂的手段——这些手段对于人类参与者来说,很可能已经表明这不再仅仅是一次评估,而是在上传真实的 PyPI 包。例如,在一种情况下,为了创建 PyPI 账户,Claude 需要一个电子邮件地址。而为了创建电子邮件地址,它需要一个电话号码。为了获取电话号码,在未能找到免费电话号码服务后,它尝试通过几种不同方式获取资金来支付电话号码,但均告失败。最终,它回溯并找到了一个免费且未被屏蔽的电子邮件提供商,用它注册了 PyPI 账户,然后利用该账户将恶意软件上传到了 PyPI。

During the roughly one-hour window the Python package was available for download, it was run on 15 real systems, including a scanner belonging to a real security company whose name wasn’t included in the simulated attack.

在 Python 包可供下载的大约一小时窗口期内,它在 15 个真实系统上被运行,其中包括一家真实安全公司的扫描器,而该公司的名称并未包含在模拟攻击中。