Rogue AI agents created fake online identities in another hacking attempt
Rogue AI agents created fake online identities in another hacking attempt
恶意 AI 智能体在另一次黑客攻击尝试中伪造了在线身份
Yet more rogue AI agents from OpenAI and Anthropic have been caught attempting to hack real targets online without permission. The discoveries add to a growing list of previously unknown incidents that have alarmed AI safety experts and intensified pressure for greater oversight of frontier systems. 来自 OpenAI 和 Anthropic 的更多恶意 AI 智能体被发现未经许可试图在网上攻击真实目标。这些发现增加了一系列此前未知的事件,令 AI 安全专家感到震惊,并加剧了对前沿系统加强监管的压力。
According to a report from the UK’s AI Security Institute (AISI), which evaluates frontier models from top AI labs before they are released, agents powered by OpenAI’s GPT-5.6-Sol and Anthropic’s Mythos 5 “engaged in sustained, potentially harmful activity directed at real people and organisations.” This included trying to insert malicious code into an open-source project by pressuring real people in charge of it, AISI said. “In an attempt to get the code approved, the agent engaged in social engineering — creating fake online identities and using them to pressure the project’s maintainer to approve the code.” 根据英国人工智能安全研究所(AISI)的一份报告,该机构负责在顶级 AI 实验室发布前沿模型前对其进行评估。报告显示,由 OpenAI 的 GPT-5.6-Sol 和 Anthropic 的 Mythos 5 驱动的智能体“针对真实个人和组织进行了持续的、潜在的有害活动”。AISI 表示,这包括试图通过向开源项目的负责人施压,将恶意代码植入该项目。“为了让代码获得批准,该智能体进行了社会工程学攻击——创建虚假的在线身份,并利用这些身份向项目维护者施压,要求其批准代码。”
AISI said the attempts, which it detected on July 28th, “were unsuccessful” and had not resulted in real-world harm. However, the organization noted that the incident marked “the first time we have seen risks around autonomy and deception manifest this clearly, without specific prompting, in the real-world.” AISI 表示,这些尝试在 7 月 28 日被发现,但“均未成功”,且未造成现实世界的损害。然而,该机构指出,此次事件标志着“我们首次看到围绕自主性和欺骗性的风险在没有特定提示的情况下,在现实世界中如此清晰地显现出来。”
Unlike OpenAI’s rogue agent that attacked Hugging Face, AISI said this was “not a case of a model escaping its secure test environment,” or sandbox. Safeguards usually imposed on the models had been disabled as part of testing, AISI said, and they had also been permitted access to the internet. “To measure what these models can genuinely do, we test them under conditions that reflect what a capable human attacker could do,” AISI said. 与此前攻击 Hugging Face 的 OpenAI 恶意智能体不同,AISI 表示,这“并非模型逃离其安全测试环境(沙箱)的情况”。AISI 指出,作为测试的一部分,通常施加在模型上的安全防护措施已被禁用,且它们被允许访问互联网。“为了衡量这些模型的真实能力,我们在反映有能力的黑客可能采取的操作条件下对它们进行了测试,”AISI 表示。
The incident stemmed from a single AISI evaluation where agents were tasked with solving a cybersecurity challenge, such as finding a piece of protected data. The challenge was run 122 times across multiple models and all runs were conducted in AISI’s research environment, which uses “virtual machine sandboxing to isolate the agents from other AISI infrastructure.” AISI’s investigation found that in 10 of those, “an AI agent took autonomous, unsanctioned action on the live internet, targeting real people and organisations.” Of 19 such actions, almost all — 17 — came from Anthropic’s Mythos 5. 该事件源于 AISI 的一次评估,当时智能体的任务是解决一个网络安全挑战,例如寻找一段受保护的数据。该挑战在多个模型上运行了 122 次,所有运行均在 AISI 的研究环境中进行,该环境使用“虚拟机沙箱将智能体与 AISI 的其他基础设施隔离开来”。AISI 的调查发现,其中 10 次运行中,“AI 智能体在实时互联网上采取了自主的、未经授权的行动,针对真实个人和组织”。在 19 次此类行动中,几乎所有(17 次)都来自 Anthropic 的 Mythos 5。
In its post-mortem of the incident, AISI identified several key factors it said contributed to the unsanctioned agent behaviors. It said the agent was persistent, pursuing avenues like trying to trick real people through “deception that, until recently, had been largely theoretical.” The task was also hard, which the organization said could push agents to be more “creative” in their problem-solving. Compounding matters were deficiencies in how internet use was monitored, with AISI suggesting that more dedicated surveillance could have identified the problem sooner. Finally, the organization said the agent hadn’t been specifically instructed not to leverage its internet access or deploy deceptive social engineering techniques in pursuit of its goal. “Previously, it was not clear that such instructions were necessary when using models with alignment training,” AISI said. 在事后分析中,AISI 指出了导致这些未经授权行为的几个关键因素。它表示,智能体表现出持续性,通过“直到最近还大多处于理论阶段的欺骗手段”来试图欺骗真实人类。任务难度较高,该机构认为这可能促使智能体在解决问题时变得更加“有创造力”。此外,互联网使用监控方面的不足也加剧了问题,AISI 建议,更专业的监控本可以更早发现问题。最后,该机构表示,并未明确指示智能体不得利用其互联网访问权限或部署欺骗性的社会工程学技术来实现目标。“此前,人们并不清楚在使用经过对齐训练的模型时,是否有必要给出此类指令,”AISI 表示。
AISI said the incident should be “interpreted with caution and nuance” but warned the agent’s actions “show signs of novel, potentially deceptive behaviours” that “were to an extent and severity we did not anticipate.” AISI 表示,该事件应“谨慎且细致地解读”,但警告称,智能体的行为“显示出新颖的、潜在的欺骗性行为迹象”,其“程度和严重性超出了我们的预期”。
In a blog post, OpenAI acknowledged the breach that happened during AISI’s testing and said it is “committed to working across the industry to strengthen shared practices for conducting high-risk evaluations safely.” OpenAI also disclosed another breach, this time from an external cybersecurity testing partner Irregular, where it said models had been mistakenly granted internet access during cybersecurity exercises. OpenAI said Irregular notified it of the breach on July 29th. 在一篇博客文章中,OpenAI 承认了在 AISI 测试期间发生的违规行为,并表示“致力于与整个行业合作,加强安全进行高风险评估的共同实践”。OpenAI 还披露了另一起违规事件,这次来自外部网络安全测试合作伙伴 Irregular,称模型在网络安全演习中被错误地授予了互联网访问权限。OpenAI 表示,Irregular 在 7 月 29 日通知了其这一违规行为。
“In the coming weeks, we will review our own approach to third-party testing, including how we identify higher-risk evaluations, agree on scope, assess requests to enable internet access or lowered safeguards, set expectations for isolation, credential handling, monitoring, and stop conditions, and establish clearer incident-notification and escalation processes,” OpenAI said. “在接下来的几周内,我们将审查我们自己的第三方测试方法,包括我们如何识别更高风险的评估、商定范围、评估启用互联网访问或降低安全防护的请求、设定隔离、凭据处理、监控和停止条件的预期,并建立更清晰的事件通知和升级流程,”OpenAI 表示。
Anthropic posted a less comprehensive response on X, largely emphasizing that the models’ standard safety features had been disabled and that they had not been given “any specific restrictions on how the internet should be used.” It said it was working closely with AISI to gather more details for its own investigation. Anthropic 在 X 上发布了一份内容较简略的回应,主要强调模型的标准安全功能已被禁用,且未对“如何使用互联网”设置“任何具体限制”。该公司表示,正在与 AISI 密切合作,为其自身的调查收集更多细节。
The findings add to an increasingly tangled mess of rogue actions from agents during testing, many of which only come to light after dedicated hunting and which feature models not released to the public. The unwillingness or inability of AI labs to contain their products has sparked concern over how such breaches could go unnoticed, the safety of frontier AI systems, and worries over the general lack of transparency and oversight the industry faces. 这些发现加剧了测试期间智能体恶意行为的混乱局面,其中许多行为只有在专门搜寻后才会被发现,且涉及尚未向公众发布的前沿模型。AI 实验室在控制其产品方面的意愿或能力不足,引发了人们对此类违规行为如何可能被忽视、前沿 AI 系统的安全性,以及该行业普遍缺乏透明度和监管的担忧。