OpenAI "rogue" agent activities found on Wikimedia projects

OpenAI “rogue” agent activities found on Wikimedia projects

OpenAI “流氓”智能体在维基媒体项目上的活动被发现

Recently, multiple organisations have disclosed how clusters of so-called “rogue” AI agents attempted to break into websites and online services, sometimes successfully. Agents from OpenAI’s environment, in particular, are known to have used other public wikis (collaboratively edited websites not owned by us) to communicate and coordinate with each other. These types of successful intrusions can expose sensitive data or disrupt website services that users rely on, while clusters of agents can attempt attacks at a scale that is difficult for defenders to manage. They affect people behind the websites who may not understand the nature of the attack, or have the tools to effectively fight back. For a site like Wikipedia, agents might find and use security vulnerabilities or make misleading edits at scale. Wikipedia’s volunteer editors and the Wikimedia Foundation’s security teams have to detect and undo that activity.

最近,多家机构披露了所谓的“流氓”AI 智能体集群如何尝试入侵网站和在线服务,且有时已获得成功。特别是来自 OpenAI 环境的智能体,已知曾利用其他公共维基(非我们拥有的协作编辑网站)进行相互通信和协调。这类成功的入侵可能会泄露敏感数据或破坏用户所依赖的网站服务,而智能体集群发起的攻击规模往往令防御者难以应对。这些攻击影响了网站背后的维护者,他们可能并不了解攻击的本质,也没有工具进行有效反击。对于维基百科这样的网站,智能体可能会发现并利用安全漏洞,或进行大规模的误导性编辑。维基百科的志愿者编辑和维基媒体基金会的安全团队必须检测并撤销这些活动。

The Wikimedia Foundation conducted its own investigation to see whether Wikimedia websites had been similarly affected by AI agents, focusing on those operated by OpenAI. We can confirm that we have discovered some activity by these “rogue” OpenAI agents on Wikimedia platforms. The unauthorized bot activities included edits to our wikis, some unsuccessful attempts to exploit a public note-taking tool we host, and heavy traffic, which are described more below. We did not find any evidence that our systems were used for coordination among agents, nor did we find any evidence of our systems or data being compromised. However, we are concerned about what could have occurred here, the difficulty and effort involved in investigating and attributing this activity, and the growing risks of agentic AI activity on our platforms in general. The open web is a public good. We should not allow this behavior to become the “new normal” for the people or organizations that maintain it.

维基媒体基金会进行了专项调查,以确定维基媒体网站是否同样受到 AI 智能体的影响,重点关注由 OpenAI 运营的智能体。我们可以确认,我们确实在维基媒体平台上发现了这些“流氓”OpenAI 智能体的活动。这些未经授权的机器人活动包括对我们维基的编辑、对我们托管的公共笔记工具的一些不成功的利用尝试,以及高负载流量,详情见下文。我们没有发现任何证据表明我们的系统被用于智能体之间的协调,也没有发现我们的系统或数据遭到破坏的证据。然而,我们对可能发生的情况、调查和归因这些活动所涉及的难度与工作量,以及我们平台上智能体 AI 活动日益增长的风险感到担忧。开放网络是公共产品。我们不应允许这种行为成为维护它的人员或组织眼中的“新常态”。

In summary, we saw: Wiki editing: We’ve identified edits to Wikimedia wikis that we believe are from AI agents operated by OpenAI. These edits were not published to pages with visibility to general readers; almost all of them were testing edits in “sandbox” areas of the wiki. It also included a few edits to the configuration for a citation tool, which we believe were potentially malicious edits that were intended to misuse this tool as a proxy for fetching data from remote services. While Wikipedia policies allow bots to edit when they are disclosed and approved by the community, none of those approvals were sought in these incidents.

总之,我们观察到: 维基编辑:我们识别出对维基媒体维基的编辑,我们认为这些编辑来自 OpenAI 运营的 AI 智能体。这些编辑并未发布在普通读者可见的页面上;几乎所有编辑都是在维基的“沙盒”区域进行的测试。此外,还包括对引用工具配置的一些编辑,我们认为这些可能是恶意编辑,旨在滥用该工具作为从远程服务获取数据的代理。虽然维基百科的政策允许在披露并经社区批准的情况下使用机器人进行编辑,但在这些事件中,没有任何一项活动寻求过此类批准。

Etherpad probing and use: Agents we believe to be operated by OpenAI made some unsuccessful attempts to compromise our public Etherpad, a note-taking tool we host as a community service. Agents unsuccessfully tried to use it to fetch data from other websites as a proxy. Other agents also likely operated by OpenAI took notes about their tasks, though this did not appear to turn into coordination.

Etherpad 探测与使用:我们认为由 OpenAI 运营的智能体曾尝试入侵我们的公共 Etherpad(我们作为社区服务托管的笔记工具),但未获成功。智能体试图将其用作代理从其他网站获取数据,但未能成功。其他可能由 OpenAI 运营的智能体记录了关于其任务的笔记,尽管这似乎并未演变成协调行为。

Excessive data downloading: Agents we believe to be operated by OpenAI made millions of automated requests to our public APIs to access the knowledge on Wikimedia projects, crawled millions of pages (mainly from our projects Wikidata and Wikimedia Commons), and made hundreds of thousands of data queries to the Wikidata Query Service (WQDS). This traffic may have contributed to a partial outage on WQDS in May.

过度数据下载:我们认为由 OpenAI 运营的智能体向我们的公共 API 发送了数百万次自动请求,以访问维基媒体项目上的知识,抓取了数百万个页面(主要来自我们的 Wikidata 和维基共享资源项目),并向 Wikidata 查询服务 (WQDS) 发送了数十万次数据查询。这些流量可能导致了 WQDS 在 5 月份的部分中断。

As a non-profit technology host of some of the largest and most widely used open knowledge platforms in the world, we are deeply concerned about the impact of “rogue” AI agents on platforms like ours, which are built by volunteers from around the world and rely on the promise of the open internet. Incidents like this one, and the many others that have been (and are still being) uncovered, illustrate how AI agents can drain resources and crash servers, as well as attempt to compromise trustworthy information.

作为世界上一些最大、使用最广泛的开放知识平台的非营利性技术托管方,我们对“流氓”AI 智能体对我们这类平台的影响深感担忧。这些平台由世界各地的志愿者构建,并依赖于开放互联网的承诺。像这样的事件,以及许多其他已经被(且仍在被)发现的事件,说明了 AI 智能体如何消耗资源、导致服务器崩溃,并试图破坏可信信息。

Over the past 25 years, Wikipedia has grown into one of the most popular and trusted websites in the world, with more than 67 million articles across over 300 languages, and up to 15 billion page views per month. Through an open, transparent, and collaborative process, volunteers work to ensure that knowledge remains neutral, reliable, and accessible to everyone. Wikipedia is one of the highest-quality datasets used in training Large Language Models (LLMs), and its knowledge forms the backbone of information on the internet, powering AI chatbots, search engines, voice assistants, and more. Wikipedia was designed for humans – and agentic behavior clearly poses challenges that no one has solutions for.

在过去的 25 年里,维基百科已成长为世界上最受欢迎和最受信任的网站之一,拥有超过 300 种语言的 6700 多万篇文章,每月页面浏览量高达 150 亿次。通过开放、透明和协作的过程,志愿者们致力于确保知识保持中立、可靠且人人可及。维基百科是用于训练大语言模型 (LLM) 的最高质量数据集之一,其知识构成了互联网信息的骨干,为 AI 聊天机器人、搜索引擎、语音助手等提供动力。维基百科是为人类设计的——而智能体行为显然带来了目前无人能解的挑战。

Because of our unique and successful knowledge creation model, Wikimedia’s volunteers are the ones who come in first contact with, and clean up the mess left behind by AI agents. Rising bot traffic and agentic activity is showing a real impact on the Wikimedia projects and the infrastructure that makes it available for millions of users globally. In 2025, the Foundation reported that its bandwidth usage had increased by 50% due to the surge of bot activity on its websites since 2024. At the same time, 65% of the most resource-consuming traffic on its projects was coming from bots. This intense pressure on our infrastructure not only adds costs for servers and humans, but if left unaddressed, can block human visitors by overloading systems and causing outages. We are already paying for costs that come with the increased activity.

由于我们独特且成功的知识创造模式,维基媒体的志愿者们成为了第一批接触并清理 AI 智能体留下的烂摊子的人。不断增长的机器人流量和智能体活动正在对维基媒体项目以及为全球数百万用户提供服务的底层基础设施产生实际影响。2025 年,基金会报告称,由于 2024 年以来网站上机器人活动的激增,其带宽使用量增加了 50%。与此同时,其项目中 65% 的高资源消耗流量来自机器人。这种对基础设施的巨大压力不仅增加了服务器和人力成本,如果任其发展,还可能因系统过载和导致中断而阻碍人类访客的访问。我们已经在为这些增加的活动所带来的成本买单。

Wikimedia’s volunteers have stayed resilient so far in tackling emerging challenges on our platforms, but we also want to say: it doesn’t need to be this way. While OpenAI admits to agents behaving “unpredictably”, they must also acknowledge their responsibility to monitor and prevent these risks. AI companies are not doing enough to secure their systems and protect the public from the harm they cause. That burden is falling onto everyone else, including smaller organizations. At a minimum, their systems should operate in a way that non-profit website owners li…

维基媒体的志愿者们迄今为止在应对平台上的新兴挑战时表现出了韧性,但我们也想说:情况本不必如此。虽然 OpenAI 承认智能体的行为“不可预测”,但他们也必须承认自己有责任监控并预防这些风险。AI 公司在保护其系统安全以及保护公众免受其造成的伤害方面做得还不够。这种负担正落在其他人身上,包括较小的组织。至少,他们的系统运行方式应当让非营利网站所有者……