Cloudflare's new AI traffic options for customers
Cloudflare’s new AI traffic options for customers
Cloudflare 为客户推出全新的 AI 流量管理选项
One year ago, we declared the first Content Independence Day, and we gave website owners the means to take back control of their content. The deal between crawlers and website owners that had held up for 30 years — we crawl you, and you get referrals — was no longer true. AI was taking everything and sending back nothing, presenting an existential threat to website owners. And so we launched a one-click “Block AI Bots” option, along with a Pay-Per-Crawl marketplace.
一年前,我们宣布了首个“内容独立日”,并赋予了网站所有者收回其内容控制权的手段。过去 30 年来爬虫与网站所有者之间维持的协议——即“我们抓取你的内容,你获得引流”——已不再适用。人工智能正在获取一切却不提供任何回馈,这对网站所有者构成了生存威胁。因此,我们推出了一键式“屏蔽 AI 机器人”选项,以及“按抓取付费”(Pay-Per-Crawl)的市场模式。
A lot has changed in a year. Last July, conversations around “AI bots” centered around blocking AI training without compensation, pointing to the win–lose deal where content was used for model training with no value driven back to the website owner. But a desire for more nuance has emerged: Content owners still want to be able to protect their content, and they should be compensated for the original content that they work hard to create, curate, and share. We also know that locking down content isn’t a one-size-fits-all solution; website owners want more options than resorting to “block all automation, every time.”
一年来,情况发生了很大变化。去年 7 月,围绕“AI 机器人”的讨论主要集中在如何阻止未经补偿的 AI 训练上,这指向了一种“赢家通吃”的局面:内容被用于模型训练,却未给网站所有者带来任何价值。但现在,人们对精细化管理的需求日益增长:内容所有者仍希望能够保护其内容,并且他们理应因自己辛苦创作、整理和分享的原创内容而获得补偿。我们也深知,封锁内容并非万能方案;网站所有者希望拥有更多选择,而不是每次都只能诉诸于“屏蔽所有自动化程序”。
If you run a small site, the problem isn’t just that someone could train models on your content — it’s that nobody can find you in the first place. So you have to make a Faustian bargain: either show up in search and let AI train on you, or risk losing discoverability. This unfairly advantages incumbent search providers if they use the same bots for both search and training; and this unfair advantage incentivizes new players to be evasive as they try to close the competitive gap.
如果你经营一个小网站,问题不仅在于有人可能利用你的内容训练模型,更在于根本没人能找到你。因此,你不得不做出一种“浮士德式的交易”:要么出现在搜索结果中并允许 AI 对你进行训练,要么冒着失去曝光度的风险。如果现有的搜索服务商同时使用相同的机器人进行搜索和训练,这将使他们获得不公平的优势;而这种不公平的优势也会促使新入局者在试图缩小竞争差距时采取规避手段。
Now, AI can be anything. Today, AI can be in anything. Google search has changed from being sorted by AI to being a full answer engine that answers your question directly on the results page. And Google is not unique in this position — this is the direction in which “search” is moving.
现在,AI 可以是任何事物。如今,AI 可以存在于任何地方。谷歌搜索已经从由 AI 进行排序,转变为一个直接在结果页面回答问题的完整答案引擎。谷歌并非个例——这就是“搜索”正在演进的方向。
We could debate the cutoff for what qualifies as “AI” today, just to find that the standard changes tomorrow. So, instead of defining a bot primarily as “AI” or not, our updated approach to classification will ask deeper questions about bot or agent behavior: What are they doing on my site? What are they storing? And how will they reshare my content?
我们可以争论今天什么是“AI”的界限,但明天标准可能就会改变。因此,我们不再仅仅将机器人定义为“AI”与否,而是采取了一种更新的分类方法,通过深入探究机器人或代理的行为来提出问题:它们在我的网站上做什么?它们在存储什么?它们将如何重新分享我的内容?
A pragmatic taxonomy. To address these questions, we need a more nuanced view — a pragmatic taxonomy that aligns with the AI use cases our customers care about. So we are opening the discussion beyond AI training alone and focusing on three AI use cases that we want all customers to be able to manage:
务实的分类法。为了解决这些问题,我们需要一种更细致的视角——一种与客户关心的 AI 用例相一致的务实分类法。因此,我们将讨论范围从单纯的 AI 训练扩展开来,重点关注我们希望所有客户都能管理的三个 AI 用例:
- Search: any behavior that collects or indexes your content, so it can answer questions about it later. The key is that Search is proactively building a database of your site to later respond to queries with. Site owners should expect to get referral traffic or other equitable compensation as a result. 搜索 (Search): 任何收集或索引你内容的行为,以便日后回答相关问题。关键在于,搜索是主动建立你的网站数据库,以便后续响应查询。网站所有者理应因此获得引流或其它公平的补偿。
- Agent: automated behavior that is acting, usually in real time, on a person’s behalf, to get something done right now. This includes chat fetch bots (e.g., ChatGPT-User) and browser-use agents (e.g., Gemini or Claude driving Chrome). The key is that it visits your web application in order to complete a job, and often there’s a human waiting on the other end. 代理 (Agent): 通常代表个人实时执行任务的自动化行为,旨在立即完成某事。这包括聊天抓取机器人(如 ChatGPT-User)和浏览器使用代理(如 Gemini 或 Claude 操作 Chrome)。关键在于,它访问你的 Web 应用程序是为了完成一项工作,且通常另一端有用户在等待。
- Training: a crawler taking your content to train or fine-tune a model. The key is that your data is permanently absorbed into the underlying architecture of the AI to improve its capabilities. 训练 (Training): 爬虫获取你的内容以训练或微调模型。关键在于,你的数据被永久吸收到 AI 的底层架构中,以提升其能力。
Many popular crawlers on the web fall into one of the classifications above; some fall into multiple. We classify plenty of other behaviors beyond the three above — including ads verification, feed fetching, and agentic transactions. But we believe it should be simple for all website owners to manage access for these three AI-centered use cases. We believe that bot operators should separate their crawlers because that creates more transparency for website owners: allowing them to better understand why a given crawler is visiting them, as well as to better manage the access they extend to that crawler. If a company runs automation that builds Search indexes, acts as an Agent, and collects data to Train their models, then we strongly encourage that company to separate the automation into three separate crawlers.
网络上许多流行的爬虫都属于上述分类之一;有些则属于多个。除了上述三种,我们还对许多其它行为进行了分类——包括广告验证、订阅源抓取和代理交易。但我们认为,所有网站所有者都应能简单地管理这三种以 AI 为中心的用例的访问权限。我们认为机器人运营商应该将其爬虫分开,因为这能为网站所有者创造更高的透明度:让他们更好地理解特定爬虫访问的原因,并更好地管理他们授予该爬虫的访问权限。如果一家公司运行的自动化程序既构建搜索索引、充当代理,又收集数据来训练模型,那么我们强烈建议该公司将这些自动化程序拆分为三个独立的爬虫。
We want a classification system that is scalable and representative of the world of automated traffic as it evolves. Tracking a bot’s purposes is nothing new, but our new taxonomy involves a few updates that better represent the state of bot traffic today. Most notably, we want to recognize that bots that have multiple purposes should be tracked with all purposes, not just one of them.
我们希望建立一个可扩展的分类系统,能够代表不断演变的自动化流量世界。追踪机器人的目的并非新鲜事,但我们的新分类法包含了一些更新,能更好地反映当今机器人流量的状态。最值得注意的是,我们希望明确:具有多种用途的机器人应该被追踪其所有用途,而不仅仅是其中之一。
New options to manage AI traffic. We want to provide more options for managing different kinds of AI traffic, to all website owners on the Cloudflare network. The managed preset to “Block AI bots” that we’ve announced in the past included single-purpose bots that crawled data for model training. But not all AI use is the same, and we want our customers to have the controls they need. So, we’re launching the ability to manage AI traffic based on three major use cases: Search, Agent, and Training crawlers. With these new options, our customers can more finely tune how they manage AI bot traffic — including customers on our Free tier.
管理 AI 流量的新选项。我们希望为 Cloudflare 网络上的所有网站所有者提供更多管理不同类型 AI 流量的选项。我们过去宣布的“屏蔽 AI 机器人”托管预设,涵盖了用于模型训练的单一用途爬虫。但并非所有的 AI 用途都相同,我们希望客户拥有他们所需的控制权。因此,我们推出了基于三大主要用例(搜索、代理和训练爬虫)来管理 AI 流量的功能。通过这些新选项,我们的客户可以更精细地调整他们管理 AI 机器人流量的方式——包括使用免费套餐的客户。
Setting new defaults. On September 15, 2026, we’ll be setting new defaults for each of these three classifications. For all new domains onboarding to Cloudflare, the categories of Training and Agent will be blocked by default on the pages that display ads, while Search will remain allowed by default. An ad is a signal that a website owner meant for a person to land there and see it — something monetizable that fuels the business. So, on those pages, we treat human attention as the end goal, and keep away the bots that may prevent this attention (i.e., Training and Agent bots).
设置新默认值。2026 年 9 月 15 日,我们将为这三类分类设置新的默认值。对于所有新接入 Cloudflare 的域名,在展示广告的页面上,“训练”和“代理”类别将被默认屏蔽,而“搜索”将保持默认允许。广告是一个信号,表明网站所有者希望用户访问并观看——这是推动业务发展的可变现内容。因此,在这些页面上,我们将人类的注意力视为最终目标,并屏蔽那些可能阻碍这种注意力的机器人(即训练和代理机器人)。