What happened to TheNumbers.com
What happened to TheNumbers.com
TheNumbers.com 到底发生了什么?
What just happened to TheNumbers.com should worry us all. The inside story of how one of film data’s most trusted sites vanished overnight, and why every website you rely on is more fragile than you think. TheNumbers.com 最近的遭遇应该引起我们所有人的警惕。这是一个关于电影数据领域最受信任的网站之一如何在一夜之间消失的内幕故事,也揭示了为什么你所依赖的每一个网站都比你想象中更加脆弱。
If you work in or around the film industry, there is a decent chance you have used the work of The Numbers this month, whether you realise it or not. Its hand-researched data is the highest quality, tracking box office grosses, budgets, home video and streaming across more than 78,000 films and 236,000 people. It gets north of eight million visitors a year, and is treated as THE definitive authority by journalists, academics, filmmakers, prediction markets, and even Guinness World Records. 如果你在电影行业工作或与之相关,那么无论你是否意识到,你很有可能在本月使用过 The Numbers 的数据。其经过人工核实的数据质量极高,涵盖了超过 7.8 万部电影和 23.6 万名从业者的票房收入、预算、家庭视频及流媒体数据。该网站每年拥有超过 800 万访客,被记者、学者、电影制作人、预测市场甚至吉尼斯世界纪录视为“权威中的权威”。
And it was this GOAT status which caused the catastrophic events of March this year. On the 5th March 2026, TheNumbers.com website vanished. The site was down for over a week, without explanation. A week later, it resurfaced at a fraction of its former size. Gone were the historical charts, the individual movie pages, and even the much-loved Report Builder. 正是这种“史上最佳”(GOAT)的地位,导致了今年三月份那场灾难性的事件。2026 年 3 月 5 日,TheNumbers.com 网站消失了。网站在没有任何解释的情况下宕机了一周多。一周后,它以仅存往日一小部分规模的状态重新上线。历史图表、独立电影页面,甚至备受喜爱的“报告生成器”(Report Builder)都消失了。
With only a generic “we’re rebuilding, please bear with us” message to go on, the internet responded as it always does - with confusion, anger, and conspiracy theories. One Reddit theory even suggested it was a deliberate rug pull designed to cripple the free site to push people towards paid products. 面对网站上仅有的一条“我们正在重建,请耐心等待”的通用信息,互联网的反应一如既往——充满了困惑、愤怒和阴谋论。Reddit 上甚至有一种理论认为,这是一场蓄意的“撤资跑路”(rug pull),旨在瘫痪这个免费网站,从而迫使人们转向付费产品。
Three months on, I spoke at length with Bruce Nash, founder and CEO of The Numbers, about what happened. He describes quite an unpleasant and eventful experience: “We got a lot of angry emails from people who are like, ‘Where’s this page that you used to have and you don’t have anymore?’” 三个月后,我与 The Numbers 的创始人兼首席执行官 Bruce Nash 进行了深入交谈,了解了事情的经过。他描述了一段相当不愉快且充满波折的经历:“我们收到了很多愤怒的邮件,人们问:‘你们以前有的那个页面去哪了?为什么现在找不到了?’”
Within his tale are a number of things that should worry anyone who runs, relies on, or simply appreciates the internet. 在他的讲述中,有许多事情值得每一个运营、依赖或仅仅是欣赏互联网的人感到担忧。
First, some background
首先,一些背景信息
On Friday 17 October 1997, mathematician and former IBM software developer Bruce Nash launched a Geocities site that tracked 300 films. Bruce described the launch in a 20th anniversary essay (which now survives only in the Internet Archive, for reasons that will become clear): “I hit a button in an Access database, uploaded some HTML pages to Geocities, and made a brief announcement on the Hollywood Stock Exchange message boards to let people know that I was starting to analyze box office for films to help them pick MovieStocks to trade on HSX.” 1997 年 10 月 17 日星期五,数学家兼前 IBM 软件开发人员 Bruce Nash 在 Geocities 上推出了一个追踪 300 部电影的网站。Bruce 在一篇 20 周年纪念文章中描述了这次发布(由于原因显而易见,该文章目前仅存于互联网档案中):“我按下了 Access 数据库中的一个按钮,将一些 HTML 页面上传到 Geocities,并在好莱坞证券交易所(HSX)的留言板上发布了一个简短的公告,告诉大家我开始分析电影票房,以帮助他们选择在 HSX 上交易的电影股票。”
From those humble beginnings, Bruce and the team he built around the site turned The Numbers into the film industry’s most reliable financial source. At the start of 2026, the database tracked 78,396 movies, 178,375 theatrical release records, and 236,176 people. 从这些卑微的起点开始,Bruce 和他围绕该网站组建的团队将 The Numbers 打造成了电影行业最可靠的财务来源。截至 2026 年初,该数据库追踪了 78,396 部电影、178,375 条院线发行记录以及 236,176 名从业者。
The robots arrive
机器人大军来袭
During its lifetime, the challenges The Numbers has faced have changed immensely. For its first quarter century or so, the traffic was manageable and mostly polite. As Bruce puts it: “Pre-AI, we got human traffic, mostly well-behaved search engine crawlers, and a few people crawling the site for personal projects. If someone got too greedy, we could spot them and block them.” 在网站运营期间,The Numbers 面临的挑战发生了巨大变化。在最初的四分之一个世纪里,流量是可控的,而且大多是“礼貌”的。正如 Bruce 所言:“在 AI 时代之前,我们面对的是人类流量、大多表现良好的搜索引擎爬虫,以及少数为了个人项目抓取网站的人。如果有人抓取过于贪婪,我们很容易就能发现并封锁他们。”
Over the past couple of years, website owners the world over have seen their web traffic change. What was initially only people browsing gave way to an ever-increasing number of bots. By 2024, automated traffic had surpassed human traffic, and just last month, Cloudflare announced that bots had reached 57.5% of web page requests. 在过去几年里,全球的网站所有者都目睹了网络流量的变化。最初仅有的人类浏览行为,逐渐被数量不断增加的机器人所取代。到 2024 年,自动化流量已经超过了人类流量;就在上个月,Cloudflare 宣布机器人流量已占到网页请求的 57.5%。
The Numbers felt this shift in two distinct waves. The first started around 2024: “We saw a big increase in crawls as AI training joined the search engine crawlers. The AI crawlers are generally less well-behaved than the search engines, which increased the management tasks for us to keep the site running smoothly.” The Numbers 在两波明显的浪潮中感受到了这种转变。第一波始于 2024 年左右:“随着 AI 训练加入搜索引擎爬虫的行列,我们发现抓取行为大幅增加。AI 爬虫通常比搜索引擎爬虫更‘不守规矩’,这增加了我们维持网站平稳运行的管理负担。”
And the second wave was stronger and more damaging: “Around December 2025, we saw another big spike in traffic which I attribute to agentic AI: a combination of AI agents that scrape sites in response to prompts, and people being able to write agents that scrape sites.” 第二波浪潮则更猛烈、更具破坏性:“大约在 2025 年 12 月,我们看到了流量的又一次激增,我将其归因于‘代理 AI’(agentic AI):即响应提示词抓取网站的 AI 代理,以及人们能够编写用于抓取网站的代理程序的结合体。”
Like every data-rich site, by early 2026 The Numbers was being hammered hard by AI bots scraping its pages over and over at an industrial scale. Bruce says that only 10% of their traffic is from humans browsing the site, with the rest coming from AI bots and automated traffic. 像每一个数据丰富的网站一样,到 2026 年初,The Numbers 遭到了 AI 机器人的猛烈攻击,它们以工业级的规模反复抓取页面。Bruce 表示,只有 10% 的流量来自人类浏览,其余部分均来自 AI 机器人和自动化流量。
Websites try to adapt to the new robots
网站试图适应新机器人
This put enormous strain on the site, but Bruce and his team were able to take measures to mitigate the worst of it. One of the cleverest was talking to the robots in their own language: “There’s stuff on the site which is designed for an LLM to read, so that it can tell somebody ‘here’s how you licence the data’ rather than ‘here’s how you scrape the website’. It’s had a huge effect. We’re now getting probably ten times the volume of licensing enquiries.” 这给网站带来了巨大的压力,但 Bruce 和他的团队采取了措施来缓解最严重的影响。其中最聪明的方法之一是用机器人自己的语言与它们对话:“网站上有一些专门为大语言模型(LLM)设计的内容,这样它就能告诉对方‘这是你如何授权获取数据的方法’,而不是‘这是你如何抓取网站的方法’。这产生了巨大的效果。我们现在收到的授权咨询量大约是以前的十倍。”
But mitigation is not the same as escape. From December through early March, the team struggled to keep the site alive under the load. Bruce estimates that: “Around 90% of our time was spent keeping the existing site running while we spent our spare moments working on a new and improved system.” 但缓解并不等于逃脱。从 12 月到 3 月初,团队一直在努力维持网站在重压下的运行。Bruce 估计:“我们大约 90% 的时间都花在维持现有网站的运行上,而只能利用空闲时间开发一套新的、改进后的系统。”
The problem was compounded by the site’s age: thirty years old, with approximately 160,000 source files serving around 2 million pages. Then, in the early hours of Thursday 5 March, the servers collapsed. 网站的“高龄”加剧了这一问题:它已经 30 岁了,拥有大约 16 万个源文件,支撑着约 200 万个页面。随后,在 3 月 5 日星期四的凌晨,服务器崩溃了。
The team scrambled to understand what had happened, initially assuming it was the sheer weight of AI traffic. It seems AI was to blame… but possibly not only in the way they first thought. 团队匆忙排查原因,最初认为是 AI 流量过大导致的。看起来确实是 AI 的错……但可能不仅仅是以他们最初认为的那种方式。
Buried in the flood of agentic AI traffic, the site’s logs showed something more pointed than scraping. As Bruce describes it: “Some of these used the site using legitimate URLs, others were looking for back doors, most likely so they could get to the data before it appeared on the site, or to manipulate the data presented to users.” 在海量的代理 AI 流量中,网站日志显示出了一些比单纯抓取更具针对性的行为。正如 Bruce 所描述的:“其中一些使用了合法的 URL 访问网站,另一些则在寻找后门,很可能是为了在数据出现在网站上之前就获取它,或者为了篡改呈现给用户的数据。”
On the advice of a friend who works in cybersecurity, the old server stayed off. For good. Restoring the backups and nursing the thirty-year-old site back online would have meant defending 160,000 legacy files against attackers who had spent months probing them. 在一位从事网络安全的朋友建议下,旧服务器被永久关闭了。恢复备份并让这个 30 年历史的网站重新上线,意味着要保护 16 万个遗留文件,抵御那些已经花了数月时间进行探测的攻击者。