“Google and Reddit do not own the Internet," web scraper says after court win
“Google and Reddit do not own the Internet,” web scraper says after court win
“谷歌和 Reddit 并不拥有互联网,”网络爬虫公司在胜诉后表示
After a big court loss last week, Google has confirmed that it won’t give up its fight to block AI bots from scraping its search results. And Reddit is weirdly along for the ride. 在上周遭遇重大法庭失利后,谷歌已确认不会放弃阻止 AI 机器人抓取其搜索结果的斗争。而 Reddit 也奇怪地卷入了这场纷争。
Curiously invoking the Digital Millennium Copyright Act (DMCA), Google sued SerpApi last December. The search giant accused the web scraper of circumventing its anti-scraping technology and then selling content scraped from Google search results through an unauthorized “Google Search API” software service. 去年 12 月,谷歌援引《数字千年版权法案》(DMCA)起诉了 SerpApi。这家搜索巨头指控该网络爬虫公司规避其反爬虫技术,并通过未经授权的“谷歌搜索 API”软件服务出售从谷歌搜索结果中抓取的内容。
According to Google, the anti-scraping tech was in place to protect copyrighted content in search results. Allegedly, SerpApi’s circumvention threatened to disrupt Google’s relationships with rights holders, including some who license content to Google to appear in so-called “knowledge panels” that are displayed in some search results for well-known people or entities. 据谷歌称,其反爬虫技术旨在保护搜索结果中的版权内容。据称,SerpApi 的规避行为威胁到了谷歌与版权持有者之间的关系,其中包括一些将内容授权给谷歌,以便在针对知名人物或实体的搜索结果中显示“知识面板”的权利人。
It was an odd use of the DMCA, since Google search results can’t be copyrighted. But Google was apparently emboldened to explore the legal theory after Reddit filed a very similar lawsuit in October, accusing SerpApi and Google-rival Perplexity of scraping Reddit content that appears in Google results. 这是一种对 DMCA 的奇怪用法,因为谷歌搜索结果本身无法获得版权保护。但在 Reddit 于 10 月提起类似诉讼,指控 SerpApi 和谷歌竞争对手 Perplexity 抓取出现在谷歌结果中的 Reddit 内容后,谷歌显然受到了鼓舞,开始探索这一法律理论。
In a blog, Google cited Reddit’s lawsuit when announcing its own challenge, which it said it filed as a “last resort” to block “malicious scraping” that violates rights holders’ choices over who can access their content. Specifically, Google alleged that SerpApi’s circumvention violated its terms and made it impossible to profit from—or offset the cost of—“billions” of bot searches. 谷歌在一篇博客中提到了 Reddit 的诉讼,并宣布了其自身的挑战。谷歌称这是作为阻止“恶意抓取”的“最后手段”,因为这种抓取行为侵犯了权利人对其内容访问权限的选择权。具体而言,谷歌指控 SerpApi 的规避行为违反了其服务条款,使其无法从“数十亿次”机器人搜索中获利或抵消相关成本。
And before it, Reddit claimed that SerpApi was evading two levels of security: Reddit’s own controls blocking scraping on its platform and Google controls blocking scraping of Reddit content in search results. 在此之前,Reddit 也声称 SerpApi 规避了两层安全防护:一是 Reddit 自身在其平台上阻止抓取的控制措施,二是谷歌阻止抓取搜索结果中 Reddit 内容的控制措施。
Meredith Rose, a senior policy counsel with expertise in the DMCA for a nonprofit public interest group called Public Knowledge, told Ars that Google and Reddit seem to be “sort of grasping at whatever tool is available” in the face of the sudden, continuous rise of AI scraping over the past three years. And while the way they’re using the DMCA is “bizarre”—and “not what the law had sort of contemplated as a use case”—she says it’s not “surprising.” 非营利性公共利益组织 Public Knowledge 的高级政策顾问、DMCA 专家 Meredith Rose 对 Ars 表示,面对过去三年 AI 抓取技术的突然且持续的兴起,谷歌和 Reddit 似乎在“试图抓住任何可用的工具”。虽然他们使用 DMCA 的方式很“离奇”,且“并非法律所预期的用例”,但她认为这并不“令人惊讶”。
Historically, the DMCA has been an effective tool to quickly stop disfavored content uses and force discussions around licensing, so turning to it may have been an obvious starting point, given Google’s goals. But Google’s and Reddit’s unusual DMCA arguments don’t seem to be winning ones. 从历史上看,DMCA 一直是快速阻止不受欢迎的内容使用并推动许可谈判的有效工具,因此考虑到谷歌的目标,转向该法案似乎是一个显而易见的起点。但谷歌和 Reddit 这种不同寻常的 DMCA 论点似乎并不奏效。
Last week, a court took the somewhat rare step of granting SerpApi’s motion to dismiss very early on in Google’s lawsuit. In that case, the judge found that Google had no DMCA standing to sue SerpApi, since it didn’t own any of the content in the search results and has not shown that it’s acting on behalf of any rights holders. 上周,法院采取了较为罕见的举措,在谷歌诉讼案的早期阶段就批准了 SerpApi 的驳回动议。在该案中,法官认定谷歌没有提起 DMCA 诉讼的法律地位,因为它并不拥有搜索结果中的任何内容,也未证明其是代表任何权利人行事。
“That does not happen terribly often,” Rose told Ars. “It really boiled down to Google didn’t allege enough about what it was protecting that was copyrighted.” “这种情况并不常见,”Rose 对 Ars 说,“归根结底,谷歌未能充分说明它所保护的内容中哪些是受版权保护的。”
Likely the timing of that decision wasn’t great for Reddit, which faced a hearing on SerpApi’s motion to dismiss its lawsuit last Thursday. It’s unclear which way the court will rule in that case, but Rose told Ars that the Google ruling doesn’t bode well for Reddit since Reddit can’t claim that it is the content owner or exclusive licensee of content in search results. 这一裁决的时机对 Reddit 来说可能并不理想,因为 Reddit 上周四刚刚就 SerpApi 驳回其诉讼的动议进行了听证。目前尚不清楚法院将如何裁决该案,但 Rose 对 Ars 表示,谷歌案的裁决对 Reddit 来说不是好兆头,因为 Reddit 无法声称自己是搜索结果中内容的拥有者或独家被许可人。
“The judge in the Google case said, ‘Well, in order to have standing to bring a lawsuit under the DMCA, you can be the copyright owner or the exclusive licensee or the person who is deploying and manufacturing the technological protection measure at issue,’” Rose told Ars. “Reddit is none of those things.” “谷歌案的法官说,‘好吧,为了具备根据 DMCA 提起诉讼的资格,你必须是版权所有者、独家被许可人,或者是部署和制造相关技术保护措施的人,’”Rose 对 Ars 说,“而 Reddit 并不符合其中任何一项。”
SerpApi is hoping that the fight will be over soon, telling Ars that the costly legal battle is worth sticking it out to defend the open web. “The bottom line is that both Google and Reddit appear to be engaged in attempts to use the DMCA to wall off the open Internet by retroactively claiming control over content that they didn’t author and don’t own,” SerpApi told Ars. SerpApi 希望这场斗争能尽快结束,并告诉 Ars,这场昂贵的法律战值得坚持下去,以捍卫开放的网络。“底线是,谷歌和 Reddit 似乎都在试图利用 DMCA 来封锁开放的互联网,通过追溯性地声称对他们未创作且不拥有的内容拥有控制权,”SerpApi 对 Ars 表示。
Google’s last chance to keep fight alive
谷歌维持诉讼的最后机会
Although Rose agreed with SerpApi that, in granting the motion to dismiss, the court gave SerpApi a big win, the fight is not over yet, as Google has a narrow path forward to keep its war against web scraping alive. 尽管 Rose 同意 SerpApi 的观点,即法院批准驳回动议是 SerpApi 的一场重大胜利,但斗争尚未结束,因为谷歌仍有一条狭窄的途径来维持其针对网络抓取的战争。
Google acknowledged that search results can’t be copyrighted but argued that “knowledge panels” sometimes include copyrighted content that Google licenses from rights holders. If Google can amend its complaint to argue that rights holders directly authorized Google to use its anti-scraping technology to prevent unauthorized access to content, then Google may be able to block a very limited amount of SerpApi’s scraping. 谷歌承认搜索结果本身无法获得版权,但辩称“知识面板”有时包含谷歌从权利人处获得许可的版权内容。如果谷歌能够修改诉状,主张权利人直接授权谷歌使用其反爬虫技术以防止未经授权的内容访问,那么谷歌或许能够阻止 SerpApi 的极小部分抓取行为。
Google’s spokesperson, José Castañeda, told Ars that Google plans to amend the complaint and is “pleased to see that the Court rejected nearly all of SerpApi’s legal arguments” otherwise attempting to dispute Google’s standing. “We look forward to filing an amended complaint, as the Court invited us to do, and we remain committed to protecting our services and partners from unauthorized access,” Castañeda said. 谷歌发言人 José Castañeda 对 Ars 表示,谷歌计划修改诉状,并“很高兴看到法院驳回了 SerpApi 几乎所有其他试图质疑谷歌诉讼资格的法律论点”。Castañeda 说:“我们期待按照法院的邀请提交修改后的诉状,我们仍然致力于保护我们的服务和合作伙伴免受未经授权的访问。”
However, Rose told Ars that Google has somewhat “talked themselves into a little bit of a corner here, both in this litigation and historically.” For Google, it could be “very dangerous” to argue that the knowledge panel is “chock full of copyrighted material,” Rose suggested. Since the search giant doesn’t license all the content in the knowledge box, Google could risk future lawsuits if the act of algorithmically creating the knowledge box without licenses suddenly becomes viewed as infringement, Rose said. 然而,Rose 对 Ars 表示,谷歌在这次诉讼以及历史上都“把自己逼进了一个死胡同”。Rose 认为,对于谷歌来说,辩称知识面板“充斥着版权材料”可能是“非常危险的”。她说,由于这家搜索巨头并没有获得知识面板中所有内容的许可,如果这种通过算法创建知识面板且未获得许可的行为突然被视为侵权,谷歌可能会面临未来的诉讼风险。