OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web

OpenAI and Microsoft knew they were starting a ‘doom loop’ for the web

OpenAI 和微软早就知道他们正在为互联网开启一个“厄运循环”

Recently unsealed court documents in the New York Times’ case against OpenAI and Microsoft are pretty damning. The companies’ own documentation warned that it was starting a “doom loop” that would damage the web, characterized its scraping of data to train its models as the “largest theft of labor in human history,” and that it made a “complete mockery of the idea of fair use.”

在《纽约时报》起诉 OpenAI 和微软的案件中,最近解封的法庭文件内容相当令人震惊。这些公司自己的内部文档曾警告称,他们正在开启一个会损害互联网的“厄运循环”(doom loop),并将抓取数据来训练模型的行为描述为“人类历史上最大规模的劳动窃取”,并称其“彻底嘲弄了合理使用(fair use)的概念”。

Many of the most eye-catching quotes from the document come from Microsoft’s Director of Applied Science, Brent Hecht. Though, the company has tried to distance itself from Hecht’s assertions. Microsoft spokesperson Alex Haurek told The Verge that “These comments reflect one employee’s individual perspective, are not a legal analysis, and do not represent the company’s views.”

文件中许多最引人注目的引语来自微软应用科学总监 Brent Hecht。不过,微软公司已试图与 Hecht 的言论撇清关系。微软发言人 Alex Haurek 对《The Verge》表示:“这些评论仅反映了一名员工的个人观点,并非法律分析,也不代表公司的立场。”

In a separate court filing, Jordan Usdan, GM for Data Strategy and Ops at Microsoft AI, characterized Hecht’s role as adversarial. He said that Hecht “holds divergent, academic, and forward-looking views about how data ecosystems for AI should operate and is employed at Microsoft to bring asymmetrical, futuristic, and academic points of view … nor is he someone who speaks for Microsoft specifically as to his theoretical views on AI’s potential effect on content creators.”

在另一份法庭文件中,微软 AI 数据战略与运营总经理 Jordan Usdan 将 Hecht 的角色描述为“对抗性”的。他表示,Hecht “对于 AI 数据生态系统应如何运作持有不同、学术且前瞻性的观点,他在微软受雇是为了提供非对称的、未来主义的和学术性的视角……他也不是那种能代表微软就 AI 对内容创作者潜在影响发表理论观点的人。”

But whether or not Microsoft wants to own these comments, it’s clear that this came true. Google Zero is real! AI is eating the web!

但无论微软是否愿意承认这些言论,事实已经显而易见:Google Zero(零点击搜索)已经成为现实!AI 正在吞噬互联网!

There are plenty more wild statements in NYT’s filing from a variety of figures, including Satya Nadella, Sam Altman, and other OpenAI employees. Here are some highlights from the 92 page document.

在《纽约时报》提交的文件中,还有许多来自萨提亚·纳德拉(Satya Nadella)、萨姆·奥特曼(Sam Altman)以及其他 OpenAI 员工等各界人士的惊人言论。以下是这份 92 页文档中的一些重点内容。

“An astonishing theft” / “令人震惊的窃取”

The introduction quotes Hecht and OpenAI’s Head of ChatGPT (presumably Nick Turley) in a way that seems to show the companies knew they posed an “existential threat” to publishers like the New York Times. Hecht calls ChatGPT and Copilot’s harvesting of data the “largest theft of labor in human history” and says that Microsoft’s defense makes a “complete mockery of the idea of ‘fair use.’”

引言中引用了 Hecht 和 OpenAI ChatGPT 负责人(推测为 Nick Turley)的话,似乎表明这些公司早就知道他们对《纽约时报》等出版商构成了“生存威胁”。Hecht 将 ChatGPT 和 Copilot 的数据采集称为“人类历史上最大规模的劳动窃取”,并表示微软的辩护“彻底嘲弄了‘合理使用’的概念”。

It’s a “doom loop” / 这是一个“厄运循环”

Satya Nadella admits that chatbots have basically replaced search and removed the need to go straight to the source for info. But perhaps more damning is an internal Microsoft document that says, “Our AI content strategy has started a ‘doom loop’ that will hurt the performance of our models and the entire web at the same time: It is highly unusual that an end-product threatens the economic foundations of its essential suppliers, but that is the situation we have created for our LLM business with respect to its ‘content supply chain.’”

萨提亚·纳德拉承认,聊天机器人基本上已经取代了搜索,消除了直接访问信息源的需求。但更具破坏性的是一份微软内部文档,其中写道:“我们的 AI 内容策略已经开启了一个‘厄运循环’,它将同时损害我们模型的性能和整个互联网:终端产品威胁到其核心供应商的经济基础是非常罕见的,但这正是我们为大语言模型业务在‘内容供应链’方面所造成的局面。”

That’s not even a real number / 那甚至不是一个真实的数字

Don’t be fooled by OpenAI or Microsoft’s claims of altruistic intent. OpenAI cofounder Greg Brockman is more interested in the “gazillions” of dollars it he could potentially make through commercial AI.

不要被 OpenAI 或微软关于利他意图的声明所迷惑。OpenAI 联合创始人 Greg Brockman 对通过商业 AI 可能赚取的“天文数字”般的金钱更感兴趣。

Paywall shmaywall / 付费墙算什么

Despite Nadella later being quoted as saying, “anything that is paywalled should be licensed,” An OpenAI representative admitted that he was “unaware” of any effort to detect or remove paywalled content from training data.

尽管纳德拉后来曾表示“任何有付费墙的内容都应该获得授权”,但一名 OpenAI 代表承认,他“不知道”有任何旨在从训练数据中检测或删除付费墙内容的努力。

“Insanely good at regurgitation” / “极其擅长复读”

Internally, it seems that OpenAI was well aware of ChatGPT’s tendency to simply reproduce copyrighted material “verbatim.” Even though it acknowledged that the “prevention of memorization” was important to “minimize copyright violations,” employees admitted that GPT-4 “memorized a ton of data and therefore will be insanely good at regurgitation.”

在内部,OpenAI 似乎非常清楚 ChatGPT 倾向于简单地“逐字”复制受版权保护的材料。尽管该公司承认“防止记忆”对于“最大限度地减少版权侵权”很重要,但员工们承认 GPT-4 “记忆了大量数据,因此在复读方面会极其出色。”

The filing then goes on to cite several examples of ChatGPT outputting long strings of copy straight from articles in the Times, Mercury News, The Denver Post, LifeHacker, and Eurogamer in response to queries.

随后,该文件列举了几个例子,显示 ChatGPT 在回答查询时,直接输出了来自《纽约时报》、《水星报》、《丹佛邮报》、LifeHacker 和 Eurogamer 文章中的长段文字。

“‘Hoovering up’ all their work” / “‘吸尘器式’地搜刮他们所有的劳动成果”

Microsoft knew how its wholesale scraping of the internet would be perceived and admitted that “almost no one intended for they [sic] content they created to be used in this fashion, nor are they compensated for its use.”

微软知道其对互联网的大规模抓取会被如何看待,并承认“几乎没有人打算让他们创造的内容以这种方式被使用,他们也没有因其使用而获得补偿。”

A “substitute for the labor of people” / “人类劳动的替代品”

OpenAI Policy Director Jack Clark saw the writing on the wall, saying that it was “creating systems that substitute for the labor of the people that define the ‘culture’ of society.” Internal documents described ChatGPT as “the modern newsstand.” OpenAI’s Nick Turley is later quoted as saying that once you get an answer from its chatbot, there is “no good reason to click” on a link to the source.

OpenAI 政策总监 Jack Clark 看到了大势所趋,他表示,公司正在“创建能够替代那些定义社会‘文化’的人类劳动的系统”。内部文档将 ChatGPT 描述为“现代报刊亭”。OpenAI 的 Nick Turley 后来被引用称,一旦你从聊天机器人那里得到了答案,就“没有充分的理由去点击”指向来源的链接了。

Destroying their own supply chain / 摧毁自己的供应链

Microsoft is quoted as admitting that “LLMs are a product that destroys its own supply chain” because it’s a substitute for its own training data in many cases.

微软被引用承认,“大语言模型是一种摧毁其自身供应链的产品”,因为在许多情况下,它本身就是其训练数据的替代品。

OpenAI knows its killing referral traffic / OpenAI 知道它正在扼杀引流流量

OpenAI’s own media and economic experts attributed the drop in referral traffic for sites like the Times directly to AI summaries like Google’s AI Overviews. They’ve speculated that search referrals may be down as much as 60 percent.

OpenAI 自己的媒体和经济专家将《纽约时报》等网站引流流量的下降直接归因于谷歌 AI 概览(AI Overviews)等 AI 摘要。他们推测,搜索引流可能下降了多达 60%。

Microsoft spokesperson Haurek cautioned that “Satya’s testimony and Microsoft’s position in this case are perfectly consistent. He spoke to broad principles and changes underway in how people find and consume information. Those observations should not be confused with conclusions about copyright questions before the Court, which Microsoft addresses in its filings.”

微软发言人 Haurek 提醒道:“萨提亚的证词与微软在本案中的立场完全一致。他谈论的是广泛的原则以及人们寻找和消费信息方式正在发生的变化。这些观察不应与法庭审理的版权问题结论相混淆,微软已在提交的文件中对此进行了说明。”

But it seems pretty clear based on this newly unsealed document that both Microsoft and OpenAI knew they were going to irreparably harm the publishing industry, the “millions of people” it employs, and, by extension, damage their own product, but carried forward anyway in pursuit of “gazillions” of dollars — doom loop be damned.

但根据这份新解封的文件,显而易见的是,微软和 OpenAI 都知道他们将对出版业及其雇佣的“数百万人”造成不可挽回的伤害,并进而损害他们自己的产品,但为了追求“天文数字”般的金钱,他们还是执意推进——哪怕厄运循环在所不惜。