Microsoft says virtually nobody was grabbing NYT articles through its chatbot
Microsoft says virtually nobody was grabbing NYT articles through its chatbot
微软表示:几乎没有人通过其聊天机器人获取《纽约时报》的文章
Microsoft’s Copilot rarely reproduces even full sentences from news articles and books, let alone substantive chunks that could substitute for the original, the company says in new legal filings as it fights copyright claims from publishers including The New York Times and book authors. 微软在应对包括《纽约时报》及多位图书作者在内的版权诉讼时,在新的法律文件中表示,其 Copilot 极少复述新闻文章或书籍中的完整句子,更不用说提供能够替代原文的实质性内容片段了。
As part of the lawsuit’s discovery, Microsoft provided 8.2 million Copilot chat logs to an expert hired by news publishers. The logs, it claims, were specifically chosen “because they hit on keywords implicating use of News Plaintiffs’ websites, and therefore the most likely to contain News Plaintiffs’ works.” It says the resulting analysis shows that 59,545 of these contained at least 16 words in common with news content used to ground the AI model. 作为诉讼取证的一部分,微软向新闻出版商聘请的专家提供了 820 万条 Copilot 聊天记录。微软声称,这些记录是经过专门挑选的,“因为它们命中了涉及新闻原告网站的关键词,因此最有可能包含原告的作品。”微软表示,随后的分析显示,其中有 59,545 条记录与用于训练 AI 模型的新闻内容有至少 16 个单词的重合。
An expert for the Center for Investigative Reporting found 51 instances of “substantial overlap” with CIR work in the dataset, Microsoft says. Similarly, an expert in the authors’ suit found that the 8.2 million conversations with Copilot only had 24 responses that contained at least 30 matching words. Only 10 of the 212 books evaluated had any matches, Microsoft claims. 微软称,调查报道中心(CIR)的一位专家在数据集中发现了 51 例与 CIR 作品存在“实质性重叠”的情况。同样,在针对作者的诉讼中,一位专家发现,在 820 万次 Copilot 对话中,只有 24 条回复包含至少 30 个匹配单词。微软声称,在评估的 212 本书中,只有 10 本书出现了匹配内容。
The Times disagreed with Microsoft’s conclusions. “The documents and testimony uncovered during discovery lead to only one conclusion: Microsoft and OpenAI stole from The New York Times to make commercial products that substitute for its journalism, threaten its business, and undermine its industry,” the Times’ lead counsel Ian Crosby said in a statement. “We look forward to Microsoft and OpenAI being held accountable for their theft.” CIR, and Authors Guild did not immediately respond to requests for comment. 《纽约时报》不同意微软的结论。《纽约时报》首席律师伊恩·克罗斯比(Ian Crosby)在一份声明中表示:“取证过程中发现的文件和证词只能得出一个结论:微软和 OpenAI 窃取了《纽约时报》的内容,用以制造替代其新闻报道、威胁其业务并破坏其行业的商业产品。我们期待微软和 OpenAI 为其盗窃行为承担责任。”调查报道中心和作者协会未立即回应置评请求。
Microsoft argues that the numbers bolster its case that using copyrighted content for AI training datasets should be considered fair use. While systems like Copilot rely on using copyrighted material, it says, the resulting systems are used for significantly different purposes than the original. The fact that they sometimes reproduce sections of text, it concludes, “hardly undermines the transformative purpose of LLM training.” 微软辩称,这些数据支持了其观点,即使用受版权保护的内容进行 AI 训练数据集应被视为“合理使用”。微软表示,虽然像 Copilot 这样的系统依赖于受版权保护的材料,但最终生成的系统与原始材料的用途截然不同。它总结称,这些系统有时会复述部分文本的事实,“几乎不会削弱大语言模型训练的变革性目的。”
The filings were made as part of a legal battle brought by publishers and authors against Microsoft and OpenAI, which the plaintiffs claim built products on their works and now compete with them directly, in part by regurgitating copyrighted content. The news publishers’ and books authors’ claims were consolidated under one judge to streamline the process, despite objections from the publishers and authors. Microsoft submitted its filing on Friday as it argues for the judge to issue a summary judgement, which would end the case at an early stage. 这些文件是出版商和作者针对微软和 OpenAI 发起的法律诉讼的一部分。原告声称,这两家公司利用他们的作品构建产品,现在正通过复述受版权保护的内容直接与他们竞争。尽管出版商和作者表示反对,但为了简化流程,新闻出版商和图书作者的诉讼请求已被合并由一位法官审理。微软于周五提交了这份文件,旨在说服法官做出简易判决,从而在早期阶段结束此案。