Is it legal to train AI models on copyrighted books? It’s complicated
Is it legal to train AI models on copyrighted books? It’s complicated
使用受版权保护的书籍训练 AI 模型合法吗?情况很复杂
You probably know by now that the AI models powering ChatGPT, Gemini, Claude, and other chatbots are trained on seemingly infinite databases of published works, containing hundreds of millions of books, online articles, academic papers, and basically anything you can find on the internet. 你现在可能已经知道,驱动 ChatGPT、Gemini、Claude 和其他聊天机器人的 AI 模型,是基于看似无限的已出版作品数据库进行训练的,其中包含了数以亿计的书籍、在线文章、学术论文,以及你在互联网上能找到的几乎任何内容。
Most published authors have, without their knowledge or consent, contributed to the development of the same AI tools that threaten to undermine their livelihoods. That seems illegal, right? The reality isn’t that simple. 大多数已出版作品的作者在不知情或未同意的情况下,为这些威胁到他们生计的 AI 工具的开发做出了“贡献”。这看起来似乎是非法的,对吧?但现实并没有那么简单。
“I think one of the issues with this entire area of law and this entire area of technology is there’s a lot going on,” Cathy Gellis, an attorney with expertise in intellectual property, copyright, and technology, told TechCrunch. “It’s very complex and there are a lot of raw feelings about what is happening, both for and against.” “我认为这一法律领域和技术领域的问题之一在于,这里面涉及的事情太多了,”知识产权、版权和技术领域的专家律师 Cathy Gellis 对 TechCrunch 表示,“这非常复杂,而且对于正在发生的事情,无论是支持方还是反对方,都有很多强烈的情绪。”
Last year, in one of the first rulings of its kind, Judge William Alsup ordered Anthropic to pay a mammoth $1.5 billion copyright settlement to a group of writers whose works were used to train the company’s AI models. At face value, this seemed like a moral victory favoring authors, but Judge Alsup actually ruled that Anthropic’s AI training was lawful. What Alsup penalized Anthropic for was pirating these books from illegal online shadow libraries. 去年,在首批此类裁决之一中,法官 William Alsup 命令 Anthropic 向一群作品被用于训练其 AI 模型的作家支付高达 15 亿美元的巨额版权和解金。从表面上看,这似乎是作者们的一次道德胜利,但 Alsup 法官实际上裁定 Anthropic 的 AI 训练行为是合法的。Alsup 惩罚 Anthropic 的原因是该公司从非法的在线影子图书馆中盗版了这些书籍。
“Like any reader aspiring to be a writer, Anthropic’s LLMs trained upon works not to race ahead and replicate or supplant them — but to turn a hard corner and create something different,” the judge wrote, comparing the way an LLM ingests trillions of words to a writer’s study of literature. “就像任何渴望成为作家的读者一样,Anthropic 的大语言模型(LLM)训练所依据的作品,并非为了抢先一步去复制或取代它们,而是为了实现转折并创造出不同的东西,”法官写道,他将大语言模型吸收数万亿单词的方式比作作家对文学的研究。
Gellis thinks the ruling is more advantageous for AI companies. What’s a $1.5 billion fine to a company projecting about $200 billion in annual revenue by 2028? “I think it is generally good news for AI training that he looked at what was going on and really sort of thought it analogous to reading a copyrighted work as opposed to copying a copyrighted work,” Gellis said. “Copyright law hinges on copying, but it doesn’t hinge on using the work or experiencing the work, consuming the work, reading the work.” Gellis 认为该裁决对 AI 公司更有利。对于一家预计到 2028 年年收入将达到约 2000 亿美元的公司来说,15 亿美元的罚款算什么?“我认为,他审视了正在发生的事情,并将其类比为‘阅读’受版权保护的作品,而不是‘复制’受版权保护的作品,这对 AI 训练来说总体上是个好消息,”Gellis 说,“版权法取决于‘复制’,但并不取决于‘使用’、‘体验’、‘消费’或‘阅读’作品。”
Copyright law hasn’t been updated since 1976, which means that judges have to figure out how to interpret guidelines from 50 years ago when confronting legal questions that have the potential to shape the future of the AI industry. 版权法自 1976 年以来一直未更新,这意味着法官在面对可能塑造 AI 行业未来的法律问题时,必须设法解读 50 年前的准则。
“Everybody is very worried right now because the law is all over the place, and it’s because of this question,” Jason Henderson, Senior Attorney and Founder of the IP & Media Practice at JWL International, told TechCrunch. “They know that the AI model has been trained on so much stuff, and the law has not really caught up to that question.” “现在每个人都很担心,因为法律界对此众说纷纭,而原因就在于这个问题,”JWL International 的高级律师兼知识产权与媒体业务创始人 Jason Henderson 对 TechCrunch 表示,“他们知道 AI 模型已经接受了海量数据的训练,而法律还没有真正跟上这个问题。”
These questions often hinge on fair use law — namely, whether use of a copyrighted work is “transformative” enough to be considered legally permissible. Fair use is a carve out of copyright law that allows for the use of copyrighted materials without explicit permission, protecting the ability to comment and iterate on copyrighted works through criticism, parody, education, and other means. Judges consider specific factors when deciding if something is fair use, including the purpose and nature of the work, the amount used, and its impact on the market. 这些问题通常取决于“合理使用”(fair use)法律——即使用受版权保护的作品是否具有足够的“转换性”(transformative),从而被视为法律允许。合理使用是版权法中的一个例外条款,允许在未经明确许可的情况下使用受版权保护的材料,旨在保护通过评论、戏仿、教育和其他方式对受版权作品进行评论和迭代的能力。法官在判定某事是否属于合理使用时会考虑特定因素,包括作品的目的和性质、使用量以及对市场的影响。
“Copyright is always about protecting and growing the market,” Henderson noted. “The courts are kind of all over the place in their reasoning [in AI cases]. What’s tending to win is if what you’re doing is you’re training on somebody’s property because your purpose is to directly compete, then the courts will frown on it… If what you’re doing is not going to compete, then the courts are tending to find ways that it will be okay.” “版权法始终是为了保护和扩大市场,”Henderson 指出,“法院在(AI 案件的)推理上有些混乱。目前的倾向是,如果你训练他人财产的目的是为了直接竞争,那么法院会对此表示反对……如果你所做的事情不会构成竞争,那么法院往往会设法认定其为合理。”
Henderson is referencing a case in which the media and technology company Thomson Reuters sued the research firm Ross Intelligence for copying its content in order to build a competing, AI-based legal platform. “Ross’s use is not transformative because it does not have a ‘further purpose or different character’ than Thomson Reuters’s,” Judge Stephanos Bibas wrote last year. In that case, Judge Bibas decided that it was not fair use to train on Reuters’ content to make a new platform that would directly compete with it. Henderson 提到的案件是媒体和技术公司汤森路透(Thomson Reuters)起诉研究公司 Ross Intelligence,指控其复制内容以构建一个具有竞争力的 AI 法律平台。去年,法官 Stephanos Bibas 写道:“Ross 的使用不具有转换性,因为它与汤森路透的内容相比,没有‘进一步的目的或不同的性质’。”在该案中,Bibas 法官裁定,利用路透社的内容来制作一个直接与其竞争的新平台,不属于合理使用。
While authors could potentially argue that chatbots are competing with them by using their works to generate new, synthetic books, that argument has not yet prevailed in court. 虽然作者们可能会辩称,聊天机器人利用他们的作品生成新的合成书籍是在与他们竞争,但这一论点尚未在法庭上胜出。
When it comes to the relationship between AI and copyright, Gellis finds it helpful to narrow down what we’re actually talking about – the way we think about copyright in terms of AI training is quite different from how we think about copyrighting AI-generated content. In one case, Thaler v. Perlmutter, the court ruled that if a work is 100% AI-generated, it’s not copyrightable, which opens a whole new can of worms – how can we definitively prove whether or not a work was generated using AI, and if so, how do we know what percentage of it was created or assisted with AI? 谈到 AI 与版权之间的关系,Gellis 认为缩小讨论范围很有帮助——我们看待 AI 训练版权的方式,与我们看待 AI 生成内容的版权保护方式截然不同。在 Thaler v. Perlmutter 一案中,法院裁定如果作品是 100% 由 AI 生成的,则不受版权保护,这引发了一系列新的难题:我们如何确切证明一部作品是否由 AI 生成?如果是,我们又如何知道其中有多少比例是由 AI 创建或辅助完成的?
“If you write your novel in [Microsoft] Word and run spell check, we kind of feel comfortable with the idea of saying that Word does not own your novel,” Gellis said. “[AI] is forcing us to look at a whole bunch of decisions that we kind of ignored for a while.” “如果你在 [Microsoft] Word 中写小说并运行拼写检查,我们很自然地认为 Word 并不拥有你的小说,”Gellis 说,“[AI] 正在迫使我们重新审视许多我们此前忽略的决定。”
Most AI companies are still lodged in pending litigation over these issues, which means that we won’t have a definitive solution to these problems any time soon. “What you are seeing is that the initial opening volleys are being influential, and that influence itself could be undone if other courts decide different things, and it’ll take later states of litigation to figure out which one will prevail,” Gellis said. “But in the meantime, all these decisions are shaping everything that’s happening. It would be kind of foolish for the AI companies to ignore them.” 大多数 AI 公司仍深陷于这些问题的未决诉讼中,这意味着我们短期内不会有明确的解决方案。“你所看到的是,最初的几轮交锋正产生影响,但如果其他法院做出不同的裁决,这种影响本身也可能被推翻,我们需要等待后续的诉讼阶段才能确定哪种观点会胜出,”Gellis 说,“但与此同时,所有这些裁决都在塑造着正在发生的一切。如果 AI 公司忽视它们,那将是愚蠢的。”