Booksellers suspect AI firms are buying and then destroying rare books
Booksellers suspect AI firms are buying and then destroying rare books
书商怀疑人工智能公司正在购买并销毁珍稀书籍
If you can truly appreciate an old book—and maybe even marvel at how its fragile, yellowing pages contain some of the earliest ways that people tried to make sense of the world around them—then headlines about tech companies that are destroying books to train AI likely torture a tender part of your soul. 如果你能真正欣赏一本古籍,甚至惊叹于它那脆弱、泛黄的书页中蕴含着人类早期试图理解世界的智慧,那么关于科技公司为训练人工智能而销毁书籍的头条新闻,很可能会刺痛你内心最柔软的地方。
It’s indeed depressing to imagine piles of book spines waiting to be fed into wood chippers while torn-out pages are cropped, scanned, and trashed. But that’s the cheapest and easiest way to scan books as fast as possible, and AI companies are in a race to advance their models by training on the kind of engaging, high-quality long-form texts that can only be found in books. 想象一下,成堆的书脊等待着被送入碎木机,而撕下的书页被裁剪、扫描后丢弃,这确实令人沮丧。但这却是以最快速度扫描书籍最廉价、最简单的方法。人工智能公司正竞相通过训练那些只能在书籍中找到的、引人入胜的高质量长篇文本来提升其模型能力。
So book lovers fear it’s likely that the practice is happening on a grander scale than is currently being reported and that some physical copies of books will be lost forever. What makes this destruction extra painful, though, is that it doesn’t have to be this way. 因此,爱书人士担心这种做法的规模可能比目前报道的要大得多,一些书籍的实体版本可能会永远消失。然而,这种破坏行为之所以格外令人痛苦,是因为它本不必如此。
Google patented a non-destructive book-scanning technology in 2009 that AI firms could use to efficiently scan books—if they were willing to slow down and invest in the process. It’s not perfect, however; studies have found that the curve of the page can distort text, and pages can be missed. As Wired reported, glitches can happen when workers move too quickly, including disembodied hands obscuring pages. 谷歌在2009年申请了一项非破坏性书籍扫描技术的专利,如果人工智能公司愿意放慢速度并投入资金,他们完全可以使用这项技术高效地扫描书籍。不过,这项技术并不完美;研究发现,书页的弧度可能会导致文字变形,且可能漏扫页面。据《连线》杂志报道,当工作人员操作过快时,可能会出现故障,包括手部遮挡页面等问题。
Overall, the trade-offs in cost and speed may not appeal to AI firms looking for the cheapest way to scan millions of titles, and Google’s method may not be the best way to handle rare books anyway. The Internet Archive, which helps libraries preserve aging collections, has long understood that scanning old texts takes time and attention to limit handling, and that’s why it’s considered such a human job. 总的来说,对于那些寻求以最低成本扫描数百万册书籍的人工智能公司而言,这种在成本和速度上的权衡可能并不具备吸引力,而且谷歌的方法也未必是处理珍稀书籍的最佳方式。长期致力于帮助图书馆保存老化馆藏的“互联网档案”(Internet Archive)早就明白,扫描古籍需要时间和专注力以减少对书籍的损耗,这就是为什么这被认为是一项需要人工完成的工作。
For AI companies bent on finding shortcuts, the Internet Archive’s method may feel almost alien. But if you’re a book lover looking for a balm while parsing rumors of AI-driven book destruction, an old post from 2021 describes how the Internet Archive scans books to avoid damaging even the most precious rare books. 对于一心寻找捷径的人工智能公司来说,互联网档案的方法可能显得格格不入。但如果你是一位在解析人工智能销毁书籍的传闻时寻求慰藉的爱书人,那么2021年的一篇旧文描述了互联网档案是如何在扫描书籍时避免损坏哪怕是最珍贵的古籍的。
In it, a book scanner who has been with the Archive since 2010, Eliza Zhang, offered a moment of zen by explaining what she likes so much about scanning books “the hard way.” The Internet Archive did try going the automated route, the post said, even testing out “commercial book scanners that feature a vacuum-powered page-turning arm.” But “it turns out those automated scanners didn’t really work well for brittle books, rare volumes, and other special collections—the kinds of material our library partners ask us to digitize,” the Internet Archive said. 文中,自2010年起就在该档案库工作的书籍扫描员Eliza Zhang分享了她为何如此钟情于“笨办法”扫描书籍,这让人感到一种宁静。文章提到,互联网档案确实尝试过自动化路线,甚至测试过“带有真空动力翻页臂的商业书籍扫描仪”。但互联网档案表示:“事实证明,这些自动化扫描仪并不适合处理脆弱的书籍、珍本和其他特殊馆藏——而这些正是我们的图书馆合作伙伴要求我们数字化的材料。”
“The job requires keen concentration,” Zhang said, since the pages of “very old, fragile books” are “paper thin.” In the post, Andrea Mills, who helps lead the Archive’s book-scanning operations, explained that “clean, dry human hands are the best way to turn pages.” “这项工作需要高度集中注意力,”Zhang说,因为“非常古老、脆弱的书籍”的书页“薄如蝉翼”。在文章中,负责领导档案库书籍扫描业务的Andrea Mills解释说,“干净、干燥的人手是翻页的最佳方式。”
To ensure each rare book only has to go through the scanning process once, Zhang takes her time. She carefully raises the scanner glass with a foot pedal each time she turns a page, then adjusts the cameras and ensures the page is readable. She also takes note of any fold-outs, setting a reminder to go back and scan the inserts so that bonus materials aren’t lost while documenting the main pages of the work. 为了确保每本珍稀书籍只需经过一次扫描过程,Zhang总是从容不迫。每次翻页时,她都会用脚踏板小心地抬起扫描仪玻璃,然后调整摄像头并确保页面清晰可读。她还会留意任何折页,并设置提醒以便回头扫描这些插页,从而确保在记录书籍正文时不会遗漏这些附加材料。
The whole time, she understands that if a page is skipped or an image is too blurry, Internet Archive’s proprietary software will stop the process and prompt her to scan it again. Practice makes perfect, though, and she reports a low error rate after more than a decade of finding a rhythm in the job. 在此过程中,她深知如果漏掉一页或图像过于模糊,互联网档案的专有软件会停止进程并提示她重新扫描。熟能生巧,在经过十多年摸索出工作节奏后,她报告的错误率非常低。
At the time, Zhang had scanned “more than 3 million pages, 14,000 foldouts, and 18,000 items (mostly books),” the post said, with the goal of guaranteeing “zero errors.” Ars asked the Internet Archive for comment, but Chris Freeland, the director of library services, said the post detailing Zhang’s work is “still the best description of our scanning process today.” 文章称,当时Zhang已经扫描了“超过300万页、1.4万个折页和1.8万件物品(主要是书籍)”,目标是确保“零错误”。Ars向互联网档案寻求置评,但图书馆服务总监Chris Freeland表示,那篇详细介绍Zhang工作的文章“至今仍是对我们扫描流程最好的描述”。
The post came after a video of Zhang’s book scanning got 1.5 million views on what was then Twitter, accompanied by a caption that would resonate with book lovers appalled by AI-driven book destruction today. “At the Internet Archive, this is how we digitize a book,” the tweet said. “We never destroy a book by cutting off its binding. Instead, we digitize it the hard way—one page at a time.” 这篇文章发布于一段Zhang扫描书籍的视频在当时的Twitter上获得150万次观看之后,视频配文在今天那些对人工智能销毁书籍感到震惊的爱书人中引起了共鸣。推文写道:“在互联网档案,我们是这样数字化书籍的。我们从不通过切断装订来毁坏书籍。相反,我们用最笨的方法——一页一页地进行数字化。”
Rare booksellers flag suspicious bulk orders
珍本书商警示可疑的大宗订单
Ever since a lawsuit in summer 2025 outed Anthropic for destroying millions of print books to train its AI models, book lovers have moved to defend some of the most precious collections from what feels like AI firms’ endless quest to feed all the books in the world into their large language models. 自2025年夏季的一起诉讼揭露了Anthropic为训练其人工智能模型而销毁数百万册印刷书籍以来,爱书人士纷纷行动起来,保护一些最珍贵的馆藏,以抵御人工智能公司似乎永无止境的企图——将世界上所有的书都喂给它们的大语言模型。
The biggest fear for people who want to see books preserved through the training process is that AI firms will callously pulp rare books that can never be replaced. As The Atlantic reported last week, social media “raged” after two recent reports indicated that AI was already endangering rare books. 对于那些希望在训练过程中保护书籍的人来说,最大的恐惧是人工智能公司会冷酷地将那些永远无法替代的珍稀书籍化为纸浆。正如《大西洋月刊》上周报道的那样,在最近两份报告指出人工智能已经危及珍稀书籍后,社交媒体上引发了“愤怒”。
First, a Telegraph report accused Silicon Valley of destroying millions of rare books and “shredding the originals,” then 404 Media reported that a book-database company called ISBNdb was advertising that it could help AI firms source books in bulk. This backlash was expected, with ISBNdb reportedly warning its clients that the optics were bad. 首先,《每日电讯报》的一篇报道指责硅谷销毁了数百万册珍稀书籍并“粉碎了原件”,随后404 Media报道称,一家名为ISBNdb的书籍数据库公司正在打广告,声称可以帮助人工智能公司批量采购书籍。这种反弹在预料之中,据报道,ISBNdb甚至警告其客户,这种做法的舆论影响很差。
Anthropic started using a codename—“Project Panama”—for its destructive book-scanning in an effort to keep it hidden from the public. There is no indication that Anthropic ever destroyed rare books, and the company has denied doing so in statements. It’s also true that some AI firms are helping preserve rare books. OpenAI and Microsoft, for example, are working with Harvard librarians on an initiative to train AI models on about 1 million public-domain books dating back to… Anthropic开始为其破坏性的书籍扫描项目使用代号——“巴拿马计划”(Project Panama),试图将其对公众隐瞒。目前没有迹象表明Anthropic曾销毁过珍稀书籍,该公司也在声明中否认了这一点。同样真实的是,一些人工智能公司正在帮助保护珍稀书籍。例如,OpenAI和微软正在与哈佛大学的图书馆员合作,计划利用约100万本公有领域的书籍来训练人工智能模型,这些书籍的历史可以追溯到……