AI companies destroy physical books – let's scan rare books before it's too late

AI companies destroy physical books – let’s scan rare books before it’s too late

AI 公司正在销毁实体书——让我们在为时已晚之前扫描珍稀书籍

TL;DR: AI companies are secretly buying, scanning, and destroying millions of physical books to train their models, permanently locking human knowledge inside private corporate servers. Anna’s Archive is urgently calling on volunteers worldwide to scan and upload books before this cultural heritage disappears forever. 简而言之: AI 公司正在秘密购买、扫描并销毁数百万本实体书以训练其模型,将人类知识永久锁定在私有的企业服务器中。Anna’s Archive 紧急呼吁全球志愿者在这些文化遗产永远消失之前,抓紧扫描并上传书籍。

Several AI companies are acquiring large quantities of secondhand books through intermediaries, scanning and destroying them, all to obtain training data “untouched by machines” from before 2022. Anthropic’s “Project Panama” was exposed in a $1.5 billion copyright settlement. In early 2024, they launched this highly confidential project. The company has spent tens of millions of dollars purchasing millions of paper books, scanning them, training its Claude LLM, and then destroying them all. It’s outrageous is that it’s legally permissible, but ethically, it’s an extremely serious crime against humanity. 几家 AI 公司正通过中间商大量收购二手书,在扫描后将其销毁,目的仅仅是为了获取 2022 年之前“未受机器污染”的训练数据。Anthropic 的“巴拿马计划”(Project Panama)在一次 15 亿美元的版权和解案中被曝光。2024 年初,他们启动了这个高度机密的项目。该公司花费数千万美元购买了数百万本纸质书,扫描并训练其 Claude 大语言模型,随后将这些书全部销毁。令人愤慨的是,这种行为在法律上竟然是“允许”的,但在道德层面,这是对人类文明极其严重的犯罪。

So why destroy physical books? Behind it lies the AI race and the interests of capital: It prevents these books from being scanned and used for training by competitors. It avoids legal risks. Destroying books is cheaper than lossless scanning. After AI companies massively scan and destroy physical books, they become the only ones in the world with digital copies. Knowledge is permanently monopolized on private servers. 为什么要销毁实体书?其背后是 AI 竞赛与资本利益:这可以防止竞争对手扫描这些书籍并将其用于训练;这可以规避法律风险;销毁书籍比进行无损扫描更便宜。在 AI 公司大规模扫描并销毁实体书后,他们便成了世界上唯一拥有这些数字副本的机构。知识被永久地垄断在私有服务器中。

This battle for old books reveals a paradox: while promising to “make human knowledge accessible,” AI companies are dismantling the most solid carriers of human knowledge. The public may gain more intelligent AI assistants, but at the cost of a vast amount of knowledge resources disappearing from the public domain. 这场针对旧书的争夺战揭示了一个悖论:AI 公司在承诺“让全人类知识触手可及”的同时,却正在拆毁人类知识最坚实的载体。公众或许得到了更智能的 AI 助手,但代价是海量的知识资源从公共领域中消失。

As the world’s largest shadow library, Anna’s Archive needs a plan to combat the destruction of physical books by AI companies. After all, the emergence of shadow libraries is the greatest miracle of knowledge sharing in the 21st century. Along with other shadow libraries, we’re building a digital library of Alexandria, an inextinguishable light of humanity. 作为全球最大的影子图书馆,Anna’s Archive 需要制定计划来对抗 AI 公司对实体书的销毁。毕竟,影子图书馆的出现是 21 世纪知识共享领域最伟大的奇迹。我们正与其他影子图书馆一道,共同构建一座数字化的亚历山大图书馆,点亮人类永不熄灭的知识之光。

We need the help of volunteers worldwide to scan materials (including books, journal articles, newspapers, magazines, ancient books, rare books, and other materials) from every library and archive around the world and upload them to the shadow library for knowledge preservation, especially those that are easily lost. If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth. For small scans and uploads, we usually award recognition and lifetime membership to Anna’s Archive. For large-scale scans and uploads of books, we can help pay for the scanning fees and other rewards. 我们需要全球志愿者的帮助,从世界各地的图书馆和档案馆扫描资料(包括书籍、期刊文章、报纸、杂志、古籍、珍本及其他材料),并上传至影子图书馆以保存知识,特别是那些容易散佚的文献。如果每个人都能扫描一本书,而全球有 1000 万名志愿者,我们就能获得 1000 万份无价的财富。对于小规模的扫描和上传,我们通常会给予表彰并授予 Anna’s Archive 的终身会员资格。对于大规模的书籍扫描和上传,我们可以协助支付扫描费用并提供其他奖励。

Since the beginning of 2025, AI-generated content has accounted for more than half of newly published internet content. A frightening reality emerges: if much of the future content consists of AI-generated books and papers, will humans be able to distinguish them? Once AI has absorbed even the last sentence written by humans on paper, all that will remain on the internet will be AI’s own words. In such a world, how can human civilization be preserved? Shadow libraries offer the best answer. 自 2025 年初以来,AI 生成的内容已占互联网新增发布内容的半数以上。一个可怕的现实浮现:如果未来大部分内容都是由 AI 生成的书籍和论文,人类还能分辨真伪吗?一旦 AI 吸收了人类在纸上写下的最后一句文字,互联网上将只剩下 AI 自己的话语。在这样的世界里,人类文明该如何保存?影子图书馆提供了最好的答案。

If you want the memory of human civilization to no longer be monopolized, if you want future generations to be able to read all of humanity’s wealth for free, if you don’t want publishers making a fortune while authors receive little, then please help us. Please make any contribution you can, whether it’s scanning and uploading books, purchasing books and papers to scan and upload, or donating. With the efforts of all humanity, the monopoly on knowledge will be broken. Each of us can make history. This is a race against time. Our ideal is to scan and upload all the world’s publications before publishers completely block knowledge, and before AI companies scan and destroy all the world’s books and papers. 如果你不希望人类文明的记忆被垄断,如果你希望后代能够免费阅读全人类的财富,如果你不希望出版商赚得盆满钵满而作者却所得寥寥,那么请帮助我们。请尽你所能做出贡献,无论是扫描并上传书籍、购买书籍和论文进行扫描上传,还是进行捐赠。在全人类的共同努力下,知识的垄断终将被打破。我们每个人都可以创造历史。这是一场与时间的赛跑。我们的理想是:在出版商彻底封锁知识之前,在 AI 公司扫描并销毁全世界的书籍和论文之前,将全球所有的出版物扫描并上传。

  • Anna’s Archive volunteer “u” —— Anna’s Archive 志愿者 “u”