Why AI Firms Are Shredding Millions of Printed Books

Lawsuits and whistleblower reports reveal that AI developers are buying used books in bulk, using destructive scanning to digitize them, and then destroying the physical volumes—raising legal, ethical, and cultural alarms about access to knowledge.

Why AI Firms Are Shredding Millions of Printed Books

4 Minutes

They arrive at warehouses in unmarked trucks. Long pallets. Cardboard boxes stamped with ISBN labels. Quiet work. Faster than preservation. The scene sounds like a logistics operation, but the cargo is literature: millions of printed books bound for industrial scanners that tear them apart page by page.

What you're reading about is not a conspiracy theory. Court filings and whistleblower reports show that several AI companies have adopted destructive scanning techniques to digitize printed texts at scale. Instead of careful, shelf-preserving capture, book spines are removed and pages are fed into high-speed, industrial scanners. The tradeoff is brutally simple: speed and lower cost versus the physical survival of the printed volume.

One revealing thread runs through the filings. Former staff involved with Google Books and later projects like Anthropic's internal efforts—referred to in litigation as Project Panama—have testified that vendors offering both gentle overhead scanning and the faster, destructive service were used. For AI builders, the decision is pragmatic. Destructive scanning costs less and accelerates throughput by orders of magnitude.

Anthropic, for example, is accused of spending millions to extract training data from printed books purchased through resellers such as Better World Books and then destroying those copies after scanning. While a judge recently ruled that using copyrighted works for AI training can fall under "fair use" in some contexts, that ruling did not spare Anthropic from punitive action when private archives were found to contain millions of digitized copies taken without proper author or publisher authorization—an action that triggered fines reportedly reaching $1.5 billion.

Publishers haven't sat idle. Coalitions of publishers have filed suits alleging that other tech giants used copyrighted titles without permission to train large language models. The legal battle centers on both rights and remedies, but a larger cultural question hangs over the courtroom: should private companies be allowed to turn public and private library stock into proprietary data and then shred the originals?

So why are tech companies buying up used books by the pallet? The blunt answer is training data. High-quality, human-written texts remain among the most valuable inputs for advanced language models. The web is already awash with low-quality AI-generated content—derisively called "AI slop"—which pollutes the signal. To build models that produce coherent, accurate prose, developers want authentic sources: well-edited fiction, rigorous non-fiction, and obscure academic monographs not found in scraped web corpora.

Enter markets like ISBNdb, an online database listing more than 111 million titles. Orders placed through such platforms range wildly—anything from a thousand copies to a single transaction of a million. Sellers who once shipped a few dozen books a week now report fivefold increases, sometimes more. Buyers often pay large sums and show little interest in subject or author; if a volume has an ISBN, it is a candidate.

There are practical and ethical dimensions to this shift. On one hand, digitizing rare texts can improve discovery and ensure content survives physical decay. On the other, the wholesale removal of printed works from circulation—followed by their destruction—deprives libraries, used-book markets, and future researchers of access. Are rare or historically significant editions being filtered out before destruction? The records are murky. What is clear is that scanned texts frequently end up in private databases, accessible only to the companies that paid for them.

Imagine a generation that needs to verify a quotation or study an archival annotation, but key volumes no longer exist in physical form and the digitized copies live behind corporate walls. The tradeoffs extend beyond property law. They touch on cultural stewardship, access to knowledge, and who controls the raw materials that power emergent AI.

Destroying books to train tomorrow's intelligence raises a difficult question: is technological progress worth the permanent loss of parts of our paper-based heritage?

Policymakers, librarians, and technologists will have to decide whether to accept private hoarding of digitized literature, mandate public access, or regulate the methods used to capture printed works. The debate is only beginning, but the stacks keep moving, and the machines keep whirring. Who will speak for the books?

Leave a Comment

Comments

No comments yet. Be the first.