AI Firms Are Buying and Shredding Rare Books to Feed Training Pipelines
AI companies are purchasing physical books in bulk, running them through high-speed scanners that slice off the spines, and destroying the originals in the process. A service called ISBNdb reportedly brokers orders as large as a million volumes while keeping buyers anonymous, even offering NDAs and framing the destruction as “digital preservation.” Books printed before 2022 command a premium because their text predates the flood of AI-generated content, making them cleaner training data. A federal judge has ruled the practice qualifies as fair use, reasoning that destroying the physical copy after scanning keeps only a single copy in existence at any moment — a legal green light that will likely push more companies toward the same approach.
The piece, an opinion thread amplifying reporting from 404 Media, argues the real harm is irreversibility. Web scraping and pirated media can be re-uploaded or reprinted, but rare titles with only a handful of surviving copies cannot be recovered once they are pulped for a dataset. The author points to Anthropic hiring the former head of Google Books partnerships to acquire “all the books in the world” as a sign of how aggressively firms are pursuing physical text.
Taken as commentary rather than straight reporting, the thread’s significance is less in new facts than in framing a live tension: courts are treating destructive scanning as legally clean format-shifting, while critics see an ethical and cultural loss that no ruling addresses. It highlights how the economics of AI training data now reach into physical archives, and how services have emerged specifically to make that acquisition quiet and deniable.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.