AI's Book-Burning Problem: Rare Texts Destroyed for Training Data
TL;DR: AI companies are secretly buying, scanning, and destroying millions of physical books—including rare and out-of-print works—to train their models on pre-2022 text. This destructive practice, exposed by 404 Media and Anna's Archive, risks permanently locking human knowledge in private corporate servers. Shadow libraries are now racing to scan and preserve these books before they vanish.
The Hidden Cost of AI's Hunger for Clean Data
In the race to build ever-smarter language models, AI companies have discovered a dirty secret: the cleanest training data is found in physical books printed before 2022, untouched by the AI-generated noise that now floods the internet. But securing that data comes at a staggering price—the destruction of the books themselves.
Recent investigations have revealed that major AI companies, including Amazon and Anthropic, are purchasing massive quantities of secondhand and rare books, scanning them, and then destroying the physical copies. The practice, dubbed "destructive scanning," is turning the dream of universal knowledge access into a corporate monopoly on human culture.
Anthropic's Project Panama: A $1.5 Billion Bet
The first domino fell in January 2026 when The Washington Post broke the story of Anthropic's "Project Panama." The company had spent tens of millions of dollars purchasing millions of paper books, scanning them to train its Claude LLM, and then destroying them all. The scale of the operation was staggering, and it was only the beginning.
In a $1.5 billion copyright settlement, Anthropic admitted to scanning pirated titles as part of this project. While the company's actions were deemed legally permissible—a federal judge ruled that training AI on legally purchased books can qualify as fair use—the ethical implications are profound. As Anna's Archive volunteer "u" wrote in a guest post, "It's outrageous that it's legally permissible, but ethically, it's an extremely serious crime against humanity."
Amazon's VGT3 Warehouse: The Book Destroyer
Amazon's involvement came to light through a clever piece of investigative journalism by 404 Media. They hid an Apple AirTag in a rare book they suspected would be acquired by an AI company and traced its journey across the country. The final destination: Amazon's VGT3 warehouse in Las Vegas, Nevada.
Workers at this facility confirmed that their sole job is to receive massive shipments of printed books, cut the bindings off to feed pages through scanners faster, and discard the destroyed remains. The operation is so systematic that workers scan ISBNs and barcodes before digitizing pages, leading rare book dealers to suspect AI companies are working through lists of ISBNs in an effort to scan every printed book in existence.
Amazon confirmed it "purchases books through commercial channels" and said the operation hadn't been reported before. The company didn't directly address AI training, but the evidence is clear: this is a large-scale, industrial book destruction operation.
Why Destroy Books? The Economics of Knowledge Hoarding
The logic behind destructive scanning is cold and calculated. There are three primary reasons AI companies choose to destroy physical books:
- Competitive advantage: Destroying books prevents competitors from scanning and using them for training. Once scanned, the digital copies become proprietary.
- Legal protection: By owning the physical copies, companies can argue their scanning is transformative use, shielding them from copyright claims.
- Cost efficiency: Lossless scanning of rare books is expensive. Cutting spines and high-speed scanning is cheaper, even when factoring in the cost of the books themselves.
This creates a troubling paradox: AI companies promise to "make human knowledge accessible" while dismantling the most solid carriers of that knowledge. The public may gain smarter AI assistants, but at the cost of vast amounts of knowledge resources disappearing from the public domain forever.
The Rare Book Problem: Irreplaceable Loss
While the destruction of common paperbacks is concerning, the loss of rare and out-of-print books is catastrophic. These works are often unavailable anywhere on the internet, making them uniquely valuable for AI training. But they are also irreplaceable pieces of cultural heritage.
Rare book dealers have expressed alarm that irreplaceable works could vanish in the pursuit of training data. A single copy of a 19th-century botanical illustration or a limited-run poetry collection could be destroyed in minutes, never to be seen again. The booksellers' fear is compounded by the theory that AI companies are systematically working through ISBN lists, which would mean no book is safe.
Model Collapse: The AI Industry's Own Footgun
There's a bitter irony in this destructive scramble for clean data. AI researchers have identified a phenomenon called model collapse, where feeding AI-generated text back into a language model during training causes a gradual deterioration in output quality. This is why pre-2022 physical books are so valuable—they are guaranteed to be human-written.
But by destroying these books, AI companies are ensuring that future generations of models will have no clean data left to train on. They are sacrificing the long-term health of the AI ecosystem for short-term competitive advantage. The very thing they're trying to avoid—model collapse—may be accelerated by their own actions.
Shadow Libraries: The Digital Alexandria
In response to this crisis, shadow libraries like Anna's Archive are mobilizing. As the largest truly open library in human history, Anna's Archive sees itself as building "a digital library of Alexandria, an inextinguishable light of humanity." The organization is calling on volunteers worldwide to scan and upload books before they're destroyed.
"If every person scans a book, and there are 10 million volunteers worldwide, we can obtain 10 million pieces of invaluable wealth," wrote volunteer "u." The organization offers recognition and lifetime membership for small scans, and can help pay for scanning fees for large-scale uploads.
The urgency is real. Since the beginning of 2025, AI-generated content has accounted for more than half of newly published internet content. If much of the future content consists of AI-generated books and papers, will humans be able to distinguish them? "Once AI has absorbed even the last sentence written by humans on paper, all that will remain on the internet will be AI's own words," warns "u."
The Legal Landscape: Fair Use or Piracy?
The legal battle over destructive scanning is ongoing and complex. A federal judge ruled that training AI on legally purchased books can qualify as fair use, a partial win for Anthropic that Meta and OpenAI also claimed. However, pirated copies are a different matter entirely.
Anthropic agreed to a $1.5 billion settlement over scanned pirated titles, and Salesforce now faces a class action over alleged book piracy. The distinction between legally purchased and pirated books is becoming a key battleground in courtrooms.
But even when companies purchase books legally, the destruction of physical copies raises questions that current copyright law doesn't adequately address. As one legal expert noted, the law is playing catch-up with technology, and the consequences are being borne by our cultural heritage.
The Race Against Time
This is a race against time, and the finish line is the complete destruction of the world's physical book collection. Anna's Archive's goal is ambitious: to scan and upload all the world's publications before publishers completely block knowledge and before AI companies destroy everything.
"This is a race against time," writes "u." "Our ideal is to scan and upload all the world's publications before publishers completely block knowledge, and before AI companies scan and destroy all the world's books and papers."
The question is whether humanity will act in time. Every book destroyed by an AI company is a piece of our collective memory lost forever. Every book scanned and uploaded to a shadow library is a piece saved. The choice is ours to make, and the clock is ticking.
Related News

AI Boosts Homework Scores but Hurts Exam Performance, Study Finds

On-Device Piano Autocomplete: 125M Model Hits 108 Notes/Sec

Don't Paste the AI: The Case for Human Answers in an Automated World

GrapheneOS confirms 2027 launch on Motorola flagships

From Joke Domain to Geopolitical Tool: The SondeHub Story

