AI's 'Book Burning': Why Companies Are Destroying Millions of Physical Books to Train Their Models

AI firms bought millions of print books, sliced off the bindings, scanned the pages, and shredded the originals — and a court ruled it was fair use. Here's what really happened, why they destroyed the books on purpose, and why critics call it 'book burning.'

By Ubedulla · 6 min read
An open antique book with its pages dissolving into streams of glowing digital data, illustrating AI companies destroying physical books to train models.
AI firms bought and destructively scanned millions of print books to build training data. Illustration.

Some of the most valuable AI models on earth were built, in part, on a quiet act of destruction: companies bought millions of physical books, sliced the bindings off, scanned the pages, and threw the originals away. This week that fact went viral all over again — reframed by critics as "AI book burning" — as people confronted the strange reality that teaching a machine to read has meant physically destroying the books it learned from.

The story is real, but it is more complicated — and more legally deliberate — than the outrage suggests. Here is what actually happened, why the books were destroyed on purpose, and why a court decided it was perfectly legal.

What actually happened

The clearest documented case is Anthropic, the company behind Claude. To assemble a massive, high-quality library of text to train its models, Anthropic didn't just scrape the web. It bought millions of physical print books and ran them through a process the industry politely calls "destructive scanning": the spines are cut off, the loose pages are fed through a scanner, the text is digitized, and the physical book — now a pile of shredded paper — is discarded.

This wasn't a rogue side project. Court filings revealed Anthropic hired an executive who had previously helped run Google's book-scanning program specifically to acquire books at enormous scale. The goal was blunt: get as many books as possible into the training set.

Here is the counterintuitive part. Anthropic didn't shred books out of carelessness — it did so because destroying them was the safer legal path.

Earlier, AI companies (Anthropic included) had trained on enormous troves of pirated books pulled from shadow libraries. That is straightforward copyright infringement. Buying a physical copy, however, changes the calculus: when you own a book, digitizing it for your own use — and crucially, not keeping both a physical and a digital duplicate — looks far more like legally permitted "format shifting" than piracy. Cutting up the book you paid for is, paradoxically, the clean version.

A federal judge agreed. In a landmark 2025 ruling in the Bartz v. Anthropic case, Judge William Alsup found that training Claude on legally purchased, destructively scanned books was fair use — a major win for the AI industry. But he drew a hard line at the pirated books: using those was infringement. That exposure was enormous, and Anthropic ultimately agreed to pay roughly $1.5 billion to settle the piracy claims — one of the largest copyright settlements on record.

So the odd lesson AI companies took away was this: if you want books in your model without legal risk, buy them and destroy them. The destruction is a feature of the legal strategy, not a bug.

Why people are calling it "book burning"

The phrase is emotionally charged — and technically inaccurate; nothing is being set on fire. But it captures a real unease. Millions of books were turned into shredded paper so their contents could live inside a chatbot. For most bulk-bought used paperbacks, that's arguably no great loss — the same text exists in countless other copies.

The sharper worry, and the reason this resurfaced now, is what happens at the edges: whether rare, antique, or out-of-print editions — books where each surviving copy actually matters — could get swept into industrial-scale buying and shredding. Once a scarce edition is cut up and tossed, that physical object is gone for good, even if its words survive as data. Critics argue that treating all books as interchangeable "training fuel" misunderstands what a book is. Defenders counter that the companies bought ordinary copies in bulk, that the text is preserved, and that a court blessed the practice.

Even rival AI leaders have used the moment to score points. Elon Musk publicly signaled that his own AI efforts wouldn't be shredding rare books to train models — a pointed jab at Anthropic that kept the story burning across social media.

The bigger picture

Strip away the "book burning" framing and you're left with something that says a lot about the AI era: the demand for training data is now so intense, and the copyright rules so unsettled, that destroying physical books became the legally cautious option. It's a vivid example of how AI's hunger for human-made content keeps colliding with the systems — copyright law, libraries, publishing — built to preserve that content in the first place. The books-into-data pipeline is unlikely to stop; the fight now is over where its limits should be.

Frequently asked questions

Did AI companies really destroy millions of books?

Yes. Court records in the Bartz v. Anthropic case established that Anthropic bought millions of physical books and destructively scanned them — cutting the bindings, scanning the pages, and discarding the physical copies — to build a training dataset for Claude. The practice is real and documented.

Is it actually "book burning"?

Not literally — no books are set on fire. "Book burning" is an emotionally charged label critics use for the mass destruction of physical books. The books are cut apart and shredded, and their text is preserved digitally in the training data.

Why would a company destroy books it paid for?

For legal safety. Buying and digitizing your own physical copy — without keeping a duplicate — is far more defensible under copyright law than using pirated files. A judge ruled this destructive scanning of purchased books was fair use, while using pirated books was infringement. Destroying the originals was part of staying on the legal side of that line.

Were rare or irreplaceable books destroyed?

The documented practice centered on buying ordinary used books in bulk, where the same text exists in many copies. The current concern — and the reason the story went viral again — is whether rare, antique, or out-of-print editions could be caught up in industrial-scale buying and shredding, since those physical objects can't be replaced once destroyed. That specific worry is the contested part of the debate, not a settled fact.

According to a 2025 U.S. federal court ruling, training on legally purchased, destructively scanned books is fair use — so that part is legal. Using pirated books was not, and Anthropic agreed to pay about $1.5 billion to settle those claims. The law here is still evolving.

Reporting based on the public record of the Bartz v. Anthropic copyright case and Judge William Alsup's 2025 fair-use ruling, plus coverage by Tom's Hardware, Futurism, Decrypt and others. The Bot Post will update this story as the copyright landscape develops.

About the author

Ubedulla

Founder & Editor

Founder and editor of The Bot Post, covering AI news and technology.

Related Articles