← Back to articles
Project Panama: How Anthropic (and Others) Buy Physical Books to Destroy Them for AI Training

Project Panama: How Anthropic (and Others) Buy Physical Books to Destroy Them for AI Training

Anthropic paid $1.5 billion to settle a class-action lawsuit over pirated books. Meanwhile, the company buys, cuts, and scans millions of physical books as part of Project Panama—an operation a federal judge deemed 'fair use'. A look into the war on printed text.

By Brice Matter··4 min read

On September 5, 2025, a U.S. federal judge approved a $1.5 billion settlement imposed on Anthropic—the largest compensation ever awarded in a copyright case in the United States. The reason: the company had used pirated copies of hundreds of thousands of books to train Claude. Approximately $3,000 per book was paid to thousands of authors.

What has since been uncovered, amid document leaks and journalistic investigations (Fortune, Futurism, Dallas Express): alongside the pirated ebooks case, Anthropic has been conducting a completely different operation for years—internally named Project Panama—with a chilling stated goal: 'destructively scan all the books in the world'.

Project Panama: The Mechanics

The process is industrial. Anthropic buys books in bulk from booksellers, resellers, and library distributors. Then:

  1. Each book goes through a hydraulic guillotine that cleanly removes the binding from the pages.
  2. The pages are scanned at an industrial pace by professional imagers.
  3. The extracted text is integrated into the training corpora of Claude models.
  4. The physical books are destroyed.

An internal document, cited in several investigations, states: 'We use a flexible code name because we don't want this to get out.'

The scale: according to revelations, we're talking about millions of volumes destroyed—including rare editions, sometimes copies of which only a few remain in the world. A Dutch bookseller told Fortune he thought it was 'spam or a phishing attempt' when he received an order for 3,000 copies of the same book.

The Judge: 'Fair Use'

Federal Judge William Alsup issued a ruling that, from the American legal perspective, changed the game. According to him, digitizing legally purchased books and using the digital copies to train an LLM constitutes fair use—as long as:

  • the digital copy replaces the paper copy (the physical book is destroyed—there is no duplication);
  • no copies are redistributed;
  • the use is transformative (training a statistical model, not making the text available).

This interpretation—challenged but upheld for now in appellate court—radically separates two cases: the pirated ebooks (clear violation, hence the $1.5 billion) and the books bought then destroyed (deemed legal).

Buy a book legally, cut it up, scan it, throw it away: that's fair use. Download it from Library Genesis: that'll cost $3,000 per book.

Anthropic Is Not Alone

The movement is industrial. Documented in the same investigations:

  • Other major AI labs reportedly conduct similar operations (names circulate, but evidence is harder to publicly establish)
  • Specialized service providers are emerging: mass purchase logistics + industrial destruction + scanning + delivery of clean text corpora
  • The second-hand book market sees abnormal price shifts in certain segments (out-of-print technical editions, university textbooks)

The Issues Raised

1. Irreversible Destruction of Heritage

Some destroyed books are rare editions of which only a few physical copies remain in the world. Scanning them does not preserve them in a heritage sense—a scan does not have the same value as an original object for a book historian, a bibliophile, or a material history researcher. Once shredded, these copies are lost forever.

2. Ethical Asymmetry Between Legal Authors and Piracy

The judge accepts as fair use the destruction of a book to train a model. Result: an author whose book was bought once for €20 and shredded is not compensated for the training. Whereas an author whose book was pirated receives $3,000. The moral coherence of this setup can be debated.

3. Lack of Consent from the Book Ecosystem

Publishers, rights holders, illustrators, preface writers, translators—nobody in the book chain was informed or consulted for what constitutes the largest transfer of text content in history to commercial systems. The logic 'I bought it, I can do what I want with it' crushes the contractual copyright ecosystem that usually governs secondary uses (translation, adaptation, excerpt).

4. Economic Concentration

Physically shredding millions of books requires considerable capital. Only a few players (Anthropic, OpenAI, Google, Meta, some Chinese labs) can afford it. This practice locks in a gap between foundational models trained on nearly complete text corpora and smaller players.

5. Impossible Traceability

Once a book is integrated into the training corpus, it becomes impossible for the author, rights holder, or even the regulator to verify that their work is in the model, or to obtain its removal. The GDPR provides a right to erasure—unachievable at this stage for a trained LLM.

6. Flaw in Legal Reasoning

American fair use relies on four criteria, including transformative nature and no economic effect on the original work. Yet Claude and its competitors now write texts that directly compete with the books that trained them (white papers, essays, practical guides). The 'no economic effect' criterion becomes difficult to defend in the medium term.

7. International Gray Area

Judge Alsup's decision applies in the United States. European copyright law (Directive 2019/790 on copyright in the digital single market) provides an explicit opt-out right for text and data mining for commercial purposes. A book bought in Paris, cut up in San Francisco to train a model sold in Berlin—which jurisdiction applies? No one really knows.

Where Are We Headed?

Three possible trajectories emerge:

  1. Prolonged Status Quo: the fair use jurisprudence holds, labs scan everything available by 2028, and regulation will be discussed when everything has been ingested.
  2. Legislative Backlash: U.S. states (led by California) and the EU impose a mandatory license with redistribution to rights holders, modeled after collective management organizations (SACD, SACEM).
  3. Explicit License Market: major publishers (Penguin Random House, Hachette) sign eight-figure deals with labs (as already seen between OpenAI and News Corp, Financial Times, Le Monde). Smaller publishers are left out.

The most likely short-term outcome: a mix of all three, with marked geographical variations. What is certain, however: the shredded books won't return. And the publishing world has yet to digest that its cumulative physical production over five centuries is literally heading to the shredder to feed statistical models.