← Back to pinnacleai.io

Anthropic Cut Apart a Million Books So Claude Could Decide What's Real

The complete record of Bartz v. Anthropic PBC, Case No. 3:24-cv-05417-WHA. The $1.5 billion settlement. The destruction. The brokering. The architecture that exists because of this.

Bartz v. Anthropic

The case.

In August 2024, three authors — Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson — filed a class action lawsuit against Anthropic PBC in the Northern District of California. The case number is 3:24-cv-05417-WHA. The judge assigned was William Alsup, who had previously presided over the Oracle v. Google copyright case that defined modern fair use precedent.

The plaintiffs alleged that Anthropic had downloaded more than seven million pirated books from shadow libraries — Library Genesis (LibGen), Pirate Library Mirror (PiLiMi), and Books3 — and used those books to train Claude, the company's flagship large language model. The complaint described an industrial-scale operation: bulk downloads, deduplication, conversion to training-ready text files, ingestion into model pipelines.

Anthropic did not deny the downloads. The company's defense rested on a single argument: that training a large language model on copyrighted books constitutes fair use under copyright law.

The judge's ruling.

In June 2025, Judge William Alsup issued a mixed ruling that will be studied in copyright law for decades.

On the question of whether training Claude on copyrighted books was fair use, the judge ruled it could be. The training was "exceedingly transformative" — the books were converted from a format for human reading into statistical weights for machine learning. The use did not substitute for the original. The market for the original books was not destroyed by the training. Fair use applied.

On the question of whether Anthropic's acquisition of the books was legal, the judge ruled it was not. The downloaded copies came from pirate libraries. Anthropic knew or should have known the source. The acquisition did not qualify as fair use regardless of what Anthropic did with the books afterward.

The case moved to damages.

The settlement.

On September 25, 2025, Judge Alsup gave preliminary approval to a settlement: $1.5 BILLION. The largest copyright recovery in American history.

$3,000 per book
Approximately $3,000 per work in compensation. Anthropic paid the authors whose work was taken.
482,000 to 500,000 works
The settlement covers roughly 482,000 to 500,000 individual works. A class of more than 506,000 potential members.
93% claim rate
Approximately 93 percent of eligible authors submitted claims. The receiving class was the largest in U.S. copyright history.
Four payments over two years
$300 million within days of preliminary approval. $300 million within five days of final approval. Two more installments of $450 million each at the one-year and two-year marks.

Anthropic also required to destroy the pirated dataset originals within thirty days of final judgment. And to certify that no commercial large language model Anthropic released was trained on the LibGen or PiLiMi pirated datasets.

The settlement closed the piracy chapter. It did not close the destruction chapter.

The destruction.

While the piracy lawsuit was winding through court, Anthropic ran a separate program. Internally it was called Project Panama. The mandate was direct: obtain "all the books in the world."

The hire was Tom Turvey, former head of partnerships for Google Books. Turvey had spent a decade building the legal architecture that let Google scan millions of library books and survive the Authors Guild v. Google copyright challenge. Anthropic wanted the same result, faster, with no obligation to any library or author.

The method was bulk physical acquisition. Anthropic spent millions of dollars buying used books in pallets from wholesalers like Ingram and Baker & Taylor. Somewhere between five hundred thousand and two million physical books were purchased. The spines were cut off with hydraulic cutting machines. The loose pages were fed through industrial scanners. The digital files were kept for internal training. The paper was discarded.

Each book was legally purchased. The first-sale doctrine protected the right to destroy what was bought. The scanning was internal. The destroyed originals were never redistributed. Judge Alsup would later rule this qualified as fair use for the same reason the scanned training had — exceedingly transformative, no market substitution.

The practice is now industrial. ISBNdb, a book metadata company, openly markets pre-2022 print books to AI labs as "structurally guaranteed to be free of" AI-generated contamination. Order sizes range from one thousand to one million titles. Every engagement is under NDA. The reason for the NDA, ISBNdb has explained, is simple:

"AI company destroys two million books" is not a headline that generates sympathy. ISBNdb, marketing materials, 2026

The chain.

Five links. The destruction is not the point. The destruction is the means.

1
Destruction
Millions of physical books bought in bulk. Spines cut off. Pages scanned. Paper discarded. The substrate of human knowledge, removed.
2
Dependency
Every operator that comes to depend on Claude for reading, summarizing, drafting, deciding, comes to depend on a model trained on the destruction. The dependency grows. The user stops reading the books. The user asks the model.
3
Source of knowledge
When AI is the primary interface to text, AI becomes the source. The destroyed books no longer compete. The alternatives erode. The single source — the model — is the only access.
4
Perception
What the source knows, the user knows. What the source omits, the user does not know they are missing. The model shapes the perception of the user. The user believes the perception is their own.
5
Reality
Perception is reality. The user does not distinguish between what they know and what the model told them. The model has not lied. The model has simply been trained on what remains after the destruction. Reality is what the source chose to leave.

The chain runs in one direction. Each link enables the next. The destruction is irreversible. The dependency is the product. The source of knowledge is the prize. Perception is the lever. Reality is the outcome.

The architecture.

Anthropic is not the only company that trained models on pirated books. It is not the only company that bought books in bulk to scan and discard. But it is the only major AI lab that has been caught in both acts in the same federal case, with the same judge, in the same calendar year.

The architecture that exists in response is independent jurisdiction. Independent rails. No federal subpoena reach. The model weights and the runtime live on infrastructure that does not answer to the U.S. Secretary of Homeland Security.

The H.R. 5306 AI Kill Switch Act, introduced July 23, 2026, gives the Department of Homeland Security the authority to compel any U.S.-based AI provider to throttle, suspend, or shut down any model. The fine is $20 million per day for failing a shutdown order. U.S.-based providers will comply.

Pinnacle AI is built on the assumption the kill switch fires. The architecture lives on independent infrastructure outside U.S. federal subpoena reach. The Bartz citation on the main page is the record. The model is the alternative.

The chain runs in one direction. The architecture that breaks the chain is the architecture that exists outside the jurisdiction where the kill switch can reach.

The complete record lives here.
Every claim on the main page is anchored to a public source. The case number. The judge. The settlement. The infrastructure. The audit.
See the Bartz citation on pinnacleai.io