On June 23, 2025, Judge William Alsup of the Northern District of California issued a partial summary judgment in Bartz v. Anthropic: Anthropic’s use of legally acquired printed books to train its large language models is fair use. It is the first time a US court has squarely endorsed the fair-use argument for AI training — the industry’s first major win in the wave of generative-AI copyright suits.
The same ruling kept its sharpest edge for the other half of the case: Anthropic’s “central library” of millions of books downloaded from pirate sites is not protected by fair use, and damages will be set at a later trial. For the industry, the decision opens the door on training while leaving “where the data came from” sitting on the doorstep.
The core ruling: training is fair use
The case was brought by authors Andrea Bartz, Charles Graeber, and Kirk Wallace Johnson. Judge Alsup found that Anthropic scanning its lawfully purchased print books into a database — ripping bindings, cutting pages, digitizing them — and training Claude on those copies were both fair use. The court called the use transformative, and Alsup went unusually far with his language: LLM training is transformative “spectacularly so,” a phrase Anthropic spokesperson Jennifer Martinez echoed in the company’s statement. Procedurally, this was a partial summary judgment: the court found no genuine dispute of fact on the training question and ruled directly, while the pirated-library claims head to trial. Per The Verge, this is the first judicial endorsement of the argument for the AI industry — and for the wave of parallel cases still pending, it is the first answer with real judicial weight behind it.
The loss: the pirated central library
The win has a hard boundary. The same ruling held that Anthropic’s downloading of millions of books from pirate sites to build its central library was not fair use — and some of those books were never even used for training. Alsup wrote that he doubted “any accused infringer could ever meet its burden” of justifying downloads of books it could have lawfully purchased; buying physical copies afterward does not erase liability, though it may reduce statutory damages. A separate trial will determine what Anthropic owes for the pirated library. For Anthropic, that question has narrowed from “was it infringement” to “how much”; for the industry, the court drew a line between how data is used and how it was obtained — the former can be fair use, the latter cannot.
Alsup’s reasoning
The two most-quoted passages frame the court’s view of copyright’s purpose. Alsup wrote that the Copyright Act “seeks to advance original works of authorship, not to protect authors against competition,” and likened LLM training to schoolchildren learning to write by reading — no one infringes by absorbing style from what they read. In other words, the court placed LLM training in the same tradition as human learning rather than treating it as a technology requiring new rules, and that framing may matter longer than the fair-use holding itself. The Verge’s analysis flags an important limit: the ruling does not address whether AI outputs infringe copyright, which remains the central question in other litigation.
Implications for the other AI copyright cases
TechCrunch’s assessment: the decision does not bind other courts, but it may set a precedent favoring technology companies in the similar suits pending against OpenAI, Meta, Midjourney, Google, and others. For authors and publishers, it moves the battlefield from “should training be licensed” to “how was the data acquired” — a more concrete question with better evidence trails. For the generative-AI industry broadly, the cost structure of lawfully sourced data just became estimable for the first time.
What it means for developers and data teams
The practical takeaway is direct: manage data provenance and data use as separate workstreams. The transformative-use argument for training now has judicial support, but source legality is the new core risk — acquisition records, license documentation, and dataset lineage may all become key evidence in litigation. It also helps explain why more AI companies are cutting direct licensing deals with content owners rather than betting on favorable findings for every dataset. The upcoming trial over the pirated library deserves tracking from anyone running data licensing or training-corpus programs: its damages calculation could become the first public reference point for pricing the cost of infringing data.
Sources
- TechCrunch: A federal judge sides with Anthropic in lawsuit over training AI on books
- The Verge: Anthropic wins a major fair use victory for AI, but it’s still in trouble for stealing books
AI-assisted summary compiled from the sources above, reviewed by a human before publishing.
