Who Owns the Words That Trained Your AI?
The copyright fight over training data isn’t a sideshow to the AI boom — it’s the bill arriving for it. Here is what the courts have actually decided, and what they pointedly haven’t.
Every large AI model you have used was trained on an enormous pile of human work — books, articles, photographs, code, songs, forum posts — and almost none of the people who made that work were asked, paid, or told. For a few years the industry treated this as a settled non-issue: the data was “publicly available,” the use was “transformative,” and anyone who objected simply didn’t understand how machine learning works. That confidence is now meeting the one institution that doesn’t care how the technology works, only whether it broke the law: a courtroom.
The short answer, as of now, is uncomfortable for everyone. Training on copyrighted material is not automatically illegal — and it is not automatically fine either. The line the courts are drawing runs straight through the middle of how these models were actually built, and it turns on details the marketing never mentions: where the data came from, whether it was paid for, and whether the model competes with the thing it learned from.
This is the explainer we wish existed when the first “AI stole my work” headline landed. No cheerleading, no doom — just what has actually been decided, what is still open, and what it means for you whether you use these tools, make the work that trains them, or both.
The one question under all the lawsuits
Strip away the specifics and nearly every AI copyright case asks the same thing: when a company copied millions of works to train a model, was that copying ‘fair use’?
Fair use is the American doctrine that lets you use copyrighted material without permission in certain circumstances — quoting a book in a review, parodying a song, and so on. Courts weigh four factors, but in the AI cases two do most of the work. The first is how transformative the use is: are you making something genuinely new, or just repackaging the original? The second is market harm: does your use substitute for the original and cost its author sales or licensing income? AI companies lean hard on the first factor — a model is not a library, they argue, it is a new kind of thing that learns statistical patterns. Authors and artists lean on the second: if the model can spit out work that competes with mine, trained on mine, I have been harmed.
Both arguments are stronger than their opponents like to admit. That is exactly why the rulings have been split down the middle rather than handed cleanly to one side.
The Anthropic ruling: the clearest line yet
The case that has told us the most is Bartz v. Anthropic. Three authors sued Anthropic for training Claude on their books. In June 2025 the judge, William Alsup, issued a decision that neither side got to celebrate for long.
On the core question, he sided with Anthropic: training on the books to build Claude was, in his words, “exceedingly transformative,” and therefore fair use. That single sentence is the most AI-friendly ruling the industry has. But he did not stop there. Anthropic hadn’t only trained on books it bought — it had also downloaded more than seven million pirated books from shadow libraries such as Library Genesis to assemble its training corpus. Building a permanent library out of pirated copies, the judge held, was not fair use, and that half of the case was headed for trial, where willful infringement can carry statutory damages of up to $150,000 per work.
Do the arithmetic on seven million books at those numbers and you understand why Anthropic settled. The agreement — roughly $1.5 billion, about $3,000 for each of an estimated 500,000 covered books — was described in the filing as the largest publicly reported copyright recovery in history, and a US judge signed it off in 2026. Anthropic, then valued at $183 billion after a fresh funding round, could absorb it. The precedent is harder to absorb: you may be allowed to train on a book, but you are not allowed to pirate it to do so, and the difference is worth over a billion dollars.
It is tempting to read the settlement as a defeat for AI. It isn’t, quite. The transformative-use finding is a genuine win for the labs, and one lawyer close to the case put it bluntly: at least in northern California, companies now have a defensible legal path to train on copyrighted work — so long as they obtain it legally. The fight is shifting from “can you train on this at all” to “did you pay to get a clean copy first,” which is a very different, and far more manageable, question for a well-funded company. That is also why the honest framing is not “creators won” but “the free-piracy era ended.”
Why Meta walked and the New York Times didn’t
If Anthropic shows where the line is, two other cases show how much still depends on how well the case is argued. In a parallel suit against Meta, brought by authors including Ta-Nehisi Coates and Sarah Silverman, a different judge granted Meta summary judgment — effectively ending it. But he was careful to say he was ruling on this lawsuit, not blessing AI training in general: the authors, he found, simply hadn’t produced the evidence of market harm that a stronger case might. That is not “training is fair use.” It is “these plaintiffs lost,” with an all-but-explicit invitation for better-prepared ones to try again.
Meanwhile the highest-profile case of all — The New York Times against OpenAI and Microsoft — grinds on, years after it was filed, with the paper amending its complaint to sharpen the focus on the infrastructure used to build the models. Newspapers can show something book authors struggle to: an AI answer that reproduces their reporting is a direct substitute for visiting their site. That is the same substitution problem we keep coming back to, the one hollowing out the economics of the open web. A model that answers in the publisher’s own words, trained on the publisher’s own work, is the market-harm argument made flesh.
Europe started from the opposite end
The United States is arguing its way toward a rule case by case. Europe wrote one down first. Under the 2019 Digital Single Market directive, “text and data mining” — which covers training — is permitted, but rightsholders can opt out by reserving their rights in a machine-readable way. Flip the American default and you have the European one: mining is allowed unless you said no, rather than forbidden unless a court says it’s fair.
The EU AI Act then bolts transparency onto that. Providers of general-purpose models have to respect those opt-outs and publish a summary of the data used to train the model. In principle this is exactly what creators have been asking for: a way to say “not my work” in advance, and a paper trail to check. In practice the opt-out only bites going forward, the training-data summaries are high-level rather than itemised, and enforcement is young and untested. It is a better default than “sue and find out.” It is not the same as control.
Whose side is the government on? Nobody’s, reliably
You might expect the referees to bring clarity. Instead, the institution meant to advise on exactly this — the US Copyright Office — spent the period producing a careful, multi-part study on AI and copyright while its own leadership was engulfed in a political fight, with the head of the Office removed and the fallout dragging through the courts and Congress. When the body that is supposed to say what the law should be is itself the subject of a power struggle, “wait for guidance” stops being useful advice. The Office’s analysis leans against treating wholesale commercial training as automatically fair, but analysis is not law, and the law is being made in courtrooms faster than anyone can codify it.
This is the part worth sitting with. There is no settled answer coming from on high any time soon. What we have instead is a scatter of rulings, a record settlement, one continent’s opt-out regime, and a lot of companies quietly signing licensing deals rather than betting the business on a judge agreeing with them next time.
And someone pays for all of it. Licensing agreements, billion-dollar settlements and standing legal teams are not run for free; their cost lands where every cost eventually lands, in the price of the product. The tidy future the labs now sketch — models trained on properly licensed data, creators compensated, everything above board — is also a materially more expensive business to run than the scrape-first version it replaces. That expense does not evaporate because the story got more respectable. It arrives, in the end, on someone’s invoice, which is the quiet way a dispute that looks like it is purely between authors and AI companies loops back to the person paying $20 a month. The bill for how these systems were built is still being totted up, and consumers are on the distribution list whether or not they ever read a word about the litigation.
What this actually means for you
Whether you use these tools, make the work that feeds them, or both, a few practical truths fall out of the mess:
- A licensing market is forming in real time. The settlements and the risk of trial have made “just scrape it” look reckless, so labs are increasingly paying publishers, stock libraries and forums for training rights. That is good for large rightsholders and roughly useless for the individual whose blog was ingested in 2022.
- Opt-outs are forward-looking and often voluntary. You can add “no training” signals now, and some crawlers honour them. Almost nothing you published before the fight began is covered, and honouring the signal is frequently a courtesy rather than an obligation.
- How the data was obtained is the whole ballgame. The Anthropic split shows courts distinguishing sharply between lawfully acquired and pirated material. Expect “we bought a clean copy” to become a standard corporate defence — and expect it to work.
- Image and music makers have a sharper claim. When a model can reproduce a recognisable style or near-copies on demand, the market-harm argument is easier to prove than it is for prose. Those cases are worth watching most closely.
- “Publicly available” was always a slogan, not a licence. The same instinct that treats your conversations as free training data treats your creative work the same way. We wrote about the data side of this in why every AI wants your data; copyright is that appetite meeting an actual legal system.
The honest state of play
Here is the fair summary, opinion clearly flagged as opinion. The industry’s strongest claim — that training is genuinely transformative — has real support in at least one significant ruling, and pretending otherwise is its own kind of hype. But the years of “move fast, scrape everything, argue fair use later” are ending, and they are ending because a court put a price on the difference between learning from a book and stealing it. The models that define this moment were, in large part, built on a foundation their makers are now paying to have legitimised after the fact.
None of this makes the tools useless or their makers villains — the same models are genuinely helpful, and the law is genuinely unsettled. But it does puncture the founding myth that the training data was free for the taking, the same way the confident claim that these systems don’t really copy sits awkwardly beside the fact that they can be made to reproduce their source material when pushed. Who owns the words that trained your AI? Increasingly, the answer is: someone who is finally getting paid — and a great many people who never will. Knowing which group you are in is worth more than any launch-day benchmark.
Frequently asked questions
Is training an AI on copyrighted material illegal?
Not categorically, at least not in the United States. A federal judge has held that training a model on books the company lawfully acquired can be fair use because the use is highly transformative. But the same ruling found that downloading millions of pirated books to build the training set was infringement. So legality currently hinges less on the training itself and more on how the data was obtained and whether it competes with the original work.
What was the Anthropic settlement about?
Authors sued Anthropic for training Claude on their books. The judge split the case: training on legitimately bought copies was fair use, but Anthropic’s use of pirated libraries was not, and that part was headed to trial with potentially ruinous damages. Anthropic settled it for roughly $1.5 billion — about $3,000 for each of an estimated 500,000 books — in what was described as the largest copyright recovery on record.
Can I stop my work being used to train AI?
Going forward, sometimes. Europe’s text-and-data-mining rules let rightsholders opt out, and many sites now publish machine-readable ‘no training’ signals that some crawlers honour. But opt-outs rarely reach models already trained, honouring them is often voluntary, and enforcement is thin. Realistically, anything you published before the fight began was very likely already ingested.
Does the law treat AI images and text differently?
The underlying question — was copying the training data fair use — is the same for text, images, audio and code. But image and music cases add a second problem: models can reproduce a recognisable style or near-copies of specific works, which makes the ‘market harm’ argument easier for plaintiffs to run. Several image and music suits are still working through exactly that point.
Sources
- Anthropic settles with authors in first-of-its-kind AI copyright infringement lawsuit ($1.5bn; Alsup fair-use ruling; Meta summary judgment) — NPR
- Copyright and Artificial Intelligence — the U.S. Copyright Office study on training and fair use — U.S. Copyright Office
- Directive (EU) 2019/790 — the text-and-data-mining exception and opt-out (Articles 3–4) — EUR-Lex
- Regulation (EU) 2024/1689 (the AI Act) — general-purpose model transparency and copyright duties — EUR-Lex