Is It Legal for AI to Train on Your Data?
Your posts, code, photos and prose almost certainly helped train someone's model. The law's answer to whether that was allowed is a lawyerly ‘mostly yes’ — with the fight now over how the data was collected, not whether it should have been.
Somewhere in the training data of the model you used this morning, there is a decent chance you are in there. Not you by name, necessarily, but your Reddit comments, the code you pushed to a public repository, the blog you kept in 2019, the photographs you posted before you thought to wonder where they would end up. The models that answer your questions and finish your sentences were built by reading an enormous amount of what people put on the internet, and “what people put on the internet” includes a great deal of what you put on the internet.
The natural question — the one that arrives the moment this stops being abstract — is whether any of that was allowed. Did an AI company need your permission to learn from your work? Did it need to pay you? Was it, in a word, legal? The honest answer in 2026 is that it is mostly legal, in most places, but the law is still being written in courtrooms as we go, and “legal” has quietly come apart from “something you agreed to”. This piece is a map of where the line currently sits, from the point of view of the person whose data it is.
We should be precise about scope first, because two very different questions get muddled together. One is about public, published, copyrightable work — your writing, art, music and code — and whether copyright law lets a company train on it. The other is about private data you type into a product — your chatbot conversations, your work documents — and whether the product’s terms let it learn from you. This article is mostly about the first. The second is governed less by copyright and more by a settings toggle, and we have written before about how those toggles tend to be set; the short version is that the default is often to train on your work unless you find the switch.
What “training on your data” actually means
A model is not a filing cabinet with your blog post in a drawer. Training is the process of adjusting billions of numerical weights so the system gets better at predicting the next token, pixel or frame; your work is one of the examples it learns the patterns from. This is why the companies argue that training is not “copying” in the everyday sense — the finished model does not, in the normal case, contain a retrievable copy of your article. It is also why the critics counter that the whole edifice was nonetheless built on copying your article at least once, to feed it in, and that a machine which can sometimes reproduce a work near-verbatim has not fully “forgotten” it.
Both things are true at once, and that tension is exactly what the courts are trying to resolve. It matters because the legal treatment turns on framing. If training is fundamentally about learning uncopyrightable patterns, it looks like fair use. If it is fundamentally about ingesting and storing millions of protected works to build a commercial product that competes with them, it looks like infringement at industrial scale. The same activity supports both stories, and which one prevails is being decided piece by piece.
The US answer: probably fair use — unless you stole the books
In the United States, the relevant test is fair use, and the clearest signal so far came from the litigation against Anthropic over its use of books. The court found the training itself to be transformative — “spectacularly so”, in the judge’s words — on the reasoning that a model learning from books to produce new text is doing something different from the books’ original purpose. That is about as favourable a statement as the AI industry could have hoped for on the core question.
And yet Anthropic still agreed to pay roughly US$1.5 billion, in a settlement given approval in 2026, at an estimated US$3,000 per work. The reason is the crucial distinction the case drew: training on the books might have been fair use, but obtaining them by downloading pirated copies from shadow libraries was a separate wrong that fair use did not excuse. The lesson is not “training is illegal”. It is “the training may be defensible; the piracy on the way in is what costs you”. How the data was acquired has become its own front in the war.
The picture is not uniform, which is the honest and slightly uncomfortable part. In the Thomson Reuters case against Ross Intelligence, a court found that copying legal headnotes to build a rival research tool was not fair use — though it stressed that Ross’s product was a search tool, not a generative model, which may limit how far that ruling travels. In the authors’ case against Meta, a different judge found the training defensible but pointedly criticised the plaintiffs for making a “half-hearted” argument on the one factor that may matter most: market harm. And the highest-profile fight, The New York Times against OpenAI, is still grinding on. The through-line is that American courts are converging not on a yes-or-no rule but on a fact-specific question: does the model’s output substitute for, and erode the market for, the thing it learned from?
Britain and Europe: an exception with an opt-out most people can’t use
Cross the Atlantic and the legal furniture changes. There is no broad fair-use doctrine; instead there are specific copyright exceptions, and the relevant one is for text and data mining. Under the EU’s copyright directive, mining copyrighted material — including for commercial AI training — is permitted unless the rightsholder has expressly reserved their rights in a machine-readable way. The AI Act layers transparency duties on top, requiring general-purpose model providers to respect those reservations and publish summaries of what they trained on.
Read that carefully and the catch appears. The opt-out is real, but it is built for entities that control a website or a dataset and can attach a machine-readable “do not mine” flag: a stock-photo library, a news publisher, a big platform. It is not built for an individual whose photographs live on someone else’s social network and whose comments are scattered across a dozen forums. You cannot realistically reserve rights you do not technically control, which means the European “opt-out” is, for most ordinary people, a right they have no practical way to exercise.
The UK, meanwhile, gave the world its first substantial courtroom test. In Getty Images v Stability AI, decided by the High Court in November 2025, Getty ended up abandoning its main training-related claims during the trial and the court rejected the remaining secondary-infringement argument, holding that a model’s weights are not themselves “infringing copies” because they do not store the original images. Getty won only a narrow, historic trademark point about watermarks appearing in outputs. It was widely read as a win for the AI developer — but a narrow, technical one that left the biggest questions, including whether the initial training in the UK infringed, deliberately unanswered.
The Copyright Office’s quiet warning
Amid the litigation, the US Copyright Office weighed in with the third part of its report on AI and copyright in May 2025, and it is worth reading because it is more sceptical than the early court wins suggest. Its position, in essence, was that using vast quantities of copyrighted work to build systems that generate content competing with that work — particularly where the material was accessed illegally — goes beyond the established boundaries of fair use. It spent considerable effort dismantling the strongest arguments the industry had been making, and floated the idea of “market dilution”: the harm not of copying one book, but of flooding the market with cheap machine-made substitutes for a whole category of human work.
A report is not a statute and not a court ruling, and its release was politically bruising. But it is a signal that the official copyright authority does not regard the fair-use question as comfortably settled in the industry’s favour, and that the market-harm argument the Meta plaintiffs fumbled is the one that could, argued properly, change the outcome.
Where you actually stand
Strip away the case names and the position of an ordinary person is fairly clear, if not especially comforting. If you have posted something publicly, it has very likely already been used in training, and the current weight of the law does not treat that as wrong in itself. The places the law bites are around the edges — pirated acquisition, near-verbatim regurgitation, provable market harm — and those are fights conducted by well-resourced plaintiffs, not by you as an individual whose comment history became one ten-millionth of a dataset.
Two further realities sharpen the point. First, the terms of service you accept increasingly grant training rights up front, which is why a story about a platform quietly opting you into AI training by default is now a genre rather than an aberration. Consent obtained through a pre-ticked box and a settings page three menus deep is consent in the legal sense and nowhere near it in the meaningful one. Second, the tools that would let you claw data back barely exist. We have written about how hard it is to get your data out of an AI tool once it is in, and the same asymmetry applies here: getting your work into a training set takes a crawler a fraction of a second; getting it out is, for practical purposes, impossible once a model is trained.
It is worth separating this cleanly from the related question of what you can own at the other end of the pipe — whether the outputs a model produces from all that training can themselves be copyrighted. That is a distinct problem with its own emerging answer, which we covered in our piece on whether you can copyright what AI makes. Training is about the inputs; authorship is about the outputs; the law is unsettled at both ends.
What you can, and can’t, do about it
None of this leaves you completely without options, but honesty requires being clear about how limited they are:
- If you control a site, use the machine-readable signals. A
robots.txtdisallow for the named AI crawlers, and the emerging opt-out conventions, are the one place an individual’s wishes are actually legible to the systems doing the scraping. They only work for content on infrastructure you own, and only for crawlers that choose to honour them. - Use the account-level opt-outs where they exist — and check the default. Several providers now offer a “don’t train on my data” toggle for what you type into their products. Assume it is off until you have turned it on, and assume it applies only to future training, not to models already built.
- Read what you grant when you sign up. The most consequential decision about your data is usually the licence you accept when you join a platform, not anything you can do afterwards. That is the moment to pay attention, however tedious.
- Keep your expectations calibrated. For material already public, there is currently no reliable way to compel its removal from an existing model. Treat “anything I’ve made public may have been trained on” as the working assumption, because it is the accurate one.
The law here is not static, and it is not obviously heading in the companies’ favour forever. The market-harm argument is getting sharper, regulators are more sceptical than the first headlines implied, and every settlement redraws the line a little. But the gap the reader should hold onto is the one between legal and consented to. Right now, a great deal of training on your work is probably lawful, and almost none of it was anything you would recognise as a choice you made. Closing that gap — making legality depend on genuine, informed permission rather than on a crawler’s head start — is the actual argument, and it is far from over.
Frequently asked questions
Can AI companies legally train on my public posts, photos and code?
In most cases, and in most places, yes — with caveats. In the United States, courts have so far treated the training itself as likely fair use because the model learns statistical patterns rather than reproducing your work. In the UK and EU, a text-and-data-mining exception permits it unless rights have been reserved in a machine-readable form. That does not mean every use is lawful — pirated acquisition, verbatim regurgitation and market harm are all still contested — but the default lean of the current law is that publicly posted material can be used to train models.
Does it matter how the AI company got the data?
It matters a great deal. The clearest US ruling so far, in the Anthropic case, found the training use itself defensible but treated the downloading and storage of pirated books as a separate wrong — which is what the company’s US$1.5 billion settlement addressed. The lesson emerging from the cases is that courts are more comfortable with training on lawfully obtained data than with training on material a company took from pirate libraries. How the data was collected can be the difference between a defensible use and an expensive one.
Can I opt out of having my data used to train AI?
Sometimes, partially, and rarely in a way that reaches models already trained. In the EU, the opt-out is real but machine-readable and site-level — it suits a stock-photo library or a publisher who controls a website, not an individual whose words are scattered across other people’s platforms. Some AI firms and platforms now offer account toggles or web-form opt-outs, but they are inconsistent, often off by default, and generally apply only to future training. There is currently no universal, person-level ‘do not train on me’ switch.
Is AI training copyright infringement or fair use?
It depends on the jurisdiction and the facts, and it is genuinely unsettled. Several US courts have leaned towards fair use for the training step while leaving room to punish piracy and market harm; one, in the Thomson Reuters case, found copying that was not fair use. UK and EU law frames the same question as an exception-with-opt-out rather than fair use. Anyone who tells you the matter is definitively resolved in either direction is ahead of the courts.
What about my private messages and documents, not just public posts?
That is a different question, governed more by the product’s terms of service and privacy law than by copyright. Whether a chatbot or workplace tool trains on what you type into it usually comes down to a setting and a default — and those defaults have repeatedly been set to ‘on’. The copyright cases above are mostly about public and published material; your private data is protected less by copyright law and more by whether you noticed the toggle.
Sources
- Copyright and Artificial Intelligence (Part 3: Generative AI Training), prepublication report — US Copyright Office
- Getty Images (US) Inc & Ors v Stability AI Ltd [2025] EWHC 2863 (Ch) — Courts and Tribunals Judiciary (England & Wales)
- AI in litigation series: an update on AI copyright cases in 2026 — Norton Rose Fulbright
- Anthropic's landmark $1.5B copyright settlement is approved — TechCrunch
- Stability AI defeats Getty Images copyright claims in first-of-its-kind dispute before the High Court — Bird & Bird
- EU AI Act's opt-out trend may limit data use for training AI models — Greenberg Traurig
