The AI DownsideDocumenting AI's downsides

Privacy

Why Every AI Wants Your Data

The default setting is almost always “yes, train on this,” and the opt-out is almost always somewhere you will not look.

Abstract editorial illustration for “Why Every AI Wants Your Data”.

There is a reason nearly every AI product is free or cheap to start, generous with its limits, and eager for you to paste in your documents, your code, your half-formed ideas and your most specific questions. You are not only the customer. You are, quietly, the supply chain. The thing that makes the next model better is data that sounds like real people solving real problems, and there is no cheaper source of that than the people already typing into the box.

The default tells you everything

You can learn a company's real intentions by reading its defaults, not its mission statement. Across much of the industry, the default for consumer products has been that your inputs may be used to improve the models unless you go and turn it off. The setting exists. The opt-out is real. But it lives a few screens deep, described in calm language, switched to the position that benefits the vendor, waiting for the overwhelming majority of users who will never open that menu.

This is the oldest trick in the data-collection book, and it works because defaults are destiny. Study after study on everything from organ donation to pension enrolment shows that whatever the default is, that is what most people “choose.” Setting the default to “train on my data” is a decision dressed up as a non-decision.

Consent that requires you to go looking for the setting is not really consent. It is a treasure hunt the vendor is betting you will lose.

The data you never typed at all

It is not only the words you paste in. The valuable signal increasingly includes the metadata around them: which answers you accepted and which you regenerated, how you rephrased a request, what you copied out, how long you lingered, which suggestion you clicked. This behavioural exhaust is, for training purposes, gold — it is direct human feedback about what a good answer looks like, produced by millions of people for free, simply by using the product. You cannot opt out of generating it, because generating it is indistinguishable from using the thing at all.

This is why the framing of “we only use your data to improve the product” is so slippery. Improvement, here, means the model learning from your judgement — your corrections, your preferences, your patience running out. You are not merely supplying raw text to be digested; you are, unpaid and mostly unaware, performing the labour of teaching the system what humans want, one thumbs-up and one frustrated re-prompt at a time. The product gets better, the company's core asset appreciates, and your contribution is folded silently into a model that will one day be sold back to you at a higher tier. It is one of the great uncompensated labour arrangements of the age, and its genius is that it does not feel like labour.

The enterprise gets the privacy the consumer does not

Here is the detail that gives the game away. The privacy protection that consumers have to dig for is frequently the default for enterprise customers. Business and API tiers routinely promise that inputs are not used for training, because businesses have lawyers, procurement teams and the leverage to demand it. The individual has none of those, so the individual gets the setting flipped the other way.

The technology is identical. The difference is entirely who is asking and how much they can push back. That two-tier arrangement — data dignity for those who can negotiate it, data extraction for those who cannot — is the part worth being angry about, calmly and on the record.

“We use it to improve the product” is doing a lot of work

The stated justification is always improvement, and improvement is real: models do get better with more human data. But “improve the product” is a phrase broad enough to cover a great deal. Your prompt might be used to train the next model, reviewed by a human rater, retained for a period you were not told, or bundled into a dataset that outlives the account you eventually delete. Each of those is a different thing, and the single reassuring phrase flattens them all into one friendly blur.

Retroactive scraping and the consent that came too late

There is also the matter of the material that was already out there. A great deal of what trained these systems was collected from the open web, forums, social platforms and archives long before anyone imagined it would become model fuel. Some platforms have since changed their terms to permit training on their users' historical posts, presented as an update rather than a request. Consent obtained after the fact, for data you posted under different assumptions, is a strange kind of consent.

What you can actually do

You cannot fix the incentives, but you can decline to be the path of least resistance:

  • Open the data controls on every AI tool you use and turn training off where the option exists. It usually does; it is just not where you would look.
  • Do not paste anything into a consumer chatbot you would mind seeing surface elsewhere — secrets, other people's personal information, anything under NDA.
  • Prefer the tiers and providers that make no-training the default, and reward the ones that treat it as a right rather than a setting.

Deletion is not what you think it is

People comfort themselves with the delete button. It is less reassuring than it looks. Deleting a conversation removes it from your view; it does not necessarily unwind everything the data has already done. If your input was used to train a model, that influence is baked into the model's weights and cannot be surgically extracted by deleting the source — the lesson has been learned and does not un-learn on request. Copies may persist in backups, in logs, in datasets already assembled, or with human reviewers who saw it, for retention periods you were never shown. “Delete” governs the visible record far more than the actual footprint.

This is the quiet asymmetry of donating data you cannot recall. The moment of contribution is frictionless and reversible-looking; the consequences are durable and largely irreversible. It is the opposite of how the interface makes it feel, and the gap is not accidental — a delete button that honestly explained how little it undoes would make people think twice before typing, which is precisely the hesitation the frictionless design is built to avoid.

The regulation the defaults are racing

None of this exists in a vacuum. Data-protection law in several regions already leans toward exactly the principles the industry's defaults flout — that consent should be freely given, specific and informed, and that privacy should be the default setting rather than the buried one. Regulators have begun to take an interest in how training data is gathered and whether “opt-out, if you can find it” clears the bar. The aggressive defaults are, in part, a bet that the value extracted now will outweigh whatever the rules eventually require, and that it is easier to ask forgiveness than permission.

That is a revealing bet to be making. A company confident that its users would gladly consent has no reason to hide the switch; you bury a setting only when you suspect that asking plainly would cost you the answer you want. The most damning evidence that the current arrangement is not in the user's interest is simply that it is arranged to avoid the user's explicit choice. Design the flow so people must actively opt in, and the ones who value the improvement will still say yes — you will just have to earn it rather than assume it.

The trust you extend is not returned in kind

There is a lopsidedness at the heart of the arrangement that deserves naming. You are asked to be transparent — to hand over your questions, your documents, your half-formed thoughts, in some cases your most private worries — to a system whose own workings are almost entirely opaque to you. You do not know exactly what is retained, for how long, who or what reviews it, whether it trained a model, or what the model quietly inferred about you along the way. The flow of candour runs one direction. You are legible to the company; the company is a black box to you.

That asymmetry is the actual product design, not an unfortunate side effect. A relationship where one party must be open and the other may stay closed is a relationship of power, and the defaults, the buried switches and the reassuring-but-vague language all work to preserve it. The remedy is not to stop using the tools — they are useful, and the data genuinely does improve them. It is to withhold the reflexive trust the interface invites, to treat the friendly box as the data-collection instrument it also is, and to extend candour in proportion to the transparency you get back. On current terms, that proportion is low, and behaving accordingly is not paranoia. It is just declining to be the most trusting party in a deal you were not allowed to see the other half of.

The models are useful and the data genuinely makes them better. The complaint is not that companies want data; it is that they arranged the defaults so that most people give it without ever deciding to. Move the switch to “off” by default and let people opt in, and the whole objection evaporates. That they have not is the answer to why every AI wants your data: because asking plainly would mean hearing “no” far more often.

Related grievances

All articles →