RC RANDOM CHAOS

Keep the two claims apart

Consumer AI trains on your chats by default. How to tell the real consent problem from unprovable 'secret breakthrough' claims - and what you can control.

· 7 min read
Keep the two claims apart

By default, a free ChatGPT account feeds your typed conversations back into OpenAI’s training pipeline. There is a setting that turns this off - Settings, then Data Controls, then ‘Improve the model for everyone.’ Most people have never opened that menu. That single design choice - collection on by default, refusal available but buried three taps deep - is the real story underneath almost every viral claim that a lab ‘trained on private conversations.’ The breakthrough accusation is a separate argument, and welding the two together is where the analysis usually falls apart.

Two claims are hiding in one headline

‘Another researcher says OpenAI trained on conversations, then claimed a breakthrough’ packs two assertions into one sentence, and they are not equally supported.

The first - that a consumer AI product uses customer conversations as training data - is documented, not leaked. OpenAI’s own help pages and usage policies say so. It needs no whistleblower and no researcher. The second - that a specific capability gain was secretly and mostly the product of ingesting private chats, then repackaged as a novel method - is a causal claim about internal engineering that almost nobody outside the lab can verify, and that the person making the claim usually can’t either.

Keep them apart. The first is a consent-and-policy problem you can act on today. The second is an attribution problem that, without internal documents, stays unproven. Systems thinking means refusing to let an outrage-friendly second claim borrow credibility from a boring, true first one. When the two travel together in a headline, the true part smuggles the unproven part past your skepticism.

What ‘trained on your chats’ actually means

When a lab trains on a conversation, your words are tokenized and used to nudge billions of model weights by tiny amounts. The model doesn’t file your chat in a lookup table it can query later. But nudged weights still memorize - researchers have repeatedly extracted verbatim training strings, including names, phone numbers, and addresses, out of large models by prompting them the right way. ‘It’s just statistics’ is not a privacy guarantee.

The tier you’re on decides your exposure, and the tiers are not the same:

  • Consumer ChatGPT, free and Plus: conversations may be used for training unless you opt out. On by default.
  • API access: OpenAI states inputs and outputs are not used to train its models by default, and are retained roughly 30 days for abuse monitoring.
  • Enterprise, Team, and Edu: contractually excluded from training.

The default also varies by region and product version, which means the honest answer to ‘is it on for me’ is ‘check your own account,’ not ‘trust a general claim.’ The people most exposed are the ones least likely to open a data-controls menu - individuals typing medical symptoms, legal trouble, and work secrets into a free chatbot because it’s faster than a search engine. Exposure tracks inversely with the effort required to reduce it, which is the opposite of how a fair system would work.

Clicking ‘I Agree’ on a terms-of-service wall is not the same as consenting to have your 2 a.m. therapy-adjacent chat become training data. Regulators have already said close to that. Italy’s data-protection authority, the Garante, forced ChatGPT offline in the country for several weeks in spring 2023, citing the absence of a legal basis for processing personal data to train the model. The U.S. Federal Trade Commission opened an inquiry the same year into how the company handles personal information. Under GDPR, ‘we buried an opt-out in a submenu’ is not an obvious lawful basis for secondary use of personal data.

The design pattern is older than AI. Default-on collection with a technically-available opt-out is the same move used for cookie tracking and app telemetry for two decades. It works because it converts inertia into consent. If a small fraction of users find and flip the toggle, the lab keeps the overwhelming majority of the data and still gets to say the choice was there all along. That’s not a UX accident. It’s the intended yield, and it’s why ‘you could have opted out’ is a deflection rather than a defense.

Deleted isn’t deleted

There’s a second-order problem people miss: opting out or deleting a chat going forward does nothing about copies already retained. In 2025, a court in the New York Times’ copyright suit against OpenAI ordered the company to preserve output logs it would otherwise have deleted - including content from deleted and temporary chats - as potential evidence. A litigation hold overrode the delete button for millions of users who are not parties to the case. Retention you were promised can be suspended by a court you’ve never heard of.

And there’s no delete for the model itself. Once your words have shifted the weights, removing the original chat doesn’t un-train the parameters it already influenced. ‘Machine unlearning’ is an open research problem, not a feature you can request. So when you evaluate any ‘we don’t keep your data’ claim, the real question is never the headline policy - it’s every legal and operational exception stacked on top of it, plus the fact that the trained model is downstream of data you can no longer reach.

Why the ‘breakthrough’ gets attached

Here’s the mechanism behind the second half of these accusations. Model capability improves from three things at once: more and better data, architectural changes, and training method. When a lab ships a jump in performance, it has a strong incentive to credit method and architecture - those look like durable intellectual property - rather than ‘we simply had more of your conversations this round.’ Data-driven gains are harder to defend as a moat and easier to attack on privacy grounds.

So the accusation writes itself: the real driver was the data, and the ‘breakthrough’ framing was cover. It’s a plausible story. It’s also, without training-set documentation and ablation studies from inside the lab, unfalsifiable from the outside. The honest position is that user conversations plausibly contribute to capability, and that no external party can currently quantify how much. Anyone announcing a precise causal split - ‘it was mostly the private chats’ - is guessing with a confident voice. That doesn’t make them wrong. It makes them unverifiable, which for your decision-making should carry the same weight.

How to read a ‘researcher says’ headline

Because these claims spread faster than they can be checked, keep a short filter for them:

  • Is there a primary source - a paper, a repository, a dataset - or just a screenshot of a post? Screenshots are not evidence.
  • Is the researcher named and reachable, with a track record you can look up, or anonymous?
  • Does the claim require internal access to be true? If so, an outsider asserting it with certainty is telling you about their confidence, not the facts.
  • Can it be reproduced? ‘I got the model to output my own earlier conversation’ is testable. ‘They trained on chats and hid it in a breakthrough’ usually is not.
  • Who benefits from the framing? Outrage is a currency. A claim engineered to travel is worth more scrutiny, not less.

Run those five and most ‘researcher says’ posts collapse into one verifiable part and one unprovable part - the same split as the headline we started with.

What you can actually check and control

You don’t need to settle the internal-attribution debate to cut your own exposure. Concrete moves, in rising order of effort:

  1. Open Settings, then Data Controls, in ChatGPT and turn off ‘Improve the model for everyone.’ It governs future chats only.
  2. Use Temporary Chat for anything sensitive; those sessions are excluded from training and from your history - though the retention exceptions above still apply.
  3. If you run a business, keep sensitive work off consumer accounts. Move to the API, Team, or Enterprise, where training exclusion is contractual rather than a toggle you hope stays flipped.
  4. In the EU or UK, file a GDPR data-subject access or deletion request. It forces a documented answer about what’s actually held on you.
  5. Read the usage policy and model card of any tool before you paste real data into it - not the marketing page. The answer to ‘is my input training data’ is almost always written down somewhere; it’s just not on the landing page.

None of this resolves the harder questions - whether opt-out-by-default is legitimate consent at population scale, and whether the capability gains were honestly attributed. Those get decided by regulators and courts, not by a settings menu.

The durable lesson isn’t about one researcher or one company. For consumer AI, ‘trained on private conversations without consent’ sits closer to the documented default than to a scandal - and the scandal, if there is one, is that default-on collection got normalized so thoroughly that a true statement about it can be dressed up as a breakthrough exposé and still move faster than the correction. Check the toggle. Read the policy. Treat anything you type into a free model as a contribution to it until the product proves otherwise in writing.

Share

Keep Reading

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.