RC RANDOM CHAOS

In January 2024, OpenAI admitted it can't train without copyright

OpenAI told the UK Parliament it can't build models without copyrighted work. Here's what that admission means for the fair use cases now in court.

· 7 min read
In January 2024, OpenAI admitted it can't train without copyright

In January 2024, OpenAI told a committee of the UK House of Lords that “it would be impossible to train today’s leading AI models without using copyrighted materials.” That sentence was not leaked. OpenAI wrote it down and submitted it as formal written evidence. It is the cleanest admission in the entire debate, and it frames everything that has happened in U.S. courts since.

When people say OpenAI “admitted two things,” they are compressing a more precise picture. One of those things really is admitted. The other is the contested heart of the litigation. Keeping them separate is the whole job, because the legal outcome, and the money, turns on which is which.

The dependency is not in dispute

The first claim is that OpenAI’s products depend on ingested work the company did not pay for. This is not a hostile characterization. It is OpenAI’s own position, stated to defend itself. The argument runs: the models need to read at scale, the scale required is effectively the whole public internet plus large book and article corpora, and licensing every rights-holder in advance is not commercially feasible. So the company treats ingestion as fair use rather than a purchase.

You can see the dependency in the training sets that keep surfacing in filings and research: Common Crawl (a bulk scrape of the web), and book collections like the one commonly called Books3, which was assembled from a pirate library. In the Authors Guild case and the parallel New York Times case (filed in the Southern District of New York in December 2023, now before Judge Sidney Stein), plaintiffs allege their specific works sat inside those corpora. OpenAI’s response has generally not been “we never used it.” It has been “using it was lawful.” That is a fight about permission and payment, not about whether the ingestion happened.

So treat the first point as settled by the defendant’s own mouth: the product does not exist without the uncompensated corpus. Hold that.

Substitution is the real fight

The second claim, that the output substitutes for the original work, is not something OpenAI admits. It is what plaintiffs must prove, and it is the factor most likely to decide the cases.

U.S. fair use lives in 17 U.S.C. § 107, which lists four factors. Three of them get most of the press. The fourth, “the effect of the use upon the potential market for or value of the copyrighted work,” is the one courts weigh most heavily in commercial disputes. If the new product competes in the same market as the thing it learned from, that factor swings hard against fair use.

This is why the New York Times complaint spends so many pages on regurgitation. It shows the model reproducing long passages of Times articles close to verbatim, including paywalled content. The legal point is not “the model memorized text.” The point is market substitution: a reader who gets the article’s substance from a chatbot did not visit the site, see the ads, or pay the subscription. That is factor four in a single screenshot.

OpenAI’s counter is that those examples were coaxed with adversarial prompting, are not how normal users interact with the product, and do not represent a real substitute for journalism. Both sides are arguing about the same factor. Neither has conceded it.

Why “transformative” carries the defense

The entire fair use case leans on factor one, the purpose and character of the use, and specifically on the word “transformative.” OpenAI argues that learning statistical patterns from text to build a general reasoning tool is a fundamentally different purpose than the original writing, the way a search index is different from the pages it indexes.

That analogy worked before. Google won the book-scanning case (Authors Guild v. Google, decided at the Second Circuit in 2015) partly because Google Books showed snippets and pointed users back to the source. It expanded the market for books instead of replacing it. The tell was that the tool sent readers toward the original.

Generative models invert that. They do not point back. They answer in place. A model that produces a usable summary, a working code function, or a passable article in the style of a named writer is not sending anyone to the source. That is the structural reason the Google precedent is a weaker shield than it looks, and it is why the substitution question keeps eating the transformative question alive.

The courts are no longer speaking with one voice

For a while, defendants could assume that “transformative” would win by default. That assumption is cracking.

In February 2025, in Thomson Reuters v. Ross Intelligence, Judge Stephanos Bibas ruled that a company which used Westlaw headnotes to train a competing legal-research AI was not protected by fair use. The decisive factor was the fourth one: Ross was building a market substitute for the very product it trained on. It was a non-generative model and the facts are narrower than OpenAI’s, but the reasoning is exactly the argument aimed at generative systems. When the output competes with the input, transformation stops rescuing you.

The money moved too. In 2025, Anthropic reached a settlement in the authors’ class action over pirated training books, reported at roughly $1.5 billion. A settlement is not a verdict, and Anthropic admitted no liability. But you do not agree to a number like that unless your lawyers think the uncompensated-corpus position could lose in front of a jury. That figure is the market pricing the risk that the first admission, the dependency, becomes a bill.

Read it as a supply chain with an unpriced input

Strip the legal vocabulary and this is an ordinary systems problem. OpenAI’s supply chain has one input priced at zero that, by the company’s own account, cannot be removed without collapsing the product. Every other input, GPUs, electricity, salaries, data-center leases, is paid at market rate. Training data is the exception, and it is load-bearing.

An unpriced, non-optional input is unstable. Either the price stays zero, which requires the fair use argument to hold across many courts and many plaintiffs, or the price gets set, by settlement, by statute, or by a licensing market. There is no durable third state where a critical input stays free forever while the company built on it is valued in the hundreds of billions. Systems do not leave that kind of gap open. Something fills it.

The ethical version of the same observation is simpler. The value the model captures did not come from nowhere. It came from specific people who wrote specific things, and the company has said in writing that it could not have built the product without them. “We could not exist without your work” and “we owe you nothing for it” are difficult to hold in the same hand once the first half is a matter of record.

What a priced market actually looks like

The practical answer is not “delete the models.” It is metering. Some of it is already being built:

  • Direct licensing. OpenAI has signed content deals with the Associated Press, Axel Springer, News Corp, and others. Each deal is an admission that the input has a price when the rights-holder is large enough to force the conversation.
  • Collective licensing. The music industry solved a similar problem with performing-rights organizations that collect and distribute royalties at scale. Text has no equivalent yet, which is why individual writers are stuck litigating one at a time.
  • Provenance and opt-out infrastructure. Standards like the Coalition for Content Provenance and Authenticity (C2PA) and machine-readable training-permission signals let a market form at all. You cannot pay for what you cannot identify.

The direction of travel is toward paying for training data the way you pay for cloud compute: as a line item, metered, with contracts. The open question is only how much and to whom, not whether.

What to watch next

Three signals will tell you which way this settles, without any legal training required.

First, whether courts keep giving factor four the weight the Ross ruling gave it. If market substitution stays central, the dependency admission gets expensive across the board.

Second, whether licensing deals reach past the handful of large publishers that can negotiate. If individual authors and small outlets get a mechanism to be paid, a real market exists. If only the big players get checks, the fair use fight over everyone else continues.

Third, whether any AI company stops describing training data as free and starts booking it as a cost. The day that line item appears on a balance sheet, the argument is effectively over, and the only remaining question is the size of the number.


Contains a referral link.

Share

Keep Reading

Latest on the Wire

Full wire →

New signal daily · RSS

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.