Skip to content
Citations Citations
Blog About Request Access

AI Training, Competing Products, and Copyright: What the Ross Appeal Means for Publishers

Francois-Xavier Bioul
Francois-Xavier Bioul · CCO at Citations LLC
7 min read

AI Training, Competing Products, and Copyright: What the Ross Appeal Means for Publishers

On June 11, 2026, the Third Circuit heard oral arguments in Thomson Reuters v. Ross Intelligence. No ruling has been issued yet. If it comes, it will be the first federal appellate word on whether training an AI system on copyrighted editorial content can be fair use.

But the question worth answering now isn't whether "AI training" is allowed or forbidden as a category. It's narrower, and more useful: does the economic function of an AI use — improving a licensed product, or building a substitute for one — change what that use should cost, and how it should be governed.

In short

Ross Intelligence built a legal research tool. To train it, Ross used materials derived from Westlaw headnotes — editorial summaries Thomson Reuters writes for every reported case. In February 2025, the District of Delaware found that Ross had copied 2,243 of the 2,830 headnotes it examined, and rejected Ross's fair use defense on summary judgment.

Two of the four fair use factors decided the outcome: the purpose and character of the use, and the effect on the market. The court found Ross's use commercial and non-transformative, and found that Ross's product was built to compete directly with Westlaw's. The Third Circuit heard argument on both fair use and the copyrightability of the headnotes on June 11, 2026. No decision has been issued.

The case draws a line publishers should already be drawing in their contracts. Content accessed to improve a licensor's own product does not perform the same economic function as content accessed to help a third party build a substitute for it — even when the underlying content is identical.

At Citations Logic, we describe this through usage evidence: records that capture not just which content was accessed, but under which rights, in which workflow, and for what declared purpose. Without that record, a publisher cannot tell which of these two events happened until the competing product is already in the market.

The commercial value of an AI usage event depends not only on what content was used, but on what market function that use performed.

Why Thomson Reuters v. Ross is different from the generative AI cases

Most AI copyright litigation involves large language models trained across the open web to generate new text. Ross is not that. Ross built a non-generative search tool: it locates relevant case law based on a query. It doesn't produce prose, summaries, or anything resembling the headnotes it trained on.

That distinction matters for how far the case travels. The court wasn't ruling on generative AI trained broadly for research or education. It was ruling on a commercial product, built from a competitor's editorial corpus, aimed at the same customers. Publishers watching this case for a blanket answer on "AI training and fair use" are watching the wrong frame. The frame that transfers is narrower and more useful: proximity between the source and the product, and what market the product enters.

Why the ruling turned on purpose and market harm — not training alone

Fair use in the U.S. weighs four factors. In Ross's case, two decided it: the purpose and character of the use, and the effect on the market for the original.

On purpose, the court found Ross's use commercial and not transformative — Ross used the headnotes to build a product that performed close to the same function they did, rather than to create something with new meaning or message. On market effect, the court found that Ross's tool was designed to compete directly with Westlaw, and that this harmed the market Thomson Reuters had built around its editorial content.

The other two factors — the nature of the copyrighted work and the amount used — favored Ross. That split matters for how publishers should read this case. It wasn't decided by the fact that training happened. It was decided by what the training was for, and what it put into the market once it was done. A non-infringing output doesn't automatically remove the competitive significance of the training use behind it.

Not every AI use performs the same economic function

This is the distinction that should be in every AI license before it's signed — not after a dispute forces it.

Type of AI use

Primary economic function

Publisher's main concern

Evidence needed

Authorized RAG

Retrieve licensed content to answer a permitted query

Attribution, retrieval scope, compensation

Content ID, passage, rights, query, timestamp

Internal consultation

Improve an organization's own internal work

Scope of users, permitted reuse

User class, purpose, content accessed, retention

General-purpose training

Improve a broad model

Training rights, corpus inclusion, downstream use

Dataset provenance, licensed works, training scope

Product enhancement

Improve the publisher's own service

Internal governance, partner controls

Workflow, feature, permitted purpose, usage record

Competing-product training

Build or improve a substitute service

Market displacement, licensing value

Content used, model purpose, product function, competitive field

Substitutable output

Replace demand for the source content or product

Revenue loss, market substitution

Output-source relationship, frequency, user workflow

The same content — the same headnote, the same paragraph, the same dataset — can sit in any row. What changes the license value isn't the content. It's which row the use is actually in.

"But the training was intermediate, and the output doesn't reproduce the source"

This is Ross's strongest argument, and publishers should take it seriously rather than dismiss it. Ross's system doesn't output headnotes. It trained on them to learn relationships between legal questions and case law, then returns judicial opinions — not Thomson Reuters's editorial text. Ross argues that use was intermediate, and that the appeal should turn on whether the headnotes were copyrightable at all.

That argument doesn't resolve the commercial question, even if it succeeds on the legal one. A license needs to separate four things that a single "AI use" clause collapses into one: whether the output copies the source, whether the training used the source, what the resulting product does, and what market that product enters. Ross may be right that its output doesn't reproduce Thomson Reuters's text. That's a separate question from whether the product built from that training competes for Westlaw's customers. Publishers who license on the first question alone are leaving the second one unpriced.

What to require before the next AI license

  • Name the permitted product, not just "AI use." A license that says "for AI training" doesn't say whether the resulting product can compete with the licensor.

  • Separate internal use from market-facing use, and require a new authorization when a use moves from one to the other.

  • Require purpose-linked usage records — a query count says nothing about which economic function was served.

  • Set an escalation trigger. A change in product, workflow, or audience should require re-authorization, not a retroactive dispute.

  • Keep the link auditable between content, granted right, and declared purpose — the record that would have told Thomson Reuters, in 2020, what Ross's use was actually for.

The renewal question this case is really asking

Whatever the Third Circuit decides, publishers already have an operating answer: "AI use" is not a precise enough licensing category to build a contract on. The next agreement should distinguish use by product, purpose, audience, market exposure, and the evidence available to verify all four.

That's the layer Citations Logic is built for: connecting authorized AI access to content-level usage evidence, so a publisher can verify whether content was used within the product, workflow, and purpose that were actually licensed — not assume it, after the fact, from a court docket.

Frequently asked questions

What is the difference between AI training use and training for a competing product?

Both start with the same act: a model learns from copyrighted content. The difference the court drew was in the fair use factors. Ross's use was found commercial, non-transformative, and aimed at a product built to compete directly with Westlaw. Content used to improve an authorized product does not perform the same economic function as content used to help a third party build a substitute for it — even when the underlying content is identical.

Does Thomson Reuters v. Ross apply to generative AI models trained on the open web?

Not directly. Ross built a non-generative legal search tool trained on a narrow, editorial corpus close to its own market. Most pending AI copyright cases involve generative models trained broadly across the web. The frame that transfers from Ross isn't "AI training is infringement" — it's the weight courts place on proximity between the source and the resulting product.

How many Westlaw headnotes did the court find infringed?

The District of Delaware found actual copying of 2,243 of the 2,830 headnotes it examined, on summary judgment, in its February 11, 2025 order.

Has the Third Circuit ruled on the appeal yet?

No. The court heard oral argument on June 11, 2026, on both the fair use question and whether the headnotes are copyrightable. No decision had been issued as of this article's publication.

What should publishers require before signing an AI licensing deal?

A license should name the permitted product rather than just "AI use," separate internal use from market-facing use, require purpose-linked usage records, set a re-authorization trigger for changes in product or audience, and keep an auditable link between content, granted right, and declared purpose.


Sources

This analysis draws on the primary record of Thomson Reuters Enterprise Centre GmbH v. Ross Intelligence, Inc.: the District of Delaware summary judgment order, February 11, 2025, the interlocutory appeal certification, May 23, 2025, and reporting on the Third Circuit oral argument, June 11, 2026.


Continue the evidence chain

AI Usage Evidence for Publishers

Proof of AI Content Usage

One-Time vs Recurring AI Licensing Revenue

EU AI Act and Content Licensing

Book an AI usage evidence assessment