Skip to content
Citations Citations
Blog About Request Access

When AI Licensing Deals Become Evidence in Court: The Mosaic Discovery Fight

Francois-Xavier Bioul
Francois-Xavier Bioul · CCO at Citations LLC
8 min read

When AI Licensing Deals Become Evidence in Court: The Mosaic Discovery Fight

An AI licensing agreement is normally read as a commercial instrument: what content, for what use, at what price, and under what reporting terms. A discovery order and a fresh amicus brief in a closely watched authors' case against Databricks and MosaicML show it can also become part of the evidentiary record a court uses to decide whether a market for that content actually exists.

That shift matters before any dispute starts, not after. It changes what a licensing negotiation should leave behind: not just a signed price, but a documented trail of what was discussed, rejected, and assumed.

In short

In In re Mosaic LLM Litigation, Magistrate Judge Lisa J. Cisneros ordered Microsoft, on August 8, 2025, to produce not only its executed AI training-data agreements but also communications reflecting the negotiations behind them — limited to counterparties that actually signed a deal, and only where the communications discuss potential or proposed terms. On August 27, 2026, the Association of American Publishers, the News/Media Alliance and STM filed an amicus brief in the same case urging the court to deny summary judgment and reject the defendants' fair use defense, arguing that unlicensed training usurps markets that include AI licensing itself.

Neither document decides the fair use question. The discovery order only governs what evidence Microsoft must hand over; the amicus brief is one side's argument, filed with a motion for leave the court still has to grant. Together, though, they show the same market-existence question moving from a negotiating point into something both sides now have to document and prove.

What the Microsoft discovery order actually compels

The case was brought in March 2024 by a group of book authors against Databricks and its subsidiary MosaicML, alleging the companies trained large language models on datasets that included unauthorized copies of their work. It proceeds before District Judge Charles R. Breyer, with discovery disputes referred to Magistrate Judge Cisneros.

Plaintiffs subpoenaed Microsoft for records of agreements under which it had licensed text data for AI training. Microsoft had already produced the executed agreements. Plaintiffs also wanted the negotiation history behind them — draft terms, rejected structures, pricing discussions.

Judge Cisneros granted that request in part. Production was limited to counterparties that had actually signed a deal, and to communications that reflect potential or proposed terms — not every email in a negotiation file. The reasoning is procedural, but it is telling: a final price shows one outcome; the negotiation record shows what the market considered, priced, and rejected along the way.

What the publishers' amicus brief argues — and what it doesn't decide

Filed three weeks later in the same litigation, the AAP / News/Media Alliance / STM brief takes a different angle. Its central argument is that large language models trained on copyrighted text can generate substitutes for that text — verbatim or near-verbatim passages, summaries, derivative-like outputs — and that this “problem of substitution” should weigh heavily against a fair use defense.

Licensing enters that argument directly. The brief states that unauthorized training usurps markets that rightfully belong to copyright owners, including markets for derivative works and AI licensing, and that licensing gives AI developers authorized access to quality content while avoiding the risk of training on unverified sources. That is a narrower and more specific claim than “any licensing deal counts as market proof” — it treats the AI licensing market itself, alongside the market for the underlying works, as one of the markets a court should weigh under the fourth fair use factor.

Two limits are worth stating plainly. The brief is an argument, not a finding — it required a motion for leave, and the court has not ruled on the underlying summary judgment motion. And its market-substitution theory is contested inside the same case: other amici, including EFF and industry groups, argue that treating a licensing market's mere existence as proof of harm would let rights holders claim a veto over any use that might conceivably compete with their work — the same “market dilution” objection that has surfaced in other AI training litigation.

Why the fourth factor turns on more than the final price

Section 107 asks courts to weigh the effect of the use upon the potential market for or value of the copyrighted work. Courts have been wary of letting that factor become circular: if any unlicensed use can retroactively be called “licensable,” almost anything can be said to harm a market that was only theoretical.

The Cisneros order sidesteps that problem procedurally rather than doctrinally. By limiting discovery to real negotiations behind deals that actually closed, it draws a line between a market that exists — because parties were negotiating and some signed — and a market a rights holder merely wishes existed. That is close to the logic the Supreme Court applied in Andy Warhol Foundation v. Goldsmith: whether a use serves as a market substitute depends on what the market for licensing that kind of use actually looks like, not on an abstract entitlement to license everything.

What to document in every AI licensing negotiation now

Neither document requires publishers to turn every negotiation into a litigation file. It does mean treating the negotiation record as an asset with a life beyond the signature — useful to both sides of a future dispute.

Record to keep

Why it matters if a dispute follows

The precise use requested

Separates “AI training” in general from the specific product, workflow and audience actually negotiated

Rights included and excluded

Stops broad deal language from being read backward to cover scope it never granted

Composition of the content set

Ties the deal to a specific catalog, not a category — the difference between a comparable and a guess

Alternatives discussed and rejected

Documents what the market actually considered substitutable — the core question under Warhol

Assumptions behind the valuation

Distinguishes a defensible benchmark from a number nobody can explain later

Reasons for the principal tradeoffs

Shows why a deal looks the way it does, not just what it says

The rights holder gets a record that can show a market existed and what it valued. The licensee gets a record of exactly what it bought, so a broadly worded clause cannot be stretched past what was actually priced. Both outcomes depend on the same discipline: writing the negotiation down before a docket — not a memory — has to reconstruct it.

Where this connects to usage evidence

A negotiation record answers what a market considered. It does not answer what actually happened to the content once a license was signed — which model, which product, which workflow, how often. That is a separate evidentiary gap, and it is the one AI usage evidence for publishers is built to close.

The two records reinforce each other at renewal. A publisher who can show what the market negotiated, and what actually happened under the license that resulted, is arguing from documentation on both fronts — not from a recollection of what a deal was supposed to mean. That is the same discipline behind how early AI licensing deals become market benchmarks, and the same gap identified in why legal presumptions still need usage records.

It also sits next to the market-effect question in what the Ross appeal means for AI training and competing products — Ross turned on whether a specific product competed with Westlaw; Mosaic turns on whether the licensing market itself is part of the evidentiary record. Different questions, same requirement: a documented trail, not an assumption.

This is the layer Citations Logic is built for: helping publishers and rights organizations turn licensing negotiations and content usage into records that hold up outside the room where the deal was made.

Frequently asked questions

What did the Microsoft discovery order in In re Mosaic LLM Litigation actually decide?

It resolved a discovery dispute, not the underlying fair use question. Magistrate Judge Cisneros ordered Microsoft to produce negotiation communications behind its AI training-data agreements, limited to counterparties that signed a deal and to communications reflecting potential or proposed terms.

Does the AAP, News/Media Alliance and STM brief prove AI training is not fair use?

No. It is an amicus argument urging the court to deny summary judgment and reject the fair use defense — not a ruling. Amicus briefs require the court's leave to be considered, and the underlying motion has not been decided.

Why do licensing negotiation records matter more than the final signed price?

A price shows one outcome. A negotiation record shows what the market considered, priced, and rejected along the way — the difference between a market that demonstrably existed and one asserted only after the fact.

What should companies document during AI licensing negotiations now?

At minimum: the precise use requested, rights included and excluded, the composition of the content set, alternatives discussed and rejected, the assumptions behind the valuation, and the reasons for the main tradeoffs.

Does this mean every AI licensing deal will end up in litigation?

No. Most won't. But a deal negotiated without a documented record leaves a publisher unable to show what the market considered if a dispute ever does reach discovery — a cost that only becomes visible after the fact.

Sources

In re Mosaic LLM Litigation, No. 3:24-cv-01451 (N.D. Cal.) — case docket, including the order of August 8, 2025 regarding Plaintiffs' subpoena to Microsoft
CourtListener docket

Association of American Publishers — press release and full brief
Book, News, and Journal Publishers File Amicus Brief in In Re Mosaic LLM Litigation, August 27, 2026

17 U.S.C. § 107 — statutory fair use factors
Section 107 of the Copyright Act

Andy Warhol Foundation for the Visual Arts, Inc. v. Goldsmith, 598 U.S. 508 (2023)
Supreme Court opinion

Continue the evidence chain

For how the same fourth fair use factor played out in a different case, read what the Ross appeal means for AI training and competing products.

AI content licensing benchmark covers what happens when early deals set the market's reference price.

For the operational gap this leaves unresolved, see why legal presumptions still need usage records.

To understand the broader category, read AI usage evidence for publishers.

Book a licensing documentation and evidence review