Skip to content
Citations Citations
Blog About Request Access

AI Content Licensing Benchmark: Why Usage Data Decides Value

Francois-Xavier Bioul
Francois-Xavier Bioul · CCO at Citations LLC
9 min read

AI Content Licensing Benchmark: Why Usage Data Decides Value

AI licensing is creating a market.

But a market does not only need prices. It needs reference points.

The first AI content licensing deals are now becoming those reference points. News Corp, Axel Springer, the Associated Press, The New York Times and other major publishers have entered the AI licensing debate through agreements, lawsuits, or both. Each new deal moves the market further from theory into precedent.

A recent Brookings Institution analysis, based on work with the Open Markets Institute, describes the structural problem clearly: the companies that reshaped the economics of content distribution are now helping shape how publishers are compensated for AI use.

That raises the question every rights leader will face next:

What is the AI content licensing benchmark — and who gets to set it?

In short

An AI content licensing benchmark is the reference point negotiators use to value content when no public market price exists. Today, early benchmarks come mainly from two flawed sources: confidential bilateral deals and litigation pressure. Both matter. Neither is enough.

A deal-based benchmark reflects what one publisher was able to negotiate at one point in time. It does not necessarily show how often that publisher’s content was retrieved, surfaced, cited, summarized, or reused by AI systems. It does not reveal the scope of rights granted. It does not show whether the deal covers training, grounding, display, summaries, attribution, or all of them together.

A usage-based benchmark starts somewhere else. It asks what actually happened to the content.

At Citations Logic, we use the term usage evidence to describe the independent record of AI content use: which works were retrieved, surfaced, cited, summarized, reused, and connected to a specific commercial context.

For publishers, collective management organizations, and rights leaders, the practical conclusion is uncomfortable but simple:

A benchmark that only the buyer can measure is not a market reference. It is the buyer’s number.

Why early AI licensing deals become benchmarks

Early deals do more than create revenue. They create expectations.

When a major publisher signs with a major AI company, the market watches. Even when the financial terms remain confidential, the structure of the deal starts to influence the next negotiation. Other publishers ask whether they should sign. Collective management organizations ask whether similar terms could be generalized. AI companies use previous agreements to frame what is “market standard.”

That is how a benchmark forms.

The problem is that most AI licensing benchmarks are forming in opacity. Public announcements usually describe partnership, access, quality journalism, attribution, innovation, or responsible AI. They rarely disclose the exact rights granted, the usage measurement method, the reporting obligations, or the renewal logic.

That makes the benchmark unstable.

A price without its scope is not a benchmark.

A deal without usage data is not a market reference.

It is a signal of bargaining power.

The News Corp–OpenAI agreement, for example, gives OpenAI permission to display content from News Corp mastheads in response to user questions and access current and archived content from major publications including The Wall Street Journal, Barron’s, MarketWatch, Investor’s Business Daily, the New York Post, The Times, The Sunday Times, The Australian and others.

The public announcement is significant. But like most such announcements, it does not create a public unit of measurement for actual AI usage.

The same pattern applies more broadly. The agreements keep coming. So do the lawsuits. But the market still lacks a shared reference point for how content contribution should be measured.

The valuation problem is deeper than lost traffic

A weak AI licensing benchmark often begins with the wrong proxy.

In the search era, publishers were trained to think in traffic. Referrals. Clicks. Pageviews. Subscriptions. Advertising conversion.

That logic made sense when users moved from search results to publisher websites.

AI changes the sequence.

An AI system can retrieve information, synthesize it, summarize it, ground an answer, or surface a citation without generating a meaningful visit back to the publisher. The content may still create value. The click may disappear.

This is why referral loss is too narrow as a benchmark.

Publisher content can contribute to AI systems in several ways:

  • training and fine-tuning;

  • factual grounding;

  • retrieval-augmented generation;

  • citation and attribution;

  • answer display;

  • summarization;

  • reasoning support;

  • temporal freshness;

  • domain authority;

  • risk reduction in professional workflows.

A benchmark based only on traffic loss misses most of this value. It prices the visible damage, not the actual contribution.

That is the trap.

If publishers accept traffic as the main reference point, they allow AI companies to define value using the metric AI itself is weakening.

Deal-based benchmarks vs. usage-based benchmarks

Deal-based benchmark

Usage-based benchmark

What it measures

What another publisher was able to negotiate

What an AI system retrieved, surfaced, cited, summarized, or reused

Primary data source

Announced, leaked, or reported deal figures

Independent usage records

Who controls the evidence

Usually the parties to someone else’s contract

The rights holder, a collective body, or a neutral evidence layer

Comparability

Weak, because scope and terms are often confidential

Stronger, if the same usage units are applied across deals

Renewal value

Anchors the next negotiation to the previous deal

Adjusts the next negotiation to documented usage

Commercial risk

Imports another publisher’s bargaining position

Creates negotiation leverage from observed contribution

Failure mode

Turns opacity into market standard

Requires measurement infrastructure to exist

This distinction matters because benchmarks harden.

Once a deal structure becomes “standard,” it becomes harder to challenge. Once a reporting format becomes normal, it becomes harder to demand better records. Once the buyer controls the only usage data, the seller’s renewal position is already weakened.

This is why usage reporting in AI licensing deals is not an operational detail. It is the mechanism through which a benchmark becomes verifiable.

Without usage reporting, a publisher gets a number.

With usage evidence, a publisher gets a market reference.

Why collective management organizations should care first

For collective management organizations, the benchmark question comes before the mandate question.

A single publisher can decide to sign a bilateral agreement quickly. That may be rational. The publisher may want revenue, visibility, strategic proximity, or litigation leverage.

A collective management organization faces a different responsibility. It does not negotiate for one catalog. It aggregates rights across many rightsholders. It must defend not only the total amount collected, but the way that amount is allocated.

That changes the standard of proof.

A collective AI licensing market cannot rely only on brand power, catalog size, or historical revenue. Those inputs may matter, but they are not enough. They say who is big. They do not say whose content was used.

A fair collective benchmark needs to answer harder questions:

  • Which works were used?

  • How often were they retrieved?

  • Were they cited, summarized, displayed, or only used as background grounding?

  • Did the AI system attribute the source?

  • Was the usage linked to a consumer query, an enterprise workflow, an educational use, or a professional task?

  • Can the record be audited independently?

  • Can the same measurement logic apply across licensees?

Without these answers, collective negotiation starts from someone else’s number.

And if that number is set by private AI deals, the imbalance is already baked in.

The strongest objection: sign now, benchmark later

The strongest objection is speed.

The market is moving now. AI companies are signing deals now. Courts are hearing cases now. Publishers that wait too long may lose leverage. Collective bodies that spend too much time defining a perfect benchmark may watch the market consolidate without them.

That objection is serious.

But it does not lead to the conclusion that rights leaders should sign first and measure later.

It leads to a sharper rule:

Sign early if needed. But never sign without a measurement clause.

A deal signed today with usage reporting obligations can produce, within one contract cycle, the data needed for a stronger benchmark. A deal signed today without usage evidence produces only a confidential number. At renewal, the publisher still depends on the buyer’s view of value.

That is not speed.

That is dependency with a deadline.

The goal is not to slow the market down. The goal is to prevent the first market reference points from being built on data that only one side can see.

What an AI content licensing benchmark should measure

A serious AI content licensing benchmark should not begin with the headline price. It should begin with the unit of value.

The minimum measurable units should include:

  1. Work-level identification
    Which title, article, chapter, entry, image, dataset, or source was used?

  2. Usage type
    Was the content retrieved, summarized, cited, displayed, used for grounding, or used in another AI workflow?

  3. Frequency
    How often did the usage occur over a defined period?

  4. Context
    Was the usage consumer-facing, educational, enterprise, professional, legal, scientific, medical, or internal?

  5. Attribution status
    Was the publisher or source visible to the user?

  6. Rights status
    Was the use authorized, restricted, excluded, or outside the scope of the license?

  7. Commercial effect
    Did the usage support a paid product, subscription feature, enterprise workflow, answer engine, API call, or agentic task?

Only after these units exist can the market argue intelligently about price.

Otherwise, the negotiation starts in the wrong place.

It asks: “What did someone else get?”

The better question is: “What was actually used?”

From benchmark to renewal leverage

The first AI licensing deal matters.

The renewal matters more.

At signing, both sides can work from assumptions. At renewal, assumptions should no longer be enough. The publisher should know which content was used, how usage evolved, where attribution failed, which works created repeated value, and whether the license scope matched real behavior.

That is where usage evidence changes the negotiation.

A publisher with usage records can say:

  • this content was repeatedly retrieved;

  • these works were surfaced in high-value contexts;

  • these sources were used without visible attribution;

  • these categories drove more AI interactions than expected;

  • this license should be repriced;

  • this reporting obligation should be tightened;

  • this distribution key should be adjusted.

A publisher without usage records can only say:

  • we believe our content matters;

  • the market has moved;

  • another publisher received more;

  • the buyer should report more.

That is not a renewal strategy.

That is a weaker version of the first negotiation.

What rights leaders should require now

Publishers, collective management organizations, and rights leaders do not need to wait for perfect regulation before acting. But they should not allow the AI licensing benchmark to form without them.

Three requirements should be on the table now.

First, every AI licensing framework should include a usage record clause. Not just a quarterly declaration. Not just aggregate reporting. A record of retrieval, citation, display, summarization, and reuse.

Second, rights holders should push for common measurement units across licensees. Without common units, deals cannot be compared. If deals cannot be compared, benchmarks cannot be defended.

Third, collective organizations should define benchmark governance before benchmark numbers. The industry should agree on what gets measured before arguing about what it is worth.

This is where Citations Logic operates: producing independent usage evidence that connects AI usage to works, rights, sources, and commercial contexts.

The most important negotiation in AI licensing may not be the next agreement.

It may be deciding what everyone will use as the reference point.

Because the next AI licensing market will be shaped by the benchmark.

And the benchmark will be shaped by whoever controls the usage data.

Continue the evidence chain

To understand the broader category, read AI usage evidence for publishers.

Usage reporting in AI licensing deals decides renewal leverage.

To understand the regulatory pressure behind documentation, read EU AI Act content licensing.

To connect valuation to collective revenue allocation, read AI copyright collective management distribution key.

To assess your own licensing evidence position, Book an AI usage evidence assessment.

Sources and further reading

Brookings Institution — Same Gatekeepers, New Tollbooths in the AI Content Licensing Market, Courtney C. Radsch, June 9, 2026.
https://www.brookings.edu/articles/same-gatekeepers-new-tollbooths-in-the-ai-content-licensing-market/

Open Markets Institute — Same Gatekeepers, New Tollbooths: Mapping the AI Content Licensing Market, 2026.
https://www.openmarketsinstitute.org/publications/report-mapping-the-ai-content-licensing-market

News Corp — News Corp and OpenAI Sign Landmark Multi-Year Global Partnership, May 22, 2024.
https://investors.newscorp.com/news-releases/news-release-details/news-corp-and-openai-sign-landmark-multi-year-global-partnership

The Wall Street Journal — Amazon to Pay New York Times at Least $20 Million a Year in AI Deal, 2025.
https://www.wsj.com/business/media/amazon-to-pay-new-york-times-at-least-20-million-a-year-in-ai-deal-66db8503

Axios — AP strikes news-sharing and tech deal with OpenAI, July 13, 2023.
https://www.axios.com/2023/07/13/ap-openai-news-sharing-tech-deal

Associated Press — OpenAI to start using news content from News Corp. as part of a multiyear deal, May 22, 2024.
https://apnews.com/article/a49144d381796df5729c746f52fbef19

European Commission — AI Act: regulatory framework for artificial intelligence.
https://digital-strategy.ec.europa.eu/en/policies/regulatory-framework-ai