AI Search Controls: What Publishers Must Record Themselves
Two things happened this summer that point in opposite directions.
On 3 June 2026, the UK Competition and Markets Authority imposed a conduct requirement on Google search. Publishers gained, in the CMA's words, a world-first ability to keep their content out of AI features such as AI Overviews.
On 21 August 2026, wikiHow filed a complaint against OpenAI in the Southern District of New York. According to the filing, wikiHow had disallowed OpenAI's GPTBot in its robots.txt since August 2023, added directives against two further OpenAI crawlers in January 2025, notified OpenAI's legal department in December 2024 — and still logged 148,529 crawler visits between May and July 2026.
One event gives publishers controls. The other shows what a control is worth when the only party who can see whether it was honoured is the publisher itself.
In short
AI Search governance now has two layers that are easy to confuse.
Controls are what a platform agrees to offer: opt-out switches, attribution links, reporting dashboards. They are declared, administered and revoked by the counterparty.
Evidence is what remains on your own servers when a control turns out to have been ignored, changed or narrowed. It is retained by you.
The CMA requirement is a control regime. It obliges Google to give publishers choices and metrics. It does not oblige anyone to hand you a record you can still produce in a dispute two years from now.
So the retention question is not "what does the platform report?" It is: of the six things you might want to prove about AI Search, which ones can only you produce?
What the CMA requirement actually delivers
The conduct requirement follows the CMA's designation of Google with strategic market status in general search services. Four obligations matter for rights teams.
Publishers get controls over whether their search content powers Google's generative AI features. After consultation feedback, the CMA extended this to cover fine-tuning of models, not just grounding — so display rights and training rights become separately controllable.
Google must attribute publisher content with clear links in AI-generated search results, and must publish comprehensible explanations of how search content feeds its generative AI.
Critically, exercising the opt-out must not cost a publisher its classic search ranking. That protection is what makes the control usable at all — and it is also the first thing that will need evidence if it is ever tested, because only the publisher holds a before-and-after ranking series.
Publishers get engagement metrics on their content in generative search features. The Publishers Association's reading of the requirement lists impressions, clicks, click-through rate and the ability to calculate click quality.
And Google has nine months to implement, though the CMA expects important parts of the controls to reach publishers well before that deadline. Compliance reports are due every six months for the first year.
Read that list again with a rights lawyer's eye. Every item is administered by the counterparty. The opt-out is a switch Google builds. The metrics are a dashboard Google populates. The attribution is a link Google renders. None of it is custody.
The wikiHow filing is a lesson about custody
The wikiHow complaint contains allegations, not findings, and discovery will test them. But the structure of the pleading is instructive regardless of outcome.
Its strongest components are the ones wikiHow could evidence unilaterally: 1,211 registered copyrights with documented chain of title, its own web server logs showing crawler identity and volume, and a dated notice to OpenAI's legal department. Those facts do not depend on OpenAI's cooperation, disclosure or record-keeping.
Its most contestable components are the inferences about what happened inside the counterparty — whether crawler activity proves training, whether a new model necessarily re-ingested the corpus, whether engineered verbatim prompts reflect ordinary user behaviour.
That asymmetry is the transferable lesson, and it applies well beyond litigation. It applies to a licensing renewal, a regulatory complaint, an audit clause and a board discussion about whether an AI channel is worth anything.
Publisher-side records carry weight. Counterparty-side inference does not. This is the same problem as retrieved is not cited, arriving through the governance door rather than the measurement door.
Six objects, ranked by who holds the record
Most AI Search checklists list controls in the order a platform presents them. That is the wrong order for a retention policy.
Rank by custody instead: keep first what only you can produce, because that is what survives a dashboard being redesigned, a reporting obligation being narrowed to one jurisdiction, or a counterparty rotating its logs.
Priority | Object | Who holds the evidence | What to retain on your side |
|---|---|---|---|
1 | Crawler identity and purpose | You, entirely | Raw access logs with user-agent, IP and timestamp, retained well beyond your default rotation window — plus a dated archive of each vendor's published crawler documentation, since declared purposes are revised without notice |
2 | Opt-out state over time | Split: you declare, they honour | Dated snapshots of every robots.txt version and every platform control setting, plus the notices you sent and when |
3 | Reproduction in outputs | You, if you test for it | Timestamped prompt-and-output captures against your own registered works, with the model version recorded |
4 | Substitution effect | You, if you baselined early | Traffic, conversion and revenue series per content class, established before the controls take effect — recorded as a series, not as a causal claim |
5 | Distribution metrics | The platform | Monthly exports archived on your own storage — a dashboard you do not control is not a record |
6 | Attribution rate | The platform | Sampled, dated captures of AI answers that draw on your catalogue, whether or not they link back |
The ranking is the argument. Items 1 to 4 cost you almost nothing today and cannot be reconstructed later. Items 5 and 6 arrive by regulation, which means they arrive on someone else's schedule and in someone else's format.
Note what the ordering does to a common assumption. Crawler logs sit at the top not because they prove the most, but because they are the only object in the list where the publisher holds complete, contemporaneous, unilateral evidence. That is exactly the object most publishers rotate out after thirty days.
The objection: the regulator is already forcing disclosure
A fair challenge: if the CMA now compels metrics and attribution, why duplicate the work?
Three reasons.
Scope. The requirement binds one company, in one jurisdiction, for search-integrated features. It says nothing about the crawler that visited you this morning on behalf of an assistant that is not Google and does not operate a search index.
Form. Regulatory reporting is built for operational use and compliance monitoring, not for evidentiary use years later. Access to a report is not custody of a record, and the difference only becomes visible at the moment you need to produce something.
Timing. Nine months of implementation means the controls describe the future. Every dispute currently on file concerns a period during which no control existed at all. Retention is retrospective insurance, and its value is set by what you kept before you knew you needed it.
The same logic governs usage reporting in AI licensing deals, and it is why the EU AI Act's transparency obligations improve a publisher's position without solving it.
What to change before the controls arrive
Three decisions, in order.
Extend log retention now, before December. This is an infrastructure ticket, not a strategy project, and it is the only item on this list whose window closes every day.
Snapshot control state with dates. Every robots.txt change, every platform toggle, every notice sent. A control you cannot date is a control you cannot enforce.
Build the substitution baseline while there is still a "before". Once opt-out is exercised, traffic moves for two reasons at once, and separating them afterwards is guesswork.
None of this requires a view on whether AI Search is good or bad for publishing. It requires only a view on who will be holding the records when the question is asked.
The distinction that decides it
A control tells you what a platform agreed to do.
A log tells you what happened.
Most of the time they say the same thing, and the distinction looks academic. In wikiHow's account, they diverged for three years, and only one of the two was still producible in August 2026.
This is the layer Citations Logic is built for: turning AI usage into records a publisher holds, rather than assurances a publisher receives. That is what AI usage evidence for publishers means in practice.
Before the CMA controls reach your teams, one question is worth putting to whoever owns your infrastructure: how many days of crawler logs do we currently keep?
If the answer is thirty, the controls will arrive before the evidence does.
Book an AI Search controls and evidence review
Read next
Retrieved Is Not Cited: Why AI Visibility Fails Publishers
Usage Reporting in AI Licensing Deals: Who Audits the Number?
Sources
Competition and Markets Authority — "CMA secures fairer deal for publishers and improves Google search services in UK", 3 June 2026
CMA secures fairer deal for publishers and improves Google search services in UK
Publishers Association — "CMA conduct requirements for Google search: Publishers Association statement", 3 June 2026
CMA conduct requirements for Google search: Publishers Association statement
IPWatchdog — "WikiHow Files Lawsuit Against OpenAI Over ChatGPT's Use of Its How-To Library", 25 August 2026
WikiHow Files Lawsuit Against OpenAI Over ChatGPT’s Use of Its How-To Library
US District Court, Southern District of New York — wikiHow, Inc. v. OpenAI, Inc. et al, case 1:26-cv-07171, filed 21 August 2026
wikiHow, Inc. v. OpenAI, Inc. et al, case 1:26-cv-07171