Skip to content
Citations Citations
Blog About Request Access

Who Can Upload Content to AI? Input Authority for Publishers and RROs

Francois-Xavier Bioul
Francois-Xavier Bioul · CCO at Citations LLC
7 min read

Who Can Upload Content to AI? Input Authority for Publishers and RROs

For two years, the content industries have built their AI defences around the crawler: robots.txt, technical blocking, licensing talks with AI platforms. That perimeter leaves another door open, one nobody forces because it opens from the inside.

In April 2026, The Bookseller reported that some editors had been uploading submitted manuscripts into consumer AI tools to read them faster. On April 16, the Authors Guild answered with a statement and a new model contract clause. Nobody in that story was attacking anything.

In short

Access to a file is not permission to put it into an AI system. An editor may read a manuscript without being entitled to send it to an AI provider; an RRO may administer member data without its mandate covering AI processing of that data.

Publishers and RROs need a rule between access and use: input authority. It defines who may bring which content into which system, for which operation, and what record shows the operation stopped where it should. It rests on four decisions (authority, operation, environment, lifecycle) and a register that records them. Neither a training opt-out nor a secure environment replaces them.

The breach that looks like good work

The scale is no longer anecdotal. Cyberhaven's 2026 AI Adoption & Risk Report, based on data movements observed across 222 companies, found that 39.7% of AI interactions involve sensitive data, including prompts, copy-paste actions and file uploads. The same report finds that about a third of employees reach AI tools through personal accounts, which sit outside single sign-on, central logging, retention policies and training controls.

These are cross-sector figures from a security vendor, not a publishing survey. They still show the mechanics: the upload is ordinary, frequent and mostly invisible.

For a publisher, the exposure sits upstream of publication: manuscripts, author contracts, reader reports. For an RRO, it covers the repertoire, licence agreements, member data and distribution files. In both cases, the material is valuable partly because it has not left the building.

Access is not input authority

Three permissions are routinely treated as one. Access lets a person open a file; use covers what may be done with the content under a licence or mandate. Input authority sits between them: the documented permission for a named person to bring specific content into a specific AI system for a specific operation. It is not a legal category but a governance decision, and most organisations have not written it down.

The publishing sector already has a template. The Authors Guild asks that no manuscript or author's personal information be uploaded to a consumer-facing AI system without the author's written permission. Where permission is granted, the work should be opted out of training, and internal AI systems should be sandboxed so that manuscripts cannot become training inputs. The Guild's model clauses also separate AI training rights, other AI uses and paid AI licences treated as subsidiary rights.

These clauses are neither law nor universal practice. They do show what an upload rule needs: named content, a permitted operation, an admitted environment, an accountable party.

Four decisions an upload policy must make

Authority. Who can authorise the input: the author, the rights holder, the mandating member, the business owner, legal or security? Access to the file cannot stand in for that authorisation.

Operation. Summarising, translating, indexing, querying through retrieval (RAG), fine-tuning, training and producing derivative content are not equivalent acts. A generic permission for "internal AI use" leaves every one of them open to interpretation.

Environment. A consumer app, a business plan, an API, a private cloud and an on-premise model carry different terms on retention, provider access, location and reuse. A training opt-out is a useful control where it exists. It does not answer retention, subprocessing, deletion or transfer.

Lifecycle. The organisation must be able to say which file went where, sent by whom, for what purpose and for how long. It also needs a procedure for deletion requests and for an upload made by mistake.

French regulators point in the same direction. The CNIL recommends that organisations deploying generative AI define authorised and prohibited uses, train end users and run regular checks. ANSSI, the national cybersecurity agency, recommends prohibiting internet-based generative AI tools for professional work involving sensitive data.

The input authority register

The four decisions only hold if they are recorded per type of content. The table below is an illustrative starting grid for publishers and RROs; each line has to be checked against the organisation's own contracts and mandates.

ContentWho can authorise inputOperations typically admissibleEnvironment admittedRecord to keep
Unpublished manuscriptAuthor, through the contracts or rights teamSummary or reader report, only with written consentBusiness tier or sandboxed internal system, training disabledConsent reference, file identifier, tool, date, deletion confirmation
Author contract, agent correspondenceLegal or contracts directorClause extraction, comparisonInternal or private deploymentRequester, purpose, retention period
Published catalogueRights director, within the rights granted by authorsIndexing, retrieval, metadata enrichment; training only under an express grantContracted environment with usage loggingGrant reference per title, operation log
RRO repertoire and member personal dataManagement body under the mandate, with the DPOMatching, deduplicationPrivate or on-premise; impact assessment where requiredLegal basis, impact assessment reference, access log
Distribution filesDistribution committee or general managementAnomaly checks, simulationsOn-premise or private cloud, no provider reuseFile version, run date, outputs retained

The last column matters most, because it is the only one that can be checked afterwards. For RROs, the distribution-file line carries particular weight, since distribution keys increasingly depend on usage evidence rather than on declarations.

Segmentation does not replace authorisation. A secure environment can still receive content the organisation had no right to use for that purpose, and a contractual authorisation can coexist with an environment that offers no guarantees. The register has to join both layers.

The register covers what enters the system. What the provider then logs, retains and deletes is the other half of the same question, set out in the ten-item AI vendor checklist libraries now use in procurement.

The objection: "Our contract already covers AI use"

Many large publishers already include a clause in author contracts that allows AI use for business operations. The objection follows: input authority is already granted, so a register adds process without adding rights.

The clause settles whether AI use is permitted at all. It does not say which operation, in which environment, decided by whom, or what record will exist. An RRO mandate raises the same issue: authority to manage members' rights is not automatically authority to feed them into a new AI operation.

The objection holds in one case: where the contract is explicit and the environment already controlled, the register is record-keeping, not a new negotiation.

What to decide before the next AI deployment

The next incident will probably look like good work: someone summarising faster or answering a client sooner. That is why the answer cannot rest on individual caution.

Before connecting any AI tool to a corpus, an organisation should be able to answer one question: who has the right to move which content across which boundary, for which operation, and how will anyone show that the operation stopped where it should? A policy that answers in prose is a declaration, not an audit trail.

The limit is scope: a register records authorised inputs and will not detect an upload through a personal account. Its first value is to make the rule explicit enough to train on and check. Citations Logic works on the record side of that problem: AI usage evidence that shows what was done with content, not only what was allowed.

Review your AI input authority with Citations Logic

Where to go next

The external vector, what AI search systems take and what publishers must log themselves, is covered in AI search controls and the evidence publishers must record.

For the question that follows every authorised input, what the provider keeps and for how long, start with how ALA turned AI procurement into an audit.

Sources

Authors Guild — "Use of AI in Publishing and New Model Contract Clause," April 16, 2026 (updated April 22, 2026)
Authors Guild statement on uploading manuscripts to AI

Authors Guild — "AI-Related Model Publishing Contract Clauses"
Authors Guild AI model clauses

Jane Friedman — "Authors Guild adds a new AI clause to their model contract regarding publishers' use of AI on manuscripts," April 2026
Jane Friedman on the Authors Guild upload clause

Cyberhaven Labs — "2026 AI Adoption & Risk Report"
Cyberhaven 2026 AI Adoption & Risk Report

CNIL — "Comment déployer une IA générative ? La CNIL apporte de premières précisions"
CNIL guidance on deploying generative AI

ANSSI — "Recommandations de sécurité pour un système d'IA générative"
ANSSI security recommendations for a generative AI system