Verified as of 27 September 2026. If anything has changed, let me know and I will update it with the date of the change.
If you publish work online, European law lets you reserve certain uses for text and data mining. That mining may play a part in building datasets and in some stages of developing artificial intelligence systems. The reservation is not a lock: it is a legal expression of objection that compliant systems must be able to detect.
This guide explains the legal basis, the available signals and their limits. It is not legal advice. Before changing a live site, check its configuration and keep a copy.
The legal basis
Article 4 of Directive (EU) 2019/790 allows reproductions and extractions for text and data mining of lawfully accessible works, provided that the rights holder has not expressly reserved those uses in an appropriate manner. When the content is publicly available online, the reservation must be expressed by machine-readable means.
The Directive does not impose a single protocol. Recital 18 gives machine-readable means as examples, including metadata and the terms and conditions of a website or service. Whether a particular signal is effective will depend on whether it expresses the reservation, can be detected in the relevant context and can be linked to the person entitled to make it.
The Article 4 reservation does not exclude the separate Article 3 exception for research organisations and cultural heritage institutions carrying out mining for the purposes of scientific research on material to which they have lawful access. Nor does it revoke licences already granted or erase copies made before the reservation.
What Kneschke v LAION showed
The Higher Regional Court of Hamburg ruled on 10 December 2025 that the temporary download of a photograph while creating a dataset could be covered by the German mining exceptions. The agency hosting the image had included a prohibition in natural language. The court did not say that every written clause is useless in principle: it found that, for the relevant point in 2021, it had not been shown that the reservation was machine-readable in the sense required for content available online.
The ruling has been appealed. The Federal Court of Justice (Bundesgerichtshof) held the hearing on 3 September 2026 and has scheduled its decision for 17 December 2026. Until then, the case should not be presented as settling the European debate.
The practical lesson is limited but important: a human-readable legal notice may not be enough on its own. It is worth pairing it with consistent technical signals and keeping evidence of when they were active.
Layered signals
No single tool prevents every use. The most prudent option is to combine a clear legal reservation with technical mechanisms suited to where you publish.
1. robots.txt
robots.txt is a file at the root of the domain that tells compliant crawlers which paths they may visit. Each operator publishes the identifiers it uses and may separate crawling for training, search or user-initiated actions.
An illustrative example that blocks four identifiers linked to training or large corpora:
User-agent: GPTBot
Disallow: /
User-agent: Google-Extended
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: CCBot
Disallow: /
Always check the names against each operator's current documentation before publishing the file.
GPTBotis OpenAI's identifier for crawling that may contribute to the development of generative models.Google-Extendedis a control token, not a separate HTTP crawler. It lets you limit certain Google AI uses without blocking ordinary indexing by Google Search.ClaudeBotis Anthropic's crawler for content that could contribute to training.CCBotis Common Crawl's crawler; its open corpus may be reused by third parties.
Do not block search identifiers by mistake if you want assistants to be able to find and link to your pages. OAI-SearchBot and Claude-SearchBot, for example, serve a different purpose from training crawlers.
robots.txt does not technically prevent access. It gives instructions to crawlers that choose to follow them. Nor does it control copies of your work hosted on other domains.
2. TDMRep
The TDM Reservation Protocol is designed specifically to express reservations relating to mining. It was published on 10 May 2024 as the final report of a W3C community group. It is not a W3C standard and is not on the official W3C standards track.
TDMRep lets you carry the signal through:
/.well-known/tdmrep.jsonon the server;- the
TDM-ReservationHTTP header; - HTML metadata;
- metadata in EPUB and PDF.
In HTML, the minimal declaration is:
<meta name="tdm-reservation" content="1">
The value 1 means that mining rights are reserved. You can add a policy explaining how to request a licence:
<meta name="tdm-reservation" content="1">
<meta name="tdm-policy" content="https://ejemplo.com/licencias-tdm.json">
For a server-wide policy, the exact syntax of the JSON file must follow the current report. Do not improvise its structure: a valid HTML tag is better than a malformed file.
TDMRep expresses a policy, not an access control. It only has a technical effect on systems that read and process it.
3. CAWG's training and data mining assertion
The Creator Assertions Working Group published version 1.1 of its training and data mining assertion on 16 May 2025. It lets a person add to a C2PA manifest information on permitted or restricted uses for mining and training.
The institutional distinction matters. On 22 January 2026 C2PA clarified that its core specification does not contain a standard mining assertion or a rights management technology. The assertion comes from CAWG and can be carried through the extensible architecture of C2PA manifests.
Its advantage is that the signal travels with the file. Its weakness is persistence: it can be lost if a platform strips metadata, generates a new copy or does not keep the manifest. Check the file after every upload and download.
4. ai.txt, noai and noimageai
ai.txt is a private proposal linked to Spawning. noai and noimageai are conventions used by some platforms and creative communities. They are not universally adopted legal or technical standards, and I have not found case law recognising them on their own as effective reservations under Article 4.
They may strengthen the evidence of a consistent intention, but they should not be your only signal when you control the domain and can use more specific mechanisms.
A human-readable clause
Also add a clear statement to your terms or rights notice. For example:
The rights holder expressly reserves the uses of the works and content on this site for text and data mining, including their use in datasets and in processes for training, fine-tuning or evaluating artificial intelligence systems, in accordance with Article 4(3) of Directive (EU) 2019/790 and its applicable transposition. This reservation does not affect uses authorised under licence or legal exceptions that cannot be reserved.
Adapt the clause to the rights you actually control. Do not reserve on behalf of photographers, collaborators or other rights holders without authorisation. A broader clause is no substitute for a machine-readable technical signal.
If you publish on third-party platforms
On a social network, agency, gallery, repository or shop you do not normally control the robots.txt, the HTTP headers or the retention of metadata. The reservation on your website only covers content served from your domain; it does not automatically govern copies hosted by third parties.
Before publishing, check:
- the terms of the licence granted to the platform;
- its policy on training and mining;
- whether it offers an opt-out;
- whether it keeps Content Credentials and other metadata;
- whether it allows linking or embedding from your domain instead of re-uploading the file.
Keep a canonical copy on your website with an identifier, a date and the reservation. This does not control other copies, but it improves traceability.
How to keep evidence
A reservation does not usually have retroactive effect. To prove when it was active:
- keep dated copies of
robots.txt, the HTML and any TDMRep file; - keep deployment logs, version history or receipts from your provider;
- capture the full HTTP response when you use headers;
- check regularly that the public site shows the signal;
- keep the original files with their manifests and metadata;
- archive the applicable terms of the platforms you use.
An external web archive can provide context, but it does not guarantee that it captures headers, restricted files or embedded metadata. Do not rely on a single piece of evidence.
What no reservation guarantees
None of these measures:
- prevents a non-compliant actor from downloading the content;
- deletes earlier copies or datasets already built;
- requires a third party to remove a work it hosts;
- proves on its own that a work was used to train a model;
- excludes scientific mining covered by Article 3;
- replaces a claim, a licence or a technical access control measure;
- settles which law applies when the acts take place in several countries.
A reservation strengthens the legal and documentary signal. It does not guarantee actual exclusion.
What to do today
- Identify which rights you control and on which domains the work is hosted.
- Separate crawling for training, search and user-initiated access.
- Set up
robots.txtwith the current identifiers you want to exclude. - Add TDMRep through HTML, a header or a well-formed file.
- Publish a legal clause consistent with the technical signals.
- Add the CAWG assertion when your workflow preserves C2PA manifests.
- Check the terms of external platforms.
- Keep dated evidence and check the configuration regularly.
On the open conversation
This guide draws on Directive 2019/790, the Code of Practice for General-Purpose AI Models, TDMRep, the CAWG and C2PA documentation and the Hamburg ruling in Kneschke v LAION. The technical landscape will keep changing. It should be reviewed after the German decision scheduled for 17 December 2026 and whenever new recognised solutions or judicial criteria emerge.
Sources
- Directive (EU) 2019/790, Arts. 3 and 4 and Recital 18.
- European Commission, Code of Practice for General-Purpose AI Models.
- W3C TDM Reservation Protocol Community Group, TDMRep final report, 10 May 2024.
- Creator Assertions Working Group, Training and Data Mining Assertion v1.1.
- C2PA, clarification on TDM assertions, 22 January 2026.
- Federal Court of Justice, I ZR 281/25, ruling scheduled for 17 December 2026.
- Hanseatic Higher Regional Court of Hamburg, 5 U 104/24, 10 December 2025.
- OpenAI, crawler documentation.
- Anthropic, crawler documentation.
- Google, Google-Extended.
- Common Crawl, CCBot.
