AI preferences and opt-out signals

Publishers have at least ten independent technical mechanisms for signalling that they do not want their images and articles used to train generative-AI models. They sit at different layers (site-wide files, HTTP headers, embedded image metadata) and were proposed by different bodies over a short period — so adoption is uneven and the signals do not always agree. Article 4 of the EU DSM Directive requires opt-outs to be expressed in a machine-readable form; this page measures which of the candidate machine-readable forms publishers actually use.

The IPTC's Generative AI Opt-Out Best Practice Recommendations (v2.0, March 2026) sets out the thirteen techniques the IPTC currently recommends to publishers wishing to express a data-mining opt-out. Most are measured here; the remainder (plain-language rights statements, per-page TDMRep meta tags, firewall-level blocking, and opt-outs embedded in epub/PDF) are noted in the methodology section below.

Based on the most recent crawl: 505 publishers probed, 6,553 images analysed.

Adoption snapshot

per site
46.3%
robots.txt — any AI bot blocked
Blocks at least one known AI/scraper UA
234 of 505 →
per site
14.3%
noarchive / nosnippet meta robots
Honoured as AI-opt-out by Bing/Copilot (noarchive) and Google (nosnippet)
72 of 505 →
per site
3.0%
IPTC PLUS:DataMining
XMP-plus:DataMining on ≥1 sampled image
15 of 505 →
per site
1.6%
robots.txt — Content-Signal
Cloudflare Content Signals (contentsignals.org)
8 of 505 →
per site
1.4%
TDMRep — per-page <meta>
<meta name="tdm-reservation" content="1"> on ≥1 article
7 of 505 →
per site
1.2%
/.well-known/tdmrep.json
TDM Reservation Protocol (EU DSM Art. 4)
6 of 505 →
per site
0.8%
/ai.txt
Spawning AI consent proposal
4 of 505 →
per site
0.4%
RSL — License: in robots.txt
Really Simple Licensing (rslstandard.org)
2 of 505 →
per site
0.4%
trust.txt — datatrainingallowed=no
JournalList trust.txt opt-out directive
2 of 505 →
per site
0.2%
CAWG Training and Data Mining Assertion
cawg.training-mining assertion in ≥1 sampled image's C2PA manifest
1 of 505 →
per site
0.2%
noai / noimageai meta robots
AI-specific tokens on X-Robots-Tag or <meta name="robots">
1 of 505 →

robots.txt — AI bot block matrix

For each known AI/scraper user-agent, the share of crawled sites whose robots.txt disallows that UA at the root. List is versioned in known_uas.yaml; we welcome pull requests adding bots we have missed.

User-agent Operator % sites blocking  
CCBot Common Crawl 38.4%
GPTBot OpenAI 36.0%
Bytespider ByteDance 34.6%
ClaudeBot Anthropic 34.1%
anthropic-ai Anthropic 32.7%
PerplexityBot Perplexity 29.5%
cohere-ai Cohere 29.3%
omgilibot Webhose 29.3%
Google-Extended Google 28.7%
Claude-Web Anthropic 28.3%
Amazonbot Amazon 28.0%
Diffbot Diffbot 27.8%
omgili Webhose 26.4%
Applebot-Extended Apple 25.8%
Meta-ExternalAgent Meta 25.8%
ChatGPT-User OpenAI 24.8%
YouBot You.com 23.0%
FacebookBot Meta 22.8%
OAI-SearchBot OpenAI 19.9%
Timpibot Timpi 19.9%
Claude-SearchBot Anthropic 18.7%
Meta-ExternalFetcher Meta 18.7%
Perplexity-User Perplexity 17.9%
Claude-User Anthropic 17.7%
Scrapy Open-source scraper 17.5%
ImagesiftBot ImageSift 17.1%
magpie-crawler Brandwatch 17.1%
TurnitinBot Turnitin 16.3%
AI2Bot AI2 16.1%
DuckAssistBot DuckDuckGo 16.1%
cohere-training-data-crawler Cohere 15.2%
MistralAI-User Mistral 14.2%
PanguBot Huawei 13.8%
DataForSeoBot DataForSEO 13.6%
DeepSeekBot DeepSeek 13.0%
FriendlyCrawler Webis 13.0%
AwarioRssBot Awario 12.6%
AwarioSmartBot Awario 12.6%
img2dataset Open-source scraper 12.6%
ia_archiver Internet Archive 11.8%
NewsNow NewsNow 11.8%
Google-CloudVertexBot Google 11.4%
archive.org_bot Internet Archive 10.8%
BLEXBot WebMeUp 10.8%
SeekrBot Seekr 8.9%
news-please news-please (open-source) 8.5%
peer39_crawler Peer39 8.5%
Gemini-Deep-Research Google 8.1%
Quora-Bot Quora 7.3%
quillbot.com QuillBot 7.1%
MyCentralAIScraperBot MyCentral 6.9%
EchoboxBot Echobox 6.7%
Grok xAI 6.7%
Poseidon Research Crawler Poseidon Research 6.7%
TaraGroup Intelligent Bot TaraGroup 6.1%
AliyunSecBot Alibaba Cloud 5.7%
AudigentAdBot Audigent 5.5%
SeznamHomepageCrawler Seznam 5.5%
ViennaTinyBot Vienna Tiny 5.5%
GoogleOther Google 5.1%
Jetslide Jetslide 5.1%
Feedfetcher-Google Google 3.7%
bingbot Microsoft 1.4%

How many AI UAs does each site block?

A site can block zero, one, or many of the 63 tracked AI UAs. The distribution reveals whether publishers are picking individual bots to refuse or applying a blanket policy.

UAs blocked Sites % of sites  
0 (no AI blocks) 274 53.9%
1–5 40 7.9%
6–10 27 5.3%
11–20 46 9.1%
21–508 121 23.8%

Image-level signals

Three of the eight mechanisms live inside the image file (or its HTTP response), not on a site-wide file. They travel with the image when it is copied or shared.

Do the signals agree?

A site that takes AI opt-out seriously might use several mechanisms together. In practice the overlap is patchy — partly because the signals serve different audiences (robots.txt for crawlers, TDMRep for the EU legal framework, IPTC PLUS:DataMining for downstream image consumers).

If a site… …does it also… % overlap N
blocks GPTBot blocks ClaudeBot 82% 183
blocks GPTBot has tdmrep.json 3% 183
has tdmrep.json blocks ≥1 AI UA 83% 6
has ai.txt has tdmrep.json 0% 4
has RSL License blocks ≥1 AI UA 100% 2

What we measure (methodology)

The "publishers probed" denominator on this page (505) is wider than the "Sites crawled" figure on the homepage. The homepage counts publishers where the crawler completed a full pass through to fetching images. The denominator here also includes publishers whose robots.txt disallowed our user-agent at the root, whose sitemap or feed discovery was blocked, or who returned no articles — because for those publishers we still read robots.txt, probed the well-known files, and so on, so we can legitimately answer "did they express an AI preference". Only publishers we couldn't reach at all (DNS / TCP failures) are excluded. Narrowing the denominator to the homepage's figure would drop precisely the publishers most likely to opt out, biasing every percentage upward.