AI preferences and opt-out signals

Publishers have at least ten independent technical mechanisms for signalling that they do not want their images and articles used to train generative-AI models. They sit at different layers (site-wide files, HTTP headers, embedded image metadata) and were proposed by different bodies over a short period — so adoption is uneven and the signals do not always agree. Article 4 of the EU DSM Directive requires opt-outs to be expressed in a machine-readable form; this page measures which of the candidate machine-readable forms publishers actually use.

The IPTC's Generative AI Opt-Out Best Practice Recommendations (v2.0, March 2026) sets out the thirteen techniques the IPTC currently recommends to publishers wishing to express a data-mining opt-out. Most are measured here; the remainder (plain-language rights statements, per-page TDMRep meta tags, firewall-level blocking, and opt-outs embedded in epub/PDF) are noted in the methodology section below.

Based on the most recent crawl: 502 publishers probed, 6,697 images analysed.

Adoption snapshot

per site
48.2%
robots.txt — any AI bot blocked
Blocks at least one known AI/scraper UA
242 of 502 →
per site
15.5%
noarchive / nosnippet meta robots
Honoured as AI-opt-out by Bing/Copilot (noarchive) and Google (nosnippet)
78 of 502 →
per site
4.2%
robots.txt — Content-Signal
Cloudflare Content Signals (contentsignals.org)
21 of 502 →
per site
2.8%
IPTC PLUS:DataMining
XMP-plus:DataMining on ≥1 sampled image
14 of 502 →
per site
1.4%
TDMRep — per-page <meta>
<meta name="tdm-reservation" content="1"> on ≥1 article
7 of 502 →
per site
1.2%
/.well-known/tdmrep.json
TDM Reservation Protocol (EU DSM Art. 4)
6 of 502 →
per site
0.6%
/ai.txt
Spawning AI consent proposal
3 of 502 →
per site
0.4%
RSL — License: in robots.txt
Really Simple Licensing (rslstandard.org)
2 of 502 →
per site
0.4%
trust.txt — datatrainingallowed=no
JournalList trust.txt opt-out directive
2 of 502 →
per site
0.2%
CAWG Training and Data Mining Assertion
cawg.training-mining assertion in ≥1 sampled image's C2PA manifest
1 of 502 →
per site
0.0%
noai / noimageai meta robots
AI-specific tokens on X-Robots-Tag or <meta name="robots">
0 of 502 →

robots.txt — AI bot block matrix

For each known AI/scraper user-agent, the share of crawled sites whose robots.txt disallows that UA at the root. List is versioned in known_uas.yaml; we welcome pull requests adding bots we have missed.

User-agent Operator % sites blocking  
CCBot Common Crawl 40.3%
GPTBot OpenAI 37.7%
Bytespider ByteDance 36.8%
ClaudeBot Anthropic 36.2%
anthropic-ai Anthropic 31.6%
Google-Extended Google 31.6%
PerplexityBot Perplexity 29.2%
Amazonbot Amazon 28.7%
omgilibot Webhose 28.3%
Applebot-Extended Apple 27.7%
Claude-Web Anthropic 27.7%
cohere-ai Cohere 27.5%
Meta-ExternalAgent Meta 27.3%
Diffbot Diffbot 26.3%
omgili Webhose 25.5%
ChatGPT-User OpenAI 24.3%
FacebookBot Meta 21.9%
YouBot You.com 20.8%
Meta-ExternalFetcher Meta 20.0%
OAI-SearchBot OpenAI 19.6%
Timpibot Timpi 18.4%
Claude-SearchBot Anthropic 16.6%
Scrapy Open-source scraper 16.4%
ImagesiftBot ImageSift 16.2%
magpie-crawler Brandwatch 16.0%
TurnitinBot Turnitin 15.6%
Perplexity-User Perplexity 15.4%
Claude-User Anthropic 15.0%
DuckAssistBot DuckDuckGo 14.6%
AI2Bot AI2 13.8%
cohere-training-data-crawler Cohere 13.8%
DataForSeoBot DataForSEO 12.8%
FriendlyCrawler Webis 12.3%
PanguBot Huawei 12.3%
AwarioRssBot Awario 11.7%
AwarioSmartBot Awario 11.7%
img2dataset Open-source scraper 11.5%
MistralAI-User Mistral 11.5%
DeepSeekBot DeepSeek 11.3%
ia_archiver Internet Archive 11.3%
Google-CloudVertexBot Google 10.9%
NewsNow NewsNow 10.5%
BLEXBot WebMeUp 10.3%
archive.org_bot Internet Archive 8.9%
peer39_crawler Peer39 8.1%
SeekrBot Seekr 8.1%
news-please news-please (open-source) 7.7%
Feedfetcher-Google Google 7.5%
Gemini-Deep-Research Google 6.7%
Quora-Bot Quora 6.3%
MyCentralAIScraperBot MyCentral 6.1%
quillbot.com QuillBot 6.1%
EchoboxBot Echobox 5.9%
Poseidon Research Crawler Poseidon Research 5.9%
AliyunSecBot Alibaba Cloud 5.1%
SeznamHomepageCrawler Seznam 5.1%
TaraGroup Intelligent Bot TaraGroup 5.1%
AudigentAdBot Audigent 4.7%
Grok xAI 4.7%
ViennaTinyBot Vienna Tiny 4.7%
Jetslide Jetslide 4.3%
GoogleOther Google 3.8%
bingbot Microsoft 0.4%

How many AI UAs does each site block?

A site can block zero, one, or many of the 63 tracked AI UAs. The distribution reveals whether publishers are picking individual bots to refuse or applying a blanket policy.

UAs blocked Sites % of sites  
0 (no AI blocks) 264 52.2%
1–5 39 7.7%
6–10 39 7.7%
11–20 54 10.7%
21–506 110 21.7%

Image-level signals

Three of the eight mechanisms live inside the image file (or its HTTP response), not on a site-wide file. They travel with the image when it is copied or shared.

Do the signals agree?

A site that takes AI opt-out seriously might use several mechanisms together. In practice the overlap is patchy — partly because the signals serve different audiences (robots.txt for crawlers, TDMRep for the EU legal framework, IPTC PLUS:DataMining for downstream image consumers).

If a site… …does it also… % overlap N
blocks GPTBot blocks ClaudeBot 84% 191
blocks GPTBot has tdmrep.json 3% 191
has tdmrep.json blocks ≥1 AI UA 83% 6
has ai.txt has tdmrep.json 0% 3
has RSL License blocks ≥1 AI UA 100% 2

What we measure (methodology)

The "publishers probed" denominator on this page (502) is wider than the "Sites crawled" figure on the homepage. The homepage counts publishers where the crawler completed a full pass through to fetching images. The denominator here also includes publishers whose robots.txt disallowed our user-agent at the root, whose sitemap or feed discovery was blocked, or who returned no articles — because for those publishers we still read robots.txt, probed the well-known files, and so on, so we can legitimately answer "did they express an AI preference". Only publishers we couldn't reach at all (DNS / TCP failures) are excluded. Narrowing the denominator to the homepage's figure would drop precisely the publishers most likely to opt out, biasing every percentage upward.