About our crawler
If you have seen Metawatch/2.0 in your logs, this page is for you. Metawatch is
a research crawler run by the IPTC, the global standards
body for news media. Once a month it visits a few hundred news publishers, reads the
embedded metadata of a small sample of the photographs they publish, and reports how much
survives. It does not collect content for AI training or any other reuse.
At a glance
| Operator | IPTC (International Press Telecommunications Council) |
|---|---|
| Purpose | Measuring whether photo metadata (Exif, IPTC, C2PA) survives publication. Results are public on this site. |
| User-Agent | Metawatch/2.0 |
| From | metadata-crawler@iptc.org |
| Signature-Agent | "https://iptc.org" (how to verify) |
| Robots token | Metawatch |
| Schedule | Monthly, starting 02:00 UTC on the 1st; occasional re-crawls of single sites |
| Volume | Typically 40–60 requests per publisher per month |
| Runs from | GitHub Actions (Microsoft Azure, United States) |
| Source code | github.com/iptc/metawatch (MIT) |
| Contact | Feedback form; removal requests to office@iptc.org |
What it fetches
For each publisher, once a month:
robots.txt.- An RSS feed or news sitemap to find recent articles. If we don't already know yours,
we look for one linked from your homepage or in
robots.txt, then try a few common paths such as/feedand/sitemap.xml, stopping at the first that works. - Up to 20 articles published in the last 30 days.
- One image per article: the lead photograph the publisher declares in its
JSON-LD or
og:imagemarkup, up to 20 MB. - A few site-wide policy files, so we can report publishers' AI and text-and-data-mining
preferences:
/.well-known/tdmrep.json,/ai.txtand/.well-known/trust.txt.
It does not run JavaScript, follow links beyond the sampled articles, submit forms, log in, or fetch anything behind a paywall.
What we keep
From each article: its URL, headline, publication date, language, keywords and its
structured-data (NewsArticle) markup. Not the article text. From each image:
its URL, dimensions, file size, which metadata fields are present and their values, and
any C2PA manifest summary. The image file itself is deleted as soon as its metadata has
been read. Everything we keep is published as open data on the
Dataset page, so you can see exactly what we recorded
about your site.
How it behaves
- One request at a time per site, at least one second apart plus a
random delay, or longer if your
robots.txtsets aCrawl-delay. If your crawl delay is long, we take fewer articles rather than staying longer: each site is capped at five minutes. robots.txtis honoured for our token (Metawatch) and forUser-agent: *. Article URLs it disallows are dropped. Where we use a feed a publisher has listed for us, we fetch that feed beforerobots.txt, because some firewalls block an address outright after it readsrobots.txt; the article URLs from the feed are still checked against your rules. Ifrobots.txtcan't be read at all, we treat the site as allowing us.- No disguise. We always send the same
User-AgentandFromheaders, never a browser's. We don't rotate addresses to get round blocks, solve CAPTCHAs, or try to defeat bot-management rules. When a site blocks us, we record it as blocked and stop.
Verifying that a request is really ours
The crawler runs on GitHub-hosted runners, whose IP addresses change from run to run, so we can't publish an IP list. Instead each request is signed with Web Bot Auth (HTTP Message Signatures, RFC 9421). Each signed request carries:
Signature-Agent: "https://iptc.org"Signature-Inputcovering@authorityandsignature-agent, withkeyid="rtvhLrHEsgX7dk-Bj5Fb9Zzi2CjxEccPinja47gY6yk",alg="ed25519"andtag="web-bot-auth", valid for 60 secondsSignature: the Ed25519 signature
Our public key is published at https://iptc.org/.well-known/http-message-signatures-directory. A request
that claims to be Metawatch but doesn't verify against that key isn't from us.
Blocking or limiting us
To stop Metawatch crawling your site, add this to your robots.txt:
User-agent: Metawatch
Disallow: / To slow it down instead:
User-agent: Metawatch
Crawl-delay: 10
Changes take effect at the next monthly crawl. You can also ask us to remove your site, or
correct its entry, at office@iptc.org. We honour
removal requests without question. A blocked site still appears on Metawatch with the
status robots_disallow, since honouring that choice is part of what we measure.
Seen something wrong?
If the crawler misbehaves (too many requests, ignoring your robots.txt, anything
else) please tell us via this short
form, with the time, your domain and, if you have them, the request headers. We read
every response.