← About Metawatch

About our crawler

If you have seen Metawatch/2.0 in your logs, this page is for you. Metawatch is a research crawler run by the IPTC, the global standards body for news media. Once a month it visits a few hundred news publishers, reads the embedded metadata of a small sample of the photographs they publish, and reports how much survives. It does not collect content for AI training or any other reuse.

At a glance

OperatorIPTC (International Press Telecommunications Council)
PurposeMeasuring whether photo metadata (Exif, IPTC, C2PA) survives publication. Results are public on this site.
User-AgentMetawatch/2.0
Frommetadata-crawler@iptc.org
Signature-Agent"https://iptc.org" (how to verify)
Robots tokenMetawatch
ScheduleMonthly, starting 02:00 UTC on the 1st; occasional re-crawls of single sites
VolumeTypically 40–60 requests per publisher per month
Runs fromGitHub Actions (Microsoft Azure, United States)
Source codegithub.com/iptc/metawatch (MIT)
ContactFeedback form; removal requests to office@iptc.org

What it fetches

For each publisher, once a month:

It does not run JavaScript, follow links beyond the sampled articles, submit forms, log in, or fetch anything behind a paywall.

What we keep

From each article: its URL, headline, publication date, language, keywords and its structured-data (NewsArticle) markup. Not the article text. From each image: its URL, dimensions, file size, which metadata fields are present and their values, and any C2PA manifest summary. The image file itself is deleted as soon as its metadata has been read. Everything we keep is published as open data on the Dataset page, so you can see exactly what we recorded about your site.

How it behaves

Verifying that a request is really ours

The crawler runs on GitHub-hosted runners, whose IP addresses change from run to run, so we can't publish an IP list. Instead each request is signed with Web Bot Auth (HTTP Message Signatures, RFC 9421). Each signed request carries:

Our public key is published at https://iptc.org/.well-known/http-message-signatures-directory. A request that claims to be Metawatch but doesn't verify against that key isn't from us.

Blocking or limiting us

To stop Metawatch crawling your site, add this to your robots.txt:

User-agent: Metawatch
Disallow: /

To slow it down instead:

User-agent: Metawatch
Crawl-delay: 10

Changes take effect at the next monthly crawl. You can also ask us to remove your site, or correct its entry, at office@iptc.org. We honour removal requests without question. A blocked site still appears on Metawatch with the status robots_disallow, since honouring that choice is part of what we measure.

Seen something wrong?

If the crawler misbehaves (too many requests, ignoring your robots.txt, anything else) please tell us via this short form, with the time, your domain and, if you have them, the request headers. We read every response.