About Metawatch
Metawatch is a project of the IPTC. It periodically scans a curated list of major news publishers worldwide, samples the photographs they publish, and reports how much embedded metadata survives the journey from photographer's camera to rendered article page. We check each image for Exif, IPTC (in IIM and XMP format) and C2PA Content Credentials.
Why this matters
Photographs carry information in their files: who took them, when and where, the licensing terms, captions, credits, alt text for accessibility and more. The IPTC Photo Metadata Standard exists so that this information travels with the image, anywhere the image goes. In practice much of it is stripped before readers see it: sometimes by publishers, sometimes by Content Delivery Networks (CDNs) as part of automatic resizing, sometimes by intermediate image-processing pipelines.
The cost falls on photographers (whose attribution is lost), agencies (whose licensing terms become invisible), and readers (who lose the provenance signals that would help them judge what they're looking at). Metawatch is meant to be a long-running, public, repeatable way to measure the problem and watch whether it improves.
What inspired this, and what has changed since
Metawatch has two ancestors. The first is IMATAG's State of Image Metadata, published in 2018 and updated for the press in 2019. The second is IPTC's own response: a prototype crawler built in 2019 and presented at that year's Photo Metadata Conference (slides, PDF). Metawatch is the third generation of that idea — the same question, asked every month rather than once.
The 2018 picture, and now
IMATAG's 2018 report split editorial images three ways. Recomputing those buckets on their
own definition of credit — an image counts if any of
Xmp.photoshop.Credit, Iptc.Application2.Credit,
Xmp.dc.rights, Iptc.Application2.Copyright or
Exif.Image.Copyright is filled — gives this:
| Editorial images | IMATAG 2018 | Metawatch, Sept 2026 |
|---|---|---|
| No embedded metadata at all | 80% | 89.3% |
| Some metadata, but no credit or copyright | 12% | 3.0% |
| Metadata including credit or copyright | 8% | 7.7% |
The credit-bearing slice has not moved in eight years: 8% then, 7.7% now. What has changed is the middle. The band of images carrying something but not a credit has collapsed from 12% to 3%, and total stripping is up nine points. Images increasingly arrive either properly credited or completely bare, with less and less in between.
The same publishers, seven years on
IMATAG's 2019 update named individual titles and published its method, so this is the closest thing to a like-for-like re-run that exists. Their figure is the share of a site's images carrying credit metadata; ours is the same test on the lead photograph of each sampled article.
| Publisher | IMATAG 2019 | Metawatch 2026 | Direction |
|---|---|---|---|
| Spiegel Online | 73% | 95% | improved |
| Le Monde | 55% | 88% | improved |
| Le Figaro | 45% | 60% | improved |
| Huffington Post UK | 40% | 0% | fell |
| Politico | 38% | 83% | improved |
| Washington Post | 31% | 0% | fell |
| stern.de | 10% | 0% | fell |
| Les Echos | 6% | 80% | improved |
| El Pais | 4% | 0% | fell |
| The Guardian | 3% | 0% | unchanged |
| Die Zeit | 1% | 0% | unchanged |
| L'Express | 1% | 0% | unchanged |
| NY Times | 0% | 0% | unchanged |
| La Vanguardia | 0% | 0% | unchanged |
| New York Post | 0% | 0% | unchanged |
| Le Point | 0% | 0% | unchanged |
| Wired | 0% | 0% | unchanged |
| La Presse | 0% | 0% | unchanged |
| USA Today | 0% | 95% | improved |
Read the two directions differently. IMATAG sampled every image wider than 400 pixels on a site's home page and article pages — between 1,000 and 7,000 images per site. We sample one lead photograph per article. A lead photograph is the image on a page most likely to have come from an agency with its credit intact, so our percentages should sit above theirs for the same publisher, whatever the truth. That asymmetry is useful: increases here may be partly our sampling, but decreases happened despite it, and a zero cannot be manufactured by a favourable sample.
So the reliable findings are the falls and the flats. Four publishers that carried credit in 2019 now carry none: the Huffington Post UK, the Washington Post, stern.de and El Pais. In each case we found not merely a missing credit but no embedded metadata of any kind on any image we sampled. And eight titles sat at or near zero in 2019 and are still there — the New York Times, Wired, Le Point, La Vanguardia, the New York Post, Die Zeit, L'Express and the Guardian. Seven years, no movement. Whichever side of the line a newsroom was on in 2019, it is almost certainly still on it.
Who has actually fixed it
The falls are the easy story. The rises are harder to trust, because our lead-image sampling flatters them — but not all of them, and the exception matters. IMATAG's method captured every image over 400 pixels on an article page, which necessarily includes the lead photograph. So a publisher whose lead images carried credit in 2019 could not have scored zero in their study. Any title that reads zero for them and high for us has genuinely changed.
USA Today is the clearest case. IMATAG measured it at 0% in 2019. We measure it at 75%, 74%, 80% and 95% across our four 2026 runs — consistently, not as a one-month blip — and the images carry a credit line, a source, a creator and a copyright notice, with a licensor URL on nearly half. Somewhere between 2019 and now, somebody there fixed the pipeline and it has stayed fixed. Les Echos is the same shape: 6% then, 75–90% across our runs now.
Within our own short history, the cleanest recent change is the Chicago Tribune: zero in June, July and August, then 47% in September, with captions, creators, credit lines and even a PLUS data-mining preference appearing together. That is one month of data and could yet reverse, but it has the shape of a deliberate change rather than noise. Diena in Latvia made a similar jump earlier and has held it.
We would rather report this honestly than cheerfully: across the same four runs, nine publishers went from nothing to something, and eighteen went from something to nothing. Improvement is real, it is achievable, and it is currently outnumbered two to one.
Against IPTC's own 2019 crawler
The 2019 IPTC prototype checked 200 news feeds across 35 countries. Across the 30 of those we still cover, the mean share of images carrying any IPTC field has risen from 4.2% to 10.2% (8.2% counting only fields that no Exif tag can satisfy, which is closer to what that tool measured). Treat it gently: the 2019 run reported exactly 0.00% for the Netherlands, Norway, Ireland, Finland and Switzerland, which is far more likely to reflect a prototype's blind spots — it read only RSS and Atom feeds and skipped images under 150 pixels — than five countries embedding nothing at all. A baseline biased low flatters its successor.
Every figure here is recomputed from the published Parquet files, so anyone preferring a different definition can redo it. If anyone at IMATAG would like to compare methodologies properly, we would welcome it — office@iptc.org.
The score
For every image we sample, we check the "Four Cs" of news photo provenance — fields drawn from IPTC IIM and IPTC XMP that long-standing wire-service training treats as universally applicable. A field counts as present if it appears in either family with a non-empty value. The image's score is the sum of weights for fields present, expressed as a percentage of 100. A site's score is the mean across all images we sampled for it; a country's score is the equal-weighted mean across that country's sites.
| Scored field | Weight |
|---|---|
Creator | 25 |
Copyright | 25 |
CaptionDescription | 25 |
CreditLine | 25 |
| Total | 100 |
A further 19 IPTC fields are tracked but do not
affect the score: we record their presence and report it on
the per-field breakdown, but absence isn't penalised. Many of these
are legitimately omitted — LocationCreated can endanger sources, studio
shoots and archival images often genuinely lack DateCreated or
Keywords, and so on. We don't want to confuse "didn't supply" with "supplied
badly".
Tracked fields: AIPromptInformation , AIPromptWriterName , AISystemUsed , AISystemVersionUsed , AltTextAccessibility , DataMining , DateCreated , DigitalSourceType , ExtendedDescriptionAccessibility , Genre , Keywords , LicensorName , LicensorURL , LocationCreated , LocationShown , ObjectName , Source , UsageTerms , WebStatement .
How a site is sampled
For each site in our list we try a small chain of discovery sources, in order, and use the first one that yields articles within the last 30 days:
- RSS feed(s) found on previous attempts.
- RSS feed(s) we auto-discover via
<link rel="alternate">on the homepage. - Common RSS paths we guess (
/feed,/rss.xml, etc.). - XML sitemap declared in
robots.txt— preferring news sitemaps, then any non-image sitemap. Image-only sitemaps are a last resort because they carry no article URLs. - Common sitemap paths we guess (
/sitemap-news.xml,/sitemap.xml, etc.). Gzipped sitemaps (.xml.gz) are supported.
From the chosen source we take up to 20 of the most recent articles, fetch each one, and
extract a single "lead" image per article — either from JSON-LD NewsArticle /
Article markup or from <meta property="og:image">. Inline
body images and related-story thumbnails are deliberately skipped: they often pick up
site chrome, ads, or lazy-load placeholders rather than the photographs the article was
built around. Also, we encourage site owners to keep metadata on higher-resolution "lead"
images, but we understand if site owners want to strip metadata from smaller images for
bandwidth optimisation purposes.
Each image is then fetched and analysed with ExifTool; the field presence and CDN attribution flow into the per-image, per-article and per-site database tables.
What each site status means
ok— we found articles, fetched them, and analysed at least one image.robots_disallow— the site'srobots.txtdisallows our user-agent. We record the entry but don't crawl.no_articles_found— discovery succeeded but no articles in the 30-day window were returned (often a sitemap that points at an empty news feed).no_feed_found— every feed and sitemap address we tried simply wasn't there (a 404, or an ordinary web page where a feed should be). The publisher has no feed or sitemap we know of; we track them until one appears.unreachable— every feed and sitemap we tried failed at the network layer. This usually means that a CDN or WAF is rejecting our requests. We make no assumptions about whether the site is healthy for human readers.timeout— the crawl of that one site exceeded a 5-minute internal cap (rare; usually means a stuck CDN downstream).
Politeness
Metawatch identifies itself as Metawatch/2.0 with a
From: metadata-crawler@iptc.org header. It honours robots.txt,
respects Crawl-delay, and waits at least one second between requests to
the same domain. If a publisher declares a long crawl delay we shrink the article sample so
the whole site finishes within a fixed time budget, rather than hammering them at the edge
of their allowance. We do not bypass paywalls, fingerprints, or CDN bot-management rules.
Site operators can find everything about the crawler — exactly what it fetches, how to verify its requests, and how to block or slow it — on About our crawler.
How often we run
Metawatch runs automatically once per month on the first of the month at 02:00 UTC.
Each run commits its Parquet output and the exported JSON for this site to the project
repository. We may also trigger ad-hoc runs while we are iterating on the crawler — those
appear as additional run directories in data/runs/.
Data and source code
Every run's full Parquet output is published under data/runs/ in the
project repository. The
Dataset page lists every run with file sizes, direct
download links, and a quick-start Python snippet. The data is licensed
CC BY 4.0; the
crawler source code is MIT-licensed. Citations and academic use are welcome — please
credit "IPTC Metawatch" with a link back.
For publishers
If you would like to be removed from the crawled list, or to correct your site's entry — a better feed URL, a missing sitemap, a name fix — please email office@iptc.org. We honour removal requests without question.
Caveats and limitations
- We sample, we don't crawl exhaustively. A 20-article window may be unrepresentative for very large publications, and the sample is necessarily skewed towards whatever a site happens to publish in any given month.
- Scoring rewards presence, not correctness. A photo with a
Creatorfield of"AP"scores the same as one naming the actual photographer. - "Stripped by CDN" versus "never embedded by the publisher" is genuinely hard to distinguish from a single end-of-pipeline fetch. The CDN page gives one view of this; Phase 2 will add a cross-check using publishers who distribute the same image through more than one CDN.
- Some publishers serve content differently to crawlers than to human readers (paywalls, soft blocks, geolocation rules). A low score for those sites may reflect what their bot-facing surface looks like, not what their readers see.
- C2PA detection is presence-first: we report whether a manifest is embedded and
the signing certificate's issuer and we look for DigitalSourceType in c2pa.actions
assertions, but do not yet parse individual
c2pa.metadataorcawg.metadataassertions.
Latest run: 2026-10-01T023219Z — 425 of
508 sites successfully crawled,
6,580 images analysed.