About Metawatch

Metawatch is a project of the IPTC. It periodically scans a curated list of major news publishers worldwide, samples the photographs they publish, and reports how much embedded metadata survives the journey from photographer's camera to rendered article page. We check each image for Exif, IPTC (in IIM and XMP format) and C2PA Content Credentials.

Why this matters

Photographs carry information in their files: who took them, when and where, the licensing terms, captions, credits, alt text for accessibility and more. The IPTC Photo Metadata Standard exists so that this information travels with the image, anywhere the image goes. In practice much of it is stripped before readers see it: sometimes by publishers, sometimes by Content Delivery Networks (CDNs) as part of automatic resizing, sometimes by intermediate image-processing pipelines.

The cost falls on photographers (whose attribution is lost), agencies (whose licensing terms become invisible), and readers (who lose the provenance signals that would help them judge what they're looking at). Metawatch is meant to be a long-running, public, repeatable way to measure the problem and watch whether it improves.

What inspired this, and what has changed since

Metawatch has two ancestors. The first is IMATAG's State of Image Metadata, published in 2018 and updated for the press in 2019. The second is IPTC's own response: a prototype crawler built in 2019 and presented at that year's Photo Metadata Conference (slides, PDF). Metawatch is the third generation of that idea — the same question, asked every month rather than once.

The 2018 picture, and now

IMATAG's 2018 report split editorial images three ways. Recomputing those buckets on their own definition of credit — an image counts if any of Xmp.photoshop.Credit, Iptc.Application2.Credit, Xmp.dc.rights, Iptc.Application2.Copyright or Exif.Image.Copyright is filled — gives this:

Editorial imagesIMATAG 2018Metawatch, Sept 2026
No embedded metadata at all80%89.3%
Some metadata, but no credit or copyright12%3.0%
Metadata including credit or copyright8%7.7%

The credit-bearing slice has not moved in eight years: 8% then, 7.7% now. What has changed is the middle. The band of images carrying something but not a credit has collapsed from 12% to 3%, and total stripping is up nine points. Images increasingly arrive either properly credited or completely bare, with less and less in between.

The same publishers, seven years on

IMATAG's 2019 update named individual titles and published its method, so this is the closest thing to a like-for-like re-run that exists. Their figure is the share of a site's images carrying credit metadata; ours is the same test on the lead photograph of each sampled article.

PublisherIMATAG 2019Metawatch 2026Direction
Spiegel Online73%95%improved
Le Monde55%88%improved
Le Figaro45%60%improved
Huffington Post UK40%0%fell
Politico38%83%improved
Washington Post31%0%fell
stern.de10%0%fell
Les Echos6%80%improved
El Pais4%0%fell
The Guardian3%0%unchanged
Die Zeit1%0%unchanged
L'Express1%0%unchanged
NY Times0%0%unchanged
La Vanguardia0%0%unchanged
New York Post0%0%unchanged
Le Point0%0%unchanged
Wired0%0%unchanged
La Presse0%0%unchanged
USA Today0%95%improved

Read the two directions differently. IMATAG sampled every image wider than 400 pixels on a site's home page and article pages — between 1,000 and 7,000 images per site. We sample one lead photograph per article. A lead photograph is the image on a page most likely to have come from an agency with its credit intact, so our percentages should sit above theirs for the same publisher, whatever the truth. That asymmetry is useful: increases here may be partly our sampling, but decreases happened despite it, and a zero cannot be manufactured by a favourable sample.

So the reliable findings are the falls and the flats. Four publishers that carried credit in 2019 now carry none: the Huffington Post UK, the Washington Post, stern.de and El Pais. In each case we found not merely a missing credit but no embedded metadata of any kind on any image we sampled. And eight titles sat at or near zero in 2019 and are still there — the New York Times, Wired, Le Point, La Vanguardia, the New York Post, Die Zeit, L'Express and the Guardian. Seven years, no movement. Whichever side of the line a newsroom was on in 2019, it is almost certainly still on it.

Who has actually fixed it

The falls are the easy story. The rises are harder to trust, because our lead-image sampling flatters them — but not all of them, and the exception matters. IMATAG's method captured every image over 400 pixels on an article page, which necessarily includes the lead photograph. So a publisher whose lead images carried credit in 2019 could not have scored zero in their study. Any title that reads zero for them and high for us has genuinely changed.

USA Today is the clearest case. IMATAG measured it at 0% in 2019. We measure it at 75%, 74%, 80% and 95% across our four 2026 runs — consistently, not as a one-month blip — and the images carry a credit line, a source, a creator and a copyright notice, with a licensor URL on nearly half. Somewhere between 2019 and now, somebody there fixed the pipeline and it has stayed fixed. Les Echos is the same shape: 6% then, 75–90% across our runs now.

Within our own short history, the cleanest recent change is the Chicago Tribune: zero in June, July and August, then 47% in September, with captions, creators, credit lines and even a PLUS data-mining preference appearing together. That is one month of data and could yet reverse, but it has the shape of a deliberate change rather than noise. Diena in Latvia made a similar jump earlier and has held it.

We would rather report this honestly than cheerfully: across the same four runs, nine publishers went from nothing to something, and eighteen went from something to nothing. Improvement is real, it is achievable, and it is currently outnumbered two to one.

Against IPTC's own 2019 crawler

The 2019 IPTC prototype checked 200 news feeds across 35 countries. Across the 30 of those we still cover, the mean share of images carrying any IPTC field has risen from 4.2% to 10.2% (8.2% counting only fields that no Exif tag can satisfy, which is closer to what that tool measured). Treat it gently: the 2019 run reported exactly 0.00% for the Netherlands, Norway, Ireland, Finland and Switzerland, which is far more likely to reflect a prototype's blind spots — it read only RSS and Atom feeds and skipped images under 150 pixels — than five countries embedding nothing at all. A baseline biased low flatters its successor.

Every figure here is recomputed from the published Parquet files, so anyone preferring a different definition can redo it. If anyone at IMATAG would like to compare methodologies properly, we would welcome it — office@iptc.org.

The score

For every image we sample, we check the "Four Cs" of news photo provenance — fields drawn from IPTC IIM and IPTC XMP that long-standing wire-service training treats as universally applicable. A field counts as present if it appears in either family with a non-empty value. The image's score is the sum of weights for fields present, expressed as a percentage of 100. A site's score is the mean across all images we sampled for it; a country's score is the equal-weighted mean across that country's sites.

Scored fieldWeight
Creator 25
Copyright 25
CaptionDescription 25
CreditLine 25
Total100

A further 19 IPTC fields are tracked but do not affect the score: we record their presence and report it on the per-field breakdown, but absence isn't penalised. Many of these are legitimately omitted — LocationCreated can endanger sources, studio shoots and archival images often genuinely lack DateCreated or Keywords, and so on. We don't want to confuse "didn't supply" with "supplied badly".

Tracked fields: AIPromptInformation , AIPromptWriterName , AISystemUsed , AISystemVersionUsed , AltTextAccessibility , DataMining , DateCreated , DigitalSourceType , ExtendedDescriptionAccessibility , Genre , Keywords , LicensorName , LicensorURL , LocationCreated , LocationShown , ObjectName , Source , UsageTerms , WebStatement .

How a site is sampled

For each site in our list we try a small chain of discovery sources, in order, and use the first one that yields articles within the last 30 days:

  1. RSS feed(s) found on previous attempts.
  2. RSS feed(s) we auto-discover via <link rel="alternate"> on the homepage.
  3. Common RSS paths we guess (/feed, /rss.xml, etc.).
  4. XML sitemap declared in robots.txt — preferring news sitemaps, then any non-image sitemap. Image-only sitemaps are a last resort because they carry no article URLs.
  5. Common sitemap paths we guess (/sitemap-news.xml, /sitemap.xml, etc.). Gzipped sitemaps (.xml.gz) are supported.

From the chosen source we take up to 20 of the most recent articles, fetch each one, and extract a single "lead" image per article — either from JSON-LD NewsArticle / Article markup or from <meta property="og:image">. Inline body images and related-story thumbnails are deliberately skipped: they often pick up site chrome, ads, or lazy-load placeholders rather than the photographs the article was built around. Also, we encourage site owners to keep metadata on higher-resolution "lead" images, but we understand if site owners want to strip metadata from smaller images for bandwidth optimisation purposes.

Each image is then fetched and analysed with ExifTool; the field presence and CDN attribution flow into the per-image, per-article and per-site database tables.

What each site status means

Politeness

Metawatch identifies itself as Metawatch/2.0 with a From: metadata-crawler@iptc.org header. It honours robots.txt, respects Crawl-delay, and waits at least one second between requests to the same domain. If a publisher declares a long crawl delay we shrink the article sample so the whole site finishes within a fixed time budget, rather than hammering them at the edge of their allowance. We do not bypass paywalls, fingerprints, or CDN bot-management rules.

Site operators can find everything about the crawler — exactly what it fetches, how to verify its requests, and how to block or slow it — on About our crawler.

How often we run

Metawatch runs automatically once per month on the first of the month at 02:00 UTC. Each run commits its Parquet output and the exported JSON for this site to the project repository. We may also trigger ad-hoc runs while we are iterating on the crawler — those appear as additional run directories in data/runs/.

Data and source code

Every run's full Parquet output is published under data/runs/ in the project repository. The Dataset page lists every run with file sizes, direct download links, and a quick-start Python snippet. The data is licensed CC BY 4.0; the crawler source code is MIT-licensed. Citations and academic use are welcome — please credit "IPTC Metawatch" with a link back.

For publishers

If you would like to be removed from the crawled list, or to correct your site's entry — a better feed URL, a missing sitemap, a name fix — please email office@iptc.org. We honour removal requests without question.

Caveats and limitations

Latest run: 2026-10-01T023219Z — 425 of 508 sites successfully crawled, 6,580 images analysed.