Per-field breakdown

The headline "% with any IPTC" hides a lot of detail. Below is the population rate of each IPTC field across the 6,680 images analysed in the latest crawl. Field names link to their definitions in the IPTC Photo Metadata Standard. The "scored" badge marks the four fields that contribute to a site's overall score (the "Four Cs" of news photo provenance); the others are tracked for context but don't affect the score — see the methodology for why.

Field Type Present Total Population Score weight
Creator scored 513 6,653 7.7% 25%
CreditLine scored 402 6,653 6% 25%
DateCreated tracked 401 6,653 6%
Copyright scored 369 6,653 5.5% 25%
ObjectName tracked 369 6,653 5.5%
CaptionDescription scored 323 6,653 4.9% 25%
Source tracked 289 6,653 4.3%
LocationCreated tracked 256 6,653 3.8%
Keywords tracked 188 6,653 2.8%
WebStatement tracked 107 6,653 1.6%
LicensorURL tracked 97 6,653 1.5%
UsageTerms tracked 47 6,653 0.7%
DataMining tracked 37 6,653 0.6%
Genre tracked 17 6,653 0.3%
DigitalSourceType tracked 7 6,653 0.1%
LicensorName tracked 3 6,653 0%
AIPromptInformation tracked 0 6,653 0%
AIPromptWriterName tracked 0 6,653 0%
AISystemUsed tracked 0 6,653 0%
AISystemVersionUsed tracked 0 6,653 0%
AltTextAccessibility tracked 0 6,653 0%
ExtendedDescriptionAccessibility tracked 0 6,653 0%
LocationShown tracked 0 6,653 0%
Total 100%

Bars are scaled to the most populated field, not to 100%, so differences within the long tail stay visible. The score weight column reflects the per-field weights in config/scoring.yaml: the four scored fields are weighted 25% each. Tracked fields show "—" because their presence does not contribute to the score.

What this tells us

Even the best-populated field — Creator at 7.7% — appears on fewer than one in ten images we fetched. The fields most relevant to licensing (WebStatement, LicensorURL) are below 3%. The dataset as a whole tells a consistent story: publishers and CDNs are discarding metadata that the photographers and agencies almost certainly embedded upstream.