Last reviewed: Next review due:
What is document metadata and why does it matter?
Every digital file contains more than just the content you can see. Photos embed GPS coordinates, camera serial numbers, and timestamps. Word documents embed the author name, organisation, revision history, and comments. PDFs can embed all of the above plus creation software details and tracked changes.
For journalists, metadata is a double-edged tool: you can use it to verify documents and photographs from sources, but you must also strip it from documents you receive and publish, to prevent it from identifying your source.
Key tools
Free, command-line tool (GUI wrappers available). Run exiftool -all= filename to strip all metadata from a file. The definitive tool for metadata analysis.
Freedom of the Press Foundation tool. Converts documents to image-based PDFs in an isolated container, stripping all metadata and active content. Recommended for documents from unknown sources.
Tools > Redact > Mark for Redaction > Apply Redaction. This properly removes the text layer. Print-to-PDF does not redact — it only covers the text visually.
When saving as PDF, check the "Document properties" export option — disable including author, organisation, and creation date if available.
Red flags
- Publishing a photo from a source without checking for GPS coordinates in EXIF data.
- Drawing a black box over text in a PDF without using a proper redaction tool.
- Publishing a Word document (.docx) directly — it contains author and revision metadata.
- Sending a source photograph without stripping EXIF data that includes location or device serial number.
- Using screenshot as a "cleaning" method — screenshots can still embed device metadata.
Document cleaning checklist
- I have run ExifTool on all images to check for GPS coordinates before publishing or sharing.
- I have stripped all EXIF metadata from images that came from a sensitive source.
- I have processed documents from unknown sources through Dangerzone before opening them on my main device.
- For redacted PDFs: I have used Adobe Acrobat's Redact tool (not a black box drawn over text) to remove underlying text.
- I have checked Word document properties (File > Info > Properties) and removed author, organisation, and company fields before exporting to PDF.
- I have verified that the published version of a document does not contain hidden tracking pixels or watermarking patterns.
Source protection tools
Metadata cleaning is part of source protection. Assess your full source protection workflow.
Source Protection ChecklistCommon mistakes
- Assuming that cropping or resizing a photo removes EXIF data — it often does not.
- Using a black rectangle drawn in a PDF viewer as redaction — the text remains selectable.
- Opening untrusted documents on your main device before sanitising them.
- Forgetting that Google Docs and other cloud editors add their own metadata.
- Not checking for hidden tracked changes in received Word documents before relying on their content.
Related guides
Primary sources
Frequently asked questions
What is EXIF data and how can it expose a source?
What is the most common redaction mistake journalists make?
What is Dangerzone and how does it help?
Does publishing a Word document expose metadata about the author?
Related guides
Primary sources
- ExifTool — Metadata Reader & Editor— Phil Harvey
- Dangerzone — Document Sanitisation Tool— Freedom of the Press Foundation
- How to Remove Metadata from Your Files— Electronic Frontier Foundation
- Security Training for Journalists— Freedom of the Press Foundation
- Reducing Risk from Malicious Documents— National Cyber Security Centre
- Redacting Sensitive Content in PDFs— Adobe
- Journalist Security Guide— Committee to Protect Journalists