Lyonite
PDF metadata guide

What Is PDF Metadata?

PDF metadata is information saved with a PDF, such as its title, author name, dates and the software used to make it. You usually won’t see it on the page.

What PDF metadata tells you

A PDF can record a title, an author name, dates and the software used to create it. Those fields are metadata: information about the document, rather than the words and pictures on its pages.

For example, a proposal might show Author “Sample Team” and Creator “Microsoft Word”. That can help you find the right file or understand how it was made. It can also reveal a name you did not mean to share. A template can supply those values, so the Author field does not necessarily identify the person who wrote the proposal.

Why the same PDF can show different properties

Document properties (Info dictionary)

The traditional properties record stores fields such as Title, Author and dates. Software can also add custom fields, such as a company name. A PDF does not have to contain every field.

/Author (Jane Doe)
/Creator (Microsoft Word)
/Producer (PDF engine)

An additional metadata record (XMP)

XMP can record the same title and author, plus other details such as document IDs, language and rights information. It uses XML, a text format software can read and extend.

If a program updates only one record, the two can disagree. That is why a useful checker identifies where each value came from instead of quietly combining them.

Common PDF metadata fields

Field
Typical meaning
Title
A descriptive title for the document.
Author
The person or organization recorded as the document author.
Subject
A short description or subject classification.
Keywords
Terms used for search, categorization, or document management.
Creator
Usually the application that created the source document before PDF conversion.
Producer
Usually the application or PDF engine that generated the PDF.
Creation date
A recorded document creation timestamp.
Modification date
A recorded modification timestamp.

Can you trust the author and dates?

Read them as recorded values, not verified facts. Names can come from an account or a template. Dates may reflect an export, a software clock or a later rewrite. The fields can also be changed directly.

Consider them alongside the document’s contents and how you received it. If two records disagree, keep that disagreement visible. A digital signature can provide a different kind of check, but whether the signed bytes match and whether you trust the signer are separate questions.

Clearing metadata does not remove a name printed on a page. Comments, filled forms, photographs and attached files may also reveal information. Metadata removal is not content redaction.

What else can be inside a PDF?

Document properties are only part of a file. A photo can contain its own camera or location metadata. A comment can record a reviewer’s name. A PDF can also contain attachments, scripts and earlier saved versions. Those are other file contents; calling all of them “metadata” hides useful differences.

In our published test of seven metadata removers, the six external free tools left the GPS data in the sample’s embedded photograph. That result concerns the tested file and workflows, not every file those tools can process.

The broader question is: what information will travel with this copy? What’s hidden inside a PDF? explains the different places to look, with examples and links to the relevant checks.

How Lyonite inspects PDF metadata

Lyonite’s PDF Metadata Checker reads your file on your device. Summary shows common properties first. Details also lets you check comments, photos, attachments and saved versions, with PDF objects and raw data available under Advanced.

The report keeps current properties separate from older or unreferenced records. It also tells you which checks could not finish, so missing results do not look like a clean bill of health.

Inspect PDF metadata in your browser →

Keep reading

Technical references