Skip to main content

Module pdf

Module pdf 

Source
Expand description

PDF.

The hardest of the Phase 1 formats, and the one where hidden data is least likely to be where you look for it. The Document Information Dictionary is the easy part and the part every tutorial covers; the leaks that matter live in XMP packets, per-object metadata, annotation authorship, and — above all — in objects left behind by incremental updates, which are physically present in the file and reachable with a hex editor long after the “current” version of the document stopped referring to them.

§Full rewrite, not incremental patching (ADR-0020)

A PDF can be edited by appending: the original bytes stay, and a new cross-reference section at the end says which objects supersede which. Nulling the Info dictionary with another such append is easy, fast, and preserves the file almost perfectly — and it leaves every previous author name exactly where it was, four kilobytes up the file. For this tool that is not a lesser fix, it is a silent failure: the user is told the document is clean and publishes it.

So the document is parsed into its object graph, scrubbed, pruned to what the catalogue can actually reach, renumbered, and written out fresh. Everything unreachable — every superseded revision — is gone because it is never written, not because it was overwritten.

The cost is honest and worth stating: the output is not byte-comparable with the input, object numbering changes, and files using features the rewrite cannot faithfully reproduce are refused rather than mangled. Refusing is the correct half of that trade.

Structs§

PdfHandler
Removal of metadata from PDF documents.