Expand description
PDF text extraction across classic xref, xref streams, cmap and CID fonts.
Part of the pith zero-dependency hashing suite: this crate depends
only on pith-digest and pith-inflate, so the whole suite resolves
without a single registry package.
§Scope
Opens a PDF, resolves objects through the cross-reference (classic
tables, xref streams with /W field widths and /Index runs,
/Prev incremental-update chains, and /ObjStm object streams), walks
the page tree and extracts text from content streams:
- text operators
Tj,TJ,',"plus the positioning opsTd,TD,Tm,T*;Dorecurses into Form XObjects (own resources override the page’s per name); - font decoding:
/ToUnicodeCMaps (bfchar,bfrangeincl. array destinations, multi-char strings and surrogate pairs), the five predefined encodings plus/Differences(AGL names +uniXXXX/uXXXXXX), and CID-keyed fonts (/Type0,/Identity-H/-Vand embedded CMap streams); - stream filters
FlateDecode(zlib + raw-deflate fallback),ASCII85,ASCIIHexwith/DecodeParmsPNG/TIFF predictors.
§Extraction rules (deterministic, documented)
- A
Td/TDwithty != 0, anyTmmoving vertically,T*,',"end the current line.Td/Tmwith a pure horizontal move of= 2 text units emits one space.
TJnumbers below-250(over a quarter em leftward gap) start a new word.- Pages join with
\x0c(form feed). Unmappable codes decode as U+FFFD (the glyph exists but cannot be mapped), never silently dropped.
§Refusals, never guesses
- Encrypted documents (
/Encryptin the trailer or xref stream) refuse extraction per page with object context. - CID fonts without
/ToUnicoderefuse — CID numbers are glyph ids, not codepoints. - Unsupported stream filters (LZW, DCT, JBIG2, JPX, CCITT, RunLength, Crypt) refuse naming the filter.
- A corrupt xref is rebuilt by scanning for
N G objheaders when possible (Document::xref_was_rebuiltreports it); only a file with no findable objects fails. - Every page failure carries the page number (
Error::Page); every object-level refusal carries the object (Error::Object).
Modules§
- cmap
- ToUnicode and encoding CMaps (PDF 32000-1 9.10.3).
- ffi
- The C ABI surface of
pith-pdf: the entry points the Python (ctypes), Node (koffi) and Go (cgo) SDKs bind through.
Structs§
Enums§
- Error
- The crate’s public error: wraps the suite
pith_digest::Errorwith the page or object where the failure happened. - Obj
- A parsed PDF object.
Dictkeeps file order: dictionary keys repeat legally and/Index-style array payloads rely on position, not sorting.
Functions§
- decode_
stream - Apply a stream object’s
/Filterchain to its raw bytes. - extract_
text - Extract all text from a PDF byte buffer (pages joined by
\x0c).
Type Aliases§
- Result
- The crate’s result type.