pith_pdf/lib.rs
1//! PDF text extraction across classic xref, xref streams, cmap and CID fonts.
2//!
3//! Part of the `pith` zero-dependency hashing suite: this crate depends
4//! only on `pith-digest` and `pith-inflate`, so the whole suite resolves
5//! without a single registry package.
6//!
7//! # Scope
8//!
9//! Opens a PDF, resolves objects through the cross-reference (classic
10//! tables, xref **streams** with `/W` field widths and `/Index` runs,
11//! `/Prev` incremental-update chains, and `/ObjStm` object streams), walks
12//! the page tree and extracts text from content streams:
13//!
14//! - text operators `Tj`, `TJ`, `'`, `"` plus the positioning ops `Td`,
15//! `TD`, `Tm`, `T*`; `Do` recurses into Form XObjects (own resources
16//! override the page's per name);
17//! - font decoding: `/ToUnicode` CMaps (`bfchar`, `bfrange` incl. array
18//! destinations, multi-char strings and surrogate pairs), the five
19//! predefined encodings plus `/Differences` (AGL names + `uniXXXX`/
20//! `uXXXXXX`), and CID-keyed fonts (`/Type0`, `/Identity-H/-V` and
21//! embedded CMap streams);
22//! - stream filters `FlateDecode` (zlib + raw-deflate fallback), `ASCII85`,
23//! `ASCIIHex` with `/DecodeParms` PNG/TIFF predictors.
24//!
25//! # Extraction rules (deterministic, documented)
26//!
27//! - A `Td`/`TD` with `ty != 0`, any `Tm` moving vertically, `T*`, `'`,
28//! `"` end the current line. `Td`/`Tm` with a pure horizontal move of
29//! >= 2 text units emits one space.
30//! - `TJ` numbers below `-250` (over a quarter em leftward gap) start a new
31//! word.
32//! - Pages join with `\x0c` (form feed). Unmappable codes decode as U+FFFD
33//! (the glyph exists but cannot be mapped), never silently dropped.
34//!
35//! # Refusals, never guesses
36//!
37//! - Encrypted documents (`/Encrypt` in the trailer or xref stream) refuse
38//! extraction per page with object context.
39//! - CID fonts without `/ToUnicode` refuse — CID numbers are glyph ids, not
40//! codepoints.
41//! - Unsupported stream filters (LZW, DCT, JBIG2, JPX, CCITT, RunLength,
42//! Crypt) refuse naming the filter.
43//! - A corrupt xref is **rebuilt by scanning** for `N G obj` headers when
44//! possible ([`Document::xref_was_rebuilt`] reports it); only a file with
45//! no findable objects fails.
46//! - Every page failure carries the page number ([`Error::Page`]); every
47//! object-level refusal carries the object ([`Error::Object`]).
48
49#![cfg_attr(not(feature = "std"), no_std)]
50// `unsafe` is denied everywhere except `ffi`, the C ABI surface the
51// language SDKs bind through: raw pointers exist only at that boundary,
52// and every exported function is a documented `unsafe extern "C"` fn.
53#![deny(unsafe_code)]
54#![deny(missing_docs)]
55
56extern crate alloc;
57
58pub mod cmap;
59mod content;
60mod document;
61pub mod ffi;
62mod font;
63mod lex;
64mod object;
65mod tables;
66mod xref;
67
68pub use document::{Document, Error, Result};
69pub use object::{Obj, Ref, decode_stream};
70
71/// Extract all text from a PDF byte buffer (pages joined by `\x0c`).
72///
73/// Equivalent to [`Document::open`] + [`Document::text`].
74pub fn extract_text(data: &[u8]) -> Result<alloc::string::String> {
75 Document::open(data)?.text()
76}