Skip to main content

pgdumpx/
lib.rs

1//! Bounded, byte-oriented access to PostgreSQL custom-format (`pg_dump -Fc`) archives.
2//!
3//! `pgdumpx` opens supported archive versions 1.14 through 1.16, parses the header and
4//! table of contents eagerly, and keeps entry payload access lazy. Selected entries are
5//! validated before they are exposed and are decompressed as streams through
6//! [`EntryDataReader`]. The default feature set supports PostgreSQL's none, gzip, LZ4,
7//! and Zstandard archive compression modes.
8//!
9//! # Byte-oriented contract
10//!
11//! Archive metadata and COPY fields are bytes first. [`ArchiveString::as_bytes`], table
12//! lookup, column lookup, [`Row`], and [`FieldRef`] do not require UTF-8. Callers opt into
13//! text conversion with fallible helpers such as [`ArchiveString::to_str`] and
14//! [`Column::name_str`]. [`FieldRef::Bytes`] and [`OwnedField::Bytes`] contain logical
15//! field bytes *after* PostgreSQL COPY-text backslash decoding; they are not the escaped
16//! on-wire spelling. `\N` is represented as [`FieldRef::Null`] rather than as bytes.
17//!
18//! # Lending rows
19//!
20//! [`CopyRowReader`] and [`TableRowReader`] reuse their current-row storage. A [`Row`]
21//! therefore borrows from the reader and remains valid only until that reader is mutably
22//! borrowed again. The row readers intentionally expose `next_row(&mut self)` instead of
23//! implementing `Iterator`. [`OwnedRow`] is used when a matched row must outlive reader
24//! advancement; normal iteration stays borrowed.
25//!
26//! # Three independent limit classes
27//!
28//! Resource controls are intentionally separated because they protect different work:
29//!
30//! - [`Limits`] contains finite structural bounds used while opening archives and parsing
31//!   individual COPY rows/column layouts.
32//! - [`ScanLimits`] bounds total rows and parser-consumed decompressed COPY bytes for a row
33//!   scan. These limits do not measure decoder read-ahead that the COPY parser has not
34//!   consumed.
35//! - [`EntryReadLimits`] bounds decompressed bytes returned by raw selected-entry access.
36//!   Crossing a raw-output limit is an error, never successful truncation.
37//!
38//! [`Limits::default_compatible`] supplies the finite `Limits` default. In contrast,
39//! [`ScanLimits::unlimited`] and [`EntryReadLimits::unlimited`] are also their respective
40//! library defaults so trusted callers can choose operation policy explicitly. The
41//! `pgdumpx extract` CLI applies its own finite 1 GiB default when its raw-output option
42//! is omitted.
43//!
44//! # Sequential row search
45//!
46//! [`CopyRowReader::find_first`] and [`TableRowReader::find_first`] are sequential scans
47//! from the reader's current position. [`TableRowReader::find_first_equal`] adds a reusable
48//! named-column exact-equality convenience layer without changing that predicate API or
49//! scan path. Column names and [`FieldRef::Bytes`] targets are exact bytes, targets are
50//! compared after COPY-text unescaping, and [`FieldRef::Null`] remains distinct from empty
51//! or literal `b"\\N"` bytes. No collation, SQL coercion, or typed-value comparison occurs.
52//!
53//! A freshly created [`TableRowReader`] starts at the beginning of the selected
54//! `TABLE DATA` entry, but there is no row-level value index in the archive. An early match
55//! terminates immediately; a late or absent match can process the rest of the selected
56//! entry unless [`ScanLimits`] stops it. Worst-case unrestricted work is proportional to
57//! the remaining selected table-data size.
58//!
59//! # Raw extraction and partial output
60//!
61//! [`Archive::copy_entry_to`] streams decompressed bytes incrementally. If a later input,
62//! decompression, limit, or destination error occurs, bytes already accepted by the
63//! destination cannot be rolled back. The operation still returns an error and never
64//! reports a partial stream as successful extraction.
65//!
66//! [`ExtractionPlan::execute`] extends that same bounded raw path to multiple tables while
67//! preserving the archive's single mutable seek invariant. It completes metadata preflight
68//! for every selector before requesting the first destination, then executes targets in
69//! deterministic plan order. If a target fails, earlier completed outcomes are retained,
70//! the current target may already have partial output, and later targets are not started.
71//!
72//! # Typed errors
73//!
74//! [`PgDumpError`] is the detailed error type. [`PgDumpError::category`] exposes a stable
75//! high-level [`ErrorCategory`], while helpers such as [`PgDumpError::dump_id`],
76//! [`PgDumpError::row_number`], [`PgDumpError::byte_offset`], and
77//! [`PgDumpError::limit_context`] expose machine-readable context without parsing display
78//! text. Variants wrapping lower-level I/O, decompression, output, COPY-input, or UTF-8
79//! failures preserve those failures through [`std::error::Error::source`].
80//!
81//! # CLI encoding boundary
82//!
83//! The Rust API remains byte-oriented. The `pgdumpx find` and `pgdumpx extract` CLI
84//! commands accept UTF-8 schema/table arguments and use an exact `SCHEMA.TABLE` selector:
85//! exactly one ASCII `.` separator, two non-empty components, and no SQL identifier
86//! quoting. `find` also requires UTF-8 column/value arguments and compares their UTF-8
87//! bytes with logical post-unescape field bytes. Raw `extract` output remains binary-safe.
88//!
89//! # Example: production row path
90//!
91//! The following compiles against the same public path used by the CLI. The search is a
92//! bounded sequential scan of one selected table-data entry.
93//!
94//! ```no_run
95//! use pgdumpx::{Archive, ColumnEqualityResult, FieldRef, ScanLimits};
96//! use std::{fs::File, io::BufReader};
97//!
98//! # fn main() -> Result<(), Box<dyn std::error::Error>> {
99//! let file = File::open("backup.dump")?;
100//! let mut archive = Archive::open(BufReader::new(file))?;
101//! let mut rows = archive.table_rows(b"public", b"orders")?;
102//!
103//! let limits = ScanLimits::unlimited()
104//!     .with_max_rows(100_000)
105//!     .with_max_decompressed_bytes(64 * 1024 * 1024);
106//! let matched = rows.find_first_equal_with_limits(
107//!     limits,
108//!     b"order_number",
109//!     FieldRef::Bytes(b"123456"),
110//! )?;
111//!
112//! if let ColumnEqualityResult::Match(row) = matched {
113//!     assert!(!row.fields().is_empty());
114//! }
115//! # Ok(())
116//! # }
117//! ```
118
119#![forbid(unsafe_code)]
120#![deny(rustdoc::broken_intra_doc_links)]
121
122mod archive;
123mod copy;
124mod copy_metadata;
125mod custom;
126mod entry;
127mod error;
128mod error_taxonomy;
129mod extraction_plan;
130mod file;
131mod io;
132mod limits;
133mod metadata_budget;
134mod metadata_filter;
135mod model;
136mod raw_entry;
137mod selector;
138mod table_rows;
139
140#[cfg(test)]
141mod archive_primitives_tests;
142#[cfg(test)]
143mod copy_tests;
144#[cfg(test)]
145mod error_tests;
146#[cfg(test)]
147mod extraction_plan_tests;
148#[cfg(test)]
149mod metadata_budget_tests;
150#[cfg(test)]
151mod metadata_open_tests;
152#[cfg(test)]
153mod raw_entry_tests;
154#[cfg(test)]
155mod scan_limits_tests;
156#[cfg(test)]
157mod table_rows_tests;
158
159pub use archive::Archive;
160pub use copy::{CopyRowReader, FieldRef, OwnedField, OwnedRow, Row};
161pub use copy_metadata::{Column, TableDataRepresentation};
162pub use entry::EntryDataReader;
163pub use error::PgDumpError;
164pub use error_taxonomy::{ErrorCategory, LimitContext, ResourceLimit};
165pub use extraction_plan::{
166    ExtractionExecutionError, ExtractionOutcome, ExtractionPlan, ExtractionPlanError,
167    ExtractionTarget, ResolvedExtractionPlan,
168};
169pub use limits::{Limits, ScanLimits};
170pub use metadata_filter::{MetadataFilter, MetadataMatch};
171pub use model::{
172    ArchiveHeader, ArchiveString, ArchiveTimestamp, ArchiveVersion, Compression, DataLocation,
173    DumpId, Section, TableRef, TocEntry,
174};
175pub use raw_entry::{BoundedEntryDataReader, EntryReadLimits};
176pub use selector::TableSelector;
177pub use table_rows::{ColumnEqualityResult, TableRowReader};