Expand description
Bounded, byte-oriented access to PostgreSQL custom-format (pg_dump -Fc) archives.
pgdumpx opens supported archive versions 1.14 through 1.16, parses the header and
table of contents eagerly, and keeps entry payload access lazy. Selected entries are
validated before they are exposed and are decompressed as streams through
EntryDataReader. The default feature set supports PostgreSQL’s none, gzip, LZ4,
and Zstandard archive compression modes.
§Byte-oriented contract
Archive metadata and COPY fields are bytes first. ArchiveString::as_bytes, table
lookup, column lookup, Row, and FieldRef do not require UTF-8. Callers opt into
text conversion with fallible helpers such as ArchiveString::to_str and
Column::name_str. FieldRef::Bytes and OwnedField::Bytes contain logical
field bytes after PostgreSQL COPY-text backslash decoding; they are not the escaped
on-wire spelling. \N is represented as FieldRef::Null rather than as bytes.
§Lending rows
CopyRowReader and TableRowReader reuse their current-row storage. A Row
therefore borrows from the reader and remains valid only until that reader is mutably
borrowed again. The row readers intentionally expose next_row(&mut self) instead of
implementing Iterator. OwnedRow is used when a matched row must outlive reader
advancement; normal iteration stays borrowed.
§Three independent limit classes
Resource controls are intentionally separated because they protect different work:
Limitscontains finite structural bounds used while opening archives and parsing individual COPY rows/column layouts.ScanLimitsbounds total rows and parser-consumed decompressed COPY bytes for a row scan. These limits do not measure decoder read-ahead that the COPY parser has not consumed.EntryReadLimitsbounds decompressed bytes returned by raw selected-entry access. Crossing a raw-output limit is an error, never successful truncation.
Limits::default_compatible supplies the finite Limits default. In contrast,
ScanLimits::unlimited and EntryReadLimits::unlimited are also their respective
library defaults so trusted callers can choose operation policy explicitly. The
pgdumpx extract CLI applies its own finite 1 GiB default when its raw-output option
is omitted.
§Sequential row search
CopyRowReader::find_first and TableRowReader::find_first are sequential scans
from the reader’s current position. TableRowReader::find_first_equal adds a reusable
named-column exact-equality convenience layer without changing that predicate API or
scan path. Column names and FieldRef::Bytes targets are exact bytes, targets are
compared after COPY-text unescaping, and FieldRef::Null remains distinct from empty
or literal b"\\N" bytes. No collation, SQL coercion, or typed-value comparison occurs.
A freshly created TableRowReader starts at the beginning of the selected
TABLE DATA entry, but there is no row-level value index in the archive. An early match
terminates immediately; a late or absent match can process the rest of the selected
entry unless ScanLimits stops it. Worst-case unrestricted work is proportional to
the remaining selected table-data size.
§Raw extraction and partial output
Archive::copy_entry_to streams decompressed bytes incrementally. If a later input,
decompression, limit, or destination error occurs, bytes already accepted by the
destination cannot be rolled back. The operation still returns an error and never
reports a partial stream as successful extraction.
ExtractionPlan::execute extends that same bounded raw path to multiple tables while
preserving the archive’s single mutable seek invariant. It completes metadata preflight
for every selector before requesting the first destination, then executes targets in
deterministic plan order. If a target fails, earlier completed outcomes are retained,
the current target may already have partial output, and later targets are not started.
§Typed errors
PgDumpError is the detailed error type. PgDumpError::category exposes a stable
high-level ErrorCategory, while helpers such as PgDumpError::dump_id,
PgDumpError::row_number, PgDumpError::byte_offset, and
PgDumpError::limit_context expose machine-readable context without parsing display
text. Variants wrapping lower-level I/O, decompression, output, COPY-input, or UTF-8
failures preserve those failures through std::error::Error::source.
§CLI encoding boundary
The Rust API remains byte-oriented. The pgdumpx find and pgdumpx extract CLI
commands accept UTF-8 schema/table arguments and use an exact SCHEMA.TABLE selector:
exactly one ASCII . separator, two non-empty components, and no SQL identifier
quoting. find also requires UTF-8 column/value arguments and compares their UTF-8
bytes with logical post-unescape field bytes. Raw extract output remains binary-safe.
§Example: production row path
The following compiles against the same public path used by the CLI. The search is a bounded sequential scan of one selected table-data entry.
use pgdumpx::{Archive, ColumnEqualityResult, FieldRef, ScanLimits};
use std::{fs::File, io::BufReader};
let file = File::open("backup.dump")?;
let mut archive = Archive::open(BufReader::new(file))?;
let mut rows = archive.table_rows(b"public", b"orders")?;
let limits = ScanLimits::unlimited()
.with_max_rows(100_000)
.with_max_decompressed_bytes(64 * 1024 * 1024);
let matched = rows.find_first_equal_with_limits(
limits,
b"order_number",
FieldRef::Bytes(b"123456"),
)?;
if let ColumnEqualityResult::Match(row) = matched {
assert!(!row.fields().is_empty());
}Structs§
- Archive
- A read-only PostgreSQL custom-format archive with eagerly parsed metadata.
- Archive
Header - Metadata parsed from a supported PostgreSQL custom archive header.
- Archive
String - An owned, byte-oriented archive metadata string.
- Archive
Timestamp - PostgreSQL’s raw broken-down archive creation timestamp fields.
- Archive
Version - The custom-archive format version stored in a PostgreSQL archive header.
- Bounded
Entry Data Reader - A streaming selected-entry reader with an optional decompressed-byte budget.
- Column
- One ordered PostgreSQL COPY column name derived from TOC metadata.
- Copy
RowReader - A lending, byte-oriented parser for PostgreSQL COPY text rows.
- DumpId
- A positive PostgreSQL dump identifier parsed from the archive TOC.
- Entry
Data Reader - A streaming, decompressed view of one validated custom-archive entry.
- Entry
Read Limits - Optional decompressed-byte budget for one raw selected-entry read.
- Extraction
Outcome - Successful result for one target in an
ExtractionPlanexecution. - Extraction
Plan - An owned reusable plan describing ordered table extraction intent.
- Extraction
Target - One archive-specific table-data target resolved for plan execution.
- Limit
Context - Typed limit, configured bound, and consumed/observed work for one failure.
- Limits
- Finite structural limits used while opening archives and parsing COPY metadata/rows.
- Metadata
Filter - An owned reusable filter over already-parsed archive TOC metadata.
- Metadata
Match - One metadata-only result produced by
Archive::filter_metadata. - Owned
Row - One owned COPY text row that can outlive its streaming reader.
- Resolved
Extraction Plan - Metadata resolved for one
ExtractionPlanpreflight against a specific archive. - Row
- A borrowed COPY text row backed by reusable parser storage.
- Scan
Limits - Optional total-work budgets for streaming COPY row scans.
- Table
Ref - A metadata-only view of a
TABLEand its optional relatedTABLE DATAentry. - Table
RowReader - A lending COPY-text row stream for one selected
TABLE DATAentry. - Table
Selector - An owned exact-byte selector for one PostgreSQL table identity.
- TocEntry
- Metadata for one archive table-of-contents entry.
Enums§
- Column
Equality Result - Result of a named-column exact-equality row search.
- Compression
- Compression algorithm declared by a supported custom archive header.
- Data
Location - The custom-archive data-location state associated with a TOC entry.
- Error
Category - Stable high-level categories for programmatic error handling.
- Extraction
Execution Error - Failure while executing an
ExtractionPlanagainst one archive. - Extraction
Plan Error - Errors produced while constructing a reusable
ExtractionPlan. - Field
Ref - A borrowed logical field from a PostgreSQL COPY text row.
- Owned
Field - An owned logical field copied from one matched COPY text row.
- PgDump
Error - Errors produced while reading PostgreSQL dump archives.
- Resource
Limit - A typed resource protected by a finite limit.
- Section
- The restore section assigned to a TOC entry by PostgreSQL.
- Table
Data Representation - The logical representation recorded for one table-data entry.