Skip to main content

Crate pgdumpx

Crate pgdumpx 

Source
Expand description

Bounded, byte-oriented access to PostgreSQL custom-format (pg_dump -Fc) archives.

pgdumpx opens supported archive versions 1.14 through 1.16, parses the header and table of contents eagerly, and keeps entry payload access lazy. Selected entries are validated before they are exposed and are decompressed as streams through EntryDataReader. The default feature set supports PostgreSQL’s none, gzip, LZ4, and Zstandard archive compression modes.

§Byte-oriented contract

Archive metadata and COPY fields are bytes first. ArchiveString::as_bytes, table lookup, column lookup, Row, and FieldRef do not require UTF-8. Callers opt into text conversion with fallible helpers such as ArchiveString::to_str and Column::name_str. FieldRef::Bytes and OwnedField::Bytes contain logical field bytes after PostgreSQL COPY-text backslash decoding; they are not the escaped on-wire spelling. \N is represented as FieldRef::Null rather than as bytes.

§Lending rows

CopyRowReader and TableRowReader reuse their current-row storage. A Row therefore borrows from the reader and remains valid only until that reader is mutably borrowed again. The row readers intentionally expose next_row(&mut self) instead of implementing Iterator. OwnedRow is used when a matched row must outlive reader advancement; normal iteration stays borrowed.

§Three independent limit classes

Resource controls are intentionally separated because they protect different work:

  • Limits contains finite structural bounds used while opening archives and parsing individual COPY rows/column layouts.
  • ScanLimits bounds total rows and parser-consumed decompressed COPY bytes for a row scan. These limits do not measure decoder read-ahead that the COPY parser has not consumed.
  • EntryReadLimits bounds decompressed bytes returned by raw selected-entry access. Crossing a raw-output limit is an error, never successful truncation.

Limits::default_compatible supplies the finite Limits default. In contrast, ScanLimits::unlimited and EntryReadLimits::unlimited are also their respective library defaults so trusted callers can choose operation policy explicitly. The pgdumpx extract CLI applies its own finite 1 GiB default when its raw-output option is omitted.

CopyRowReader::find_first and TableRowReader::find_first are sequential scans from the reader’s current position. TableRowReader::find_first_equal adds a reusable named-column exact-equality convenience layer without changing that predicate API or scan path. Column names and FieldRef::Bytes targets are exact bytes, targets are compared after COPY-text unescaping, and FieldRef::Null remains distinct from empty or literal b"\\N" bytes. No collation, SQL coercion, or typed-value comparison occurs.

A freshly created TableRowReader starts at the beginning of the selected TABLE DATA entry, but there is no row-level value index in the archive. An early match terminates immediately; a late or absent match can process the rest of the selected entry unless ScanLimits stops it. Worst-case unrestricted work is proportional to the remaining selected table-data size.

§Raw extraction and partial output

Archive::copy_entry_to streams decompressed bytes incrementally. If a later input, decompression, limit, or destination error occurs, bytes already accepted by the destination cannot be rolled back. The operation still returns an error and never reports a partial stream as successful extraction.

ExtractionPlan::execute extends that same bounded raw path to multiple tables while preserving the archive’s single mutable seek invariant. It completes metadata preflight for every selector before requesting the first destination, then executes targets in deterministic plan order. If a target fails, earlier completed outcomes are retained, the current target may already have partial output, and later targets are not started.

§Typed errors

PgDumpError is the detailed error type. PgDumpError::category exposes a stable high-level ErrorCategory, while helpers such as PgDumpError::dump_id, PgDumpError::row_number, PgDumpError::byte_offset, and PgDumpError::limit_context expose machine-readable context without parsing display text. Variants wrapping lower-level I/O, decompression, output, COPY-input, or UTF-8 failures preserve those failures through std::error::Error::source.

§CLI encoding boundary

The Rust API remains byte-oriented. The pgdumpx find and pgdumpx extract CLI commands accept UTF-8 schema/table arguments and use an exact SCHEMA.TABLE selector: exactly one ASCII . separator, two non-empty components, and no SQL identifier quoting. find also requires UTF-8 column/value arguments and compares their UTF-8 bytes with logical post-unescape field bytes. Raw extract output remains binary-safe.

§Example: production row path

The following compiles against the same public path used by the CLI. The search is a bounded sequential scan of one selected table-data entry.

use pgdumpx::{Archive, ColumnEqualityResult, FieldRef, ScanLimits};
use std::{fs::File, io::BufReader};

let file = File::open("backup.dump")?;
let mut archive = Archive::open(BufReader::new(file))?;
let mut rows = archive.table_rows(b"public", b"orders")?;

let limits = ScanLimits::unlimited()
    .with_max_rows(100_000)
    .with_max_decompressed_bytes(64 * 1024 * 1024);
let matched = rows.find_first_equal_with_limits(
    limits,
    b"order_number",
    FieldRef::Bytes(b"123456"),
)?;

if let ColumnEqualityResult::Match(row) = matched {
    assert!(!row.fields().is_empty());
}

Structs§

Archive
A read-only PostgreSQL custom-format archive with eagerly parsed metadata.
ArchiveHeader
Metadata parsed from a supported PostgreSQL custom archive header.
ArchiveString
An owned, byte-oriented archive metadata string.
ArchiveTimestamp
PostgreSQL’s raw broken-down archive creation timestamp fields.
ArchiveVersion
The custom-archive format version stored in a PostgreSQL archive header.
BoundedEntryDataReader
A streaming selected-entry reader with an optional decompressed-byte budget.
Column
One ordered PostgreSQL COPY column name derived from TOC metadata.
CopyRowReader
A lending, byte-oriented parser for PostgreSQL COPY text rows.
DumpId
A positive PostgreSQL dump identifier parsed from the archive TOC.
EntryDataReader
A streaming, decompressed view of one validated custom-archive entry.
EntryReadLimits
Optional decompressed-byte budget for one raw selected-entry read.
ExtractionOutcome
Successful result for one target in an ExtractionPlan execution.
ExtractionPlan
An owned reusable plan describing ordered table extraction intent.
ExtractionTarget
One archive-specific table-data target resolved for plan execution.
LimitContext
Typed limit, configured bound, and consumed/observed work for one failure.
Limits
Finite structural limits used while opening archives and parsing COPY metadata/rows.
MetadataFilter
An owned reusable filter over already-parsed archive TOC metadata.
MetadataMatch
One metadata-only result produced by Archive::filter_metadata.
OwnedRow
One owned COPY text row that can outlive its streaming reader.
ResolvedExtractionPlan
Metadata resolved for one ExtractionPlan preflight against a specific archive.
Row
A borrowed COPY text row backed by reusable parser storage.
ScanLimits
Optional total-work budgets for streaming COPY row scans.
TableRef
A metadata-only view of a TABLE and its optional related TABLE DATA entry.
TableRowReader
A lending COPY-text row stream for one selected TABLE DATA entry.
TableSelector
An owned exact-byte selector for one PostgreSQL table identity.
TocEntry
Metadata for one archive table-of-contents entry.

Enums§

ColumnEqualityResult
Result of a named-column exact-equality row search.
Compression
Compression algorithm declared by a supported custom archive header.
DataLocation
The custom-archive data-location state associated with a TOC entry.
ErrorCategory
Stable high-level categories for programmatic error handling.
ExtractionExecutionError
Failure while executing an ExtractionPlan against one archive.
ExtractionPlanError
Errors produced while constructing a reusable ExtractionPlan.
FieldRef
A borrowed logical field from a PostgreSQL COPY text row.
OwnedField
An owned logical field copied from one matched COPY text row.
PgDumpError
Errors produced while reading PostgreSQL dump archives.
ResourceLimit
A typed resource protected by a finite limit.
Section
The restore section assigned to a TOC entry by PostgreSQL.
TableDataRepresentation
The logical representation recorded for one table-data entry.