Skip to main content

Module diff

Module diff 

Source
Expand description

Content verification primitives for faucet verify (#701).

A pipeline can report green while its destination quietly diverges from the source: a missed change event, a hand edit downstream, a retried page that landed twice. Row counts (reconcile:) cannot see a row that exists on both sides with different values. This module holds the pure building blocks the verifier composes:

  • KeyRange — a half-open slice of an integer key space, bisectable, and convertible into the PK-range ShardSpec every SQL source already understands (so a range read needs no new connector code);
  • Normalizer + row_hash — a canonical, type-tolerant fingerprint of one row’s compared columns, so 1 and 1.0, or two spellings of the same instant, hash alike;
  • DigestAccumulator / ContentDigest — an order-independent digest of a set of rows (count + folded row hashes + observed key bounds), computed client-side while streaming;
  • ServerDigest — the same shape computed inside the backend (only comparable when both sides report the same algorithm);
  • diff_rows — the row-level comparison of two leaf ranges, keyed.

Nothing here performs I/O; the CLI (cli/src/verify/) drives the sources.

Structs§

ContentDigest
An order-independent digest of a set of rows.
Difference
One differing key.
DigestAccumulator
Streams rows into a ContentDigest.
KeyRange
A half-open integer key range [lo, hi); None = unbounded on that side.
Normalizer
How values are canonicalised before hashing and comparing, so two backends’ spellings of the same value agree.
ServerDigest
A digest a backend computed itself, without shipping rows. Two server digests are comparable only when their algorithm ids match — each backend hashes its own text rendering of a row, so a Postgres digest says nothing about a SQLite table. When the algorithms differ the verifier streams both sides and uses ContentDigest instead.
VerifyReport
Summary of one verification, serialisable for --json and the HTTP API.

Enums§

DifferenceKind
Why one key differs between the two sides.
Side

Functions§

diff_rows
Compare two leaf ranges row by row. Returns every differing key, ordered by key text so the report is stable. Rows on either side are matched on the canonical key text; a key present more than once on a side is reported as a DifferenceKind::Duplicate and not matched.
is_excluded
Whether a column is excluded: an exact name, or a prefix* glob.
key_int
The integer value of a single-column key, when it is one.
key_text
The canonical text of a record’s key tuple (k1=…;k2=…), used to match rows across the two sides.
parse_timestamp_micros
UTC microseconds for a timestamp string in RFC 3339 form or the YYYY-MM-DD HH:MM:SS[.fff][Z|±hh:mm] form SQL backends emit (a naive timestamp is read as UTC). None when the string is not a timestamp.
plan_ranges
Split [min, max] into target contiguous ranges tiling the whole key space (the first is unbounded below, the last unbounded above), via the same planner PK-range sharding uses.
row_hash
A 128-bit fingerprint of record’s compared columns: the key columns plus columns (every non-excluded field when columns is None), each canonicalised through norm. Column order does not matter — the fields are hashed in sorted name order.