Expand description
Content verification primitives for faucet verify (#701).
A pipeline can report green while its destination quietly diverges from the
source: a missed change event, a hand edit downstream, a retried page that
landed twice. Row counts (reconcile:) cannot see a row that exists on both
sides with different values. This module holds the pure building blocks
the verifier composes:
KeyRange— a half-open slice of an integer key space, bisectable, and convertible into the PK-rangeShardSpecevery SQL source already understands (so a range read needs no new connector code);Normalizer+row_hash— a canonical, type-tolerant fingerprint of one row’s compared columns, so1and1.0, or two spellings of the same instant, hash alike;DigestAccumulator/ContentDigest— an order-independent digest of a set of rows (count + folded row hashes + observed key bounds), computed client-side while streaming;ServerDigest— the same shape computed inside the backend (only comparable when both sides report the samealgorithm);diff_rows— the row-level comparison of two leaf ranges, keyed.
Nothing here performs I/O; the CLI (cli/src/verify/) drives the sources.
Structs§
- Content
Digest - An order-independent digest of a set of rows.
- Difference
- One differing key.
- Digest
Accumulator - Streams rows into a
ContentDigest. - KeyRange
- A half-open integer key range
[lo, hi);None= unbounded on that side. - Normalizer
- How values are canonicalised before hashing and comparing, so two backends’ spellings of the same value agree.
- Server
Digest - A digest a backend computed itself, without shipping rows. Two server
digests are comparable only when their
algorithmids match — each backend hashes its own text rendering of a row, so a Postgres digest says nothing about a SQLite table. When the algorithms differ the verifier streams both sides and usesContentDigestinstead. - Verify
Report - Summary of one verification, serialisable for
--jsonand the HTTP API.
Enums§
- Difference
Kind - Why one key differs between the two sides.
- Side
Functions§
- diff_
rows - Compare two leaf ranges row by row. Returns every differing key, ordered by
key text so the report is stable. Rows on either side are matched on the
canonical key text; a key present more than once on a side is reported as a
DifferenceKind::Duplicateand not matched. - is_
excluded - Whether a column is excluded: an exact name, or a
prefix*glob. - key_int
- The integer value of a single-column key, when it is one.
- key_
text - The canonical text of a record’s key tuple (
k1=…;k2=…), used to match rows across the two sides. - parse_
timestamp_ micros - UTC microseconds for a timestamp string in RFC 3339 form or the
YYYY-MM-DD HH:MM:SS[.fff][Z|±hh:mm]form SQL backends emit (a naive timestamp is read as UTC).Nonewhen the string is not a timestamp. - plan_
ranges - Split
[min, max]intotargetcontiguous ranges tiling the whole key space (the first is unbounded below, the last unbounded above), via the same planner PK-range sharding uses. - row_
hash - A 128-bit fingerprint of
record’s compared columns: the key columns pluscolumns(every non-excluded field whencolumnsisNone), each canonicalised throughnorm. Column order does not matter — the fields are hashed in sorted name order.