vole-document 0.1.0-alpha.20

Persistent procedural document runtime: byte-exact reconstruction plus a content-addressed procedural seed DAG, queryable observations with provenance, and selective late materialization.
Documentation
# ADR-0018: Partial materialization is a scoped random-access query-cost result, measured against generic compressors

- **Status:** Accepted — scoped positive on decode CPU, with an
  explicit no-I/O-win caveat (Phase 7.3)
- **Date:** 2026-10-05

## Context

Whole-file compression loses to generic tools on every file tested (ADR-0017),
so VOLE's pivoted claim is *random-access / query cost*: serving one narrow output
range without materializing the whole document. A descriptor may carry an
optional `OBSERVATION_INDEX` record (`observation_index_v1`) and the `view` CLI
serves a byte range, a PDF indirect object, an encoded stream, or a revision,
reporting measured `ObservationStats` (`src/materialize/observation.rs`).

The claim is falsifiable and must be measured on a large, multi-object document
against the honest alternative: a sequential decoder must inflate the whole
prefix to reach a late output offset.

## Decision

- **Adopt the mechanism for query serving; claim only decode work.**
  `materialize_observation` skips ops that do not intersect the requested range
  when the program is a sequence of linear, independent block producers (the
  replay lane's `INLINE · DEFLATE_REPLAY · INLINE` shape qualifies), and decodes
  only the referenced entropy channels. The index is advisory and re-checked
  against the authoritative program at parse time; every served slice is
  byte-compared against the full materialization.
- **The measurement corpus is a deterministic, large, multi-object PDF**
  (`pdf-make-large DIR [OBJECTS]`, default 800 pages/streams, ≥ 32 MiB, distinct
  per-stream zlib content). The regenerable bytes are gitignored; the ledger is
  committed.
- **Two axes, two verdicts.** Whole-file size (ADR-0017) is reported separately
  and remains a loss. Only the query-cost court is a potential win.
- **Do not claim an I/O win while v1 reads the whole descriptor.** The CLI reads
  the entire `.voldoc` into RAM (`fs::read`) and parses it, so
  `descriptor_bytes_traversed` is a **CPU-side approximation**, not a bytes-read
  figure. Until an mmap/seek reader lands, the honest metric is decode CPU, and
  the receipt must say so.

## Consequences

- **Measured result (Phase 7.3, receipt
  `evidence/campaigns/2026-10-05-phase7-partial-a5764c9/`):** on a 33,789,340 B
  (32.22 MiB), 800-stream, 800-replayed PDF, all **18/18** pre-registered queries
  are byte-exact. For mid/late queries the indexed lane touches ~0.41–0.43 MB
  (`descriptor_bytes_traversed` alone — its `entropy_bytes_decoded` breakdown is
  a subset already counted there and must not be added) regardless of offset,
  versus gzip inflating `a+len`: in the late region (≥ 50 % in) it is
  **1.4–2.3 %** of gzip's bytes and **~2–5× faster** than gzip / **~4–13×**
  faster than xz on CPU. Contract hypothesis H1 is met vs gzip and xz (8/8 late
  queries) and **not** met vs zstd.
- **Where it loses (recorded, not hidden):** in the early region (≤ ~8–16 MiB)
  the constant ~0.02–0.03 s / ~38 MB descriptor parse loses to gzip/xz/zstd, which
  process only 256 B–8 MB; **zstd's raw decompressor is faster than VOLE at every
  point** (≤ 0.01 s), at ~36 MB RSS; VOLE's peak RSS is ~38 MB versus gzip's
  ~1.2 MB; and whole-file size is still **2.98×** xz.
- **The decisive v1 caveat:** on-disk I/O is **not** reduced. At 31 MiB gzip
  reads a ~9.8 MB compressed prefix while VOLE reads its entire **17.5 MB**
  `.voldoc`. There is **no bytes-read win**. `descriptor_bytes_traversed` tracks
  bytes the path *needed* on the CPU side, not bytes fetched from storage.
- **Prerequisite for the next claim:** an **mmap/seek descriptor reader** so the
  index's record directory can be walked without materializing the whole
  container, making `descriptor_bytes_traversed` a real I/O figure. Only then can
  a "bytes read to serve a query" win be claimed; until then any such statement
  would be a CPU-cost claim in I/O clothing.
- **Neutral/degenerate regime acknowledged:** the stress corpus uses level-0
  (stored) zlib streams, so corrections are tiny (28 B/stream) and the corpus is
  the weak-producer regime. Distinct per-stream plaintext is what makes the
  query court non-degenerate; it is not a survey of strong real-world DEFLATE
  (cf. ADR-0017's producer corpus).

## References

- `src/materialize/observation.rs` (`materialize_observation`,
  `ObservationStats`), `src/container/observation.rs` (`observation_index_v1`),
  `src/adapter/pdf/replay.rs` (`propose_pdf_deflate_replay_rans_indexed`)
- `src/main.rs` (`pdf-make-large`, `view`), `tools/partial-court.sh`,
  `tools/partial-table.jq`
- Receipt `evidence/campaigns/2026-10-05-phase7-partial-a5764c9/`
- `docs/evidence/phase7-partial-report.md`
- ADR-0017: generic lossless compressors are the whole-file comparator