VOLE-Document
A persistent procedural document runtime with byte-exact reconstruction.
VOLE-Document inverse-proceduralizes a document into a bounded deterministic
reconstruction description β reconstruction structure, parameters/state, typed
residual channels, and typed rANS channels β persists the recovered procedural
state, and lets you query it directly (page text, structure, preview, streams,
objects, revisions, exact byte ranges) with a typed observation API, provenance,
and EXPLAIN/EXPLAIN ANALYZE. The exact original bytes are always recoverable.
The governing invariant of the exact profile is uncompromising:
descriptor: materialize(descriptor) == original_bytes
persistent field: materialize(field_root) == original_bytes
Parsing successfully, producing "the same" text, the same object graph, the same pages, the same rendering, or a canonical re-save are not substitutes.
Compression is an implementation detail here, not the product. rANS is the entropy substrate beneath the representation, never the procedural model. The point is a document that remains a queryable, reconstructible computational field rather than a re-opened file. See
SPEC.md,PROJECT_STATE.md, anddocs/.
π Findings
Phase 11 changed the axis. On whole-file size VOLE still loses to generic compressors and is not a compressor (0/27, ADR-0017). But the persistent procedural field now has scoped, receipted measured wins: a warm, cache-served narrow observation reads 8.4β8.9 KB / ~99 Β΅s versus the fair preprocessed SQLite baseline's 24,393 B / ~1 ms and Poppler's 136,128 B / ~8 ms; lifetime wall beats raw PDF tooling on 5/5 documents; and after deleting the source a new process still rematerializes the exact original (length + SHA-256 +
cmp). Cold narrow observations still lose on bytes, cross-document sharing loses to CDC/tar|xz, and an earlier warm-byte claim was falsified by the independent skeptic (the cache-served answer read had been excluded from the byte count) β seedocs/reviews/phase-11-skeptic-review.md. ReadFINDINGS.mdand ADRs 0017β0028.
Status
Current release: 0.1.0-alpha.13 (Phase 11 β persistent procedural document
field). Phase 11 turns the byte-exact representation into a persistent,
queryable procedural field: inverse-proceduralize early (object/stream/page
tree, including /ObjStm), persist the recovered state as a content-addressed
seed DAG (filesystem or EntropyFS backend), navigate it with a bounded
hierarchical observation index, and materialize as late as possible β a
narrow observation resolves its minimum dependency closure and, when warm, never
opens the descriptor at all.
Exactness is unchanged and remains the invariant, not a competitive win.
materialize(descriptor) == original_bytes and materialize(field_root) == original_bytes hold byte-for-byte (length + SHA-256 + cmp), including after the
source is deleted and across process restarts; the derived cache is disposable
and off the authority path.
Phase 11's measured wins are scoped and receipted (see
docs/phases/phase-11-results.md):
- Warm narrow observation (agent-interactive): 0 descriptor bytes read, 8.4β8.9 KB total overhead, ~99 Β΅s wall β versus the fair preprocessed-SQLite baseline (24,393 B / ~1 ms) and Poppler page text (136,128 B / ~8 ms) on the small/medium cases. This is an overhead + wall win; it requires a warm derived cache and it is reported honestly as such.
- Lifetime wall vs raw PDF tooling: win on 5/5 documents.
- Cross-process reuse: a repeated observation in a new OS process executes 0 seed nodes and reads 0 seed bytes.
- Partial descriptors: a cold narrow observation reads its record closure (~0.22β0.36 MB), not the whole ~17 MB descriptor (β98%).
Phase 11's recorded losses and corrections: the cold narrow-observation
byte court still loses to SQLite; lifetime vs the preprocessed DB crosses over on
only 2/5 documents (absent for large/producer docs); whole-file size loses (0/27);
finer-than-object sharing loses to tar|xz and strongest CDC
(ADR-0028); and the independent
skeptic falsified the original warm-byte-vs-SQLite claim
(review).
The pre-Phase-11 consolidation still stands for the axes it covers: through
Phase 10, the single-file representation stack did not beat purpose-built
baselines on any measured axis (ADR-0023,
FINDINGS.md). Phase 11 does not overturn that; it adds a new axis
β a persistent procedural field β on which scoped wins are now measured. Whole-file
size remains a recorded loss against generic lossless tools (0/27, ADR-0017);
Phase 7 measured a scoped random-access decode-CPU win (ADR-0018); Phase 8
measured a scoped random-access bytes-read win versus non-seekable
sequential codecs only (ADR-0019); Phase 9 added the cross-document store and
recorded a robust loss (ADR-0020/0021); Phase 10.1 added an optional encoder-only
governor that buys no bytes (ADR-0022). See the Phase-11 subsection below for the
field results, and ADRs 0024β0028 for the Phase-11 architecture.
Phase 9 (branch phase9, ADR-0020/0021) adds a cross-document
content-addressed object store and then measures the one axis a single-file
compressor structurally cannot serve. The store is exact (a store-backed
descriptor materializes the same bytes as its standalone form) and its three
accounting universes are kept permanently distinct. The measured result is a
recorded negative: over 37 locally generated files (5,579,469 B, 10
deliberate-sharing strata) the unique-reachable universe U = 3,369,900 B loses
to per-file min LZ (1,304,307 B) and to the strongest pinned
content-defined-chunk dedup (borg 1.2.4: deterministic 771,383 B raw;
compressed non-deterministic, 210,835β210,840 B). The negative is robust β
forcing the finest candidate in the current set (--force pdf-deflate-replay)
still gives U = 2,360,054 B, and a per-stratum oracle ~2,537,730 B β but its
size is partly an artifact of externalization granularity/candidate selection:
the auto winner emits only 0β1 objects per file. Under the auto candidate only
byte-identical opaque repeats win (repeat-bin, U = 133,048 B); forcing
PDF_DEFLATE_REPLAY (one object per deflate stream) flips the shared-payload
fixture to 264,139 B, a win over raw CDC (285,257 B) that still loses to
LZ (34,591 B) and compressed CDC (28,195 B). No current candidate emits more
than one object per file (ADR-0021, docs/evidence/phase9-store-report.md,
docs/evidence/phase9-skeptic-review.md). Cross-document sharing is store
amortization, never "compression".
| Area | State | Evidence |
|---|---|---|
Exact .voldoc container (framing, header, records) |
Implemented | src/container/, unit + conformance courts |
| Typed errors + stable exit codes | Implemented | src/error.rs |
| Centralized resource limits | Implemented | src/limits.rs |
| CRC32C framing + SHA-256 archival identity | Implemented | src/integrity.rs |
| Document Reconstruction Algebra (literal subset) | Implemented | src/dra/ |
| Coverage certificate (checked invariant) | Implemented | src/dra/program.rs |
| RAW exact opaque adapter | Implemented | src/adapter/opaque/ |
| Candidate complete-cost court + decode-before-commit | Implemented | src/encode/ |
CLI (encode/decode/verify/inspect/capabilities) |
Implemented | src/main.rs |
| Exact court over a mixed corpus | Measured | evidence/campaigns/ |
| Native rANS floor (order-0 / typed byte channels) | Measured | campaign 2026-10-05-phase2-f6af30b |
RLE candidate (REPEAT_LAST run-length) |
Measured | campaign 2026-10-05-phase2-f6af30b |
| BYTE_RANS candidate (order-0 byte channel) | Measured | campaign 2026-10-05-phase2-f6af30b |
| Entropy capsule (full decoder-entry state, not a seed) | Measured | ADR-0006; src/entropy/ |
| PDF lexical span cover (Phase 3.1) | Measured | campaign 2026-10-05-phase3-486aa17 |
| PDF byte-authoritative physical scanner (Phase 3.2β3.3) | Measured | campaign 2026-10-05-phase3-486aa17 |
| PDF incremental revision map (Phase 3.4) | Measured | campaign 2026-10-05-phase3-486aa17 |
| qpdf differential oracle court (oracle, never authority) | Measured | tools/pdf-oracle.sh; campaign 2026-10-05-phase3-486aa17 |
PDF lexical channel transposition (split/join) |
Implemented | src/adapter/pdf/channels.rs; tests/pdf_channels.rs |
INTERLEAVE_CHANNELS DRA op (DRA v3) |
Measured | src/dra/op.rs; campaign 2026-10-05-phase4-3840bc4 |
| Compact entropy model wire v2 (sparse/dense, smaller chosen) | Measured | src/entropy/model.rs; campaign 2026-10-05-phase4-3840bc4 |
Forced-candidate ablation (encode --force KIND) |
Measured | tools/phase4-court.sh; campaign 2026-10-05-phase4-3840bc4 |
PDF typed channels (PDF_CHANNELS) |
Recorded (rejected on corpus) | campaign 2026-10-05-phase4-3840bc4; exact but loses to BYTE_RANS on complete cost (ADR-0010) |
Positional DRA ops (MARK_OFFSET / EMIT_OFFSET, DRA v4) |
Implemented | src/dra/op.rs; tests/pdf_layout.rs |
PDF layout candidate (PDF_LAYOUT) |
Recorded (rejected on cost) | campaign 2026-10-05-phase5-7193001; byte-exact and predicts xref offsets/startxref, but loses to RAW/BYTE_RANS on DRA framing cost (ADR-0011) |
Packed segment framing (PACK_SEGMENTS, DRA v5) |
Implemented | src/dra/op.rs; opcode 0x08, one op + a compact varint item table over a single data object, amortizing per-segment framing |
| PDF layout on packed framing (layout-v2) | Recorded (beats RAW at scale, rejected vs BYTE_RANS) |
campaign 2026-10-05-phase5-4521778; packed+coalesced layout beats RAW on many.pdf (10,069 vs 10,215) but still loses to BYTE_RANS (5,181); residual stored literally (ADR-0012) |
PACKED_CHANNELS DRA op (DRA v6) |
Implemented | opcode 0x09; reconstructs from a data channel + a plan channel (serialized item table) with a declared output length validated at eval; universe β phase5-8 |
PDF layout + rANS (PDF_LAYOUT_RANS) |
Recorded β rejected vs BYTE_RANS |
campaign 2026-10-05-phase5-8-cf8048d; byte-exact, but wins 0 / loses 8 / declines 3 head-to-head (ADR-0013) |
Phase-5 forced-candidate court (--force pdf-layout) |
Measured | tools/phase5-court.sh; campaigns 2026-10-05-phase5-7193001 and 2026-10-05-phase5-4521778 |
Phase-5.8 forced-candidate court (--force pdf-layout-rans) |
Measured | tools/phase5-8-court.sh; campaign 2026-10-05-phase5-8-cf8048d |
Lexer stream opacity (stream+EOL opaque span) |
Adopted | campaign 2026-10-05-phase6-0d0bb79; stream-data bytes are a byte-authoritative span, so /FlateDecode stream spans are exact |
DEFLATE_REPLAY DRA op (DRA v8) |
Implemented | src/dra/op.rs; opcode 0x0A, explicit replay_codec tag (preflate-0.7.6-experimental), exact raw-DEFLATE replay from (plaintext, corrections) with a declared output length statically bounded before the engine runs, catch_unwind-isolated, mandatory feature bit (opt-in deflate-replay cargo feature) |
Exact DEFLATE replay, raw plaintext (PDF_DEFLATE_REPLAY) |
Recorded β rejected vs BYTE_RANS |
campaign 2026-10-05-phase6-0d0bb79; byte-exact, but on flate.pdf 56,736 vs BYTE_RANS 49,291 (plaintext β bitstream) |
Exact DEFLATE replay, shared rANS plaintext (PDF_DEFLATE_REPLAY_RANS) |
Adopted β first structural win (over BYTE_RANS, on a self-authored fixture) |
campaign 2026-10-05-phase6-0d0bb79; flate.pdf 36,102 vs BYTE_RANS 49,291 (β13,189 B). A per-mechanism result only: not a generic-compressor win, and its enabling condition is not produced by tested transformers (ADR-0015, docs/evidence/phase7b-skeptic-review.md) |
Phase-6 replay court (--force pdf-deflate-replay[-rans]) |
Measured | tools/phase6-court.sh; campaign 2026-10-05-phase6-0d0bb79 |
Producer-stratified Flate ratio harness (deflate-stats) |
Measured | tools/pdf-corpus.sh; amendment campaign 2026-10-05-phase7-corpus-b-c4eb77e (supersedes 2026-10-05-phase7-corpus-f1f8d26); 24/24 replayed, 0 declined; diagnostic only, no new candidate |
Generator-family Flate corpus (producers: ReportLab/Cairo/LibreOffice/pdfTeX) |
Measured (claim corrected 7.0c) | campaign 2026-10-05-phase7-producers-e071250; 87/87 replayed; PDF_DEFLATE_REPLAY_RANS beats BYTE_RANS on Cairo (58,711 β 34,574, β24,137 B) but the Cairo file is a repeated-identical-bytes harness artifact (six byte-identical streams), and generic LZ does ~2Γ better (gzip -9 17,382 B; xz -9e 16,852 B); the "first authoring-generator witness" claim is withdrawn; no candidate changed |
| Generic-compressor baseline ladder (gzip/zstd/xz/brotli vs VOLE) | Measured β the honest comparison | campaign 2026-10-05-phase7-baselines-7b9f662; tools/baselines.sh + opt-in baseline image; VOLE beats gzip/zstd/xz/brotli on 0/27 files (+460,320 B vs the best generic); prior "wins" were relative to the weak order-0 BYTE_RANS lane |
Coverage-guided fuzzing (cargo-fuzz, ten targets) |
Measured | campaign 2026-10-05-phase7-fuzz-ca6a92b; pinned nightly-bookworm-slim-2026-10-04 + cargo-fuzz 0.13.2; 9/10 targets zero-crash; two upstream preflate-rs findings (F1 mitigated + regression test; F2 contained on the decode path by process isolation, ADR-0016) |
Process-isolated DEFLATE replay (__replay-worker) |
Implemented | replay_bounded runs preflate in a child under RLIMIT_AS + a wall-clock timeout; knobs VOLE_REPLAY_WORKER/VOLE_REPLAY_MEM_MB/VOLE_REPLAY_TIMEOUT_MS; a library embedder without a worker keeps the in-process residual |
Partial materialization (OBSERVATION_INDEX + view) |
Measured β scoped positive (decode CPU, no I/O win) | campaign 2026-10-05-phase7-partial-a5764c9; 18/18 queries byte-exact; mid/late queries touch ~0.41β0.43 MB (descriptor_bytes_traversed alone; its entropy_bytes_decoded breakdown is a subset already counted there and must not be added) vs gzip inflating a+len (late region 1.4β2.3 %, ~2β5Γ faster than gzip, ~4β13Γ than xz); v1 reads the whole descriptor, so on-disk I/O is not reduced and it loses in the early region (β€ ~8β16 MiB) (ADR-0018) |
Deterministic large corpus generator (pdf-make-large) |
Tooling | encode-time subcommand; classic-xref PDF of OBJECTS (default 800) distinct zlib FlateDecode streams, β₯32 MiB, correct by construction; bytes gitignored |
Seek DIRECTORY + Read + Seek reader (view) |
Measured β scoped bytes-read win vs non-seekable sequential codecs | campaign 2026-10-05-phase8-seek-08de2a9; 18/18 queries byte-exact; seeked view reads a constant ~0.44β0.46 MB (GRAPH + OBSERVATION_INDEX + DIRECTORY floor), 12β21Γ fewer bytes than sequential gzip/zstd/xz prefixes for late queries; loses at offset 0 and early (ADR-0019) |
| Seekable/blocked random-access baseline (bgzip / blocked xz / pixz) | Recorded β the honest random-access comparison (falsifies βgeneral random-access winβ) | amendment to campaign 2026-10-05-phase8-seek-08de2a9 (seekable-report.md); late query VOLE 460,713 B vs bgzip 23,808 B (~19Γ), xz-64KiB 15,344 B (~30Γ), xz-1MiB 179,892 B (~2.6Γ), xz-4MiB 708,612 B (VOLE wins), pixz 2,810,832 B (VOLE wins); BGZF whole-file 7,995,600 B < 17,566,832 B; docs/evidence/phase8-skeptic-review.md |
| PDF structural adapters (Phases 7β8) | Planned | β |
Content-addressed object store (ObjectStore/EmbeddedStore, EXTERNAL_REF) |
Implemented | src/store/, tests/store.rs; ADR-0020; Id = BLAKE3-256, mandatory FEATURE_EXTERNAL_OBJECTS, externalize/hydrate, gc |
| Three accounting universes (standalone / unique-reachable / amortized) | Measured | campaign 2026-10-05-phase9-store-fdb2845; store account; Ξ£ amortized == unique reachable by construction |
| Cross-document store court (universes vs per-file LZ + generic CDC) | Recorded β store axis is a loss | campaign 2026-10-05-phase9-store-fdb2845 (ADR-0021); U = 3,369,900 B loses to per-file min LZ (1,304,307 B) and to CDC borg 1.2.4 (771,383 B raw deterministic / 210,835β210,840 B compressed, non-deterministic); the negative is robust (forced --force pdf-deflate-replay U = 2,360,054 B; per-stratum oracle ~2,537,730 B) but partly an externalization-granularity artifact; under the auto candidate only repeat-bin wins (133,048 B), while forcing PDF_DEFLATE_REPLAY flips shared-payload to 264,139 B (win over raw CDC, loss to LZ/zstd); no current candidate emits >1 object/file; 37/37 byte-exact; docs/evidence/phase9-skeptic-review.md |
EntropyFS store backend (EntropyFsStore, optional adapter) |
Implemented (optional, not the measured backend) | non-default entropyfs-store feature; list/remove decline (no per-blob delete); ADR-0008/0020 |
| DSFB search governance (Phase 10.1) | Implemented (encoder-only, dsfb-search) / search recorded negative |
campaign 2026-10-05-phase10-governor-d2b09c9; dependency-free dsfb-search = []; guided never worse than fixed (H1) and matches exhaustive on 8/8 holdout with β€ Β½ candidates (H2), but the fixed heuristic already equals exhaustive on every workload, so the search adds zero bytes (H3); zero decode authority proven (ADR-0022) |
| Partial materialization checkpoints (beyond v1) | Planned | v1 random-access view measured in 7.3 (ADR-0018); the seek reader landed in Phase 8 (ADR-0019), leaving checkpoint bytes as future work |
"Implemented" means the mechanism exists and is tested. "Measured" means there is
a sealed campaign under evidence/. The Phase-1 core establishes exactness,
framing, integrity, bounds, and receipts before any entropy or format-aware
mechanism is allowed to compete; Phase 2 then measures entropy channels on that
same exactness floor.
Phase 2 measured results
Phase 2 is order-0 typed byte channels only β no context model, no typed
residuals, and no format awareness. On a 9-file mixed corpus the cumulative
coreβfull ladder over serialized .voldoc bytes is:
sum_source = 590081
sum_core = 464474 (RAW + RLE)
sum_full = 291304 (RAW + RLE + BYTE_RANS)
delta = 173170 (sum_core - sum_full)
The entire delta is attributed to the two files where BYTE_RANS wins
(text-256k.bin 262144 β 143746; skewed.bin 65536 β 11382). RLE wins the
long runs (zeros-64k.bin 65536 β 283; runs.bin 65536 β 3088). Negative
controls hold: BYTE_RANS never wins on random 64 KiB (stored RAW at 65536 β
65845, a 309-byte fixed framing overhead) or on empty/one-byte inputs (RLE).
The canonical model's bytes are charged like any other bytes, so on tiny or
high-entropy inputs order-0 rANS loses to RAW/RLE as required β this is a scoped
measurement on one deterministic corpus, not a general compression claim.
The entropy substrate is optional in the build: default = ["rans"]. The exact
DEFLATE replay stack is opt-in (--features deflate-replay); a
channel-bearing or replay-bearing descriptor decoded without the required
feature returns an explicit UnsupportedFeature, never a silent
reinterpretation.
Receipt:
evidence/campaigns/2026-10-05-phase2-f6af30b/.
Phase 3 measured results
Phase 3 adds a byte-authoritative PDF physical scanner: an owned lexer, a
conservative structural span cover, /Length resolution, and an append-only
revision map. The sealed campaign 2026-10-05-phase3-486aa17 runs over a
deterministic 9-item corpus (7 valid PDFs plus malformed.pdf and notpdf.bin
as negative controls):
- Coverage β
all_covered = true: 171 spans, 19 objects, and 8 revisions across the corpus, partitioned into a contiguous cover of[0, len)with no gap and no overlap. - Byte-exactness β
all_exact = true:materialize(descriptor) == original_bytesfor every item, including both negative controls through the opaque RAW lane. - Validated detection β a file is a PDF only when the bytes show a
%PDF-header and an indirect object and a%%EOF; the extension is never authority, and both controls reportis_pdf = false. - Revision map β the incremental input yields two append-only revisions with
a
/Prevchain, and/Sizeis treated as never decreasing. - qpdf oracle β 100% object-number agreement with qpdf 11.3 (classic 4/4,
two-page 6/6, incremental 5/5);
qpdf --checkreports valid;pdfinfopages 1/2/1. Divergence is expected where objects are compressed inside object streams: those have no physicalN G objmarker, so a physical scanner enumerates fewer objects than qpdf's semantic view. qpdf is an oracle, never the byte authority.
The literal PDF candidate currently loses to RAW: RAW won all 9 items and
PDF_PHYSICAL won 0. This is the expected Phase-3 result β the physical lane
persists each span as one literal INLINE op with no structural compression, so
its per-span overhead loses once complete cost is charged. Structural
compression (xref//Length proceduralization, stream replay) is Phase 5+ and is
not claimed here.
Receipt:
evidence/campaigns/2026-10-05-phase3-486aa17/.
Phase 4 measured results
Phase 4 transposes the byte-authoritative lexical cover into typed channels
(one kind id per token, one 4-byte length per token, one payload stream per
lexical kind) and reconstructs them with the bounded INTERLEAVE_CHANNELS DRA op
(DRA v3). Each channel gets its own order-0 byte-rANS model, and model wire v2
serializes to whichever of sparse or dense is smaller. The sealed campaign
2026-10-05-phase4-3840bc4 runs a forced-candidate ablation (encode --force KIND) over a deterministic 10-file corpus:
A0 RAW = 71036
A1 + RLE = 71036
A2 + BYTE_RANS = 43297
A3 + PDF_PHYSICAL = 43297
A4 + PDF_CHANNELS = 43297 leave-one-out channel delta = 0
All 10 files round-trip byte-exactly (cmp + verify) and the qpdf oracle
re-check passes. Auto winners: RAW = 8, BYTE_RANS = 2, PDF_PHYSICAL = 0,
PDF_CHANNELS = 0. On the text-heavy scale sample bigtext.pdf (65,549 B) the
forced sizes were RAW = 65,871, BYTE_RANS = 38,142, PDF_CHANNELS = 46,432:
typed channels beat RAW by ~21.6% but lose to BYTE_RANS by ~8,290 B. Compact
sparse models cut the per-channel model overhead from 7,224 B to 1,981 B, which
was not enough to close the gap. The honest conclusion is that coarse lexical
transposition plus per-channel order-0 models does not beat a monolithic
order-0 BYTE_RANS on this corpus, so the typed lexical channels were rejected
by the complete-cost court and PDF_CHANNELS is recorded (not adopted). A win
would require conditioning and ordering rather than more marginal per-kind
models β that is Phase 5+ work (ADR-0010).
Receipt:
evidence/campaigns/2026-10-05-phase4-3840bc4/.
Phase 5 measured results
Phase 5 adds positional DRA ops (MARK_OFFSET / EMIT_OFFSET, DRA v4)
and the first candidate that replaces literal structural bytes with
procedurally determined ones: the classic cross-reference layout lane
(PDF_LAYOUT), which marks each indirect object's introducer offset and the
xref section start, then regenerates the 10-digit xref entry offsets and the
startxref value from those marks. The sealed campaign
2026-10-05-phase5-7193001 runs the forced-candidate ablation
(encode --force KIND) over a deterministic 10-file corpus:
A0 RAW = 71116
A1 + RLE = 71116
A2 + BYTE_RANS = 43377
A3 + PDF_PHYSICAL = 43377
A4 + PDF_CHANNELS = 43377
A5 + PDF_LAYOUT = 43377 leave-one-out layout delta = 0
All 10 files round-trip byte-exactly through their auto winner (cmp +
verify); auto winners are RAW = 8, BYTE_RANS = 2, and PDF_PHYSICAL /
PDF_CHANNELS / PDF_LAYOUT = 0. The prediction works and the descriptor is
exact β classic.pdf regenerates 3 of 4 xref entry offsets plus the
startxref, and incremental.pdf regenerates 5 of 7 entries plus two
startxref values β but the lane still loses to RAW and BYTE_RANS on
complete cost: classic.pdf 798 vs RAW 659, and bigtext.pdf 66,066 vs RAW
65,879 / BYTE_RANS 38,150. Layout wins 0 of the 7 classic-xref files. The
reason is framing, not prediction: the DRA pays a MarkOffset per object and
per xref section plus an EmitOffset per predicted entry, and that per-segment
op framing (tag + operand length) costs more than the ~7 digits saved per
predicted offset at document scale. The predicted structure is right; the
reconstruction container is too expensive. This is recorded honestly as a
negative result (ADR-0011) and a format-design input for later phases.
Receipt:
evidence/campaigns/2026-10-05-phase5-7193001/.
Packed framing (Phase 5.7) measured results
Phase 5.7 attacks the framing root cause directly. The PACK_SEGMENTS DRA op
(opcode 0x08, bumping the DRA graph to version 5, universe
phase6-prep;β¦;dra-5;β¦+packed) amortizes per-segment framing: instead of one
tagged op per literal run, it carries one op plus a compact varint item table
(Literal varint-length / Mark / Emit) over a single data object. The layout
candidate was rebuilt on this op (layout-v2), and literal coalescing (5.7.2b)
merges adjacent literal runs: on many.pdf (200 objects, 9,881 B) the item table
fell from 1,413 to 805 items.
The sealed campaign 2026-10-05-phase5-4521778 runs the forced-candidate
ablation over an 11-file corpus:
A0 RAW = 81371
A2 + BYTE_RANS = 48598
A5 + PDF_LAYOUT = 48598 leave-one-out layout delta = 0
Forced sizes:
many.pdf layout-v2 10069 RAW 10215 BYTE_RANS 5181 (layout beats RAW by 146 B)
classic.pdf layout 711 RAW 663
bigtext.pdf layout 65929 RAW 65883 BYTE_RANS 38154
Packed framing plus coalescing makes structural layout prediction beat RAW at
scale (many.pdf 10,069 vs 10,215), which the per-segment Phase-5 framing never
managed. It is still not adopted: it loses to BYTE_RANS (5,181), wins 0 of
the 8 classic-xref samples, and the leave-one-out layout delta is 0, so layout is
never the auto winner. The honest conclusion is that the framing is fixed, but the
residual data object is stored literally, so any order-0 entropy lane
dominates it. The remaining lever is to entropy-code the residual data object β
structural prediction composed with rANS on the residual, which is the paper's
layered model β not more literal packing. This is recorded as a partial positive
(ADR-0012).
Receipt:
evidence/campaigns/2026-10-05-phase5-4521778/.
Layout + rANS residual (Phase 5.8) measured results
Phase 5.8 builds the lever ADR-0012 named: the PACKED_CHANNELS DRA op (opcode
0x09, DRA v6, universe
phase5-8;β¦;dra-6;β¦+packed+packed-channels) reconstructs from a data channel
plus a plan channel (the serialized item table) with a declared output length
validated at eval, and the PDF_LAYOUT_RANS candidate codes the layout plan's
data object and its item table each as their own order-0 rANS channel. The sealed
campaign 2026-10-05-phase5-8-cf8048d runs the forced-candidate ablation over the
11-file corpus:
A0 RAW = 81591
A2 + BYTE_RANS = 48818
A6 + PDF_LAYOUT_RANS = 48818 leave-one-out layout+rANS delta = 0
Forced sizes:
classic.pdf RAW 683 BYTE_RANS 728 PDF_LAYOUT 731 PDF_LAYOUT_RANS 883
bigtext.pdf RAW 65903 BYTE_RANS 38174 PDF_LAYOUT_RANS 38341
many.pdf RAW 10235 BYTE_RANS 5201 PDF_LAYOUT 10089 PDF_LAYOUT_RANS 5914
On many.pdf the layout+rANS size breaks down as data 7,877, plan 1,815, models
645, payload 4,775. Head-to-head against BYTE_RANS: win 0, lose 8, declined
3. The honest conclusion: layout+rANS does not beat BYTE_RANS. Channel 0
codes nearly the whole file β the same job BYTE_RANS does with one channel β so
the plan channel (1,815 B on many.pdf) plus a second model are added metadata
BYTE_RANS never pays. Three phases (4, 5, 5.7) plus this one converge: at the
tested scale, PDF structural proceduralization does not beat a whole-file order-0
rANS lane.
Receipt:
evidence/campaigns/2026-10-05-phase5-8-cf8048d/.
Exact DEFLATE replay (Phase 6) measured results
Phases 4β5.8 all proceduralize plain syntax that BYTE_RANS already models
well. Phase 6 attacks a different layer: bytes the producer has already
entropy-coded. The DEFLATE_REPLAY DRA op (opcode 0x0A, DRA v8, universe
phase6;β¦;dra-8;β¦+deflate-replay-preflate-0.7.6-experimental) reconstructs the
original raw DEFLATE bitstream of a /FlateDecode stream from (plaintext, corrections). Its replay_codec tag names the correction representation as an
experimental, version-coupled preflate-0.7.6 layout (not frozen v1), and an
unknown tag fails closed. The
byte-authoritative scanner owns stream discovery and /Filter classification;
preflate never discovers streams. A lexer fix makes stream+EOL payloads opaque
spans. Two candidates use the op: PDF_DEFLATE_REPLAY (raw, deduplicated
plaintext objects) and PDF_DEFLATE_REPLAY_RANS (each unique plaintext is one
shared order-0 byte-rANS channel). The sealed campaign
2026-10-05-phase6-0d0bb79 runs the forced-candidate ablation over the 12-file
corpus:
A0 RAW = 139950
A2 + BYTE_RANS = 98560
A6 + PDF_LAYOUT_RANS = 98560
A7 + PDF_DEFLATE_REPLAY = 98560
A8 + PDF_DEFLATE_REPLAY_RANS= 85371 leave-one-out replay-rANS delta = -13189
Forced sizes on flate.pdf (57,513 B):
RAW 57908 BYTE_RANS 49291 PDF_DEFLATE_REPLAY 56736 PDF_DEFLATE_REPLAY_RANS 36102
PDF_DEFLATE_REPLAY_RANS is the auto winner on flate.pdf and beats
BYTE_RANS by 13,189 B. The six FlateDecode streams are the content
plaintext p1 at levels 9/6/1/0 (one shared plaintext, four appearances), a
graphics stream p2 at level 6, and an incompressible stream p3 at level 6
that DEFLATE stores; the descriptor replays all six and codes their 3 unique
plaintexts as order-0 channels (streams=6 replayed=6 channels=3 objects=6).
Of p1's four appearances only the level-0 stream is weakly coded (stored);
level 1 is ~19% of the plaintext and levels 6/9 are strong, so the win needs the
shared plaintext to also have a large/weakly-coded appearance. BYTE_RANS
order-0-codes the six streams to 49,291 B; it does not carry them verbatim. The
raw-plaintext variant loses (56,736 B) because a strongly-compressed stream's
plaintext is nearly as large as the stream it replaces. Head-to-head vs
BYTE_RANS: win 1, lose 0, decline 11 (the other files have no lone
FlateDecode stream); every auto winner is exact (cmp + verify). This is
one composed sample, at commit 0d0bb79, measured on a single synthetic
fixture: the win requires a shared plaintext that also has a large/weakly-coded
appearance (unique strongly-compressed streams lose, by up to 4.46Γ), and the
losing region is unique, strongly-compressed plaintext. It is the first measured
positive for a PDF structural candidate
(ADR-0015); the plain-syntax converging negatives (ADR-0010βADR-0013) stand.
Receipt:
evidence/campaigns/2026-10-05-phase6-0d0bb79/.
Producer-stratified Flate ratio (Phase 7.0) measured results
tools/pdf-corpus.sh builds a locally-generated corpus from distinct producer
lineages (Ghostscript 10.00.0 at five /PDFSETTINGS, qpdf 11.3.0 in four modes,
a hand-written stored-block-zlib base, plus the Phase-3 synthetic set), and
vole-document deflate-stats measures the exact-replay ratio per FlateDecode
stream β correction/compressed, (plaintext+corr)/compressed,
(rANS(plaintext)+corr)/compressed β with p10/p50/p90 by producer, the
exact-replay acceptance rate, and every decline. It also reports two aggregate
complete costs so shared plaintext is not overcounted:
replayed_rans_full_bytes (naive per-stream sum) and
replayed_rans_dedup_bytes (one charge per unique plaintext + one per unique
correction, matching the shared-channel candidate).
On 24 FlateDecode streams (re-measured after the Stage-A lexer stream-boundary
fix): 24 replayed, 0 declined (acceptance 1.000). The pre-fix run saw only 17
streams (11 replayed, 6 declined); the census rose because the old over-read had
swallowed whole stream objects (every qpdf-generated file was undercounted;
qpdf-preserve-objectstreams.pdf had reported zero). The Phase-6 win region
appears only in our own hand-authored fixtures (hand-base2.pdf: deduped
rANS 55,531 vs naive 111,062; flate.pdf: 34,051 vs 89,437). The same geometry
in qpdf-preserve-objectstreams.pdf (55,531 vs 111,062) is present only
because qpdf --object-streams=preserve copied and renumbered the two
byte-identical raw streams already authored in hand-base2.pdf (confirmed via
qpdf --raw-stream-data: all four stream hashes are ec028dc1β¦); the fixture
already wins 112,011 β 56,885 B and qpdf adds +41 B, so 99.93% of the reported
qpdf win is inherited. No genuinely transformed producer output exhibits the
region. Corpus-wide correction/compressed p10/p50/p90 = 0.000320 / 0.014716
/ 0.097360; the 0.004518 median is the pdf-make-samples subset only.
The six former declines were all not_zlib because of a scanner locality
limitation β the lexer's find_endstream required an EOL before endstream,
which Ghostscript omits (spec "should", qpdf-tolerated) β not because those
streams are not zlib (they begin 78 9c, confirmed via qpdf --raw-stream-data).
That limitation is fixed (commit c4eb77e): find_endstream now locates the
endstream keyword by right-termination. This is a scoped result about
locally generated files; qpdf/Ghostscript are transformers, not authoring apps,
and browser/office/TeX families remain a recorded gap. No candidate is adopted.
Receipt:
evidence/campaigns/2026-10-05-phase7-corpus-b-c4eb77e/
(amendment; supersedes the original
2026-10-05-phase7-corpus-f1f8d26/);
report: docs/evidence/phase7-corpus-report.md.
Complete-cost court (the decisive Phase-7.0 measurement). The ratio harness
is a diagnostic; the court decides. A second sealed campaign
(2026-10-05-phase7-court-99dc72e, commit 99dc72e, driver
tools/pdf-court.sh) runs the real CLI over all 23 locally generated corpus
files β unforced and with --force raw|byte-rans|pdf-deflate-replay|pdf-deflate-replay-rans
β and compares complete serialized .voldoc sizes. PDF_DEFLATE_REPLAY_RANS vs
BYTE_RANS: win 3, lose 8, decline 12 β all 3 wins self-authored. The wins
are exactly the Phase-6 shared-plaintext geometry, and in all three the unforced
court picks PDF_DEFLATE_REPLAY_RANS:
qpdf-preserve-objectstreams.pdf BYTE_RANS 112147 -> PDF_DEFLATE_REPLAY_RANS 56980 (-55167)
hand-base2.pdf BYTE_RANS 112011 -> PDF_DEFLATE_REPLAY_RANS 56885 (-55126)
_synthetic/flate.pdf BYTE_RANS 49291 -> PDF_DEFLATE_REPLAY_RANS 36102 (-13189)
The reported "qpdf win" is not a transformed-producer result:
qpdf --object-streams=preserve copied and renumbered the two byte-identical raw
streams (ec028dc1β¦) already authored in our hand-base2.pdf fixture (the
fixture wins 112,011 β 56,885 B; qpdf adds only +41 B, so 99.93% of the
55,167 B win is inherited). Every genuinely transformed producer output
loses or declines on this corpus: all 5 Ghostscript variants and both qpdf
compression variants (unique, strongly-compressed plaintext) lose, and the 12
files with no replayable Flate lane decline (a forced lane the input does not
propose is a typed Usage error, recorded null). So on this locally generated
corpus exact replay wins 3 / loses 8 / declines 12, and all 3 wins are
self-authored (two fixtures plus a preserved copy of one). --deterministic-id
is a /ID-only normalization (55,165 B no-flag vs 55,167 B flagged): it makes
the court reproducible but neither creates nor destroys the win. The Phase-6 win
is real and byte-exact, but its enabling condition (a plaintext shared across
streams with a large/weakly-coded appearance) is not produced by the tested
transformers, which motivates Phase 7.2 (nested content proceduralization). All
23 auto winners are verify + cmp byte-exact. The corpus is locally
generated and is not a population sample; qpdf and Ghostscript are
transformers, not authoring applications, and browser/PDFium, LibreOffice,
pdfTeX and Adobe outputs remain a recorded gap. No candidate is adopted and no
wire format changed.
Receipt:
evidence/campaigns/2026-10-05-phase7-court-99dc72e/.
Generator-family Flate corpus (Phase 7.0b) measured results
The Phase-7.0 gap was that qpdf and Ghostscript are transformers and never emit
the shared plaintext the win region needs. Phase 7.0b adds a separate, opt-in
producers image (Dockerfile stage + compose service, base
debian:bookworm-slim@sha256:3783cc01β¦, the same digest as tools; ~724 MB) with
four real authoring generators: ReportLab 3.6.12, Cairo 1.20.1/libcairo 1.16.0,
LibreOffice Writer 7.4.7.2, and pdfTeX 3.141592653-2.6-1.40.24. All four ran.
tools/pdf-corpus-producers.sh generates one PDF per family from the same
deterministic content document as tools/pdf-corpus.sh, fingerprints every
/FlateDecode payload, and applies
qpdf --deterministic-id --stream-data=preserve --object-streams=preserve only when
it leaves every payload byte-identical (ReportLab, Cairo, LibreOffice; pdfTeX kept
raw). ReportLab/Cairo/pdfTeX are byte-reproducible; LibreOffice is not.
deflate-stats: 87 Flate streams, 87 replayed, 0 declined β acceptance 1.000;
corpus-wide correction/compressed p10/p50/p90 = 0.034759 / 0.047945 / 0.068028.
Complete-cost court: PDF_DEFLATE_REPLAY_RANS vs BYTE_RANS β win 1 / lose 3 /
decline 0:
cairo-vector.pdf BYTE_RANS 58711 -> PDF_DEFLATE_REPLAY_RANS 34574 (-24137) WIN
reportlab-multipage.pdf BYTE_RANS 11144 -> PDF_DEFLATE_REPLAY_RANS 14456 (+3312) lose
pdftex-doc.pdf BYTE_RANS 24435 -> PDF_DEFLATE_REPLAY_RANS 35098 (+10663) lose
libreoffice-export.pdf BYTE_RANS 72791 -> PDF_DEFLATE_REPLAY_RANS 375265 (+302474) lose
Correction (Phase 7.0c): the Cairo "win" is a harness repeated-bytes artifact.
An independent adversarial review found that tools/pdf-corpus-producers.sh draws
one identical page six times (no per-page variation), so Cairo emits six streams
whose compressed bytes are identical and whose plaintexts are identical. The
court therefore cannot distinguish plaintext-sharing from plain compressed-byte
repetition, and generic LZ captures far more of the same redundancy: on
cairo-vector.pdf, gzip -9 = 17,382 B, zlib -9 = 17,376 B and xz -9e =
16,852 B β about half the 34,574 B reported as the "winning"
PDF_DEFLATE_REPLAY_RANS size. BYTE_RANS is a weak order-0 baseline with no LZ, so
the β24,137 B delta is a win over an order-0 lane on repeated identical bytes. The
phrases "a genuine authoring application does produce the win region", "the producer
creates the geometry" and "first authoring-generator witness" are withdrawn.
Correct characterization: our deterministic generator repeated one identical page
six times; Cairo emitted six byte-identical streams (compressed bytes and plaintext
both identical); this witnesses a repeated-identical-bytes region already captured
better by generic LZ, not the shared-plaintext-vs-distinct-compression mechanism.
ReportLab and pdfTeX are the same story (identical repeats) and lose at complete
cost; LibreOffice shares nothing. The delta is conditional on repeated identical
page content and is not a population claim; prior "wins" were measured against a
weak order-0 baseline and the generic ladder (below) is the honest comparison. No
candidate or wire format changed. Review:
docs/evidence/phase7b-skeptic-review.md.
Receipt:
evidence/campaigns/2026-10-05-phase7-producers-e071250/;
script tools/pdf-corpus-producers.sh; ledger
evidence/corpus/phase7-producers/provenance.json.
Generic-compressor baseline ladder (Phase 7.0c) measured results
Phase 7.0/7.0b measured VOLE only against BYTE_RANS, a whole-file order-0
byte-rANS lane with no LZ. Phase 7.0c adds the honest comparison: a pinned,
opt-in baseline image (the pinned dev base plus gzip, zstd, xz, brotli,
jq) and tools/baselines.sh, which for every corpus file records the smallest
lossless complete-file size for gzip -9, zstd -19 --long=27, xz -9e and
brotli -q 11 (each round-trip verified) and the complete serialized .voldoc
size of every VOLE lane, then compares against the best VOLE lane.
Result: on 27 corpus files the best VOLE lane beats gzip/zstd/xz/brotli on 0
files. The best generic compressor is smaller on every file, by +460,320 B
total (phase7 +410,355 over 23 files; producers +49,965 over 4). On the Cairo file
the "winning" 34,574 B is 2.07Γ brotli's 16,670 B; on flate.pdf the 36,102 B
is 1.91Γ xz's 18,884 B:
cairo-vector.pdf source 58424 gzip 17382 zstd 16836 xz 16852 brotli 16670 BYTE_RANS 58711 best VOLE 34574 (+17904 vs brotli)
_synthetic/flate.pdf source 57513 gzip 22426 zstd 20171 xz 18884 brotli 18891 BYTE_RANS 49291 best VOLE 36102 (+17218 vs xz)
Prior "wins" were relative to a weak order-0 baseline and do not survive the
generic ladder. VOLE's byte-exact structural reconstruction is unchanged; its
compression claim does not survive on this corpus. All 27 auto winners are
verify + cmp byte-exact; no candidate or wire format changed. Review:
docs/evidence/phase7b-skeptic-review.md.
Receipt:
evidence/campaigns/2026-10-05-phase7-baselines-7b9f662/
(full per-file table baseline-table.md); driver tools/baselines.sh.
Partial materialization / observation views (Phase 7.3) measured results
Whole-file compression loses (above), so Phase 7.3 measures the pivoted axis β
random-access query cost. view serves one output range from a descriptor
carrying an optional, advisory OBSERVATION_INDEX: when the program is a
sequence of linear independent ops it evaluates only the ops intersecting
[a,b) and decodes only the referenced entropy channels, and every served slice
is cmp'd against the full materialization. The corpus is a 33,789,340 B
(32.22 MiB), 800-stream deterministic PDF from the encode-time pdf-make-large
subcommand (qpdf --check rc 0; 800 replayed / 0 declined).
Result: a scoped positive on decode CPU, byte-exact on all 18
pre-registered queries. For any query at β₯ ~2 MiB the indexed lane touches
~0.41β0.43 MB (descriptor_bytes_traversed alone; its entropy_bytes_decoded
breakdown is a subset already counted there and must not be added) regardless
of offset, while gzip must inflate a + len; in the late region (β₯ 50 % in) that
is 1.4β2.3 % of gzip's bytes, and VOLE is ~2β5Γ faster than gzip and
~4β13Γ faster than xz on CPU. It loses in the early region (β€ ~8β16 MiB),
never beats zstd's raw decompressor on wall time, uses ~38 MB peak RSS vs
gzip's ~1.2 MB, and whole-file size is still 2.98Γ xz. The decisive caveat:
view reads and parses the whole descriptor, so on-disk I/O is not
reduced (no bytes-read win yet; descriptor_bytes_traversed is a CPU-side
approximation) β an mmap/seek reader is the prerequisite for that claim.
Receipt:
evidence/campaigns/2026-10-05-phase7-partial-a5764c9/
(query-table.md, report.md); report
docs/evidence/phase7-partial-report.md;
drivers tools/partial-court.sh, tools/partial-table.jq; ADR-0018.
Seek-based partial I/O (Phase 8.3) measured results
Phase 8 supplies the mmap/seek reader ADR-0018 required. A descriptor may carry
an optional seek DIRECTORY record (first record, fixed offset 64,
FLAG_OPTIONAL, cross-checked and never trusted); the view CLI peeks only the
64-byte header and serves from a Read + Seek reader that fetches those record
classes the query needs (header, DIRECTORY, GRAPH, OBSERVATION_INDEX,
INTEGRITY, and the one referenced object/channel/model) β it never fs::reads
the whole descriptor. This is a floor, not a small read: a 256-byte request
still incurs ~440 KB (~1,700Γ). A partial read is an observation
(integrity_verified == false); materialize/decode/verify remain the
archival authority.
Result: a scoped bytes-read win versus non-seekable sequential codecs,
byte-exact on all 18 pre-registered queries. On the same 33,789,340 B
(32.22 MiB) corpus the seekable descriptor is 17,566,832 B (DIRECTORY 27,390 B).
The seeked view reads a constant 439,679β461,367 B regardless of offset
(floor = header 64 + DIRECTORY 27,390 + GRAPH 265,462 + OBSERVATION_INDEX
146,711 + INTEGRITY 52), β€ 2.6 % of the descriptor for every query. In the late
region (β₯ 50 % in, 8/8 queries) that is 4.7 %β~21Γ fewer bytes than gzip's
compressed prefix (9,764,864 B vs 460,713 B at 31 MiB) and ~12β13Γ fewer than
zstd/xz; a strace -P descriptor-file cross-check equals the instrumented count
- exactly 64 B (the header peek). CPU drops to ~0.00 s and peak RSS from ~38 MB to ~3.8 MB.
Not a general random-access-I/O win (Phase 8.4 amendment). Against
seekable/blocked formats, for the same late query: bgzip (BGZF) reads
23,808 B (~19Γ fewer than VOLE's 460,713 B), xz --block-size=64KiB
15,344 B (~30Γ fewer) and 1MiB 179,892 B (~2.6Γ fewer) β and BGZF's
whole file (7,995,600 B) is even smaller than the seekable descriptor. VOLE only
wins where the block size is large (xz --block-size=4MiB 708,612 B; pixz
2,810,832 B, 16 MiB blocks). See docs/evidence/phase8-skeptic-review.md.
Where it loses (recorded). At a = 0 the constant floor exceeds gzip's first
bytes (327,680 B) and xz's (73,728 B); at early queries (β€ ~1.7 MiB) it loses to
xz's tiny compressed prefix; and versus compact-block seekable formats it loses
across the late region. The floor is constant in the offset, so it would
dominate a descriptor smaller than ~9 MB β a large-document mechanism. Whole-file
size is still 3.01Γ xz (and 0.46Γ BGZF). One locally generated corpus;
no population claim.
Validator caveat. A referenced channel's decoded_length is cross-checked
against its record; an unreferenced CHANNEL_LENGTHS entry is not β benign,
since analyze_ops never uses an unused length, so no wrong bytes are served.
Receipt:
evidence/campaigns/2026-10-05-phase8-seek-08de2a9/
(query-table.md, report.md, seekable.jsonl, seekable-table.md,
seekable-report.md); reports
docs/evidence/phase8-seek-report.md,
docs/evidence/phase8-skeptic-review.md;
drivers tools/seek-court.sh, tools/seek-table.jq,
tools/seekable-baselines.sh; ADR-0019.
Encoder-only search governance (Phase 10.1) measured results
The brief's Phase 10 asks for a DSFB-style search governor with zero decode
authority. Empirically, the published dsfb 0.1.2 crate is Drift-Slew Fusion
Bootstrap state estimation β a Kalman-like f64 observer over sensor channels
with no candidate / residual / search-directive API β present in the lockfile only
transitively through the optional EntropyFS engine. It is therefore recorded as
unavailable-for-purpose and is not a dependency: the feature is the
dependency-free dsfb-search = [].
The governor (src/encode/governor.rs) is encoder-only and integer-only. It
produces typed residual diagnostics, proposes candidates from a tiny parametric
space over existing mechanisms (scale_bits, partition, replay, packed,
depth), and maps the incumbent's dominant residual to a directive with a pure
govern function. Every candidate it can enable is serialized, parsed,
materialized, byte-compared, and priced by the same complete-cost court; no
governance state is persisted and src/container/header.rs is untouched, so a
governer-produced descriptor decodes byte-exactly in a build without the
feature.
The court (tools/governor-court.sh, campaign
2026-10-05-phase10-governor-d2b09c9, ADR-0022) compares Exhaustive,
FixedHeuristic, and DsfbGuided over a small deterministic cohort split into
disjoint tune/holdout/control sets. Pre-registered hypotheses: H1 HELD
(DsfbGuided.final β€ FixedHeuristic.final everywhere); H2 HELD
(guided.final == exhaustive.final on 8/8 holdout with β€ Β½ the candidates);
H4 HELD (negative controls Stop(Raw) and match the RAW descriptor
byte-for-byte). But H3 also HOLDS, and it is the substantive result: the fixed
heuristic already attains the exhaustive minimum on every workload, so the
parametric search adds zero bytes. Representative rows (final bytes /
candidates): flate.pdf 36,161 (fixed 10, exhaustive 438, guided 150);
bigtext.pdf 38,274 (7/414/126); many.pdf 5,301 (7/414/126). The fixed
complete-cost court is retained; the governor remains an optional, encoder-only
feature. Small locally generated cohort; no population claim.
Receipt:
evidence/campaigns/2026-10-05-phase10-governor-d2b09c9/
(results.json, hypotheses.json, decode-proof.json, report.md); driver
tools/governor-court.sh; ADR-0022.
Persistent procedural document field (Phase 11) measured results
Phase 11 stops treating the persisted document as a re-openable archive and makes it a queryable procedural field (ADRs 0024β0028):
- Procedural seed DAG. Each recovered computation is a canonical, versioned
node (
NodeId = BLAKE3("VOLE:PSEED:v1" β state β dep ids)), stored one blob per node in a content-addressed seed store, on a plain-filesystem backend or the optional EntropyFS engine (entropyfs-store). The DAG is immutable and cycle-free; the id is the fingerprint, so a changed dependency yields a new node and reuse has no invalidation pass. - Hierarchical observation index. A bounded selector-keyed tree (fanout β€ 256, depth β€ 3, node β€ 8 KiB) navigated path-only; it accelerates and never defines truth (a lying/oversized/out-of-closure node is rejected fail-closed).
- Progressive inverse compiler. Stage A durable capture; Stage B cheap eager
inversion (physical spans, revisions, objects, streams, decoded streams, page
tree,
/ObjStm-hosted page objects); Stage C demand-driven deepening. Real producers recover their pages: Cairo 6, LibreOffice 61, pdfTeX 6, ReportLab 6. - Observation engine. Typed selectors (Document/Page/Object/Stream/Revision/
ByteRange/TextMatch) Γ representations (Metadata/Text/Structure/Operators/
Encoded/Decoded/Exact/Preview/FullDocument), each answer a
FieldAnswerwith a typedbasis(authored / directly-observed / deterministically-derived / heuristic / inferred / unresolved) andintegrity_scope, plus deterministic planning andEXPLAIN/EXPLAIN ANALYZE. - Selective late materialization, reuse, and source independence. A narrow observation reads its minimum closure; the derived cache is off-wire and closure-keyed; the source PDF can be deleted and a new process still queries the field and rematerializes the exact original. An immutable edit witness shows one page-content override sharing 318/319 index entries and the descriptor (0 bytes read) while both roots stay exact.
Measured, with the four accounting universes kept distinct (source /
archive+store / procedural field / derived cache; ADR-0027). Representative
numbers (page-1 text; evidence/campaigns/2026-10-06-phase11-desc-free-a8ad6f4/):
| lane | bytes read | wall | note |
|---|---|---|---|
| VOLE field, warm (cache-served) | 8.4β8.9 KB overhead | ~99 Β΅s | 0 descriptor bytes |
| VOLE field, cold (partial descriptor) | 0.08β0.36 MB | β | re-derives text from the page channel |
| A1 preprocessed SQLite | 24,393 B | ~1 ms | one-time extraction charged |
A0 Poppler pdftotext -f1 -l1 |
136,128 B | ~8 ms | β |
Honest correction (independent skeptic). The original warm-byte claim omitted
the cache-served answer read (measured in the derived-cache universe), so on large
documents the warm byte court is a loss; the surviving win is the overhead +
wall result and the small producer documents. The LLM token court now reports a
pinned tokenizer (bert-base-uncased, asset SHA-256 verified) and finds V wins
2 / ties 2 / loses 2 versus page-local extraction β a working-set measure, never a
text-quality claim. Finer-than-object sharing wins versus per-file LZ and frozen
CDC but loses versus tar|xz and strongest CDC (unique_bytes is a lower
bound only).
Receipts: evidence/campaigns/2026-10-06-phase11-* (plan, io, partial, desc-free,
entropyfs, lifetime, share, llm-tokens, edit, skeptic); results write-up
docs/phases/phase-11-results.md; review
docs/reviews/phase-11-skeptic-review.md.
Quick start (Docker only)
All project commands run inside pinned containers. The host only invokes Docker.
# Build the toolchain image (pinned by digest in Dockerfile)
# Build, test, lint, format
# MSRV gate (Rust 1.89)
# Generic-compressor baseline ladder (Phase 7.0c; opt-in baseline image)
# Partial-materialization query court (Phase 7.3; opt-in baseline image)
# Seek-based partial-I/O bytes-read court (Phase 8.3; opt-in baseline image)
# Seekable/blocked random-access baseline court (Phase 8.4; opt-in baseline image)
# adds tabix(bgzip)+pixz to the baseline stage, then builds bgzip/xz-blocked/pixz
# and measures the honest random-access cost (covering block(s) + index).
# Phase 1 exact court (writes an evidence receipt)
CLI surface (activated by the pipeline, not by extension β extensions are hints, never authority):
vole-document encode [--force KIND] INPUT OUTPUT.voldoc
vole-document decode INPUT.voldoc OUTPUT
vole-document materialize INPUT.voldoc OUTPUT
vole-document view INPUT.voldoc [OUTPUT] --byte-range A:L | --pdf-object N:G | --pdf-stream N:G | --pdf-revision I [--stats]
vole-document verify INPUT.voldoc
vole-document inspect INPUT.voldoc
vole-document pdf-make-large DIR [OBJECTS]
vole-document capabilities
# Phase 11 β persistent procedural field (feature `field`, on by default)
vole-document field-ingest INPUT.voldoc --store DIR [--entropyfs]
vole-document observe --store DIR --field HEX (--page N | --object N | --stream N | --revision N | --byte-range A..B) --kind KIND
vole-document find --store DIR --field HEX --text PATTERN
vole-document explain --store DIR --field HEX ... [--analyze]
vole-document preview --store DIR --field HEX --page N [--json]
vole-document materialize --store DIR --field HEX --exact --output FILE
vole-document cache --store DIR [--clear]
vole-document field-edit --store DIR --field HEX --page N --content-bytes HEX
vole-document share-account [--store DIR] INPUT.voldoc...
encode --force KIND forces the complete-cost court to consider only one
candidate family (raw, rle, byte-rans, pdf-physical, pdf-channels,
pdf-layout, pdf-layout-rans, pdf-deflate-replay, pdf-deflate-replay-rans)
for honest per-mechanism ablation; it never
bypasses exactness, and it fails with a typed usage error when the input does not
propose that kind.
Fuzzing
Two layers, both Docker-only:
-
Deterministic property/mutation courts (
tests/property.rs,tests/goldens.rs,tests/malformed.rs) plus the longer soak runtools/soak-fuzz.sh(VOLE_FUZZ_ITERS, default 200000). -
Coverage-guided libFuzzer targets (Phase 7.1) in the standalone
fuzz/cargo-fuzzpackage (excluded fromcargo package), built and run in the pinned dated-nightlyfuzzDocker service:FUZZ_SECONDS=60
Ten targets cover the .voldoc container parser/materializer, the
encodeβdecode round trip, the DRA decoder/analyzer/evaluator, the rANS model
and channel decoders, the PDF lexer and physical scanner, xref//Prev, the
DEFLATE_REPLAY wrapper, and verify. See fuzz/README.md for the pinned
toolchain, seeds, and regeneration. The sealed campaign
evidence/campaigns/2026-10-05-phase7-fuzz-ca6a92b/ observed zero crashes on
nine of ten targets and reported two upstream preflate-rs findings (one
mitigated fail-closed, one recorded upstream resource limitation).
Repository layout
src/ one crate; modules for architectural separation
tests/ exact / malformed / conformance courts
fuzz/ cargo-fuzz coverage-guided targets (excluded from the crate)
tools/ court and gate scripts (run inside Docker)
docs/ architecture, ADRs, security, phase notes
evidence/ immutable campaign receipts (machine-readable)
research/ LOCAL ONLY β gitignored (paper, snapshots, subagent findings)
research/ is intentionally excluded from version control. Durable findings that
matter to a phase are frozen into ADRs and phase notes and referenced from
receipts by hash.
Licensing
Dual-licensed under either MIT or Apache-2.0, at your option. See
LICENSE-MIT and LICENSE-APACHE.
Third-party license note. The deflate-replay feature (Phase 6) is
opt-in and depends on preflate-rs, which depends on cabac, licensed
LGPL-3.0-or-later. The default build (default = ["rans", "store", "field"])
is permissive-only and contains no LGPL code. Rust links statically by default,
so a binary built with --features deflate-replay (or --all-features)
contains LGPL code and carries the corresponding obligations. See
ADR-0014.