vole-document 0.1.0-alpha.13

Byte-exact procedural document storage: deterministic reconstruction state, typed residuals, and entropy-coded channels that materialize the exact original document bytes.
Documentation
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
405
406
407
408
409
410
411
412
413
414
415
416
417
418
419
420
421
422
423
424
425
426
427
428
429
430
431
432
433
434
435
436
437
438
439
440
441
442
443
444
445
446
447
448
449
450
451
452
453
454
455
456
457
458
459
460
461
462
463
464
465
466
467
468
469
470
471
472
473
474
475
476
477
478
479
480
481
482
483
484
485
486
487
488
489
490
491
492
493
494
495
496
497
498
499
500
501
502
503
504
505
506
507
508
509
510
511
512
513
514
515
516
517
518
519
520
521
522
523
524
525
526
527
528
529
530
531
532
533
534
535
536
537
538
539
540
541
542
543
544
545
546
547
548
549
550
551
552
553
554
555
556
557
558
559
560
561
562
563
564
565
566
567
568
569
570
571
572
573
574
575
576
577
578
579
580
581
582
583
584
585
586
587
588
589
590
591
592
593
594
595
596
597
598
599
600
601
602
603
604
605
606
607
608
609
610
611
612
613
614
615
616
617
618
619
620
621
622
623
624
625
626
627
628
629
630
631
632
633
634
635
636
637
638
639
640
641
642
643
644
645
646
647
648
649
650
651
652
653
654
655
656
657
658
659
660
661
662
663
664
665
666
667
668
669
670
671
672
673
674
675
676
677
678
679
680
681
682
683
684
685
686
687
688
689
690
691
692
693
694
695
696
697
698
699
700
701
702
703
704
705
706
707
708
709
710
711
712
713
714
715
716
717
718
719
720
721
722
723
724
725
726
727
728
729
730
731
732
733
734
735
736
737
738
739
740
741
742
743
744
745
746
747
748
749
750
751
752
753
754
755
756
757
758
759
760
761
762
763
764
765
766
767
768
769
770
771
772
773
774
775
776
777
778
779
780
781
782
783
784
785
786
787
788
789
790
791
792
793
794
795
796
797
798
799
800
801
802
803
804
805
806
807
808
809
810
811
812
813
814
815
816
817
818
819
820
821
822
823
824
825
826
827
828
# VOLE-Document

Byte-exact procedural document storage.

VOLE-Document persists a **bounded deterministic reconstruction description** of a
document β€” reconstruction structure, parameters/state, typed residual channels,
and (in the entropy phases) typed rANS channels β€” and materializes the exact
original bytes on demand.

The governing invariant of the exact profile is uncompromising:

```text
materialize(descriptor) == original_bytes
```

Parsing successfully, producing "the same" text, the same object graph, the same
pages, the same rendering, or a canonical re-save are **not** substitutes.

> This project is deliberately *not* "a PDF optimizer that happens to use rANS".
> rANS is the entropy substrate beneath the representation, never the procedural
> model. See [`SPEC.md`](SPEC.md) and [`docs/`](docs/) for the architecture.

> ### πŸ“Œ Findings
>
> **The current VOLE representation stack does not beat purpose-built baselines
> on any measured axis.** The durable results are byte-exactness, an auditable
> reconstruction representation, and a complete, receipted record of the
> negatives. Read the authoritative consolidation: **[`FINDINGS.md`](FINDINGS.md)**
> and **[ADR-0023](docs/adr/0023-consolidated-findings.md)**.

## Status

Current release: **`0.1.0-alpha.12`** (Phase 10 β€” DSFB encoder-only governance +
negative-results consolidation). The consolidated verdict (ADR-0023,
[`FINDINGS.md`](FINDINGS.md)) supersedes every earlier "win" claim: **no measured
axis beats a purpose-built baseline.** Exactness is unchanged
(`materialize(descriptor) == original_bytes`) and is stated as the invariant, not
as a competitive win; whole-file size remains a recorded loss against generic
lossless tools, **0/27** (ADR-0017). Phase 7 measured a scoped random-access
**decode-CPU** win (ADR-0018); Phase 8 measured a scoped random-access
**bytes-read** win versus **non-seekable sequential** codecs only β€” it loses to
seekable/blocked formats (ADR-0019); Phase 9 adds the cross-document
content-addressed **store** and records a robust β€” but partly
externalization-granularity-artifact β€” loss on the store axis (ADR-0020/0021);
Phase 10.1 adds an optional, **encoder-only** search governor behind the
non-default, dependency-free `dsfb-search` feature and records that the
parametric search buys **no bytes** on its cohort (ADR-0022). Nothing else wins.

Phase 9 (branch `phase9`, ADR-0020/0021) adds a cross-document
**content-addressed object store** and then measures the one axis a single-file
compressor structurally cannot serve. The store is exact (a store-backed
descriptor materializes the same bytes as its standalone form) and its three
accounting universes are kept permanently distinct. The **measured result is a
recorded negative**: over 37 locally generated files (5,579,469 B, 10
deliberate-sharing strata) the unique-reachable universe `U = 3,369,900 B` loses
to per-file min LZ (`1,304,307 B`) and to the strongest pinned
content-defined-chunk dedup (borg 1.2.4: deterministic `771,383 B` raw;
compressed non-deterministic, 210,835–210,840 B). The negative is **robust** β€”
forcing the finest candidate in the current set (`--force pdf-deflate-replay`)
still gives `U = 2,360,054 B`, and a per-stratum oracle ~`2,537,730 B` β€” but its
size is **partly an artifact of externalization granularity/candidate selection**:
the auto winner emits only 0–1 objects per file. Under the auto candidate only
byte-identical opaque repeats win (`repeat-bin`, `U = 133,048 B`); forcing
`PDF_DEFLATE_REPLAY` (one object per deflate stream) flips the `shared-payload`
fixture to `264,139 B`, a win over **raw** CDC (`285,257 B`) that still loses to
LZ (`34,591 B`) and compressed CDC (`28,195 B`). No current candidate emits more
than one object per file (ADR-0021, `docs/evidence/phase9-store-report.md`,
`docs/evidence/phase9-skeptic-review.md`). Cross-document sharing is store
*amortization*, never "compression".

| Area | State | Evidence |
|---|---|---|
| Exact `.voldoc` container (framing, header, records) | **Implemented** | `src/container/`, unit + conformance courts |
| Typed errors + stable exit codes | **Implemented** | `src/error.rs` |
| Centralized resource limits | **Implemented** | `src/limits.rs` |
| CRC32C framing + SHA-256 archival identity | **Implemented** | `src/integrity.rs` |
| Document Reconstruction Algebra (literal subset) | **Implemented** | `src/dra/` |
| Coverage certificate (checked invariant) | **Implemented** | `src/dra/program.rs` |
| RAW exact opaque adapter | **Implemented** | `src/adapter/opaque/` |
| Candidate complete-cost court + decode-before-commit | **Implemented** | `src/encode/` |
| CLI (`encode`/`decode`/`verify`/`inspect`/`capabilities`) | **Implemented** | `src/main.rs` |
| Exact court over a mixed corpus | **Measured** | `evidence/campaigns/` |
| Native rANS floor (order-0 / typed byte channels) | **Measured** | campaign `2026-10-05-phase2-f6af30b` |
| RLE candidate (`REPEAT_LAST` run-length) | **Measured** | campaign `2026-10-05-phase2-f6af30b` |
| BYTE_RANS candidate (order-0 byte channel) | **Measured** | campaign `2026-10-05-phase2-f6af30b` |
| Entropy capsule (full decoder-entry state, not a seed) | **Measured** | ADR-0006; `src/entropy/` |
| PDF lexical span cover (Phase 3.1) | **Measured** | campaign `2026-10-05-phase3-486aa17` |
| PDF byte-authoritative physical scanner (Phase 3.2–3.3) | **Measured** | campaign `2026-10-05-phase3-486aa17` |
| PDF incremental revision map (Phase 3.4) | **Measured** | campaign `2026-10-05-phase3-486aa17` |
| qpdf differential oracle court (oracle, never authority) | **Measured** | `tools/pdf-oracle.sh`; campaign `2026-10-05-phase3-486aa17` |
| PDF lexical channel transposition (`split`/`join`) | **Implemented** | `src/adapter/pdf/channels.rs`; `tests/pdf_channels.rs` |
| `INTERLEAVE_CHANNELS` DRA op (DRA v3) | **Measured** | `src/dra/op.rs`; campaign `2026-10-05-phase4-3840bc4` |
| Compact entropy model wire v2 (sparse/dense, smaller chosen) | **Measured** | `src/entropy/model.rs`; campaign `2026-10-05-phase4-3840bc4` |
| Forced-candidate ablation (`encode --force KIND`) | **Measured** | `tools/phase4-court.sh`; campaign `2026-10-05-phase4-3840bc4` |
| PDF typed channels (`PDF_CHANNELS`) | **Recorded (rejected on corpus)** | campaign `2026-10-05-phase4-3840bc4`; exact but loses to `BYTE_RANS` on complete cost (ADR-0010) |
| Positional DRA ops (`MARK_OFFSET` / `EMIT_OFFSET`, DRA v4) | **Implemented** | `src/dra/op.rs`; `tests/pdf_layout.rs` |
| PDF layout candidate (`PDF_LAYOUT`) | **Recorded (rejected on cost)** | campaign `2026-10-05-phase5-7193001`; byte-exact and predicts xref offsets/`startxref`, but loses to RAW/`BYTE_RANS` on DRA framing cost (ADR-0011) |
| Packed segment framing (`PACK_SEGMENTS`, DRA v5) | **Implemented** | `src/dra/op.rs`; opcode `0x08`, one op + a compact varint item table over a single data object, amortizing per-segment framing |
| PDF layout on packed framing (layout-v2) | **Recorded (beats RAW at scale, rejected vs `BYTE_RANS`)** | campaign `2026-10-05-phase5-4521778`; packed+coalesced layout beats RAW on `many.pdf` (10,069 vs 10,215) but still loses to `BYTE_RANS` (5,181); residual stored literally (ADR-0012) |
| `PACKED_CHANNELS` DRA op (DRA v6) | **Implemented** | opcode `0x09`; reconstructs from a data channel + a plan channel (serialized item table) with a declared output length validated at eval; universe β†’ `phase5-8` |
| PDF layout + rANS (`PDF_LAYOUT_RANS`) | **Recorded β€” rejected vs `BYTE_RANS`** | campaign `2026-10-05-phase5-8-cf8048d`; byte-exact, but wins 0 / loses 8 / declines 3 head-to-head (ADR-0013) |
| Phase-5 forced-candidate court (`--force pdf-layout`) | **Measured** | `tools/phase5-court.sh`; campaigns `2026-10-05-phase5-7193001` and `2026-10-05-phase5-4521778` |
| Phase-5.8 forced-candidate court (`--force pdf-layout-rans`) | **Measured** | `tools/phase5-8-court.sh`; campaign `2026-10-05-phase5-8-cf8048d` |
| Lexer stream opacity (`stream`+EOL opaque span) | **Adopted** | campaign `2026-10-05-phase6-0d0bb79`; stream-data bytes are a byte-authoritative span, so `/FlateDecode` stream spans are exact |
| `DEFLATE_REPLAY` DRA op (DRA v8) | **Implemented** | `src/dra/op.rs`; opcode `0x0A`, explicit `replay_codec` tag (`preflate-0.7.6-experimental`), exact raw-DEFLATE replay from `(plaintext, corrections)` with a declared output length statically bounded before the engine runs, `catch_unwind`-isolated, mandatory feature bit (opt-in `deflate-replay` cargo feature) |
| Exact DEFLATE replay, raw plaintext (`PDF_DEFLATE_REPLAY`) | **Recorded β€” rejected vs `BYTE_RANS`** | campaign `2026-10-05-phase6-0d0bb79`; byte-exact, but on `flate.pdf` 56,736 vs `BYTE_RANS` 49,291 (plaintext β‰ˆ bitstream) |
| Exact DEFLATE replay, shared rANS plaintext (`PDF_DEFLATE_REPLAY_RANS`) | **Adopted β€” first structural win (over `BYTE_RANS`, on a self-authored fixture)** | campaign `2026-10-05-phase6-0d0bb79`; `flate.pdf` 36,102 vs `BYTE_RANS` 49,291 (**βˆ’13,189 B**). A per-mechanism result only: not a generic-compressor win, and its enabling condition is not produced by tested transformers (ADR-0015, `docs/evidence/phase7b-skeptic-review.md`) |
| Phase-6 replay court (`--force pdf-deflate-replay[-rans]`) | **Measured** | `tools/phase6-court.sh`; campaign `2026-10-05-phase6-0d0bb79` |
| Producer-stratified Flate ratio harness (`deflate-stats`) | **Measured** | `tools/pdf-corpus.sh`; amendment campaign `2026-10-05-phase7-corpus-b-c4eb77e` (supersedes `2026-10-05-phase7-corpus-f1f8d26`); 24/24 replayed, 0 declined; diagnostic only, no new candidate |
| Generator-family Flate corpus (`producers`: ReportLab/Cairo/LibreOffice/pdfTeX) | **Measured (claim corrected 7.0c)** | campaign `2026-10-05-phase7-producers-e071250`; 87/87 replayed; `PDF_DEFLATE_REPLAY_RANS` beats `BYTE_RANS` on Cairo (58,711 β†’ 34,574, βˆ’24,137 B) but the Cairo file is a **repeated-identical-bytes harness artifact** (six byte-identical streams), and generic LZ does ~2Γ— better (gzip -9 17,382 B; xz -9e 16,852 B); the "first authoring-generator witness" claim is withdrawn; no candidate changed |
| Generic-compressor baseline ladder (gzip/zstd/xz/brotli vs VOLE) | **Measured β€” the honest comparison** | campaign `2026-10-05-phase7-baselines-7b9f662`; `tools/baselines.sh` + opt-in `baseline` image; **VOLE beats gzip/zstd/xz/brotli on 0/27 files** (+460,320 B vs the best generic); prior "wins" were relative to the weak order-0 `BYTE_RANS` lane |
| Coverage-guided fuzzing (`cargo-fuzz`, ten targets) | **Measured** | campaign `2026-10-05-phase7-fuzz-ca6a92b`; pinned `nightly-bookworm-slim-2026-10-04` + `cargo-fuzz 0.13.2`; 9/10 targets zero-crash; two upstream `preflate-rs` findings (F1 mitigated + regression test; F2 contained on the decode path by process isolation, ADR-0016) |
| Process-isolated DEFLATE replay (`__replay-worker`) | **Implemented** | `replay_bounded` runs `preflate` in a child under `RLIMIT_AS` + a wall-clock timeout; knobs `VOLE_REPLAY_WORKER`/`VOLE_REPLAY_MEM_MB`/`VOLE_REPLAY_TIMEOUT_MS`; a library embedder without a worker keeps the in-process residual |
| Partial materialization (`OBSERVATION_INDEX` + `view`) | **Measured β€” scoped positive (decode CPU, no I/O win)** | campaign `2026-10-05-phase7-partial-a5764c9`; 18/18 queries byte-exact; mid/late queries touch ~0.41–0.43 MB (`descriptor_bytes_traversed` alone; its `entropy_bytes_decoded` breakdown is a subset already counted there and must not be added) vs gzip inflating `a+len` (late region 1.4–2.3 %, ~2–5Γ— faster than gzip, ~4–13Γ— than xz); **v1 reads the whole descriptor, so on-disk I/O is not reduced** and it loses in the early region (≀ ~8–16 MiB) (ADR-0018) |
| Deterministic large corpus generator (`pdf-make-large`) | **Tooling** | encode-time subcommand; classic-xref PDF of `OBJECTS` (default 800) distinct zlib `FlateDecode` streams, β‰₯32 MiB, correct by construction; bytes gitignored |
| Seek `DIRECTORY` + `Read + Seek` reader (`view`) | **Measured β€” scoped bytes-read win vs non-seekable sequential codecs** | campaign `2026-10-05-phase8-seek-08de2a9`; 18/18 queries byte-exact; seeked `view` reads a constant ~0.44–0.46 MB (GRAPH + OBSERVATION_INDEX + DIRECTORY floor), 12–21Γ— fewer bytes than sequential gzip/zstd/xz prefixes for late queries; loses at offset 0 and early (ADR-0019) |
| Seekable/blocked random-access baseline (bgzip / blocked xz / pixz) | **Recorded β€” the honest random-access comparison (falsifies β€œgeneral random-access win”)** | amendment to campaign `2026-10-05-phase8-seek-08de2a9` (`seekable-report.md`); late query VOLE 460,713 B vs bgzip 23,808 B (~19Γ—), xz-64KiB 15,344 B (~30Γ—), xz-1MiB 179,892 B (~2.6Γ—), xz-4MiB 708,612 B (VOLE wins), pixz 2,810,832 B (VOLE wins); BGZF whole-file 7,995,600 B < 17,566,832 B; `docs/evidence/phase8-skeptic-review.md` |
| PDF structural adapters (Phases 7–8) | Planned | β€” |
| Content-addressed object store (`ObjectStore`/`EmbeddedStore`, `EXTERNAL_REF`) | **Implemented** | `src/store/`, `tests/store.rs`; ADR-0020; `Id = BLAKE3-256`, mandatory `FEATURE_EXTERNAL_OBJECTS`, `externalize`/`hydrate`, `gc` |
| Three accounting universes (standalone / unique-reachable / amortized) | **Measured** | campaign `2026-10-05-phase9-store-fdb2845`; `store account`; `Ξ£ amortized == unique reachable` by construction |
| Cross-document store court (universes vs per-file LZ + generic CDC) | **Recorded β€” store axis is a loss** | campaign `2026-10-05-phase9-store-fdb2845` (ADR-0021); `U = 3,369,900 B` loses to per-file min LZ (`1,304,307 B`) and to CDC borg 1.2.4 (`771,383 B` raw deterministic / 210,835–210,840 B compressed, non-deterministic); the negative is robust (forced `--force pdf-deflate-replay` `U = 2,360,054 B`; per-stratum oracle ~`2,537,730 B`) but partly an externalization-granularity artifact; under the auto candidate only `repeat-bin` wins (`133,048 B`), while forcing `PDF_DEFLATE_REPLAY` flips `shared-payload` to `264,139 B` (win over raw CDC, loss to LZ/zstd); no current candidate emits >1 object/file; 37/37 byte-exact; `docs/evidence/phase9-skeptic-review.md` |
| EntropyFS store backend (`EntropyFsStore`, optional adapter) | **Implemented (optional, not the measured backend)** | non-default `entropyfs-store` feature; `list`/`remove` decline (no per-blob delete); ADR-0008/0020 |
| DSFB search governance (Phase 10.1) | **Implemented (encoder-only, `dsfb-search`) / search recorded negative** | campaign `2026-10-05-phase10-governor-d2b09c9`; dependency-free `dsfb-search = []`; guided never worse than fixed (H1) and matches exhaustive on 8/8 holdout with ≀ Β½ candidates (H2), but the fixed heuristic already equals exhaustive on every workload, so the search adds **zero bytes** (H3); zero decode authority proven (ADR-0022) |
| Partial materialization checkpoints (beyond v1) | Planned | v1 random-access `view` measured in 7.3 (ADR-0018); the **seek reader landed in Phase 8** (ADR-0019), leaving checkpoint bytes as future work |

"Implemented" means the mechanism exists and is tested. "Measured" means there is
a sealed campaign under `evidence/`. The Phase-1 core establishes exactness,
framing, integrity, bounds, and receipts before any entropy or format-aware
mechanism is allowed to compete; Phase 2 then measures entropy channels on that
same exactness floor.

### Phase 2 measured results

Phase 2 is **order-0 typed byte channels only** β€” no context model, no typed
residuals, and no format awareness. On a 9-file mixed corpus the cumulative
core→full ladder over serialized `.voldoc` bytes is:

```text
sum_source = 590081
sum_core   = 464474   (RAW + RLE)
sum_full   = 291304   (RAW + RLE + BYTE_RANS)
delta      = 173170   (sum_core - sum_full)
```

The entire delta is attributed to the two files where `BYTE_RANS` wins
(`text-256k.bin` 262144 β†’ 143746; `skewed.bin` 65536 β†’ 11382). `RLE` wins the
long runs (`zeros-64k.bin` 65536 β†’ 283; `runs.bin` 65536 β†’ 3088). Negative
controls hold: `BYTE_RANS` never wins on random 64 KiB (stored RAW at 65536 β†’
65845, a 309-byte fixed framing overhead) or on empty/one-byte inputs (RLE).
The canonical model's bytes are charged like any other bytes, so on tiny or
high-entropy inputs order-0 rANS loses to RAW/RLE as required β€” this is a scoped
measurement on one deterministic corpus, not a general compression claim.

The entropy substrate is optional in the build: `default = ["rans"]`. The exact
DEFLATE replay stack is **opt-in** (`--features deflate-replay`); a
channel-bearing or replay-bearing descriptor decoded without the required
feature returns an explicit `UnsupportedFeature`, never a silent
reinterpretation.

Receipt:
[`evidence/campaigns/2026-10-05-phase2-f6af30b/`](evidence/campaigns/2026-10-05-phase2-f6af30b/).

### Phase 3 measured results

Phase 3 adds a **byte-authoritative PDF physical scanner**: an owned lexer, a
conservative structural span cover, `/Length` resolution, and an append-only
revision map. The sealed campaign `2026-10-05-phase3-486aa17` runs over a
deterministic 9-item corpus (7 valid PDFs plus `malformed.pdf` and `notpdf.bin`
as negative controls):

- **Coverage** β€” `all_covered = true`: 171 spans, 19 objects, and 8 revisions
  across the corpus, partitioned into a contiguous cover of `[0, len)` with no
  gap and no overlap.
- **Byte-exactness** β€” `all_exact = true`: `materialize(descriptor) ==
  original_bytes` for every item, including both negative controls through the
  opaque RAW lane.
- **Validated detection** β€” a file is a PDF only when the bytes show a `%PDF-`
  header **and** an indirect object **and** a `%%EOF`; the extension is never
  authority, and both controls report `is_pdf = false`.
- **Revision map** β€” the incremental input yields two append-only revisions with
  a `/Prev` chain, and `/Size` is treated as never decreasing.
- **qpdf oracle** β€” 100% object-number agreement with qpdf 11.3 (classic 4/4,
  two-page 6/6, incremental 5/5); `qpdf --check` reports valid; `pdfinfo` pages
  1/2/1. Divergence is expected where objects are compressed inside object
  streams: those have no physical `N G obj` marker, so a physical scanner
  enumerates fewer objects than qpdf's semantic view. qpdf is an oracle, never
  the byte authority.

The literal PDF candidate currently **loses to RAW**: RAW won all 9 items and
`PDF_PHYSICAL` won 0. This is the **expected Phase-3 result** β€” the physical lane
persists each span as one literal `INLINE` op with no structural compression, so
its per-span overhead loses once complete cost is charged. Structural
compression (xref/`/Length` proceduralization, stream replay) is Phase 5+ and is
not claimed here.

Receipt:
[`evidence/campaigns/2026-10-05-phase3-486aa17/`](evidence/campaigns/2026-10-05-phase3-486aa17/).

### Phase 4 measured results

Phase 4 transposes the byte-authoritative lexical cover into **typed channels**
(one kind id per token, one 4-byte length per token, one payload stream per
lexical kind) and reconstructs them with the bounded `INTERLEAVE_CHANNELS` DRA op
(DRA v3). Each channel gets its own order-0 byte-rANS model, and model wire **v2**
serializes to whichever of sparse or dense is smaller. The sealed campaign
`2026-10-05-phase4-3840bc4` runs a forced-candidate ablation (`encode --force
KIND`) over a deterministic 10-file corpus:

```text
A0 RAW            = 71036
A1 + RLE          = 71036
A2 + BYTE_RANS    = 43297
A3 + PDF_PHYSICAL = 43297
A4 + PDF_CHANNELS = 43297     leave-one-out channel delta = 0
```

All 10 files round-trip byte-exactly (`cmp` + `verify`) and the qpdf oracle
re-check passes. Auto winners: RAW = 8, `BYTE_RANS` = 2, `PDF_PHYSICAL` = 0,
`PDF_CHANNELS` = 0. On the text-heavy scale sample `bigtext.pdf` (65,549 B) the
forced sizes were RAW = 65,871, `BYTE_RANS` = 38,142, `PDF_CHANNELS` = 46,432:
typed channels beat RAW by ~21.6% but **lose to `BYTE_RANS` by ~8,290 B**. Compact
sparse models cut the per-channel model overhead from 7,224 B to 1,981 B, which
was not enough to close the gap. The honest conclusion is that coarse lexical
transposition plus per-channel order-0 models does **not** beat a monolithic
order-0 `BYTE_RANS` on this corpus, so the typed lexical channels were **rejected
by the complete-cost court** and `PDF_CHANNELS` is recorded (not adopted). A win
would require *conditioning and ordering* rather than more marginal per-kind
models β€” that is Phase 5+ work (ADR-0010).

Receipt:
[`evidence/campaigns/2026-10-05-phase4-3840bc4/`](evidence/campaigns/2026-10-05-phase4-3840bc4/).

### Phase 5 measured results

Phase 5 adds **positional DRA ops** (`MARK_OFFSET` / `EMIT_OFFSET`, DRA v4)
and the first candidate that replaces literal structural bytes with
*procedurally determined* ones: the classic cross-reference **layout** lane
(`PDF_LAYOUT`), which marks each indirect object's introducer offset and the
`xref` section start, then regenerates the 10-digit xref entry offsets and the
`startxref` value from those marks. The sealed campaign
`2026-10-05-phase5-7193001` runs the forced-candidate ablation
(`encode --force KIND`) over a deterministic 10-file corpus:

```text
A0 RAW            = 71116
A1 + RLE          = 71116
A2 + BYTE_RANS    = 43377
A3 + PDF_PHYSICAL = 43377
A4 + PDF_CHANNELS = 43377
A5 + PDF_LAYOUT   = 43377     leave-one-out layout delta = 0
```

All 10 files round-trip byte-exactly through their auto winner (`cmp` +
`verify`); auto winners are RAW = 8, `BYTE_RANS` = 2, and `PDF_PHYSICAL` /
`PDF_CHANNELS` / `PDF_LAYOUT` = 0. **The prediction works and the descriptor is
exact** β€” `classic.pdf` regenerates 3 of 4 xref entry offsets plus the
`startxref`, and `incremental.pdf` regenerates 5 of 7 entries plus two
`startxref` values β€” but the lane still **loses to RAW and `BYTE_RANS` on
complete cost**: `classic.pdf` 798 vs RAW 659, and `bigtext.pdf` 66,066 vs RAW
65,879 / `BYTE_RANS` 38,150. Layout wins 0 of the 7 classic-xref files. The
reason is framing, not prediction: the DRA pays a `MarkOffset` per object and
per xref section plus an `EmitOffset` per predicted entry, and that per-segment
op framing (tag + operand length) costs more than the ~7 digits saved per
predicted offset at document scale. The predicted structure is right; the
reconstruction *container* is too expensive. This is recorded honestly as a
negative result (ADR-0011) and a format-design input for later phases.

Receipt:
[`evidence/campaigns/2026-10-05-phase5-7193001/`](evidence/campaigns/2026-10-05-phase5-7193001/).

### Packed framing (Phase 5.7) measured results

Phase 5.7 attacks the framing root cause directly. The `PACK_SEGMENTS` DRA op
(opcode `0x08`, bumping the DRA graph to **version 5**, universe
`phase6-prep;…;dra-5;…+packed`) amortizes per-segment framing: instead of one
tagged op per literal run, it carries **one op plus a compact varint item table**
(`Literal` varint-length / `Mark` / `Emit`) over a single data object. The layout
candidate was rebuilt on this op (**layout-v2**), and literal coalescing (5.7.2b)
merges adjacent literal runs: on `many.pdf` (200 objects, 9,881 B) the item table
fell from **1,413 to 805** items.

The sealed campaign `2026-10-05-phase5-4521778` runs the forced-candidate
ablation over an 11-file corpus:

```text
A0 RAW            = 81371
A2 + BYTE_RANS    = 48598
A5 + PDF_LAYOUT   = 48598     leave-one-out layout delta = 0
```

Forced sizes:

```text
many.pdf     layout-v2 10069   RAW 10215   BYTE_RANS 5181   (layout beats RAW by 146 B)
classic.pdf  layout      711   RAW   663
bigtext.pdf  layout    65929   RAW 65883   BYTE_RANS 38154
```

Packed framing plus coalescing makes **structural layout prediction beat RAW at
scale** (`many.pdf` 10,069 vs 10,215), which the per-segment Phase-5 framing never
managed. It is still **not adopted**: it loses to `BYTE_RANS` (5,181), wins 0 of
the 8 classic-xref samples, and the leave-one-out layout delta is 0, so layout is
never the auto winner. The honest conclusion is that the framing is fixed, but the
residual data object is stored **literally**, so any order-0 entropy lane
dominates it. The remaining lever is to entropy-code the residual data object β€”
structural prediction **composed with** rANS on the residual, which is the paper's
layered model β€” not more literal packing. This is recorded as a partial positive
(ADR-0012).

Receipt:
[`evidence/campaigns/2026-10-05-phase5-4521778/`](evidence/campaigns/2026-10-05-phase5-4521778/).

### Layout + rANS residual (Phase 5.8) measured results

Phase 5.8 builds the lever ADR-0012 named: the `PACKED_CHANNELS` DRA op (opcode
`0x09`, DRA **v6**, universe
`phase5-8;…;dra-6;…+packed+packed-channels`) reconstructs from a **data channel**
plus a **plan channel** (the serialized item table) with a declared output length
validated at eval, and the `PDF_LAYOUT_RANS` candidate codes the layout plan's
data object and its item table each as their own order-0 rANS channel. The sealed
campaign `2026-10-05-phase5-8-cf8048d` runs the forced-candidate ablation over the
11-file corpus:

```text
A0 RAW               = 81591
A2 + BYTE_RANS       = 48818
A6 + PDF_LAYOUT_RANS = 48818     leave-one-out layout+rANS delta = 0
```

Forced sizes:

```text
classic.pdf  RAW 683  BYTE_RANS 728  PDF_LAYOUT 731  PDF_LAYOUT_RANS 883
bigtext.pdf  RAW 65903  BYTE_RANS 38174  PDF_LAYOUT_RANS 38341
many.pdf     RAW 10235  BYTE_RANS 5201  PDF_LAYOUT 10089  PDF_LAYOUT_RANS 5914
```

On `many.pdf` the layout+rANS size breaks down as data 7,877, plan 1,815, models
645, payload 4,775. Head-to-head against `BYTE_RANS`: **win 0, lose 8, declined
3**. The honest conclusion: layout+rANS does **not** beat `BYTE_RANS`. Channel 0
codes nearly the whole file β€” the same job `BYTE_RANS` does with one channel β€” so
the plan channel (1,815 B on `many.pdf`) plus a second model are added metadata
`BYTE_RANS` never pays. Three phases (4, 5, 5.7) plus this one converge: at the
tested scale, PDF structural proceduralization does not beat a whole-file order-0
rANS lane.

Receipt:
[`evidence/campaigns/2026-10-05-phase5-8-cf8048d/`](evidence/campaigns/2026-10-05-phase5-8-cf8048d/).

### Exact DEFLATE replay (Phase 6) measured results

Phases 4–5.8 all proceduralize **plain** syntax that `BYTE_RANS` already models
well. Phase 6 attacks a different layer: bytes the producer has **already
entropy-coded**. The `DEFLATE_REPLAY` DRA op (opcode `0x0A`, DRA **v8**, universe
`phase6;…;dra-8;…+deflate-replay-preflate-0.7.6-experimental`) reconstructs the
*original* raw DEFLATE bitstream of a `/FlateDecode` stream from `(plaintext,
corrections)`. Its `replay_codec` tag names the correction representation as an
experimental, version-coupled preflate-0.7.6 layout (not frozen v1), and an
unknown tag fails closed. The
byte-authoritative scanner owns stream discovery and `/Filter` classification;
`preflate` never discovers streams. A lexer fix makes `stream`+EOL payloads opaque
spans. Two candidates use the op: `PDF_DEFLATE_REPLAY` (raw, deduplicated
plaintext objects) and `PDF_DEFLATE_REPLAY_RANS` (each unique plaintext is one
shared order-0 byte-rANS channel). The sealed campaign
`2026-10-05-phase6-0d0bb79` runs the forced-candidate ablation over the 12-file
corpus:

```text
A0 RAW                      = 139950
A2 + BYTE_RANS              =  98560
A6 + PDF_LAYOUT_RANS        =  98560
A7 + PDF_DEFLATE_REPLAY     =  98560
A8 + PDF_DEFLATE_REPLAY_RANS=  85371     leave-one-out replay-rANS delta = -13189
```

Forced sizes on `flate.pdf` (57,513 B):

```text
RAW 57908   BYTE_RANS 49291   PDF_DEFLATE_REPLAY 56736   PDF_DEFLATE_REPLAY_RANS 36102
```

`PDF_DEFLATE_REPLAY_RANS` is the auto winner on `flate.pdf` and **beats
`BYTE_RANS` by 13,189 B**. The six `FlateDecode` streams are the content
plaintext `p1` at levels 9/6/1/0 (one shared plaintext, four appearances), a
graphics stream `p2` at level 6, and an incompressible stream `p3` at level 6
that DEFLATE stores; the descriptor replays all six and codes their **3 unique
plaintexts** as order-0 channels (`streams=6 replayed=6 channels=3 objects=6`).
Of `p1`'s four appearances only the level-0 stream is weakly coded (stored);
level 1 is ~19% of the plaintext and levels 6/9 are strong, so the win needs the
shared plaintext to *also* have a large/weakly-coded appearance. `BYTE_RANS`
order-0-codes the six streams to 49,291 B; it does not carry them verbatim. The
raw-plaintext variant loses (56,736 B) because a strongly-compressed stream's
plaintext is nearly as large as the stream it replaces. Head-to-head vs
`BYTE_RANS`: **win 1, lose 0, decline 11** (the other files have no lone
`FlateDecode` stream); every auto winner is exact (`cmp` + `verify`). This is
**one composed sample**, at commit `0d0bb79`, measured on a single synthetic
fixture: the win requires a shared plaintext that also has a large/weakly-coded
appearance (unique strongly-compressed streams lose, by up to 4.46Γ—), and the
losing region is unique, strongly-compressed plaintext. It is the first measured
positive for a PDF structural candidate
(ADR-0015); the plain-syntax converging negatives (ADR-0010–ADR-0013) stand.

Receipt:
[`evidence/campaigns/2026-10-05-phase6-0d0bb79/`](evidence/campaigns/2026-10-05-phase6-0d0bb79/).

### Producer-stratified Flate ratio (Phase 7.0) measured results

`tools/pdf-corpus.sh` builds a locally-generated corpus from distinct producer
lineages (Ghostscript 10.00.0 at five `/PDFSETTINGS`, qpdf 11.3.0 in four modes,
a hand-written stored-block-zlib base, plus the Phase-3 synthetic set), and
`vole-document deflate-stats` measures the exact-replay ratio per `FlateDecode`
stream β€” `correction/compressed`, `(plaintext+corr)/compressed`,
`(rANS(plaintext)+corr)/compressed` β€” with p10/p50/p90 by producer, the
exact-replay acceptance rate, and every decline. It also reports two aggregate
complete costs so shared plaintext is not overcounted:
`replayed_rans_full_bytes` (naive per-stream sum) and
`replayed_rans_dedup_bytes` (one charge per unique plaintext + one per unique
correction, matching the shared-channel candidate).

On 24 `FlateDecode` streams (re-measured after the Stage-A lexer stream-boundary
fix): **24 replayed, 0 declined (acceptance 1.000)**. The pre-fix run saw only 17
streams (11 replayed, 6 declined); the census rose because the old over-read had
swallowed whole stream objects (every qpdf-generated file was undercounted;
`qpdf-preserve-objectstreams.pdf` had reported zero). The Phase-6 win region
appears **only in our own hand-authored fixtures** (`hand-base2.pdf`: deduped
rANS 55,531 vs naive 111,062; `flate.pdf`: 34,051 vs 89,437). The same geometry
in `qpdf-preserve-objectstreams.pdf` (55,531 vs 111,062) is present **only
because qpdf `--object-streams=preserve` copied and renumbered the two
byte-identical raw streams already authored in `hand-base2.pdf`** (confirmed via
`qpdf --raw-stream-data`: all four stream hashes are `ec028dc1…`); the fixture
already wins 112,011 β†’ 56,885 B and qpdf adds +41 B, so **99.93% of the reported
qpdf win is inherited**. No genuinely transformed producer output exhibits the
region. Corpus-wide `correction/compressed` p10/p50/p90 = 0.000320 / **0.014716**
/ 0.097360; the `0.004518` median is the `pdf-make-samples` subset only.
The six former declines were all `not_zlib` **because of a scanner locality
limitation** β€” the lexer's `find_endstream` required an EOL before `endstream`,
which Ghostscript omits (spec "should", qpdf-tolerated) β€” not because those
streams are not zlib (they begin `78 9c`, confirmed via `qpdf --raw-stream-data`).
That limitation is **fixed** (commit `c4eb77e`): `find_endstream` now locates the
`endstream` keyword by right-termination. This is a **scoped** result about
locally generated files; qpdf/Ghostscript are transformers, not authoring apps,
and browser/office/TeX families remain a recorded gap. No candidate is adopted.

Receipt:
[`evidence/campaigns/2026-10-05-phase7-corpus-b-c4eb77e/`](evidence/campaigns/2026-10-05-phase7-corpus-b-c4eb77e/)
(amendment; supersedes the original
[`2026-10-05-phase7-corpus-f1f8d26/`](evidence/campaigns/2026-10-05-phase7-corpus-f1f8d26/));
report: [`docs/evidence/phase7-corpus-report.md`](docs/evidence/phase7-corpus-report.md).

**Complete-cost court (the decisive Phase-7.0 measurement).** The ratio harness
is a diagnostic; the court decides. A second sealed campaign
(`2026-10-05-phase7-court-99dc72e`, commit `99dc72e`, driver
`tools/pdf-court.sh`) runs the real CLI over all 23 locally generated corpus
files β€” unforced and with `--force raw|byte-rans|pdf-deflate-replay|pdf-deflate-replay-rans`
β€” and compares complete serialized `.voldoc` sizes. `PDF_DEFLATE_REPLAY_RANS` vs
`BYTE_RANS`: **win 3, lose 8, decline 12 β€” all 3 wins self-authored**. The wins
are exactly the Phase-6 shared-plaintext geometry, and in all three the unforced
court picks `PDF_DEFLATE_REPLAY_RANS`:

```text
qpdf-preserve-objectstreams.pdf  BYTE_RANS 112147 -> PDF_DEFLATE_REPLAY_RANS 56980  (-55167)
hand-base2.pdf                   BYTE_RANS 112011 -> PDF_DEFLATE_REPLAY_RANS 56885  (-55126)
_synthetic/flate.pdf             BYTE_RANS  49291 -> PDF_DEFLATE_REPLAY_RANS 36102  (-13189)
```

The reported "qpdf win" is not a transformed-producer result:
`qpdf --object-streams=preserve` copied and renumbered the two byte-identical raw
streams (`ec028dc1…`) already authored in our `hand-base2.pdf` fixture (the
fixture wins 112,011 β†’ 56,885 B; qpdf adds only **+41 B**, so **99.93% of the
55,167 B win is inherited**). Every **genuinely transformed** producer output
loses or declines on this corpus: all 5 Ghostscript variants and both qpdf
compression variants (unique, strongly-compressed plaintext) **lose**, and the 12
files with no replayable Flate lane **decline** (a forced lane the input does not
propose is a typed `Usage` error, recorded `null`). So on this locally generated
corpus exact replay **wins 3 / loses 8 / declines 12, and all 3 wins are
self-authored** (two fixtures plus a preserved copy of one). `--deterministic-id`
is a `/ID`-only normalization (55,165 B no-flag vs 55,167 B flagged): it makes
the court reproducible but neither creates nor destroys the win. The Phase-6 win
is real and byte-exact, but its enabling condition (a plaintext shared across
streams with a large/weakly-coded appearance) is **not produced by the tested
transformers**, which motivates Phase 7.2 (nested content proceduralization). All
23 auto winners are `verify` + `cmp` byte-exact. The corpus is **locally
generated** and is **not a population sample**; qpdf and Ghostscript are
**transformers, not authoring applications**, and browser/PDFium, LibreOffice,
pdfTeX and Adobe outputs remain a recorded gap. No candidate is adopted and no
wire format changed.

Receipt:
[`evidence/campaigns/2026-10-05-phase7-court-99dc72e/`](evidence/campaigns/2026-10-05-phase7-court-99dc72e/).

### Generator-family Flate corpus (Phase 7.0b) measured results

The Phase-7.0 gap was that qpdf and Ghostscript are **transformers** and never emit
the shared plaintext the win region needs. Phase 7.0b adds a separate, opt-in
`producers` image (`Dockerfile` stage + compose service, base
`debian:bookworm-slim@sha256:3783cc01…`, the same digest as `tools`; ~724 MB) with
four real **authoring generators**: ReportLab 3.6.12, Cairo 1.20.1/libcairo 1.16.0,
LibreOffice Writer 7.4.7.2, and pdfTeX 3.141592653-2.6-1.40.24. All four ran.
`tools/pdf-corpus-producers.sh` generates one PDF per family from the same
deterministic content document as `tools/pdf-corpus.sh`, fingerprints every
`/FlateDecode` payload, and applies
`qpdf --deterministic-id --stream-data=preserve --object-streams=preserve` only when
it leaves every payload byte-identical (ReportLab, Cairo, LibreOffice; pdfTeX kept
raw). ReportLab/Cairo/pdfTeX are byte-reproducible; LibreOffice is not.

`deflate-stats`: **87 Flate streams, 87 replayed, 0 declined β†’ acceptance 1.000**;
corpus-wide `correction/compressed` p10/p50/p90 = 0.034759 / 0.047945 / 0.068028.

Complete-cost court: `PDF_DEFLATE_REPLAY_RANS` vs `BYTE_RANS` β€” **win 1 / lose 3 /
decline 0**:

```text
cairo-vector.pdf        BYTE_RANS  58711 -> PDF_DEFLATE_REPLAY_RANS  34574  (-24137)  WIN
reportlab-multipage.pdf BYTE_RANS  11144 -> PDF_DEFLATE_REPLAY_RANS  14456  (+3312)   lose
pdftex-doc.pdf          BYTE_RANS  24435 -> PDF_DEFLATE_REPLAY_RANS  35098  (+10663)  lose
libreoffice-export.pdf  BYTE_RANS  72791 -> PDF_DEFLATE_REPLAY_RANS 375265 (+302474)  lose
```

**Correction (Phase 7.0c): the Cairo "win" is a harness repeated-bytes artifact.**
An independent adversarial review found that `tools/pdf-corpus-producers.sh` draws
**one identical page six times** (no per-page variation), so Cairo emits six streams
whose **compressed bytes are identical *and* whose plaintexts are identical**. The
court therefore cannot distinguish plaintext-sharing from plain compressed-byte
repetition, and generic LZ captures far more of the same redundancy: on
`cairo-vector.pdf`, `gzip -9` = **17,382 B**, `zlib -9` = 17,376 B and `xz -9e` =
**16,852 B** β€” about **half** the 34,574 B reported as the "winning"
`PDF_DEFLATE_REPLAY_RANS` size. `BYTE_RANS` is a weak order-0 baseline with no LZ, so
the βˆ’24,137 B delta is a win over an order-0 lane on repeated identical bytes. The
phrases "a genuine authoring application does produce the win region", "the producer
creates the geometry" and "first authoring-generator witness" are **withdrawn**.
Correct characterization: *our deterministic generator repeated one identical page
six times; Cairo emitted six byte-identical streams (compressed bytes and plaintext
both identical); this witnesses a repeated-identical-bytes region already captured
better by generic LZ, not the shared-plaintext-vs-distinct-compression mechanism.*
ReportLab and pdfTeX are the same story (identical repeats) and **lose** at complete
cost; LibreOffice shares nothing. The delta is **conditional** on repeated identical
page content and is **not** a population claim; prior "wins" were measured against a
weak order-0 baseline and the generic ladder (below) is the honest comparison. No
candidate or wire format changed. Review:
[`docs/evidence/phase7b-skeptic-review.md`](docs/evidence/phase7b-skeptic-review.md).

Receipt:
[`evidence/campaigns/2026-10-05-phase7-producers-e071250/`](evidence/campaigns/2026-10-05-phase7-producers-e071250/);
script `tools/pdf-corpus-producers.sh`; ledger
[`evidence/corpus/phase7-producers/provenance.json`](evidence/corpus/phase7-producers/provenance.json).

### Generic-compressor baseline ladder (Phase 7.0c) measured results

Phase 7.0/7.0b measured VOLE only against `BYTE_RANS`, a whole-file **order-0
byte-rANS** lane with no LZ. Phase 7.0c adds the honest comparison: a pinned,
opt-in `baseline` image (the pinned `dev` base plus `gzip`, `zstd`, `xz`, `brotli`,
`jq`) and `tools/baselines.sh`, which for every corpus file records the smallest
lossless complete-file size for `gzip -9`, `zstd -19 --long=27`, `xz -9e` and
`brotli -q 11` (each round-trip verified) and the complete serialized `.voldoc`
size of every VOLE lane, then compares against the best VOLE lane.

**Result: on 27 corpus files the best VOLE lane beats gzip/zstd/xz/brotli on 0
files.** The best generic compressor is smaller on every file, by **+460,320 B**
total (phase7 +410,355 over 23 files; producers +49,965 over 4). On the Cairo file
the "winning" 34,574 B is **2.07Γ—** brotli's 16,670 B; on `flate.pdf` the 36,102 B
is **1.91Γ—** xz's 18,884 B:

```text
cairo-vector.pdf        source 58424  gzip 17382  zstd 16836  xz 16852  brotli 16670  BYTE_RANS 58711  best VOLE 34574  (+17904 vs brotli)
_synthetic/flate.pdf    source 57513  gzip 22426  zstd 20171  xz 18884  brotli 18891  BYTE_RANS 49291  best VOLE 36102  (+17218 vs xz)
```

Prior "wins" were relative to a weak order-0 baseline and do not survive the
generic ladder. VOLE's byte-exact structural reconstruction is unchanged; its
*compression* claim does not survive on this corpus. All 27 auto winners are
`verify` + `cmp` byte-exact; no candidate or wire format changed. Review:
[`docs/evidence/phase7b-skeptic-review.md`](docs/evidence/phase7b-skeptic-review.md).

Receipt:
[`evidence/campaigns/2026-10-05-phase7-baselines-7b9f662/`](evidence/campaigns/2026-10-05-phase7-baselines-7b9f662/)
(full per-file table `baseline-table.md`); driver `tools/baselines.sh`.

### Partial materialization / observation views (Phase 7.3) measured results

Whole-file compression loses (above), so Phase 7.3 measures the pivoted axis β€”
**random-access query cost**. `view` serves one output range from a descriptor
carrying an optional, advisory `OBSERVATION_INDEX`: when the program is a
sequence of linear independent ops it evaluates only the ops intersecting
`[a,b)` and decodes only the referenced entropy channels, and every served slice
is `cmp`'d against the full materialization. The corpus is a 33,789,340 B
(32.22 MiB), 800-stream deterministic PDF from the encode-time `pdf-make-large`
subcommand (`qpdf --check` rc 0; 800 replayed / 0 declined).

**Result: a scoped positive on decode CPU, byte-exact on all 18
pre-registered queries.** For any query at β‰₯ ~2 MiB the indexed lane touches
~0.41–0.43 MB (`descriptor_bytes_traversed` alone; its `entropy_bytes_decoded`
breakdown is a subset already counted there and must not be added) regardless
of offset, while gzip must inflate `a + len`; in the late region (β‰₯ 50 % in) that
is **1.4–2.3 %** of gzip's bytes, and VOLE is **~2–5Γ— faster than gzip** and
**~4–13Γ— faster than xz** on CPU. It **loses** in the early region (≀ ~8–16 MiB),
**never beats zstd's raw decompressor** on wall time, uses ~38 MB peak RSS vs
gzip's ~1.2 MB, and whole-file size is still **2.98Γ—** xz. The decisive caveat:
`view` reads and parses the **whole** descriptor, so on-disk I/O is **not**
reduced (no bytes-read win yet; `descriptor_bytes_traversed` is a CPU-side
approximation) β€” an mmap/seek reader is the prerequisite for that claim.

Receipt:
[`evidence/campaigns/2026-10-05-phase7-partial-a5764c9/`](evidence/campaigns/2026-10-05-phase7-partial-a5764c9/)
(`query-table.md`, `report.md`); report
[`docs/evidence/phase7-partial-report.md`](docs/evidence/phase7-partial-report.md);
drivers `tools/partial-court.sh`, `tools/partial-table.jq`; ADR-0018.

### Seek-based partial I/O (Phase 8.3) measured results

Phase 8 supplies the mmap/seek reader ADR-0018 required. A descriptor may carry
an optional seek `DIRECTORY` record (first record, fixed offset 64,
`FLAG_OPTIONAL`, cross-checked and never trusted); the `view` CLI peeks only the
64-byte header and serves from a `Read + Seek` reader that fetches those record
*classes* the query needs (header, DIRECTORY, GRAPH, OBSERVATION_INDEX,
INTEGRITY, and the one referenced object/channel/model) β€” it never `fs::read`s
the whole descriptor. This is a **floor**, not a small read: a 256-byte request
still incurs ~440 KB (~1,700Γ—). A partial read is an *observation*
(`integrity_verified == false`); `materialize`/`decode`/`verify` remain the
archival authority.

**Result: a scoped bytes-read win versus non-seekable sequential codecs,
byte-exact on all 18 pre-registered queries.** On the same 33,789,340 B
(32.22 MiB) corpus the seekable descriptor is 17,566,832 B (DIRECTORY 27,390 B).
The seeked `view` reads a **constant 439,679–461,367 B** regardless of offset
(floor = header 64 + DIRECTORY 27,390 + GRAPH 265,462 + OBSERVATION_INDEX
146,711 + INTEGRITY 52), ≀ 2.6 % of the descriptor for every query. In the late
region (β‰₯ 50 % in, 8/8 queries) that is **4.7 %–~21Γ— fewer bytes than gzip's
compressed prefix** (9,764,864 B vs 460,713 B at 31 MiB) and ~12–13Γ— fewer than
zstd/xz; a `strace -P` descriptor-file cross-check equals the instrumented count
+ exactly 64 B (the header peek). CPU drops to ~0.00 s and peak RSS from ~38 MB
to **~3.8 MB**.

**Not a general random-access-I/O win (Phase 8.4 amendment).** Against
**seekable/blocked** formats, for the same late query: bgzip (BGZF) reads
**23,808 B** (~19Γ— fewer than VOLE's 460,713 B), `xz --block-size=64KiB`
**15,344 B** (~30Γ— fewer) and `1MiB` **179,892 B** (~2.6Γ— fewer) β€” and BGZF's
whole file (7,995,600 B) is even smaller than the seekable descriptor. VOLE only
wins where the block size is large (`xz --block-size=4MiB` 708,612 B; pixz
2,810,832 B, 16 MiB blocks). See `docs/evidence/phase8-skeptic-review.md`.

**Where it loses (recorded).** At `a = 0` the constant floor exceeds gzip's first
bytes (327,680 B) and xz's (73,728 B); at early queries (≀ ~1.7 MiB) it loses to
xz's tiny compressed prefix; and versus compact-block seekable formats it loses
across the late region. The floor is constant in the offset, so it would
dominate a descriptor smaller than ~9 MB β€” a large-document mechanism. Whole-file
size is still **3.01Γ—** xz (and 0.46Γ— BGZF). One locally generated corpus;
**no population claim**.

**Validator caveat.** A *referenced* channel's `decoded_length` is cross-checked
against its record; an *unreferenced* `CHANNEL_LENGTHS` entry is not β€” benign,
since `analyze_ops` never uses an unused length, so no wrong bytes are served.

Receipt:
[`evidence/campaigns/2026-10-05-phase8-seek-08de2a9/`](evidence/campaigns/2026-10-05-phase8-seek-08de2a9/)
(`query-table.md`, `report.md`, `seekable.jsonl`, `seekable-table.md`,
`seekable-report.md`); reports
[`docs/evidence/phase8-seek-report.md`](docs/evidence/phase8-seek-report.md),
[`docs/evidence/phase8-skeptic-review.md`](docs/evidence/phase8-skeptic-review.md);
drivers `tools/seek-court.sh`, `tools/seek-table.jq`,
`tools/seekable-baselines.sh`; ADR-0019.

### Encoder-only search governance (Phase 10.1) measured results

The brief's Phase 10 asks for a **DSFB**-style search governor with **zero decode
authority**. Empirically, the published `dsfb 0.1.2` crate is **Drift-Slew Fusion
Bootstrap state estimation** β€” a Kalman-like `f64` observer over sensor channels
with no candidate / residual / search-directive API β€” present in the lockfile only
transitively through the optional EntropyFS engine. It is therefore recorded as
*unavailable-for-purpose* and is **not** a dependency: the feature is the
**dependency-free** `dsfb-search = []`.

The governor (`src/encode/governor.rs`) is encoder-only and integer-only. It
produces typed residual diagnostics, proposes candidates from a tiny parametric
space over *existing* mechanisms (`scale_bits`, `partition`, `replay`, `packed`,
`depth`), and maps the incumbent's dominant residual to a directive with a pure
`govern` function. Every candidate it can enable is serialized, parsed,
materialized, byte-compared, and priced by the *same* complete-cost court; no
governance state is persisted and `src/container/header.rs` is untouched, so a
governer-produced descriptor decodes byte-exactly in a build **without** the
feature.

The court (`tools/governor-court.sh`, campaign
`2026-10-05-phase10-governor-d2b09c9`, ADR-0022) compares `Exhaustive`,
`FixedHeuristic`, and `DsfbGuided` over a small deterministic cohort split into
disjoint tune/holdout/control sets. Pre-registered hypotheses: **H1 HELD**
(`DsfbGuided.final ≀ FixedHeuristic.final` everywhere); **H2 HELD**
(`guided.final == exhaustive.final` on 8/8 holdout with ≀ Β½ the candidates);
**H4 HELD** (negative controls `Stop(Raw)` and match the RAW descriptor
byte-for-byte). But **H3 also HOLDS**, and it is the substantive result: the fixed
heuristic already attains the exhaustive minimum on **every** workload, so the
parametric search adds **zero bytes**. Representative rows (final bytes /
candidates): `flate.pdf` 36,161 (fixed 10, exhaustive 438, guided 150);
`bigtext.pdf` 38,274 (7/414/126); `many.pdf` 5,301 (7/414/126). **The fixed
complete-cost court is retained**; the governor remains an optional, encoder-only
feature. Small locally generated cohort; **no population claim**.

Receipt:
[`evidence/campaigns/2026-10-05-phase10-governor-d2b09c9/`](evidence/campaigns/2026-10-05-phase10-governor-d2b09c9/)
(`results.json`, `hypotheses.json`, `decode-proof.json`, `report.md`); driver
`tools/governor-court.sh`; ADR-0022.

## Quick start (Docker only)

All project commands run inside pinned containers. The host only invokes Docker.

```sh
# Build the toolchain image (pinned by digest in Dockerfile)
docker compose build dev

# Build, test, lint, format
docker compose run --rm --no-TTY dev cargo test  --all-features
docker compose run --rm --no-TTY dev cargo clippy --all-targets --all-features -- -D warnings
docker compose run --rm --no-TTY dev cargo fmt --all --check

# MSRV gate (Rust 1.89)
docker compose build msrv
docker compose run --rm --no-TTY msrv cargo build --locked

# Generic-compressor baseline ladder (Phase 7.0c; opt-in baseline image)
docker compose build baseline
docker compose run --rm --no-TTY dev cargo build --locked --all-features
docker compose run --rm --no-TTY -e BASELINE_CORPUS=phase7 baseline \
  sh tools/baselines.sh /tmp/baselines.json evidence/corpus/phase7

# Partial-materialization query court (Phase 7.3; opt-in baseline image)
docker compose run --rm --no-TTY dev \
  ./target/debug/vole-document pdf-make-large evidence/corpus/phase7-large 800
docker compose run --rm --no-TTY dev ./target/debug/vole-document encode \
  --force pdf-deflate-replay-rans-indexed evidence/corpus/phase7-large/large.pdf /tmp/large.voldoc
docker compose run --rm --no-TTY baseline \
  sh tools/partial-court.sh /tmp/queries.jsonl evidence/corpus/phase7-large/large.pdf \
  /tmp/large.voldoc /tmp/large.pdf.gz /tmp/large.pdf.zst /tmp/large.pdf.xz

# Seek-based partial-I/O bytes-read court (Phase 8.3; opt-in baseline image)
docker compose build baseline   # adds strace to the dev-derived toolchain
docker compose run --rm --no-TTY dev ./target/debug/vole-document encode \
  --force pdf-deflate-replay-rans-indexed evidence/corpus/phase8-large/large.pdf /tmp/large.seek.voldoc
docker compose run --rm --no-TTY baseline \
  sh tools/seek-court.sh /tmp/queries.jsonl evidence/corpus/phase8-large/large.pdf \
  /tmp/large.seek.voldoc /tmp/large.pdf.gz /tmp/large.pdf.zst /tmp/large.pdf.xz

# Seekable/blocked random-access baseline court (Phase 8.4; opt-in baseline image)
# adds tabix(bgzip)+pixz to the baseline stage, then builds bgzip/xz-blocked/pixz
# and measures the honest random-access cost (covering block(s) + index).
docker compose run --rm --no-TTY baseline \
  sh tools/seekable-baselines.sh /tmp/seekable.jsonl evidence/corpus/phase8-large/large.pdf \
  /tmp/large.seek.voldoc /tmp/seekable
jq -rs -f tools/seekable-table.jq /tmp/seekable.jsonl > /tmp/seekable-table.md

# Phase 1 exact court (writes an evidence receipt)
docker compose run --rm --no-TTY dev sh tools/phase1-court.sh
```

CLI surface (activated by the pipeline, not by extension β€” extensions are hints,
never authority):

```text
vole-document encode      [--force KIND] INPUT   OUTPUT.voldoc
vole-document decode      INPUT.voldoc   OUTPUT
vole-document materialize INPUT.voldoc   OUTPUT
vole-document view        INPUT.voldoc [OUTPUT] --byte-range A:L | --pdf-object N:G | --pdf-stream N:G | --pdf-revision I [--stats]
vole-document verify      INPUT.voldoc
vole-document inspect     INPUT.voldoc
vole-document pdf-make-large DIR [OBJECTS]
vole-document capabilities
```

`encode --force KIND` forces the complete-cost court to consider only one
candidate family (`raw`, `rle`, `byte-rans`, `pdf-physical`, `pdf-channels`,
`pdf-layout`, `pdf-layout-rans`, `pdf-deflate-replay`, `pdf-deflate-replay-rans`)
for honest per-mechanism ablation; it never
bypasses exactness, and it fails with a typed usage error when the input does not
propose that kind.

## Fuzzing

Two layers, both Docker-only:

- **Deterministic property/mutation courts** (`tests/property.rs`,
  `tests/goldens.rs`, `tests/malformed.rs`) plus the longer soak run
  `tools/soak-fuzz.sh` (`VOLE_FUZZ_ITERS`, default 200000).
- **Coverage-guided libFuzzer targets** (Phase 7.1) in the standalone `fuzz/`
  `cargo-fuzz` package (excluded from `cargo package`), built and run in the
  pinned dated-nightly `fuzz` Docker service:

  ```sh
  docker compose build fuzz
  docker compose run --rm --no-TTY fuzz cargo fuzz build
  FUZZ_SECONDS=60 docker compose run --rm --no-TTY fuzz sh tools/fuzz.sh
  ```

Ten targets cover the `.voldoc` container parser/materializer, the
`encode`β†’`decode` round trip, the DRA decoder/analyzer/evaluator, the rANS model
and channel decoders, the PDF lexer and physical scanner, xref/`/Prev`, the
`DEFLATE_REPLAY` wrapper, and `verify`. See `fuzz/README.md` for the pinned
toolchain, seeds, and regeneration. The sealed campaign
`evidence/campaigns/2026-10-05-phase7-fuzz-ca6a92b/` observed zero crashes on
nine of ten targets and reported two upstream `preflate-rs` findings (one
mitigated fail-closed, one recorded upstream resource limitation).

## Repository layout

```text
src/            one crate; modules for architectural separation
tests/          exact / malformed / conformance courts
fuzz/           cargo-fuzz coverage-guided targets (excluded from the crate)
tools/          court and gate scripts (run inside Docker)
docs/           architecture, ADRs, security, phase notes
evidence/       immutable campaign receipts (machine-readable)
research/       LOCAL ONLY β€” gitignored (paper, snapshots, subagent findings)
```

`research/` is intentionally excluded from version control. Durable findings that
matter to a phase are frozen into ADRs and phase notes and referenced from
receipts by hash.

## Licensing

Dual-licensed under either MIT or Apache-2.0, at your option. See
[`LICENSE-MIT`](LICENSE-MIT) and [`LICENSE-APACHE`](LICENSE-APACHE).

**Third-party license note.** The `deflate-replay` feature (Phase 6) is
**opt-in** and depends on `preflate-rs`, which depends on `cabac`, licensed
**LGPL-3.0-or-later**. The default build (`default = ["rans"]`) is
**permissive-only** and contains no LGPL code. Rust links statically by default,
so a binary built **with** `--features deflate-replay` (or `--all-features`)
contains LGPL code and carries the corresponding obligations. See
[ADR-0014](docs/adr/0014-lgpl-cabac-dependency.md).