# `.voldoc` wire format (provisional)
> **Status: PROVISIONAL, not frozen v1.** The layout below is implemented and
> tested, but it is not a stability commitment. Version identity is carried both
> in the header and in the universe declaration string, and the universe string
> changes whenever any opcode, coder, limit semantic, adapter meaning, or hash
> semantic changes. See `PROJECT_STATE.md` and `docs/adr/` for the freeze policy.
All multi-byte integers are little-endian. Offsets are byte offsets from the
start of the file.
## File
```text
file := header record*
header := 64 bytes (fixed, see below)
record := tag:u8 flags:u8 reserved:u16=0 length:u32 payload:[u8;length] crc32c:u32
```
`crc32c` is CRC-32C (Castagnoli) over the 8-byte record header *and* the payload.
`reserved` must be zero. The final record is always `TRAILER`.
## Header (64 bytes)
| 0 | 8 | `magic` | `56 4F 4C 44 4F 43 1A 00` = ASCII `VOLDOC` + `0x1A` + `0x00` (source of truth: `MAGIC`) |
| 8 | 2 | `major` | format major version (this build: `0`) |
| 10 | 2 | `minor` | format minor version (this build: `1`) |
| 12 | 4 | `mandatory_features` | unknown bits fail closed |
| 16 | 4 | `optional_features` | may be ignored |
| 20 | 1 | `exactness_profile` | `0 = EXACT_BYTES` (only normative value) |
| 21 | 1 | `source_format` | `0 = OPAQUE`, `1 = PDF` (Phase 3); unknown fails closed |
| 22 | 2 | `reserved_a` | must be `0` |
| 24 | 16 | `universe_id` | first 16 bytes of `SHA-256(universe_string)` |
| 40 | 8 | `declared_source_len` | exact reconstructed length |
| 48 | 12 | `reserved_b` | must be `0` |
| 60 | 4 | `header_crc32c` | CRC-32C over bytes `[0,60)` |
The magic constant is defined once in `src/container/header.rs` (`MAGIC`); this
document defers to that source of truth.
## Records
| `0x01` | `UNIVERSE` | UTF-8 universe declaration string |
| `0x02` | `FORMAT` | `source_format:u8`, `basis_len:u32`, `basis:[u8]` |
| `0x10` | `OBJECT` | raw object bytes |
| `0x20` | `GRAPH` | encoded reconstruction program |
| `0x30` | `MODEL` | canonical dense entropy model (see [Entropy records](#entropy-records-phase-2)) |
| `0x40` | `ENTROPY_CHANNEL` | typed channel capsule: 33-byte header + renorm payload |
| `0x50` | `RESIDUAL` | reserved (later) |
| `0x60` | `CHECKPOINT` | reserved (later) |
| `0x70` | `INDEX` | reserved (later) |
| `0x80` | `EXTERNAL_REF` | reserved (Phase 9+) |
| `0xF0` | `INTEGRITY` | `sha256:[u8;32]`, `source_len:u64` |
| `0xFF` | `TRAILER` | `record_count:u32`, `payload_bytes:u64`, `magic:[u8;8]` |
Rule: **unknown record whose `flags` lacks `FLAG_OPTIONAL` (`0x01`) fails closed**
with `UnsupportedFeature`. Unknown explicitly-optional records are skipped.
Requirements enforced by `Descriptor::parse`:
- exactly one `UNIVERSE`, `FORMAT`, `GRAPH`, `INTEGRITY`, `TRAILER`;
- `SHA-256(universe)[..16] == header.universe_id`;
- `FORMAT.source_format == header.source_format`;
- `INTEGRITY.source_len == header.declared_source_len`;
- `TRAILER.record_count` equals the number of records actually read;
- no record after `TRAILER`.
## Graph (reconstruction program)
```text
graph := version:u8=5 op_count:u32 op*
op := EMIT_OBJECT(0x01) u32_object_id
| INLINE(0x02) u32_len [u8;len]
| REPEAT_LAST(0x03) u32_count
| DECODE_CHANNEL(0x04) u32_channel_id
| INTERLEAVE_CHANNELS(0x05) u32_kinds_channel u32_lengths_channel \
u32_first_payload_channel u8_payload_channel_count
| MARK_OFFSET(0x06) u8_slot
| EMIT_OFFSET(0x07) u8_slot u8_width
| PACK_SEGMENTS(0x08) u32_data_object u32_item_count item*
item := LITERAL(0x01) u32_leb128_len
| MARK(0x02) u8_slot
| EMIT(0x03) u8_slot u8_width
```
Semantics:
- `EMIT_OBJECT` appends the referenced object's bytes (authority: **Literal**).
- `INLINE` appends inline bytes (authority: **Literal**).
- `REPEAT_LAST` repeats the bytes produced by the *immediately preceding literal
instruction* `count` more times (authority: **Generated**). Consecutive
`REPEAT_LAST` and a leading `REPEAT_LAST` are invalid.
- `DECODE_CHANNEL` appends the exact decoded bytes of the referenced entropy
channel (authority: **EntropyChannel**). The channel's decoded length is known
statically from its descriptor, so the op's output length is still bounded at
parse time.
- `INTERLEAVE_CHANNELS` (introduced in DRA version 3) reconstructs a byte string
from a **kind channel**, a **length channel**, and a contiguous run of
**payload channels** (authority: **EntropyChannel**):
- `kinds_channel` holds one kind byte per token, in file order;
- `lengths_channel` holds one little-endian `u32` per token, aligned with the
kind stream (its decoded length must be exactly `4 * token_count`);
- payload channel `first_payload_channel + k` carries the concatenated bytes
of every token whose kind is `k`, in file order, for
`k in 0..payload_channel_count`.
Evaluation walks the kind/length sequence, and for each token appends the next
`length` bytes of the payload channel named by its kind, advancing a per-kind
cursor. All indices, the kind range, and cursor bounds are validated before
allocation; a kind outside the payload range, misaligned kind/length counts,
a length overrun, or unconsumed payload bytes is rejected as
`InvalidGraph`/`CoverageViolation`. The op is bounded and non-Turing-complete
like the rest of the DRA.
- `MARK_OFFSET` (introduced in DRA version 4) records the current output position
(a `u64`) into the named slot and emits **no** bytes (authority: **Generated**,
zero-length). `slot` is a `u8`, so every value `0..=255` is in range; the
program models [`MAX_OFFSET_SLOTS`](src/dra/program.rs) = **256** slots. Slot
`255` is **reserved for the most recent classic `xref` section start**; slots
`0..=254` are available to the layout builder for indirect-object introducer
offsets. Marking a slot does not disturb the pending `REPEAT_LAST` block.
- `EMIT_OFFSET` (introduced in DRA version 4) emits the decimal form of the
position most recently recorded in `slot`, left zero-padded with ASCII `0` to
exactly `width` bytes (authority: **Generated**). `width` is in `1..=20`. The
value is decoded only at materialization time, so analysis charges exactly
`width` output bytes; at materialization a value that needs more than `width`
digits is rejected as `InvalidGraph`. A slot must have been marked by an
earlier `MARK_OFFSET` in program order, or the op is rejected during analysis
(before allocation) with `InvalidGraph`. Prediction is deterministic and never
invents bytes: the layout builder emits an `EMIT_OFFSET` only when it has
verified that the marked position reproduces the source digits, and otherwise
falls back to a literal `INLINE`.
- `PACK_SEGMENTS` (introduced in DRA version 5) reconstructs output from a
**compact item table over one data object** (authority: **Literal** for
`Literal` items, **Generated** for `Mark`/`Emit`), amortizing per-segment op
framing. `data_object` is an index into the descriptor's object table; the
table holds exactly `item_count` items:
- `LITERAL(0x01) { len }` copies the next `len` contiguous bytes of the data
object and advances the data cursor; `len` is a `u32` LEB128 varint;
- `MARK(0x02) { slot }` records the current output position (a `u64`) into
`slot` and emits no bytes, exactly like `MARK_OFFSET`;
- `EMIT(0x03) { slot, width }` emits the marked decimal position, left
zero-padded to `width` bytes, exactly like `EMIT_OFFSET`.
The item table is bounded by the same limits as the graph and is validated
before allocation: an unknown item tag, a truncation, a slot that was never
marked, a `width` beyond `1..=20`, a literal run that overruns the data object,
or a data object not consumed **exactly** is rejected (`InvalidGraph` /
`CoverageViolation`). Like every other op it is deterministic and
non-Turing-complete, and it never invents bytes: an `Emit` is only present when
the builder verified the marked position reproduces the source digits.
## Coverage certificate (checked invariant, not stored bytes)
The coverage map is **derived** deterministically from the program and object
lengths during `parse`, and validated before any byte materialization:
```text
union(spans) == [0, declared_source_len)
spans are contiguous (no gaps, no overlapping authorities)
```
A descriptor whose predicted length disagrees with `declared_source_len`, or
whose spans are not contiguous, is rejected with `CoverageViolation` **before**
allocation. This catches the dangerous "the parser forgot a source distinction"
class of bugs at the representation boundary.
## Universe declaration
The Phase-5.7 (Phase-6 preparation) universe string is:
```text
vole-document;universe;phase6-prep;exact-bytes;dra-5;opaque+entropy+pdf+channels+offsets
```
The header's `universe_id` is the first 16 bytes of `SHA-256` over this string.
Any change to an opcode, coder, limit semantic, adapter meaning, or hash semantic
requires a new universe string. This supersedes the Phase-5 string
(`vole-document;universe;phase-5;exact-bytes;dra-4;opaque+entropy+pdf+channels+offsets`),
which superseded the Phase-4 string
(`vole-document;universe;phase-4;exact-bytes;dra-3;opaque+entropy+pdf+channels`).
## Entropy records (Phase 2, extended in Phase 4)
Phase 2 introduces two entropy records and one graph op (`MODEL`,
`ENTROPY_CHANNEL`, `DECODE_CHANNEL`). Phase 4 extends the `MODEL` wire to
version 2 (sparse/dense) and adds the `INTERLEAVE_CHANNELS` op and the typed
PDF-channel candidate layout. All of it is implemented and measured but remains
**PROVISIONAL** (the wire format is not frozen v1).
### `MODEL` (`0x30`)
A canonical, self-describing frequency table over the fixed 256-symbol byte
alphabet. Two wire versions are accepted on decode; **version 2** is what this
build emits:
```text
model_v2 := version:u8=2 form:u8 scale_bits:u8 payload
form = 0 (SPARSE): present_count:u16 (symbol:u8 freq:u16)*
form = 1 (DENSE): count:u16=256 freq:[u16;256]
model_v1 := version:u8=1 scale_bits:u8 count:u16=256 freq:[u16;256] (legacy, dense)
```
- **v2 form selection.** The encoder serializes to whichever form is *strictly*
smaller; on a tie it picks the dense form, so the mapping from model to bytes
stays deterministic. Sparse symbols must be strictly ascending and unique with
`freq >= 1`; the dense form carries all 256 little-endian `u16` frequencies.
- **Legacy v1.** The original dense payload (`[1][scale_bits][count=256][u16 x
256]`, exactly 516 bytes) is still decodable, so older descriptors remain
readable.
- `freq` entries are little-endian and **must sum to exactly `1 << scale_bits`**
in both versions; trailing or truncated payloads are rejected.
- `scale_bits` is in `1..=15` (frequencies are stored as `u16`).
- Normalization from observed counts is integer-only, deterministic, and
tie-broken by lower symbol index; symbols seen zero times get frequency zero.
- A channel's `scale_bits` must equal the `scale_bits` of the model it names, and
a channel may not name a missing model; both are checked during `parse`.
### PDF typed-channel candidate layout (Phase 4)
The `PDF_CHANNELS` candidate (`source_format = 1`) transposes the Phase-3 lexical
cover into a fixed set of channels — the indices below are a wire contract of
this candidate. With `KIND_COUNT = 12`:
```text
channel 0 kinds: one kind byte per token, in file order
channel 1 lengths: four little-endian bytes per token, aligned with kinds
channels 2..2+KIND_COUNT payloads[k]: bytes of every token of kind k, in file order
(with KIND_COUNT = 12 this is channels 2..13)
```
Each channel is independently order-0 byte-rANS coded with its own `MODEL`
(a model id per channel, in that order), and reconstruction is a single
`INTERLEAVE_CHANNELS` op with `kinds_channel = 0`, `lengths_channel = 1`,
`first_payload_channel = 2`, `payload_channel_count = KIND_COUNT`. Every model
and payload byte is charged in the complete-cost court; the candidate is
**proposed and measured but not adopted** — it loses to `BYTE_RANS` on this
corpus (see `PROJECT_STATE.md`, ADR-0010). This section is **PROVISIONAL**.
### `ENTROPY_CHANNEL` (`0x40`)
One typed channel capsule. The payload is a fixed **33-byte header** followed by
the renormalization payload; the total record payload length must equal
`33 + payload_len` exactly (no trailing bytes, no truncation).
| 0 | 1 | `coder` | `1 = CODER_ORDER0_BYTE_RANS` |
| 1 | 2 | `coder_version` | `1` |
| 3 | 1 | `scale_bits` | must match the referenced model |
| 4 | 1 | `lane_count` | `1` (single-lane only; else `UnsupportedFeature`) |
| 5 | 4 | `model_id` | index into the descriptor's `MODEL` table |
| 9 | 8 | `symbol_count` | number of symbols encoded |
| 17 | 8 | `decoded_length` | exact decoded bytes |
| 25 | 4 | `initial_state` | scalar decoder entry state |
| 29 | 4 | `payload_len` | length of the following payload |
| 33 | `payload_len` | `payload` | renormalization bytes in forward decoder-consumption order |
A channel is never a bare seed: the model, decoder state, payload, and counts
are all required to reconstruct bytes (see [`docs/adr/0006-rans-substrate.md`](docs/adr/0006-rans-substrate.md)).
## PDF layout candidate (Phase 5, rebuilt on packed framing in Phase 5.7)
The `PDF_LAYOUT` candidate (`source_format = 1`) applies only to PDFs with a
classic cross-reference section and **no** cross-reference stream, and with at
most 255 indirect objects; anything else declines. Since Phase 5.7 it persists
`objects = [data]` (one packed data object), `models` and `channels` empty, and
reconstructs entirely from **one `PACK_SEGMENTS` op** whose item table interleaves
literals and predictions:
- every literal byte is appended to the single data object and referenced by a
`LITERAL { len }` item, with adjacent literal runs **coalesced** into one item;
- a `MARK { slot }` (slot = the object's index in the physical object table)
before each indirect object's introducer bytes, and a `MARK { slot: 255 }` at
the xref section start;
- an `EMIT { slot, width: 10 }` (followed by a `LITERAL` for the generation/status
field) for each xref entry whose 10-digit offset equals the marked offset of its
target object, and a literal item otherwise;
- an `EMIT { slot: 255, width }` for the `startxref` value when the marked xref
start reproduces it, and a literal fallback otherwise.
This replaced the per-segment `MARK_OFFSET`/`EMIT_OFFSET` program of the original
Phase-5 lane (kept in the DRA as `0x06`/`0x07` but no longer emitted by this
candidate) to amortize framing. The candidate is byte-exact by construction (the
builder verifies serialize → parse → materialize → byte-compare and declines on
any mismatch) and its analysis-only metadata (`pdf-layout;objects=…;
xref_predicted=…;xref_literal=…;startxref_predicted=…`) is deterministic.
It is **recorded, not adopted**. Packed framing plus coalescing makes it **beat
RAW at scale** (`many.pdf` 10,069 vs 10,215), which the per-segment lane never did,
but it is still dominated by `BYTE_RANS` (5,181) because the residual data object
is stored literally; layout wins 0 of the 8 classic-xref samples and the
leave-one-out layout delta is 0 (campaign `2026-10-05-phase5-4521778`). See
`PROJECT_STATE.md` and ADR-0012. This section is **PROVISIONAL**.
## Feature policy
- `default = ["rans"]`: the native scalar entropy decoder (`ryg-rans-rs`
`=0.5.1`, **safe manual** API only) is present by default.
- Built with `--no-default-features`, the exact RAW/RLE floor still compiles and
materializes channel-free descriptors byte-for-byte.
- A descriptor that declares `MODEL`/`ENTROPY_CHANNEL` records but is decoded
without the `rans` feature returns an explicit `UnsupportedFeature`
(`ErrorClass` exit code 6). It is never silently reinterpreted and never
partially materialized.
## Canonical descriptor encoding
The *source document* is never canonicalized. The *descriptor* is canonical: a
given `Descriptor` always serializes to the same bytes, and
`serialize(parse(x)) == x` holds for every descriptor this implementation
produces (verified by the conformance court). This makes receipts and hashes
stable across runs and implementations.
## Integrity levels
| framing | per-record CRC-32C; header CRC-32C |
| whole source | `INTEGRITY` SHA-256 of the reconstructed bytes |
A deep verify (`vole-document verify`) materializes the source and checks the
whole-source digest. During development, `byte_compare` is the court authority.
## PDF source format (Phase 3)
The header's `source_format` selector gains a second normative value:
```text
source_format := 0 = OPAQUE
| 1 = PDF (Phase 3)
```
An unknown class still fails closed with `UnsupportedFeature`; it is never
silently reinterpreted as opaque.
A PDF descriptor (`source_format = 1`) is produced by the byte-authoritative
physical scanner. The scanner lexes the input into a contiguous span cover and
classifies each span structurally (`%PDF-` header, `obj`/`endobj`,
`stream`/`endstream`, `xref`, `trailer`, `startxref`, `%%EOF`, comments,
whitespace, and raw stream data), resolves direct/indirect `/Length`, and builds
an append-only revision map delimited by `%%EOF` with `/Prev` links.
The Phase-3 candidate persists the exact physical partition as **literal-span DRA
ops**: one `INLINE` instruction per physical span, in ascending offset order,
with no structural compression. Because the cover is contiguous and
non-overlapping, concatenating those spans reconstructs the source byte-for-byte
by construction. The per-span **kinds are deterministic analysis metadata** —
`scan` can recompute them at any time, they are not stored as trusted semantics,
and the literal bytes are the authority.
`PDF_PHYSICAL` competes in the complete-cost court like every other candidate and
currently loses to RAW (the expected Phase-3 outcome; structural compression is
Phase 5+). This section remains **PROVISIONAL**; the wire layout is not frozen
v1.