vole-document 0.1.0-alpha.20

Persistent procedural document runtime: byte-exact reconstruction plus a content-addressed procedural seed DAG, queryable observations with provenance, and selective late materialization.
Documentation

VOLE-Document

A persistent procedural document runtime with byte-exact reconstruction: materialize(descriptor) == original_bytes, always.

VOLE-Document inverse-proceduralizes a document into a bounded deterministic reconstruction description — reconstruction structure, parameters/state, typed residual channels, and typed rANS channels — persists the recovered procedural state as a queryable document field, and lets you query it directly (page text, structure, preview, streams, objects, revisions, exact byte ranges) with a typed observation API, per-answer provenance, and EXPLAIN/EXPLAIN ANALYZE.

The forward direction is ordinary: a file is opened, parsed, and rendered. VOLE runs it backwards. It recovers the computation that would produce these exact bytes and stores that computation instead of the file — as a reconstruction program plus typed residual and entropy channels — then materializes the original bytes only when asked. Because the reconstruction is a program, the persisted document becomes a field that can answer many typed questions directly, without re-opening, re-parsing, or re-rendering the source.

The governing invariant of the exact profile is uncompromising:

descriptor:       materialize(descriptor) == original_bytes
persistent field: materialize(field_root)  == original_bytes

Parsing successfully, producing "the same" text, the same object graph, the same pages, the same rendering, or a canonical re-save are not substitutes. Every exact court requires all three of equal length, equal SHA-256, and cmp byte equality. Compression is an implementation detail here, not the product: rANS is the entropy substrate beneath the representation, never the procedural model. This project is neither a compressor nor a database — whole-file size and database style both lose, and the losses are recorded (see Findings).

Why this matters

Document-heavy AI systems repeatedly turn the same source material into temporary working representations. A document may be parsed for ingestion, extracted into text, divided into chunks, indexed for retrieval, converted for another consumer, rendered for visual inspection, cached, and then partially reconstructed again when a later task needs different information.

A typical lifetime can involve the same underlying document passing repeatedly through work such as:

parse
extract
chunk
index
convert
render
cache
re-read
re-extract
re-contextualize

Each representation is useful, but most captures only one view of the document and much of the computation that produced it is discarded. A later operation that needs a different view often starts again from the source or from another derived representation.

VOLE-Document explores a different lifetime model:

source document
      ↓
inverse once
      ↓
persistent reconstructive state
      ↓
observe only what this computation needs
      ↓
text / structure / tables / resources / provenance / exact bytes

The document is inverse-compiled into durable computational state rather than treated only as an opaque file to be repeatedly decoded. That state retains enough information to reconstruct the exact original bytes while also exposing narrower observations directly through the document field.

This creates the possibility of carrying useful work forward across the lifetime of a document. Parsing decisions, recovered structure, provenance, package relationships, native format structure, and derived observations can become persistent state rather than transient products of a single request.

The long-term question is therefore not only how cheaply a document can be stored, but how much repeated work can be avoided when the same document participates in many computations over time:

traditional lifetime

source
 ├─ parse → text
 ├─ parse → chunks
 ├─ parse → structure
 ├─ parse → tables
 ├─ render → preview
 ├─ parse → provenance
 └─ reopen → exact source


VOLE lifetime

source
   ↓
persistent DocumentField
   ├─ text
   ├─ chunks / blocks
   ├─ structure
   ├─ tables
   ├─ resources
   ├─ provenance
   ├─ previews
   └─ exact source

This matters most for workloads that repeatedly revisit heterogeneous documents and ask different questions of them: retrieval systems, document agents, research systems, technical knowledge bases, compliance and audit workflows, and long-lived document infrastructure.

The economic hypothesis is measurable: if enough useful document computation can be retained in compact procedural state, the cumulative cost of repeated parsing, extraction, materialization, I/O, and model context can fall over the lifetime of the document. VOLE-Document measures that hypothesis directly rather than assuming it. The repository records the regions where the field wins, ties, declines, or loses against direct tooling and persistent database baselines.

Exact source closure is part of that model. A narrower observation never has to become the archival authority for the document: the persistent field can answer derived questions while retaining a verified path back to the original bytes.

How it works

flowchart TD
    A["source bytes (PDF / DOCX / EPUB / ODT)"] --> B["native inverse compiler"]
    B --> C["DocumentField: seed DAG + observation index"]
    C --> D["typed observations (text, structure, bytes, ...)"]
    C --> E["materialize --exact => original bytes"]

Source bytes enter a format-native inverse compiler (a PDF physical scanner, a WordprocessingML inverse, a bounded-XHTML/OCF inverse, or a bounded OpenDocument (ODF) inverse over a shared byte-authoritative ZIP layer). The recovered state is persisted as a content-addressed procedural seed DAG plus a bounded observation index. Queries resolve their minimum dependency closure and materialize as late as possible; the exact whole document is just one observation (FullExactDocument) among many. Authority is layered and never confused: the DRA plus INTEGRITY is normative for reconstruction, the seed DAG is normative for observations, indexes are advisory, and the derived cache is disposable (ADRs 0024, 0029).

What it does

Capability What it means
Exact reconstruction materialize(descriptor) and materialize(field_root) equal the original bytes (length + SHA-256 + cmp).
Inverse proceduralization Recovers a bounded DRA reconstruction program plus typed residual and order-0 rANS channels from the document.
Persistent field Stores the recovered state as a content-addressed seed DAG that survives deletion of the source.
Typed observations Selectors (document / page / object / stream / revision / byte-range / text-match) × representations (metadata / text / structure / operators / encoded / decoded / exact / preview).
Provenance Every answer carries a typed basis, scope, dependency ids, and exact source spans.
EXPLAIN explain shows the intended plan; explain --analyze reports the actual work (bytes read by class, nodes executed vs reused, decodes, wall/CPU).
Partial materialization Serves one byte range, object, stream, or revision from an advisory seek DIRECTORY + observation index without materializing the whole document.
Multi-format One field vocabulary over PDF, DOCX, EPUB and ODT, with retained native structure and format=…;common;… provenance.
Hostile-input contract Typed errors, checked arithmetic, bounded resources, fail-closed unknowns; the decoder never executes document content.

Supported formats

Format Physical layer Native inverse Exact Common observations
PDF owned lexer + physical span scanner objects, streams, revisions, page tree, /ObjStm yes metadata, text, find
DOCX shared byte-authoritative ZIP + OPC WordprocessingML stories, paragraphs, runs, tables, notes, tracked changes yes metadata, text, heading, block, table, cell, resource, link, find
EPUB shared byte-authoritative ZIP + OCF package, manifest, spine, bounded XHTML yes metadata, text, heading, block, table, cell, resource, link, find
ODT shared byte-authoritative ZIP + ODF OpenDocument: paragraphs, headings, lists, tables, notes, tracked changes, sections yes metadata, text, heading, block, table, cell, resource, link, find
XLSX, PPTX, others — — PROPOSED —

"Universal" means the observation vocabulary is shared across the four implemented formats, not that every format is supported. Details and capability gaps: Format support.

Quick start (Docker only)

All commands run inside pinned containers; the host only invokes Docker.

# Build the pinned toolchain image, then run the gate
docker compose build dev
docker compose run --rm --no-TTY dev cargo test --all-features --locked
docker compose run --rm --no-TTY dev cargo clippy --all-targets --all-features -- -D warnings

End-to-end: ingest a document, inspect and run an observation, then reconstruct the exact bytes.

# 1. Wrap the source in the exact container (RAW accepts any bytes; the format
#    is detected from the bytes, never the extension).
docker compose run --rm --no-TTY dev \
  ./target/debug/vole-document encode --force raw report.docx report.docx.voldoc

# 2. Inverse-proceduralize into a persistent field. Prints the field id (HEX).
docker compose run --rm --no-TTY dev \
  ./target/debug/vole-document field-ingest report.docx.voldoc --store /tmp/field

# 3. Show the plan, then run it and measure the actual work.
docker compose run --rm --no-TTY dev \
  ./target/debug/vole-document explain --store /tmp/field --field "$FIELD" --block 1 --kind text --analyze

# 4. Query one observation (text, structure, table cell, …).
docker compose run --rm --no-TTY dev \
  ./target/debug/vole-document observe --store /tmp/field --field "$FIELD" --block 1 --kind text

# 5. Reconstruct the exact original bytes — length + SHA-256 + cmp all match.
docker compose run --rm --no-TTY dev \
  ./target/debug/vole-document materialize --store /tmp/field --field "$FIELD" --exact --output report.docx.out

explain --analyze for a narrow observation reports whole_source_materialized: false — the field answered from its minimum closure, not by re-materializing the document. The full CLI surface is in CLI.

Current status

Release 0.1.0-alpha.20 (Phases 13–15 complete; see Changelog). Phase 15 repaired the frozen real100-v1 court (release build; storage universes reported separately) and measured the frontier — Phase 15 results. Two structural wins: --packed cuts persistent bytes to 0.719× and file count to 0.009× at latency parity (ADR-0043), and zlib-rs inflate is 1.58× miniz_oxide at 1.00× RSS, byte-identical and meeting the pre-registered bar (recommended, not adopted; ADR-0045). Two negatives: residency wins only below ~1 MiB (7 ms cold vs 9 ms resident; ADR-0042) and adaptive promotion fails all three pre-registered falsifiers (opt-in, default-off; ADR-0046). CUDA is deferred (the bandwidth gate is unopened; the pinned Docker lanes cannot see the GPU; ADR-0048). VOLE holds pdf/text_repeat and docx/table, loses the rest to SQLite/FTS, and 3/5 >100 MiB PDFs still fail at encode (ADR-0041). Headline measurements, each with its own results doc:

  1. Exactness holds after the source is gone. PDF, DOCX and EPUB rematerialize byte-for-byte (length + SHA-256 + cmp) after the source and the descriptor are deleted, in a fresh process: removal 38/38, triplet 96/96 (Phase 12 results).
  2. Warm observations are cheap. A repeated narrow observation reads 0 descriptor bytes and adds 8.4–8.9 KB of descriptor-free overhead at ~99 µs wall — an overhead-only figure that excludes the cached answer payload it also reads (Phase 11 results).
  3. Small-document lifetime is a scoped win. On a self-authored 841 B–61 KB corpus, the field beats direct per-query tooling and the cold one-time baseline; the source-retaining SQLite+FTS5 baseline wins the large-document frontier and wall/CPU at N=1000 (Phase 12 results).
  4. Whole-file size is a recorded loss. The best VOLE lane beats gzip/zstd/xz/brotli on 0/27 files; the best generic compressor is smaller on every file (Findings).
  5. Cross-document durable work reuse is a recorded negative (N3). The warm reuse fraction 0.339907 falls to 0.0 after cache --clear; only exact representation identity is shared (Phase 12 results).

Current limitations:

  • Not a compressor. Whole-file size loses to generic lossless tools on every measured file (the representation is coarser than LZ77).

  • Not a database. A source-retaining SQLite+FTS5 baseline wins the large-document byte frontier and wall/CPU at N=1000.

  • Self-authored corpora through Phase 12; first real-corpus court run. Published performance results through Phase 12 use self-authored deterministic corpora. real100-v1 is a frozen 100-document NASA/NIST corpus selected independently of VOLE performance; its first frontier court is mixed — VOLE wins repeated observations and DOCX tables/metadata and loses cold lookups and >100 MiB PDFs. A real EPUB-content loss (the XHTML DOCTYPE the policy forbade) was found and fixed (13.7, ADR-0040); residual non-DOCTYPE EPUB declines remain (frontier report).

  • Partial reusability. Cross-document durable work reuse is a negative, and XLSX/PPTX and other adapters remain PROPOSED.

Documentation

Reproducibility

Everything runs in digest-pinned Docker services (compose.yaml); nothing runs on the host. Each sealed run under evidence/campaigns/<date>-<phase>-<gitsha>/ records the base image digest, rustc/cargo versions, Cargo.lock SHA-256, git commit and dirty state, CPU architecture, oracle versions, and the exact command. Receipts are immutable: corrections are amendments, never rewrites. The docs themselves are checked by tools/check-docs.sh.

Repository layout

src/        one crate; modules for architectural separation
tests/      exact / malformed / conformance courts
fuzz/       cargo-fuzz coverage-guided targets (excluded from the crate)
tools/      court, gate, and doc-check scripts (run inside Docker)
docs/       architecture, formats, reference, ADRs, phases, reviews, project
evidence/   immutable campaign receipts (machine-readable)
research/   LOCAL ONLY — gitignored (paper, snapshots, subagent findings)

License and citation

Dual-licensed under either MIT or Apache-2.0, at your option. See LICENSE-MIT and LICENSE-APACHE. Citation metadata is in CITATION.cff.

Third-party license note. The opt-in deflate-replay feature depends on preflate-rs, which depends on cabac (LGPL-3.0-or-later). The default build (default = ["rans", "store", "field"]) is permissive-only. A binary built with --features deflate-replay (or --all-features) links LGPL code and carries the corresponding obligations (ADR-0014).