vole-document 0.1.0-alpha.24

Persistent procedural document runtime: byte-exact reconstruction plus a content-addressed procedural seed DAG, queryable observations with provenance, and selective late materialization.
Documentation
# Multi-format adapters

Phase 12 makes the field format-universal across PDF, DOCX and EPUB: three
native inverse compilers converge on one `DocumentField` with common
observations and retained native structure. Phase 13.3 adds a fourth, ODT
(OpenDocument). “Universal” means the observation
vocabulary is shared, not that every format is supported.

## Three representational layers (ADR-0029)

They stay distinct, reference each other, and are never conflated:

1. **Exact physical source state** — PDF spans; ZIP local headers / compressed
   member spans / data descriptors / central directory / ZIP64 / EOCD; XML member
   bytes. Normative; reconstructs `original_bytes`.
2. **Format-native procedural state** — PDF object/stream/revision; OPC
   part/relationship and WordprocessingML paragraph/table/story; EPUB
   package/manifest/spine/nav/XHTML. Normative for native observations.
3. **Shared observation vocabulary** — a small common selector/representation set
   plus per-format escape hatches. Never a lossy universal AST.

A shared Rust `Selector` enum is interface reuse, not shared semantics.

## Byte-authoritative ZIP layer (ADR-0030)

DOCX (OPC) and EPUB (OCF) share one physical scanner analogous to the PDF
scanner. It covers `[0, N)` (`Prefix · LocalHeader · MemberData ·
DataDescriptor · CentralDirectory · … · Eocd · Trailing · Unclassified`) with a
`validate()` that rejects any gap/overlap/wrong total. A member's identity is
`(archive ordinal, local-header offset)`, distinct from the advisory logical part
name; duplicate names are never normalized. The exact leaf is the raw compressed
span — no unzip/rezip. The `zip`/`rawzip` crates are oracle-only. Reject-vs-
opaque split: a broken physical cover is a typed reject; a semantics/resource
problem preserves the exact bytes and declines only the decode.

## DOCX adapter (ADR-0032)

- Main part discovered **semantically** from `_rels/.rels` (the officeDocument
  relationship), never a hardcoded `/word/document.xml`.
- WordprocessingML subset: paragraphs/runs/text, styles with heading identity via
  resolved `outlineLvl`, sections, headers/footers, notes, comments, bookmarks,
  hyperlinks, fields, tracked changes, drawings/resources, numbering. Unknown
  namespaces are preserved, never interpreted.
- Stories are explicit (`Main`, `Header`, `Footer`, `Footnote`, `Endnote`,
  `Comment`, `TextBox`, `Glossary`); text/find observations are scoped to one
  story and never silently mixed. Tables are first-class.
- Versioned extraction profiles (`DocxExtractProfile`); the profile identity is
  hashed into the canonical selector.
- XML is derived-only (`quick-xml`, no DTD/entities, UTF-8); exact XML bytes stay
  `Q_ref`.

## EPUB adapter (ADR-0033)

- OCF container: `mimetype` (first, stored, exactly 20 bytes) and
  `META-INF/container.xml` → rootfile(s); the package document is located
  semantically, never from a hardcoded `OEBPS/content.opf`.
- Spine-first reading order: `SpineItem(n)` is the format-native coordinate.
  Reflowable EPUB has **no intrinsic pages**; `Page(n)` exists only where a
  page-list nav, `epub:type="pagebreak"`, or a fixed-layout viewport defines it,
  and is never synthesized.
- Bounded XHTML observations (headings, paragraphs, lists, tables, links,
  resources, fragment ids, semantic sections; SVG/MathML preserved). No script
  execution, no remote fetch.
- Versioned `EpubExtractProfile`; exact-byte preservation and EPUB conformance
  are separate outcomes.

## ODT adapter (ADR-0038)

- OpenDocument (ODF) package: the mandatory stored `mimetype`
  (`application/vnd.oasis.opendocument.text`) and `META-INF/manifest.xml`. ODF is
  **not** OPC (no `[Content_Types].xml`, no `officeDocument` relationship), so —
  like EPUB — the adapter reuses the ZIP layer and the shared XML policy but not
  the OPC graph. The main content part is located **semantically** from the ODF
  manifest (never a hardcoded `content.xml`).
- Bounded OpenDocument content model (`office:body`/`office:text`): paragraphs,
  headings (`text:h` + `text:outline-level`), spans, lists, tables
  (column/row spans, covered cells), links, bookmarks, notes, images/resources,
  tracked changes (`text:changed-region` kinds), and sections. No intrinsic pages:
  `Page(n)` is never synthesized.
- Versioned `OdtExtractProfile` (tracked changes Final/Original/All, notes
  include/exclude, hidden, tabs, breaks); common vocabulary plus native
  `odt-part`/`odt-paragraph`/`odt-heading`/`odt-table`/`odt-cell`/`odt-list`/
  `odt-find`.
- Progressive inversion: only the requested part is parsed, on demand, and the
  canonical derived model is persisted and reused; exact leaves stay the 12.2
  member raw spans. A missing/malformed manifest is a typed decline with
  exactness preserved.

## Shared vocabulary (ADR-0031)

Common selectors/representations are added additively: `metadata`, `text`,
`heading`, `block`, `table`, `cell`, `resource`, `link`, `find`. Native selectors
remain first-class peers. Widening the enums fails closed on unknown
pairs. Provenance carries through: a common observation is
`DeterministicallyDerived`/`Heuristic`, never exact, and `metadata` is a shared
selector *name* with per-format semantics. Capability gaps are explicit, not
silently approximated (see [Format support](../reference/format-support.md)).

## Format detection and ingestion

One format-agnostic `field-ingest` detects the format from bytes (never a file
name) and routes through the matching adapter. Every answer's provenance is
tagged `format=<fmt>;common;<native>`.

## Measured results (mixed)

| Court | Result | Receipt |
|---|---|---|
| Source + descriptor removal, all three formats | 38/38 byte-exact | `evidence/campaigns/2026-10-06-phase12-removal-dc4d5a3/` |
| DOCX/EPUB logical triplet equivalence | 96/96 | `evidence/campaigns/2026-10-06-phase12-triplet-dc4d5a3/` |
| PDF no-regression (`N6`, A2 vs A11) | 32/32 exact, 0 regressions | `evidence/campaigns/2026-10-06-phase12-pdf-noregression-0d23a02/` |
| Hostile ZIP/OPC/OCF/XML fixtures | 315/315 assertions | `evidence/campaigns/2026-10-06-phase12-security-33f6d04/` |
| Lifetime + ablation ladder | small-doc win; A1 wins large frontier | `evidence/campaigns/2026-10-06-phase12-lifetime-ablations-06db12a/` |

Interpretation. The ablation ladder attributes the small-document win to the
content adapters (A4→A5) and to persistent semantic reuse (A5→A6, which trades
bytes for CPU); EntropyFS (A9) is a loss and `A7`/`A8` are not separable. The
source-retaining SQLite+FTS5 baseline (A1) wins the large-document frontier and
wall/CPU at N=1000. The cross-format equivalence is a self-authored,
generator-defined triplet (adapter consistency, not third-party independence).

Limitations. All corpora are locally generated (841 B–61 KB); “large” means large
in that corpus. The demo’s A0/A1 baselines pay Python interpreter startup, which
flatters VOLE. Cross-document durable work reuse is a recorded negative
([Persistence and caching](persistence-and-caching.md)). See
[phase-12-results.md](../phases/phase-12-results.md) and the independent review
[phase-12-skeptic-review.md](../reviews/phase-12-skeptic-review.md).

## Relevant ADRs

[0009](../adr/0009-pdf-byte-authority.md),
[0024](../adr/0024-document-field-authority.md),
[0029](../adr/0029-multi-format-authority-model.md),
[0030](../adr/0030-zip-physical-layer.md),
[0031](../adr/0031-common-observation-model.md),
[0032](../adr/0032-docx-adapter-scope.md),
[0033](../adr/0033-epub-adapter-scope.md),
[0035](../adr/0035-phase12-lifetime-benchmark.md).