stella-docx-kernel 0.35.0

Bounded DOCX package projection and WordprocessingML scanning
Documentation

stella-docx-kernel

stella-docx-kernel projects bounded DOCX packages and WordprocessingML into a versioned, host-neutral document representation.

The crate supports native Rust consumers directly. Browser consumers can enable the wasm feature; @stll/docx-core/projection provides the corresponding TypeScript binding.

The artifact

The WebAssembly artifact is committed under packages/docx-core/src/generated with its own size budget. Regenerate it with bun --filter @stll/docx-core wasm:generate. The committed bytes are compared against a fresh build, and CI builds on linux/amd64, so on any other platform regenerate through scripts/regenerate-wasm-canonically.sh @stll/docx-core, which runs the same step in that image. CI is still the arbiter of the bytes: when its drift check fails it uploads what it built, and that is what to commit.

The parser treats package identifiers as document facts, not durable application identities. ZIP and XML resource limits are enforced at the input boundary; semantic scans also bound XML events and inline-context copies and report both as deterministic work units.

Projection wire schema 6 reports formatting completeness per family. Bold and superscript depend on style inheritance; their status is unknown when the style cascade cannot be resolved. Highlight reports direct run markup and remains known without a styles part. An unread formatting hierarchy makes the affected families unknown. Callers read the status for each family before interpreting its spans; the former global formatting status is removed.

project_docx_with_review_facts reuses the package-directory scan and selected part decompression for the document projection, plus attributed revisions and comments. Each review family is complete or explicitly unknown. Content and location details use the same typed known/unknown contract; they remain unknown until the canonical paragraph state machine can prove offsets after revision-view normalization, rather than exposing a second parser's guessed coordinates. Revision attribution covers every recognized revision element in the effective main-document body, including text boxes and structural or formatting changes; historical markup nested inside a change snapshot is not counted a second time.