# Gap Audit — melinoe
Audit date: 2026-06-04. Scope: correctness/soundness, performance, memory
efficiency, branding capability surface, testing, benchmarking, documentation.
## 2026-08-20 Panic-recovery assertion oracle — closed
The partition panic-recovery regression now asserts both exact panic payloads
and the empty state of the recovered payload mutex. The provider conformance
scan drops `existence_only_assertions` from 2 to 0 without changing another
class. All-feature and no-default locked checks, strict Clippy, Nextest,
doctests, and Rustdoc pass; a temporary mutation that discards the first
payload fails the focused test.
## Method
Full read of every source, test, and bench module; baseline `cargo test`,
`cargo clippy --all-targets`, and `cargo miri test` across all paths (Stacked
Borrows default + Tree Borrows on the projection/branding paths).
The 0.6.0 increment audited the full source tree again end to end; the access
core, token families, guards, atomics, and Cow/slice paths remain optimal and
unchanged. Two gaps were found and closed: a shard-count SSOT duplication in the
partition driver (`src/region/mod.rs`, `src/sync/partition/mod.rs`) and a feature-gate
defect in `examples/codegen.rs`. The prior increment audited `src/cell/cow.rs`,
`src/atomic.rs`, `src/static_assertions.rs`, `tests/conditional_cow.rs`,
`tests/conditional_atomics.rs`, and the Mnemosyne / conditional-atomic Criterion
harnesses.
The Apollo provider increment adds `tests/apollo_boundary.rs` as an explicit
consumer contract for branded scratch boundaries. It verifies that the static
`Borrowed` ZST policy returns a pointer-identical `Cow::Borrowed` with zero
element clones, while the static `Retained` ZST policy produces independent
owned storage with exactly one clone per element. Evidence tier:
value-semantic integration tests plus the existing ZST/type-level policy
surface.
### Parallel executor validation boundary — resolved (0.9.0)
The `ParallelExecutorFn` alias allowed safe registration of any raw unsafe
function pointer even though sound raw-slot access depends on exact-once index
coverage, blocking completion, and context lifetime. ADR 0001 selects a
transparent `ParallelExecutor` newtype whose unsafe constructor is the single
proof boundary. Registration remains safe because it accepts only a validated
capability. The old alias is deleted rather than retained as a compatibility
path. Evidence target: compile-time layout assertion, value-semantic partition
contracts, Miri, and the real Moirai scheduler consumer.
The compile-time pointer-layout assertion, 121 value-semantic workspace tests,
30 doctests, three focused Miri executor-path tests, Clippy, rustdoc, and semver
classification pass. The Moirai consumer now constructs the validated
capability at its scheduler bridge.
The current cleanup audited the region partitioning tree. `src/region/mod.rs`
still held module documentation, the `WriterShard` capability, and the
`ShardChunks` exact-size iterator in one file. It now acts as the documentation
and re-export root, with `region::shard` owning shard capability behavior and
`region::chunks` owning exact-size chunk iteration.
The same pass closed the cross-repo thread-local cache duplication: themis and
mnemosyne carried equivalent nightly-`#[thread_local]` / stable
`std::thread_local!` cache pairs for `Copy` values. `thread_cached!` now provides
one declaration-site TLS primitive with same-thread reuse, overwrite, and
cross-thread independence tests.
## Findings
### Halo workspace integration — consolidated into melinoe crate
The `crates/halo` sub-crate was consolidated into the root `melinoe` crate.
`BrandedVec`, `BrandedVecDeque`, `BrandedDrain`, and `BrandedVecDequeDrain`
now live in `melinoe::collections` (re-exported at the crate root, gated on
`alloc`). The halo workspace entry, Cargo.toml, and standalone crate are
removed. All tests and benchmarks are migrated into the trunk crate's
`tests/` and `benches/` directories. This eliminates the cross-crate
dependency boundary and reduces the workspace to a single member.
Evidence tier: `cargo nextest run` (121/121 pass), `cargo clippy --all-targets
--all-features -- -D warnings` clean, `cargo doc --no-deps` clean, feature
matrix (`--no-default-features`, `--features alloc`, default) all build clean.
Previously: the upstream `ryancinsight/halo` repository compiles, but the
inspected tree contained a broad independent ghost-token implementation, large
collection/graph modules, missing-doc warnings, and scratch/log/generated
artifacts. Bulk import was rejected; the staged migration through `crates/halo`
was the intermediate step before full consolidation.
### Default provider feature policy — closed
Melinoe did not expose the Atlas-wide default `parallel` and
`mnemosyne-memory` feature contract. Added zero-dependency `parallel` and a
`mnemosyne-memory` feature forwarding to `alloc`, which is the minimum memory
surface required by branded Cow/cell ownership boundaries. Evidence tier:
Cargo metadata audit across Apollo, Leto, Hermes, Mnemosyne, Moirai, Melinoe,
Themis, and Hephaestus; formatting and diff checks. Compile/test residual:
target lockfile access denied before rustc.
### Soundness — clean
The four `unsafe` sites (token minting, cell access, `from_mut`, `Send`/`Sync`
impls) each discharge their obligation through a higher-ranked invariant lifetime
or the `#[repr(transparent)]` layout chain, with inline `// SAFETY:` reasoning.
Miri reports no aliasing or data-race violations on any access, slice-view,
partition, cross-thread, or projection path. Evidence tier: **machine-checked**
(Miri) on top of type-level encoding.
### Performance / memory — already optimal on the access core
Token access lowers to a bare load/store (confirmed by `examples/codegen.rs` and
the `access` benchmarks); tokens and guards are ZST / `#[repr(transparent)]`
(pinned by `src/static_assertions.rs`). There is no synchronization instruction
to remove on the hot path. The remaining memory-efficiency lever was **reaching a
sub-field of a large branded payload through a permit** without copying the whole
payload — previously only expressible by cloning the payload out. Closed below.
### Capability surface — one gap closed
`MelinoeCell` exposed whole-value `borrow` / `borrow_mut` and whole-slice
`CellSliceExt`, but no way to project a guard onto a component (the `Ref::map`
analogue) or to derive two disjoint `&mut` from one write permit. Added
`MelinoeRef`/`MelinoeMut` `map` and `map_split` ([0.2.0]).
### Partition driver memory — gap closed
`sync::partition_map` could reserve `parts` join handles even though it can
spawn only non-empty shards. For empty input this reserved needless capacity;
for `parts > len` it could amplify allocation beyond useful work. The old
ceiling division also used `len + parts - 1`, which is overflow-prone at
`usize::MAX` inputs. Fixed in [0.2.1]: chunk size now uses
`1 + (len - 1) / requested_parts`, and the handle vector capacity is the actual
non-empty shard count. Evidence tier: value-semantic integration tests plus
Criterion scheduling benchmarks.
### Multithreading ergonomics — gap closed
The fixed `parts: usize` API forced callers to compute a worker count outside
the crate and could not express cache/tile-oriented chunk sizing directly.
Added `PartitionPlan` in [0.3.0] with fixed part count, reported hardware
parallelism, and fixed chunk-size variants. `partition_map_with` /
`partition_for_each_with` execute the same disjoint-shard engine with a typed
plan, and `partition_map_available` / `partition_for_each_available` use
`std::thread::available_parallelism()` for the common hardware-parallel case.
Evidence tier: type-level API plus value-semantic tests for plan equivalence,
chunk tiling, and platform-independent full coverage.
### Boundary policy monomorphization — gap closed
Conditional ownership and atomic ordering were previously expressed either by
ad hoc benchmark-local `Cow` branching or by runtime `Ordering` arguments.
Added [0.4.0] ZST policy surfaces: `Borrowed` / `Retained` for static
borrow-or-retain decisions, and `Relaxed` / `AcqRel` / `SeqCst` for atomic
ordering contracts. The runtime `RetainDecision` and `Ordering` APIs remain for
data-dependent cases. Evidence tier: compile-time ZST size assertions plus
value-semantic tests for borrowed pointer identity, retained copy independence,
runtime retain decisions, and ZST atomic ordering equivalence.
Refinement pass: `AtomicOrder` is sealed to the crate's audited policy set;
`BrandedAtomic::get_mut` / `into_inner` use the standard atomic unique/owned
APIs rather than pointer reads; static `Cow` policies dispatch through policy
method bodies, so the borrowed monomorph contains no clone branch. Added
read-permit-gated `BrandedAtomic::as_atomic` for zero-copy interop with raw
atomic APIs while preserving the shared-phase token proof; `as_atomic_mut` and
`into_atomic` cover unique/owned extraction. Latest patch routes
`BrandedAtomic::*_with` methods directly through the sealed atomic mediation
surface with `AtomicOrder` associated constants, avoiding runtime-ordering
wrapper calls in static policy monomorphs. Latest Cow refinement routes direct
`borrow_cow` / `retain_cow`, generic-policy `borrow_cow_with`, and runtime
`borrow_cow_if` through the same sealed `Borrowed` / `Retained` policy bodies,
so clone/no-clone behavior has one implementation source. Benchmark expansion
adds direct-vs-ZST-policy Cow rows and read-permit-gated `as_atomic` interop;
targeted Criterion reruns show static Cow policy rows match direct methods
within local run noise and shared atomic interop matches raw atomic throughput.
### Partition shard-count SSOT — gap closed (0.6.0)
The shard count was computed in two places: the private `shard_count(len, chunk)`
ceiling-division helper (used to size the worker-handle `Vec`) and, implicitly,
the `ShardChunks` iterator that actually yields the shards. The two agreed, but
the duplication was a latent SSOT/DRY hazard — a future change to chunking could
desynchronize the reserved capacity from the real yield. `ShardChunks` now
implements `ExactSizeIterator` with an exact `size_hint`
(`ceil(remaining / chunk)`, decrementing as consumed), and `partition_map_with`
reserves capacity from `chunks.len()`. The helper and the `ResolvedPartitionPlan`
struct are removed; `PartitionPlan::resolve` returns only the chunk size. The
empty/over-partitioned memory-efficiency contract is unchanged (the iterator
reports `0` for an empty region) and is pinned by both the new exact-size tests
and the `partition_driver/empty_region` benchmark (~1.0 ns with the input slice
black-boxed, no spawn). The new
`ExactSizeIterator` impl is additive public API ([minor]). Evidence tier:
value-semantic tests plus Criterion confirmation of no regression.
### Region module hierarchy — gap closed (0.6.0)
`src/region/mod.rs` mixed three responsibilities: conceptual documentation,
the zero-cost write capability, and its chunk iterator. This did not affect
runtime behavior, but it weakened SRP/SoC and made the exact-size iterator less
discoverable. Split into `region/mod.rs` (docs + public re-exports),
`region/shard.rs` (`WriterShard` and its zero-copy slice views), and
`region/chunks.rs` (`ShardChunks`, exact `size_hint`, `ExactSizeIterator`).
The public API remains `melinoe::region::{WriterShard, ShardChunks}` and
`melinoe::WriterShard`; no compatibility shim was added. Evidence tier:
type-level/API preservation plus the partition integration suite. The
partition-driver benchmark was hardened to black-box the input slice before the
refreshed run, preventing the empty-region row from collapsing to a
compile-time-known `Vec::new()` result.
### Registered partition executor aliasing — gap closed (0.7.0)
The custom `ParallelExecutorFn` path already tiled cell ranges and result slots
by task index, but each task reconstructed `&mut Context` from the same raw
executor payload. That was stronger aliasing than the code needed and invalid
for a truly concurrent executor. The task wrapper now reconstructs only a shared
read-only context, then writes through raw pointers solely to its disjoint
`MaybeUninit<R>` result slot and non-overlapping cell range. Evidence tier:
type-level/API preservation plus value-semantic partition tests; the unsafe
contract remains explicit on `ParallelExecutorFn` because external schedulers
must still invoke each task index exactly once and block until completion.
### Registered partition executor lifecycle — gap closed (0.7.0)
The registered executor is process-global. Without a reset hook, an integration
test or temporary scheduler installation could leave later partition calls on a
custom driver, creating hidden cross-test/process coupling. Added
`clear_parallel_executor`, which stores a null executor pointer and restores the
default scoped-thread driver for subsequent calls. The partition suite now
clears before/after the registered-executor test and includes value-semantic
coverage proving that after a clear the deterministic executor is not invoked
and the region contents still match the identity mapping. Evidence tier:
type-level/API addition plus value-semantic integration tests.
### Partition module hierarchy — gap closed (0.7.0)
The `std` partition driver had grown three concerns in one file: shard sizing
policy, global custom-executor registration, and execution of the scoped/default
driver. Split it into `sync::partition::{plan, executor, driver}` leaf modules
under `src/sync/partition/`. Public re-exports remain unchanged
(`melinoe::sync::{PartitionPlan, partition_map_with, register_parallel_executor,
...}`), while the file tree now matches SRP/SoC boundaries and keeps executor
unsafe state isolated from plan resolution. Evidence tier: type-level/API
preservation plus the partition integration suite.
### Existence-only assertion audit — gap closed (0.7.0)
The reentrancy panic tests and doctest examples still contained existence-only
error assertions (`is_err`). Replaced them with value-semantic assertions:
re-entrant acquisition checks compare to `Err(Reentered)`, and panic-unwind
tests downcast and compare the sentinel payload. A workspace search over `src`
and `tests` now finds no `is_err` / `is_ok` / `is_some` / `is_none` assertions.
Evidence tier: value-semantic tests plus source-wide assertion audit.
### Thread-local cache duplication — gap closed (0.7.0)
Repeated per-thread value caches in Atlas consumers used the same two-way cfg
shape: nightly `#[thread_local]` storage for the fast path and stable
`std::thread_local!` fallback. A normal generic
type cannot declare a fresh static per cache site, so the canonical abstraction
is a small declaration macro. Added `thread_cached!`, which emits one module per
cache site with `get_or_init` and `set` for `Copy` values. `build.rs` now
declares `nightly_tls_active` separately from `doc_cfg_active`, so TLS fast-path
cfg does not require enabling the nightly documentation feature. Evidence tier:
value-semantic integration tests for initialization, same-thread reuse,
overwrite, and per-thread independence. Benchmark evidence:
`thread_cached_4096x` covers cached `get`, cached `get_or_init`, `set`, and
`clear`/`set`; the stable fallback measures sub-nanosecond to ~1.1 ns per
operation on this host. Latest refinement stores both cfg paths as
`Cell<Option<T>>`, eliminating generated unsafe from the nightly TLS path.
### Feature hygiene — gap closed (0.6.0)
`examples/codegen.rs` used the alloc-gated `CellCowExt::borrow_cow` but carried no
`required-features`, so `cargo test --no-default-features` failed to compile the
example despite the checklist claiming the gate passed. Fixed by declaring
`required-features = ["alloc"]` for the example in `Cargo.toml`. The full feature
matrix (`--no-default-features`, `--no-default-features --features alloc`,
`--features std`) now builds and tests clean.
### Partition driver executor-path duplication — gap closed (M-3)
`sync::scoped::partition::driver.rs` (~531 lines) held two near-duplicate driver
bodies — `partition_map_with` (mutable `WriterShard`) and
`partition_read_map_with` (shared `&[T]`) — that shared byte-for-byte identical
executor-path scaffolding: the `MaybeUninit` out-buffer allocation, the
`ExecutorDropGuard` leak/panic handling, the panic-collection mutex, and the
`Vec::from_raw_parts` teardown on both the success and unwind paths. Extracted a
single generic engine `driver_core::drive<R, Run>(num_chunks, run)` that owns all
of that machinery; the two drivers now differ only in the `run: Fn(usize) -> R`
closure they pass (mutable-shard construction vs shared sub-slice). The
executor-path `unsafe` (out-buffer init tracking + free-backing-on-unwind) lives
in exactly one place. Split `driver.rs` into `driver_core`/`map`/`read_map` leaf
modules; public exports unchanged. The drop-only-initialized-slots + unwind
teardown behavior is preserved exactly — the existing
`custom_executor_panic_safety_drops_success_elements` /
`read_custom_executor_panic_safety_drops_success_elements` tests (asserting a
drop count of exactly 3 when task 2 of 4 panics) pass unmodified, as does the
whole partition suite. Evidence tier: value-semantic + differential
(concurrent-vs-sequential) tests preserved unmodified, plus Miri on the partition
module.
### Indexed disjoint-shard accessor — gap closed (CR-7)
`WriterShard::chunks` → `ShardChunks` is a *sequential* lending iterator (each
`next()` reborrows the remainder), unsuitable for a work-stealing pool that wants
partition `c` on demand. Consumers therefore re-derived disjoint sub-slices with
`core::slice::from_raw_parts_mut` and a hand-written SAFETY argument
(`moirai-parallel/src/melinoe_ext.rs`), and Melinoe's own driver hand-rolled the
same range math internally. Added `WriterShard::par_chunks(chunk_size) ->
ParChunks<'a, 'brand, T>` (`region::par_chunks`): an indexed view holding the base
pointer, region length, and (clamped ≥1) chunk size, exposing `len() = ceil(len /
chunk)` and `unsafe get_unchecked_chunk(index) -> WriterShard`. The
`from_raw_parts_mut` sub-slicing now lives in exactly one authoritative
`// SAFETY:` block; the disjointness contract (distinct indices → non-overlapping
ranges, each requested at most once — exactly the `ParallelExecutorFn` guarantee)
is documented in one `# Safety` section. `get` is `unsafe` because the returned
shard carries the region's `'a` (not `&self`), so a pool can hold several
partitions live at once — the type system cannot enforce single-use, so the
caller promises it. The write-side partition driver now drives through
`par_chunks`, closing the internal duplication. Evidence tier: value-semantic
tests (exact `len`, complete/disjoint partition coverage, two-index non-aliasing
mutation, single-partition, `Send`/`Sync`), a `ShardChunks` len-parity unit test,
a `compile_fail` brand-escape doctest, and **Miri** (Stacked/Tree Borrows clean)
on the two-live-index aliasing test.
## Residual risk / non-goals
- **Cross-repo CR-7 follow-up (moirai-parallel).** `moirai-parallel/src/melinoe_ext.rs`
still re-derives disjoint sub-slices via `DisjointMutPtr` +
`from_raw_parts_mut` in `par_partition_for_each` / `par_partition_map`. Melinoe
now owns the primitive (`WriterShard::par_chunks`) that replaces this, but the
consumer refactor is **deliberately not done here**: it belongs in the moirai
repo and is blocked on Melinoe publishing this change (co-evolution: upstream
commit + push → `cargo update -p melinoe` in moirai → swap the hand-rolled
ranges for `par_chunks(...).get_unchecked_chunk(c)` → verify). Filed as a moirai
backlog item; not a Melinoe defect.
- Projecting arbitrary *separate* cells (not sub-components of one payload) to
simultaneous `&mut` is intentionally **not** added: distinctness of two
independent `MelinoeCell`s is not provable to the borrow checker without a
runtime pointer check, which would violate the zero-cost invariant. The slice
(`CellSliceExt` + `split_at_mut`), `WriterShard`, and `map_split` paths cover
the disjoint-`&mut` need where disjointness is structurally provable.
- `--all-features` requires a nightly toolchain (the `nightly` feature gates
`feature(doc_cfg)`); this is by design. Stable builds use the default feature
set. Not a defect.
## Status
0.7.0 increment implemented and tracked in `checklist.md` / `backlog.md`.
Stable gates green: `fmt --check`, `clippy --all-targets -- -D warnings`,
`test`, `doc --no-deps`, and the full feature matrix (`--no-default-features`,
`--no-default-features --features alloc`). The
`partition_driver` Criterion group was rerun (fast sweep); `empty_region`
(~1.0 ns with a black-boxed input slice) confirms the no-spawn / zero-capacity
contract survives the SSOT refactor. Version is 0.7.0 ([minor], additive public
API: `thread_cached!`; internal region hierarchy split preserves public
exports). CHANGELOG synchronized.
All prior verification residuals are now resolved. **Miri** is clean across the
full suite (no UB, no data races): `projection` (6), `partition` (15, including
the new exact-size tests under real `std::thread::scope`), `threads` (6),
`conditional_atomics` (8), `conditional_cow` (5), `branding` (7), `multi_token`
(8), `slice_views` (4), `differential` (3) — evidence tier: machine-checked.
**2026-09-11 re-verification.** The suite has grown since the entry above
(`partition` is now 26 tests, not 15) and an unbounded `cargo miri test` no
longer completes: the three `partition` proptests run 256 cases each by default
and spawn real threads over up to 255 elements, which exceeds a 20-minute budget
under Miri (observed exit 143 = SIGTERM, not UB). Re-verified clean at
`PROPTEST_CASES=8` — all 26 partition tests in ~67s, full suite green under both
Stacked Borrows and Tree Borrows. `.github/workflows/ci.yml` now runs both
models in a dedicated `miri` job, so the README's Miri claim is enforced rather
than asserted.
**`cargo-semver-checks`** runs via the git-rev baseline (`--baseline-rev HEAD`):
v0.5.0 → v0.6.0 reports no semver update required, confirming the [minor]
classification; default registry comparison awaits publication, and
semver-checks 0.48.0 skips its lints against the current nightly rustdoc-JSON
format (tool/format mismatch, not a crate defect). **Nightly clippy**
`--all-targets --all-features -- -D warnings` is clean (the local MSYS2 nightly
bakes the stable channel, so the `doc_cfg` feature gate needs
`RUSTC_BOOTSTRAP=1`).
### Prior increments (historical)
Current minor increment implemented and tracked in `checklist.md` /
`backlog.md`. Stable gates green: `fmt --check`, `clippy --all-targets -D
warnings`, `test`, `doc --no-deps`, no-default feature tests, and benchmark
compilation for both Criterion harnesses. `cargo miri test --test partition`
passes under Stacked Borrows and Tree Borrows. Version bumped 0.2.1 → 0.3.0
([minor], additive public API). CHANGELOG synchronized.
Current target is now 0.4.0 ([minor], additive public API) for conditional
`Cow` boundary policies and monomorphized atomic ordering policies. Stable gates
are green: `fmt --check`, `clippy --all-targets -D warnings`, `test`,
`doc --no-deps`, no-default feature tests, and benchmark compilation for all
five Criterion harnesses. Miri passes for conditional atomic / conditional Cow
tests under Stacked Borrows and Tree Borrows.
`cargo-semver-checks` is installed, but default comparison fails because
`melinoe` is not found in crates.io; a baseline rev or registry release is
needed before tagging. Stable `--all-features` still fails at the documented
nightly `doc_cfg` feature gate.
Benchmark note: the access suite was run repeatedly here (incl. a
`--sample-size 200` sweep; the `--measurement-time 10` variant crashed on the
constant-folding Melinoe read micro-benchmark, which balloons to ~51e9 estimated
iterations — `measurement-time 5` is the safe ceiling for this suite). This
machine proved load-saturated: single-threaded Melinoe micro-figures were stable
but the multithreaded/lock-contended absolutes swung 15–60% run-to-run (the
`single_thread` partition baseline ranged 15–25 ms). Because this work was
rebased onto a parallel branch that had independently **refreshed all
benchmark numbers** on a cleaner machine, the merged BENCHMARKS.md retains that
branch's internally-consistent figures (e.g. `AtomicU64` increment ~30×) rather
than this session's noisy re-measurements; only the genuinely new rows
(`projection_1024x`, `partition_driver`, and the `PartitionPlan`
available-parallelism / fixed-chunk rows) carry this session's data, with the
core-count-dependent ones labelled "measure locally." Ratios are the durable
signal across both machines.
[0.4.0]: CHANGELOG.md
[0.3.0]: CHANGELOG.md
[0.2.1]: CHANGELOG.md
[0.2.0]: CHANGELOG.md