themis-topology 0.11.1

Typed placement law for NUMA, HBM, memory-tier, and worker locality contracts
Documentation

Themis

Themis provides typed placement law for Atlas runtime and memory crates.

The crates.io package is named themis-topology; its Rust library crate remains themis so existing imports do not change.

It is the shared source of truth for:

  • NUMA node identity
  • worker identity
  • locality domains
  • memory tiers, including HBM
  • placement hints
  • topology snapshots
  • reported CPU efficiency classes and group-aware affinity masks
  • confining the calling thread to one logical processor, with a typed failure
  • current CPU/NUMA node queries

It does not own allocation, scheduling, queues, worker loops, or thread-local allocator state. Mnemosyne remains the allocation owner. Moirai remains the execution owner. Themis supplies the law both consume.

With the default melinoe feature, Themis also exposes branded placement scopes. ThreadLocalPlacement uses Melinoe's thread-confined token for worker-local placement state. SyncRegionPlacement uses Melinoe's sync-region token for placement snapshots that may move between execution domains. Branded storage remains melinoe::MelinoeCell; Themis does not define a second cell name. NumaPinnedSlice and ConstNumaPinnedSlice construct their owned cell storage through Melinoe's collections::BrandedVec handoff, preserving the brand on every cell while keeping NUMA node identity and placement permits owned by Themis.

Splitting a region hands out several capabilities for one brand, so the NUMA node tag — not the Melinoe token — is what keeps their &mut views apart. A cell's tag therefore comes from a construction path that proves it: owning the cell, or an exclusive borrow through from_unique. There is no way to staple a caller-chosen tag onto a shared cell reference, and the PinnedCell trait family is unsafe so downstream placement types must discharge the same obligation. See ADR 0002.

Boundary

themis
├── typed placement identifiers
├── memory tier and placement hint vocabulary
├── topology snapshots and distance lookup
└── current locality query

mnemosyne -> themis   allocation placement
moirai    -> themis   worker and task placement
leto      -> themis   cache-sized tiling hints

No Themis API stores allocator state, scheduler state, or raw thread-local storage pointers.

Current locality query contract

current_numa_node() is the fast placement path: under std it caches the calling thread's last observed NUMA node in thread-local storage and falls back to NumaNodeId::ZERO when the platform cannot report locality. It does not re-probe the operating system on every call. Callers that may migrate between NUMA nodes must call refresh_current_numa_node() at an explicit scheduling boundary; the refresh probes the platform and replaces that thread's cached value. try_current_numa_node() and current_processor() are uncached, optional probes for verification or diagnostics and preserve None when the platform does not expose the requested value. In no_std builds the cached placement query has the deterministic node-zero fallback because no OS probe is available.

This distinction is intentional: Themis owns the locality vocabulary and query contract, while Moirai decides scheduling boundaries and Mnemosyne decides how a placement hint affects allocation. No consumer should add a second NUMA cache or substitute a fabricated cache/topology value.

CacheLevel is provider-reported topology law. Linux reads cache-index records from sysfs and Windows reads GetLogicalProcessorInformationEx; unavailable or malformed cache data is None, never a synthetic capacity. Leto consumes cache sizes as tiling hints, and Moirai consumes shared-processor rows as chunk-locality hints. Consumers must preserve typed cache absence.

CpuTopology::efficiency() similarly discharges efficiency-class absence once and returns total class-level queries. On Windows, ProcessorAffinityGroups owns Themis's flattened processor numbering and partitions processor sets into sorted native group masks without dropping unrepresentable ids. bind_current_thread(processor) confines the calling thread to one logical processor on Windows and Linux and reports a refusal as a typed BindError (Unsupported, OutOfRange, or Os { code }), so a scheduler can refuse to start a worker whose binding failed. Allocation and worker loops remain consumer responsibilities.

Benchmarks

Run the topology benchmark with:

cargo bench --bench topology

The benchmark exercises real CpuTopology::single_node construction and processor_node_pairs traversal. Its output is empirical local timing only; it is not a statistical baseline or a speedup claim.

Evidence

Current correctness claims rest on type-level encoding plus value-semantic unit tests. Branded placement-state claims rest on Melinoe token invariants. OS topology discovery is empirical and falls back to a single-node topology when platform data is unavailable. Cache discovery is empirical platform data with value-semantic parser tests; cache absence remains explicit until a provider reports a complete hierarchy. Leto and Moirai carry the typed absence through their public consumer surfaces. Benchmark claims currently rest on the dependency-free topology benchmark harness and must be treated as empirical local timing unless a criterion baseline is added.