Skip to main content

Module loader

Module loader 

Source
Expand description

Weight loader: CMF tensor directory → Pipeline.

Storage rule: models WITH task masks are dequantized to f32 (masked execution needs f32 row access; skill files are small by design). Models without masks keep quantized matrices zero-copy from the mmap (QTensor::Mapped) — this is what lets a 15B file run in a few GB of RSS instead of 60 GB of f32.

Layer kinds come from arch.layer_types: FullAttention loads self_attn.* (with auto-detected Qwen3.5 extras: per-head qk-norm by tensor presence, output gate by q_proj row count); LinearAttention loads the canonical core vmf_attn.* (folded at convert time).

Structs§

PerSequenceState
Fixed per-sequence state a file declares, from its header alone (no weights loaded): bytes of bounded-anchor rings, bytes of recurrent vectors (S + conv rings), and how many layers still hold a growing per-position KV. A strictly-O(1) file has growing_layers == 0.

Enums§

GrowthMode
Which expert_append records the loader mounts (CMF_GROWTH).
Overlay
Tensor source selector (spec §9): backbone, one skill’s overlay, or a soft superposition of top-m skills (claim 14 working tensors).

Constants§

DECISION_MODEL_REFUSAL
The refusal every generative entry point (Pipeline::from_model*, and through it run/chat/route/bench) gives a file carrying the DECISION feature bit. Such a file is served by the decision runtime instead.

Functions§

growth_mode
CMF_GROWTH = active (unset) | all | off.
mounted_growth_records
The expert_append records the current growth_mode mounts, with their positions in header.skills (the chain rule of the format is evaluated at that position).
per_sequence_state_bytes