Expand description
Weight loader: CMF tensor directory → Pipeline.
Storage rule: models WITH task masks are dequantized to f32 (masked
execution needs f32 row access; skill files are small by design).
Models without masks keep quantized matrices zero-copy from the mmap
(QTensor::Mapped) — this is what lets a 15B file run in a few GB
of RSS instead of 60 GB of f32.
Layer kinds come from arch.layer_types: FullAttention loads
self_attn.* (with auto-detected Qwen3.5 extras: per-head qk-norm by
tensor presence, output gate by q_proj row count); LinearAttention
loads the canonical core vmf_attn.* (folded at convert time).
Structs§
- PerSequence
State - Fixed per-sequence state a file declares, from its header alone (no
weights loaded): bytes of bounded-anchor rings, bytes of recurrent
vectors (S + conv rings), and how many layers still hold a growing
per-position KV. A strictly-O(1) file has
growing_layers == 0.
Enums§
- Growth
Mode - Which
expert_appendrecords the loader mounts (CMF_GROWTH). - Overlay
- Tensor source selector (spec §9): backbone, one skill’s overlay, or a soft superposition of top-m skills (claim 14 working tensors).
Constants§
- DECISION_
MODEL_ REFUSAL - The refusal every generative entry point (
Pipeline::from_model*, and through it run/chat/route/bench) gives a file carrying the DECISION feature bit. Such a file is served by the decision runtime instead.
Functions§
- growth_
mode CMF_GROWTH=active(unset) |all|off.- mounted_
growth_ records - The
expert_appendrecords the currentgrowth_modemounts, with their positions inheader.skills(the chain rule of the format is evaluated at that position). - per_
sequence_ state_ bytes