Skip to main content

QTensor

Enum QTensor 

Source
pub enum QTensor {
    F32 {
        data: Vec<f32>,
        rows: usize,
        cols: usize,
    },
    Mapped {
        model: Arc<CmfModel>,
        idx: usize,
        dtype: TensorDtype,
        rows: usize,
        cols: usize,
        row_scale: Vec<f32>,
        col_field: Vec<f32>,
        vbit_offsets: Vec<usize>,
        repack: Vec<u8>,
    },
}

Variants§

§

F32

Fields

§data: Vec<f32>
§rows: usize
§cols: usize
§

Mapped

Fields

§model: Arc<CmfModel>
§idx: usize

Index into the model’s tensor directory.

§rows: usize
§cols: usize
§row_scale: Vec<f32>

Per-row scales, dequantized to f32 up front (tiny).

§col_field: Vec<f32>

q8_2f column field (θ), dequantized up front; empty for q8_row.

§vbit_offsets: Vec<usize>

Vbit only: byte offset of each row’s packed data within the tensor blob ([rows + 1], computed once at load — the per- matvec prefix scan over row bit-widths was O(rows) each call).

§repack: Vec<u8>

q8-family decode repack (load-time, optional): rows in groups of 4, interleaved in 16-byte units — one 64-byte line per iteration feeds all 4 sdot lanes, ONE sequential weight stream per worker instead of four (this is where llama.cpp’s repacked Q8 kernels get their bandwidth). Empty = off (CMF_REPACK=0, non-SDOT arch, or an ineligible shape). Trades an anonymous copy of the quants for mmap pages that go cold.

Implementations§

Source§

impl QTensor

Source

pub fn from_f32(data: Vec<f32>, rows: usize, cols: usize) -> Self

Source

pub fn from_model(model: &Arc<CmfModel>, name: &str) -> Result<Self, String>

Wrap a directory tensor without dequantizing the payload. Falls back to dequantized f32 for dtypes without a fused kernel.

Source

pub fn model_dtype(&self) -> Option<TensorDtype>

The layout this tensor is stored in, when it is mapped from a model. The frames branch on it — a q2tp gate against a q4tp down is a real combination in the 2-bit profile and needs a different kernel.

Source

pub fn model_idx(&self) -> Option<usize>

The tensor’s index in the model directory, when it is mapped from one. The GPU frames bind by index rather than by name — a name lookup per layer per token is not free, and the index is what the device cache is keyed on anyway.

Source

pub fn model_arc(&self) -> Option<Arc<CmfModel>>

The model this tensor is mapped from, when it is mapped at all. The GPU frames need the container to reach the bytes; a QTensor already holds it, and threading a second handle down every call site to say the same thing invites the two to disagree.

Source

pub fn rows(&self) -> usize

Source

pub fn mapped_q4tp(&self) -> Option<(&Arc<CmfModel>, usize)>

Same slot as mapped_q4t for a q4tp tensor — the fused DiT FFN picks its kernels by which of the two answers.

Source

pub fn mapped_device_gemm(&self) -> Option<(&Arc<CmfModel>, usize)>

(model, tensor idx) for a mapped weight in ANY codec the fused device paths can run — four-bit tiled or either int8 layout.

The fused DiT chains asked for mapped_q4tp by name, so an eight-bit container never reached them and rendered through per-op GEMMs even after those kernels learned its codec. The gate is what the codec has a device GEMM for, not which codec it is.

Source

pub fn mapped_q2tp(&self) -> Option<(&Arc<CmfModel>, usize)>

(model, tensor idx) for a q2tp mapped weight — the 2-bit twin of mapped_q4tp, used by the mixed MoE profile.

Source

pub fn cols(&self) -> usize

Source

pub fn mapped_q1(&self) -> Option<(&Arc<CmfModel>, usize)>

(model, tensor idx) for a q1 mapped weight — the wgpu token graph keys its resident VRAM cache by idx. None for any other dtype/kind.

Source

pub fn graph_weight(&self) -> Option<(&Arc<CmfModel>, usize, u8, &[f32])>

(model, idx, kind, row_scale) for a graph-capable mapped weight. kind: 0=q8_row (per-row scales), 1=q1, 2=q4_block, 3=q1t (tile-embedded, no rs), 5=q4_tiled, 6=q4tp, 7=q8_2f (both scale planes live inside the tensor). None only for vbit.

The old comment here claimed q4_block was unhandled while the arm right below mapped it, and it named q8_2f as unhandled after that stopped being true — a stale comment on this function is how a model silently loses the graph, so it is worth keeping honest.

Source

pub fn as_f32(&self) -> Option<&[f32]>

Dense f32 view — only for owned tensors. Masked/sparse execution paths require it; quantized weights don’t support masks yet.

Source

pub fn row_f32(&self, r: usize, dst: &mut [f32])

Dequantize one row into dst (embedding lookup).

Source

pub fn sparse_col_ok(&self) -> bool

Can this tensor’s columns be read cheaply (for sparse down_proj)? True for F32/Q8Row/Q8_2f (per-row scale, direct strided access); false for group-packed q4/vbit (column access would unpack whole groups — sparse execution falls back to f32 for those).

Source

pub fn add_col_scaled(&self, c: usize, w: f32, out: &mut [f32])

down_proj [hidden, inter]: accumulate w · col(c) into out [hidden] — reads ONLY column c (one neuron) from the mmap, no full-matrix dequant. out[k] += w · down[k, c].

Source

pub fn prefetch_row(&self, r: usize)

Touch the head of row r so the DRAM latency of the next neuron’s weights overlaps the current one’s arithmetic.

Scattered rows are what per-token sparsity reads, and a 2 KB stride is past what the hardware prefetcher follows: without this every row starts with a cold miss that nothing hides. One touch per 512 bytes is enough — the rest of the row is a sequential run the prefetcher does pick up.

Source

pub fn add_row_scaled( &self, r: usize, w: f32, out: &mut [f32], scratch: &mut [f32], )

out += w · row(r) — the transposed twin of add_col_scaled.

A neuron’s down weights are a COLUMN of [hidden, inter], and a column is strided: reading one costs a cache line per element, so per-neuron dynamic sparsity saves arithmetic and no bytes. Stored transposed (down_proj.t.weight, [inter, hidden]) the same weights are a contiguous ROW, and this accumulate reads exactly the neurons the token asked for.

Source

pub fn row_dot(&self, r: usize, x: &[f32], scratch: &mut [f32]) -> f32

Dot of row r with x (gate/up active-neuron path). Reads only row r from the mmap — no full dequant. q4/vbit dequant the row into scratch first (rare for active-FFN weights).

Source

pub fn matvec(&self, x: &[f32], out: &mut [f32], pool: Option<&Pool>)

out = W · x (row-major). F32 delegates to the historical bit-exact path; Mapped runs the fused int8 kernel.

Source

pub fn matvec2( &self, x1: &[f32], x2: &[f32], o1: &mut [f32], o2: &mut [f32], pool: Option<&Pool>, )

Fused two-input matvec (MTP verify pair): weights streamed once.

Source§

impl QTensor

Source

pub fn q4tp_mapped(&self) -> Option<(&Arc<CmfModel>, usize)>

Batched matvec (prefill-GEMM): xs — row-major [b, cols], out — row-major [b, rows]. Element-wise semantics are IDENTICAL to b matvec calls (same dot kernels in the same order); the win — the weight row streams from DRAM once per batch, not b times. (model, index) when this is a memory-mapped q4tp tensor — the identity a device-resident chain needs to hand tp_matmat the weight without going through this struct’s own dispatch.

Source

pub fn matmat( &self, xs_all: &[f32], b: usize, out: &mut [f32], pool: Option<&Pool>, )

Source§

impl QTensor

Source

pub fn device_matmat(&self, xs: &[f32], b: usize, out: &mut [f32]) -> bool

The device GEMM this tensor would take, run once on the caller’s data — the startup parity probe’s arm, and the one place that knows which entry point each codec has.

It exists because the probe used to look for a q4tp weight by name AND dtype, and a container packed any other way was declared “host path” for the whole render even though its codec had a device GEMM of its own. A gate that only recognizes one codec is a gate that silently downgrades every other one.

Source

pub fn matvec_many<const N: usize>( ts: [&QTensor; N], x: &[f32], outs: [&mut [f32]; N], pool: Option<&Pool>, )

Multi-matrix job (roadmap §3 P0): N tensors sharing one input run under a SINGLE pool dispatch — QKV or gate+up cost one barrier instead of N. Per-row math is the exact same kernel as matvec (bit-identical outputs); only the dispatch is fused. Falls back to N sequential matvecs when the set is not a uniform q8-family/F32 group or there is no pool.

Source§

impl QTensor

Source

pub fn matvec2_many<const N: usize>( ts: [&QTensor; N], x1: &[f32], x2: &[f32], o1s: [&mut [f32]; N], o2s: [&mut [f32]; N], pool: Option<&Pool>, )

Pair-input multi-matrix job: N tensors × 2 shared inputs under a single pool dispatch — the MTP/pair decode path publishes one job for Q/K/V (and one for gate+up) instead of one per tensor. Per-row math is exactly matvec2’s kernels; bit-identical.

Source

pub fn matvec_silu_mul( gate: &QTensor, up: &QTensor, x: &[f32], out: &mut [f32], pool: Option<&Pool>, ) -> bool

Fused gate+up matvec with SiLU·mul: for each row r, computes silu(gate·x) * (up·x) and writes to out[r]. ONE pool dispatch, no intermediate g/u buffers, no separate silu pass. Falls back (returns false) for unsupported dtype combos.

Source

pub fn matvec_silu_mul_limited( gate: &QTensor, up: &QTensor, x: &[f32], out: &mut [f32], limit: f32, pool: Option<&Pool>, ) -> bool

Fused gate+up+SiLU with the GLM asymmetrical clamp. limit == 0 preserves the historical unclamped helper; a positive limit clamps up to both sides and gate only from above, matching the GLM SwiGLU reference. Keeping the limit in the row kernel avoids the two intermediate vectors and the extra combine pass on the Q2TP experts.

Source

pub fn moe_gate_up_many( pairs: &[(&QTensor, &QTensor)], x: &[f32], outs: &mut [Vec<f32>], pool: Option<&Pool>, ) -> bool

Every routed expert’s fused gate/up/SiLU under ONE pool dispatch.

The per-expert path pays a pool barrier per expert per stage: at 9 experts over 40 layers that is ~720 barriers a token, and a decode profile of Qwen3.6-35B-A3B showed the pool parked in psynch_cvwait about twice as long as it spent computing. Laying every expert’s rows end-to-end in one virtual row space collapses the stage to a single dispatch. The per-row body is the single-expert q4tp arm verbatim, so outputs are bit-identical.

false = something is outside the fused q4tp kernel (dtype, shape, or a transformed tensor); the caller walks the ordinary per-expert path. Float activations use the same exact scalar rows, still fused under one pool dispatch.

Source

pub fn moe_gate_up_many_limited( pairs: &[(&QTensor, &QTensor)], x: &[f32], outs: &mut [Vec<f32>], limit: f32, pool: Option<&Pool>, ) -> bool

Batched gate/up/SiLU with the optional GLM clamp. The public legacy helper above keeps its historical unclamped semantics; callers that implement a reference with a positive SwiGLU limit use this variant.

Source

pub fn moe_down_many( downs: &[&QTensor], gs: &[Vec<f32>], weights: &[f32], out: &mut [f32], pool: Option<&Pool>, ) -> bool

Every routed expert’s down projection, weighted and summed into out, under ONE pool dispatch.

Partitioned by OUTPUT row rather than by expert: each row is owned by a single worker, so the experts are summed in the caller’s order — the same sequence of f32 adds the serial out[i] += w·eo[i] loop performs, hence bit-identical. Partitioning by expert instead would race on the shared accumulator.

Source

pub fn moe_gate_up_rows( pairs: &[(&QTensor, &QTensor)], groups: &[Vec<usize>], xs: &[f32], outs: &mut [Vec<f32>], pool: Option<&Pool>, ) -> bool

moe_gate_up_many for SEVERAL tokens at once, decode-exact: expert e (pairs[e], q4tp) serves the tokens groups[e] (row indices into xs, each cols wide). Every (expert, token) output is bit-identical to moe_gate_up_many run on that token alone — the same int8 activation split, VNNI dots, outlier terms and inline SiLU — while each weight row is read once for all the tokens routed to its expert (the speculative verify’s expert sharing). outs is flat in (expert, token-of-group) order. False = not covered (not q4tp): the caller takes the per-token path. With float activations, the exact scalar row kernel replaces the int8 dot without changing the shared dispatch or route-order reduction.

Source

pub fn moe_down_rows( downs: &[&QTensor], group_lens: &[usize], gs: &[Vec<f32>], outs: &mut [Vec<f32>], pool: Option<&Pool>, ) -> bool

The per-(expert, token) down terms moe_down_many weights and sums, for SEVERAL tokens: outs[p][o] = down_e[o] · gs[p] (int8 split of gs[p], VNNI dot, outlier terms — bit-identical to that kernel’s d), each down row read once for its expert’s whole group. The caller sums w·d per token in its route order, which reproduces moe_down_many’s f32 sequence exactly. Layout as moe_gate_up_rows.

Auto Trait Implementations§

Blanket Implementations§

Source§

impl<T> Any for T
where T: 'static + ?Sized,

Source§

fn type_id(&self) -> TypeId

Gets the TypeId of self. Read more
Source§

impl<T> Borrow<T> for T
where T: ?Sized,

Source§

fn borrow(&self) -> &T

Immutably borrows from an owned value. Read more
Source§

impl<T> BorrowMut<T> for T
where T: ?Sized,

Source§

fn borrow_mut(&mut self) -> &mut T

Mutably borrows from an owned value. Read more
Source§

impl<T> From<T> for T

Source§

fn from(t: T) -> T

Returns the argument unchanged.

Source§

impl<T> Instrument for T

Source§

fn instrument(self, span: Span) -> Instrumented<Self> ⓘ

Instruments this type with the provided Span, returning an Instrumented wrapper. Read more
Source§

fn in_current_span(self) -> Instrumented<Self> ⓘ

Instruments this type with the current Span, returning an Instrumented wrapper. Read more
Source§

impl<T, U> Into<U> for T
where U: From<T>,

Source§

fn into(self) -> U

Calls U::from(self).

That is, this conversion is whatever the implementation of From<T> for U chooses to do.

Source§

impl<T, U> TryFrom<U> for T
where U: Into<T>,

Source§

type Error = !

The type returned in the event of a conversion error.
Source§

fn try_from(value: U) -> Result<T, !>

Performs the conversion.
Source§

impl<T, U> TryInto<U> for T
where U: TryFrom<T>,

Source§

type Error = <U as TryFrom<T>>::Error

The type returned in the event of a conversion error.
Source§

fn try_into(self) -> Result<U, <U as TryFrom<T>>::Error>

Performs the conversion.
Source§

impl<T> WithSubscriber for T

Source§

fn with_subscriber<S>(self, subscriber: S) -> WithDispatch<Self> ⓘ
where S: Into<Dispatch>,

Attaches the provided Subscriber to this type, returning a WithDispatch wrapper. Read more
Source§

fn with_current_subscriber(self) -> WithDispatch<Self> ⓘ

Attaches the current default Subscriber to this type, returning a WithDispatch wrapper. Read more