Skip to main content

Crate taconite_qwen35

Crate taconite_qwen35 

Source
Expand description

Qwen3.5-2B (Qwen/Qwen3.5-2B, its text model) chatting on an AMD XDNA NPU.

The model is a hybrid: 18 Gated DeltaNet (linear attention) layers and 6 gated softmax-attention layers, each followed by a SwiGLU MLP. The bundle iron/applications/qwen3_5/export_qwen35.py writes holds every compiled IRON kernel, the weights (the NPU’s pre-packed) and the tokenizer; this crate replays the forward the Python app runs:

NPUhost (here)
promptevery projection as an flm.GEMM over 256-row chunks, one hardware contexttokenizer, embedding, norms, the DeltaNet’s conv + recurrence, RoPE, attention
a generated tokenevery projection and the LM head as a GEMVbfp16, a second contextthe same, one row
an image (the vision tower)every projection as an flm.GEMM (K = 1024), a third contextpreprocessing (image.rs), LayerNorms, 2D RoPE, attention

Qwen35::chat answers one user turn (greedy), with or without images; Qwen35::check verifies a bundle against the references its exporter recorded.

Modules§

image
Image preprocessing, as the checkpoint’s Qwen2VLImageProcessorFast (torchvision backend) does it:
model
The forward, as iron/applications/qwen3_5/qwen35_npu.py runs it: the projections on the NPU (prefill: flm.GEMMs over the prompt’s 256-row chunks, one dispatch a projection; decode: GEMVbfp16s), the rest on the host in f32 – the norms and residual adds, the DeltaNet’s causal convolution, gates and recurrence (one step a token; the heads in parallel), partial M-RoPE, the attention layers’ KV cache and GQA attention. With a vision tower in the bundle (crate::vision), an image’s embeddings take its tokens’ rows and M-RoPE’s (t, h, w) positions.
npu
The NPU side: every kernel of the bundle – the prefill GEMMs (an instruction stream a chunk count) on one hardware context, the decode GEMVs on another – and helpers for the device buffers.
tokenizer
Qwen3.5’s tokenizer, bit-exact with HF tokenizers (tok(text, add_special_tokens=False)["input_ids"]), std-only, read straight from the checkpoint’s tokenizer.json (the bundle carries a copy).
vision
The vision tower, as iron/applications/qwen3_5/qwen35_vision.py runs it: an image’s patches -> the text model’s embeddings of its <|image_pad|> tokens.

Structs§

ChatOptions
How to answer a turn.
Qwen35
RgbImage
An 8-bit RGB image, rows top to bottom, [height, width, 3].
Stats
What a generation did.
Timing
Wall time per stage, in first-seen order; NPU dispatch time is kept under npu:<kernel>.

Enums§

Error

Constants§

VERSION