Expand description
Qwen3.5-2B (Qwen/Qwen3.5-2B, its text model) chatting on an AMD XDNA
NPU.
The model is a hybrid: 18 Gated DeltaNet (linear attention) layers and
6 gated softmax-attention layers, each followed by a SwiGLU MLP. The
bundle iron/applications/qwen3_5/export_qwen35.py writes holds every
compiled IRON kernel, the weights (the NPU’s pre-packed) and the
tokenizer; this crate replays the forward the Python app runs:
| NPU | host (here) | |
|---|---|---|
| prompt | every projection as an flm.GEMM over 256-row chunks, one hardware context | tokenizer, embedding, norms, the DeltaNet’s conv + recurrence, RoPE, attention |
| a generated token | every projection and the LM head as a GEMVbfp16, a second context | the same, one row |
| an image (the vision tower) | every projection as an flm.GEMM (K = 1024), a third context | preprocessing (image.rs), LayerNorms, 2D RoPE, attention |
Qwen35::chat answers one user turn (greedy), with or without
images; Qwen35::check verifies a bundle against the references its
exporter recorded.
Modules§
- image
- Image preprocessing, as the checkpoint’s
Qwen2VLImageProcessorFast(torchvision backend) does it: - model
- The forward, as
iron/applications/qwen3_5/qwen35_npu.pyruns it: the projections on the NPU (prefill:flm.GEMMs over the prompt’s 256-row chunks, one dispatch a projection; decode:GEMVbfp16s), the rest on the host in f32 – the norms and residual adds, the DeltaNet’s causal convolution, gates and recurrence (one step a token; the heads in parallel), partial M-RoPE, the attention layers’ KV cache and GQA attention. With a vision tower in the bundle (crate::vision), an image’s embeddings take its tokens’ rows and M-RoPE’s (t, h, w) positions. - npu
- The NPU side: every kernel of the bundle – the prefill GEMMs (an instruction stream a chunk count) on one hardware context, the decode GEMVs on another – and helpers for the device buffers.
- tokenizer
- Qwen3.5’s tokenizer, bit-exact with HF
tokenizers(tok(text, add_special_tokens=False)["input_ids"]), std-only, read straight from the checkpoint’stokenizer.json(the bundle carries a copy). - vision
- The vision tower, as
iron/applications/qwen3_5/qwen35_vision.pyruns it: an image’s patches -> the text model’s embeddings of its<|image_pad|>tokens.
Structs§
- Chat
Options - How to answer a turn.
- Qwen35
- RgbImage
- An 8-bit RGB image, rows top to bottom,
[height, width, 3]. - Stats
- What a generation did.
- Timing
- Wall time per stage, in first-seen order; NPU dispatch time is kept
under
npu:<kernel>.