Skip to main content

Crate taconite_sapiens2

Crate taconite_sapiens2 

Source
Expand description

Sapiens2-Pose (facebook/sapiens2-pose-0.4b, -1b) on an AMD XDNA NPU: a person box in an image -> 308 keypoints (body, feet, hands, face).

The bundle iron/applications/sapiens2_pose/export_sapiens2.py writes holds every compiled IRON kernel and the weights (the NPU ones pre-packed); this crate replays the forward the Python app runs (sapiens2_common.py / sapiens2_npu.py):

NPUhost (here)
backbone (ViT, 3081 tokens; 0.4b: 24 layers, 1024 wide, 1b: 40 layers, 1536 wide)every projection (flm.GEMMs: patch embedding, qkv, o, the SwiGLU gate+up, down) and the attention (the MHA operator)the box crop (preprocess.rs), RMSNorms, q / k norms, 2D RoPE, residual adds
head (2 transposed convs to 256 x 192, 3 1 x 1 convs, the predictor)every convolution as an flm.GEMM (a transposed conv as one GEMM over its input’s 2 x 2 windows)the window layout, InstanceNorm + SiLU
keypointsargmax + DARK refinement, back through the crop (post.rs)

Sapiens2::pose gives a box’s keypoints in image coordinates and their heatmap scores, as HF’s post_process_pose_estimation.

Re-exports§

pub use post::Keypoint;
pub use preprocess::BBox;

Modules§

model
The forward, as sapiens2_common.py runs it with sapiens2_npu.NpuBackend: every projection and convolution an flm.GEMM dispatch (all of its rows in one), the attention an MHA dispatch, the rest here in f32.
npu
The NPU side: the bundle’s kernels (kernels naming the same xclbin share its hardware context: a context an flm.GEMM K, and the attention’s – 6 of NPU2’s 16, shared by every process), flat bf16 device buffers, and the GEMM / MHA dispatches.
post
Heatmaps -> keypoints, as HF’s post_process_pose_estimation: each heatmap’s argmax (its value the score), refined by DARK / UDP – one Newton step on the log of the heatmap blurred by an 11 x 11 Gaussian (sigma 2, zero border, rescaled to keep the heatmap’s max) – then mapped from heatmap pixels back through the crop window.
preprocess
A person box -> the model’s input crop, as HF’s Sapiens2ImageProcessor(boxes=...) makes it: the box padded by 1.25 and widened or heightened to the crop’s aspect ratio, the region sampled onto the crop with PyTorch’s grid_sample (align_corners, zero padding; bilinear when the crop shrinks the region, bicubic when it grows it) on the float image, then ImageNet-normalized.

Structs§

Config
The model’s constants (the manifest’s params).
Sapiens2
Timing
Wall time per stage, in first-seen order; NPU dispatch time is kept under npu:<kernel>.

Enums§

Error

Constants§

VERSION
The bundle format: 2 when some GEMM leaves its bias to the host (<i>.down.bias, d<j>.bias: 1b); 1 (0.4b) is read too.

Functions§

cosine
Cosine similarity.