Skip to main content

Module kernels

Module kernels 

Source
Expand description

Portable SIMD kernels using macerator.

These kernels work across all architectures supported by macerator:

  • aarch64 (NEON)
  • x86_64 (AVX2, AVX512, SSE)
  • wasm32 (SIMD128)
  • Scalar fallback for embedded/other platforms

Functionsยง

max_f32
Find the maximum element in a f32 slice using SIMD.
min_f32
Find the minimum element in a f32 slice using SIMD.
scatter_add_batched
Batched scatter-add for middle-dim reductions. For tensors like [B, M, K] reducing dim=1.
scatter_add_f32
Scatter-add: for each row, add to corresponding output positions. Used for cache-friendly first-dim and middle-dim reductions.
sum_f32
Sum all elements in a f32 slice using SIMD with 4 accumulators.
sum_rows_f32
Sum each row, storing results in output slice. Used for last-dim reductions.