Expand description
Portable SIMD kernels using macerator.
These kernels work across all architectures supported by macerator:
- aarch64 (NEON)
- x86_64 (AVX2, AVX512, SSE)
- wasm32 (SIMD128)
- Scalar fallback for embedded/other platforms
Functionsยง
- max_f32
- Find the maximum element in a f32 slice using SIMD.
- min_f32
- Find the minimum element in a f32 slice using SIMD.
- scatter_
add_ batched - Batched scatter-add for middle-dim reductions. For tensors like [B, M, K] reducing dim=1.
- scatter_
add_ f32 - Scatter-add: for each row, add to corresponding output positions. Used for cache-friendly first-dim and middle-dim reductions.
- sum_f32
- Sum all elements in a f32 slice using SIMD with 4 accumulators.
- sum_
rows_ f32 - Sum each row, storing results in output slice. Used for last-dim reductions.