Expand description
SIMD-optimized kernels for tensor operations.
Provides portable SIMD implementations via macerator with automatic
dispatch to the best available instruction set:
- aarch64: NEON
- x86_64: AVX2, AVX512, SSE
- wasm32: SIMD128
- Other: Scalar fallback
Enable with the simd feature flag (enabled by default).
Modules§
Enums§
Functions§
- abs_
inplace_ f32 - add_
inplace_ f32 - add_
shared_ row_ inplace_ f32 - bool_
and_ inplace_ u8 - bool_
and_ u8 - bool_
not_ inplace_ u8 - bool_
not_ u8 - bool_
or_ inplace_ u8 - bool_
or_ u8 - bool_
xor_ inplace_ u8 - bool_
xor_ u8 - cmp_f32
- cmp_
scalar_ f32 - div_
inplace_ f32 - div_
shared_ row_ inplace_ f32 - mask_
fill_ f32 - Conditional fill:
out[i] = if mask[i] != 0 { fill_value } else { tensor[i] } - mask_
fill_ f64 - Conditional fill for f64.
- mask_
fill_ i64 - Conditional fill for i64.
- mask_
fill_ u8 - Conditional fill for u8 (bool tensors).
- mask_
where_ f32 - Conditional select:
out[i] = if mask[i] != 0 { value[i] } else { tensor[i] } - mask_
where_ f64 - Conditional select for f64.
- mask_
where_ i64 - Conditional select for i64.
- mask_
where_ u8 - Conditional select for u8 (bool tensors).
- mul_
inplace_ f32 - mul_
shared_ row_ inplace_ f32 - recip_
inplace_ f32 - sub_
inplace_ f32 - sub_
shared_ row_ inplace_ f32