pub fn reduce<T: Copy + MaybeSendSync, Op: ElementOp<T>, M, R, U>(
src: &StridedView<'_, T, Op>,
map_fn: M,
reduce_fn: R,
init: U,
) -> Result<U>Expand description
Full reduction with map function: reduce(init, op, map.(src)).
ยงEvaluation order
reduce_fn must be associative and commutative, and init must be an
identity of reduce_fn whenever the reduction runs on several threads.
A contiguous source is folded with sixteen independent lanes (see the
crate source for the exact lane assignment), non-contiguous sources are
traversed in a cache-friendly loop order, and parallel execution folds
chunks independently and combines them in chunk order. Floating point
results can therefore differ from a strict left to right fold in the last
bits; for a fixed shape, layout and thread count the result is repeatable.
Because map_fn and reduce_fn are opaque closures this entry cannot
select the dtype-specific SIMD kernels of crate::ErasedReducePlan;
callers that reduce with a known operation (sum, product, max, min) on a
supported dtype can use that plan for the fastest path.