pub fn reduce_axis<T, Op, M, R, U>(
src: &StridedView<'_, T, Op>,
axis: usize,
map_fn: M,
reduce_fn: R,
init: U,
) -> Result<StridedArray<U>, StridedError>where
T: Copy + MaybeSendSync,
Op: ElementOp<T>,
M: Fn(T) -> U + MaybeSync,
R: Fn(U, U) -> U + MaybeSync,
U: Clone + MaybeSendSync,Expand description
Reduce along a single axis, returning a new StridedArray.
Every output folds init once and then its reduced elements. When the
reduced axis has unit stride those elements are folded with the lane
scheme of reduce, so reduce_fn must be associative and commutative.
When the leading kept axis has unit stride, a block of adjacent outputs
is accumulated one contiguous source column at a time, which keeps the
left to right order per output. With the parallel feature, outputs are
split across threads once the number of source elements read exceeds the
threading threshold; each output is computed by one thread, so the result
does not depend on the thread count.