pub struct ErasedReducePlan { /* private fields */ }Expand description
Dtype-erased reduction wrapper.
This is the erased replay boundary for full-tensor scalar reductions and axis reductions with a fixed output layout. It supports only operations with an unambiguous identity value in the selected dtype.
§Evaluation order
The op is matched once per execution; every loop runs a kernel
specialized for one (dtype, op) pair. Integer results are exact
(wrapping) and ReduceOp::Max / ReduceOp::Min do not depend on
order, so the rules below only affect the rounding of floating point and
complex ReduceOp::Sum, ReduceOp::Product and
ReduceOp::SumSquares. Two details of float max and min follow the
visit order: a NaN input makes the result NaN, but which NaN (payload and
sign) is returned is unspecified, and when the extremum is a zero its sign
is unspecified if both +0.0 and -0.0 occur.
- Full reductions (from
Self::compile, orSelf::compile_axeswith every axis reduced) visit the source in a layout derived order: extent one axes are dropped, the rest are ordered by increasing absolute stride and contiguous axes are fused into runs. Each run is reduced by the SIMD kernel (unit stride,simdfeature), a sixteen lane kernel (other unit stride runs) or an eight lane kernel (strided runs), and runs are combined left to right. A dense column major source therefore reduces exactly like one contiguous slice. - With a parallel context and more than
MINTHREADLENGTHelements, a full reduction splits the traversal into one deterministic range per worker thread. Each range runs the same run kernel as the serial path and the partials are combined in a fixed tree. The result is bitwise repeatable for a fixed layout and thread count, but may differ between thread counts and from the serial result. - Axis reductions with kept axes compute every output independently, so the result never depends on the context or thread count. Each output folds its reduced elements sequentially in caller reduced axis order, except that a unit stride leading reduced run (after fusing adjacent reduced axes) of length at least 16 is reduced by the contiguous run kernel, with runs folded in order. When instead the leading kept axis has unit stride, adjacent outputs are accumulated as a block from one contiguous source run per reduced element; each output still folds in caller reduced axis order.
Implementations§
Source§impl ErasedReducePlan
impl ErasedReducePlan
Sourcepub fn compile(
dtype: KernelDType,
op: ReduceOp,
dims: &[usize],
src_strides: &[isize],
) -> Result<ErasedReducePlan, StridedError>
pub fn compile( dtype: KernelDType, op: ReduceOp, dims: &[usize], src_strides: &[isize], ) -> Result<ErasedReducePlan, StridedError>
Validate and store a full-reduction plan for one dtype and source layout.
Sourcepub fn compile_axes(
dtype: KernelDType,
op: ReduceOp,
src_dims: &[usize],
src_strides: &[isize],
dest_dims: &[usize],
dest_strides: &[isize],
axes: &[usize],
) -> Result<ErasedReducePlan, StridedError>
pub fn compile_axes( dtype: KernelDType, op: ReduceOp, src_dims: &[usize], src_strides: &[isize], dest_dims: &[usize], dest_strides: &[isize], axes: &[usize], ) -> Result<ErasedReducePlan, StridedError>
Validate and store an axis-reduction plan for one dtype and fixed source/output layouts.
axes names the source axes reduced away. Output dimensions must be the
remaining source dimensions in source-axis order. When all axes are
reduced, any output layout with exactly one reachable element is accepted.
pub fn dtype(&self) -> KernelDType
pub fn op(&self) -> ReduceOp
Sourcepub fn execute(
&self,
ctx: &ExecContext,
dest: &mut ErasedRawStridedMut<'_>,
src: &ErasedRawStridedRef<'_>,
) -> Result<(), StridedError>
pub fn execute( &self, ctx: &ExecContext, dest: &mut ErasedRawStridedMut<'_>, src: &ErasedRawStridedRef<'_>, ) -> Result<(), StridedError>
Execute the reduction into an erased output descriptor.
Sourcepub fn execute_uninit(
&self,
ctx: &ExecContext,
dest: &mut ErasedRawStridedUninitMut<'_>,
src: &ErasedRawStridedPtr<'_>,
) -> Result<(), StridedError>
pub fn execute_uninit( &self, ctx: &ExecContext, dest: &mut ErasedRawStridedUninitMut<'_>, src: &ErasedRawStridedPtr<'_>, ) -> Result<(), StridedError>
On success, every reachable destination slot is fully overwritten;
unreachable holes are neither read nor initialized. Validation errors
are returned before any destination write. A panic during execution
may leave a partially initialized MaybeUninit destination, which is
still safely droppable; no readable value is promised for unwritten
reachable slots.
Trait Implementations§
Source§impl Clone for ErasedReducePlan
impl Clone for ErasedReducePlan
Source§fn clone(&self) -> ErasedReducePlan
fn clone(&self) -> ErasedReducePlan
1.0.0 (const: unstable) · Source§fn clone_from(&mut self, source: &Self)
fn clone_from(&mut self, source: &Self)
source. Read more