Expand description
§strided-basic
Shared typed CPU primitives and concrete copy, concatenation and reduction
execution. Generic map/zip, SIMD, indexing plans and their validation stay with
the traversal implementation. Concrete ordinary arithmetic/indexing dispatch
lives in strided-kernel; runtime-DAG fusion lives in strided-fused.
Neither of those packages is a dependency of this crate.
use strided_basic::{copy_into, StridedArray};
let src = StridedArray::<f64>::from_fn_col_major(&[2, 3], |i| (i[0] + 2*i[1]) as f64);
let mut dst = StridedArray::<f64>::col_major(&[2, 3]);
copy_into(&mut dst.view_mut(), &src.view()).unwrap();
assert_eq!(dst.get(&[1, 2]), 5.0);copy_into_uninit copies into StridedViewMut<MaybeUninit<T>> without first
initializing the destination. Success initializes every logical element, not
unreachable holes. It preserves the shared bounded-thread policy.
simd is enabled by default. parallel opts into the shared execution policy;
ExecContext::serial() requests the serial path. Features select implementation
capabilities, not operation families.
The execution module is a documented low-level kernel-extension contract.
Prevalidated entry points are unsafe, and a layout marker alone does not satisfy
their requirements. Application code should use the checked root APIs.
See kernel boundaries and the packaged
NOTICE / THIRD-PARTY-LICENSES for the Strided.jl lineage.
This new package has not yet been published; use the matching workspace checkout. Cache-optimized kernels for strided multidimensional array operations.
This crate is a Rust port of Julia’s Strided.jl and StridedViews.jl libraries, providing efficient operations on strided multidimensional array views.
§Core Types
StridedView/StridedViewMut: Dynamic-rank strided views over existing dataStridedArray: Owned strided multidimensional arrayElementOptrait and implementations (Identity,Conj,Transpose,Adjoint): Type-level element operations applied lazily on accessExecutionPolicy/with_execution_policy: optional bounds on strided-owned CPU fanout without creating a Rayon pool
§Primary API (view-based, Julia-compatible)
§Map Operations
map_into: Apply a function element-wise from source to destinationzip_map2_into,zip_map3_into,zip_map4_into: Multi-array element-wise operations
§In-place Update Operations
map_update_into,zip_update2_into,zip_update3_into:dest[i] = f(dest[i], inputs...), reading the old destination without a second view of it
§Reduce Operations
reduce: Full reduction with map functionreduce_axis: Reduce along a single axis
§Basic Operations
copy_into: Copy array contentsadd,mul: Element-wise arithmeticaxpy: y = alpha*x + y (array version)sum,dot: Reductionssymmetrize_into,symmetrize_conj_into: Matrix symmetrization
§Example
use strided_basic::{StridedView, StridedViewMut, StridedArray, Identity, map_into};
// Create a column-major array (Julia default)
let src = StridedArray::<f64>::from_fn_col_major(&[2, 3], |idx| {
(idx[0] * 10 + idx[1]) as f64
});
let mut dest = StridedArray::<f64>::col_major(&[2, 3]);
// Map with view-based API
map_into(&mut dest.view_mut(), &src.view(), |x| x * 2.0).unwrap();
assert_eq!(dest.get(&[1, 2]), 24.0); // (1*10 + 2) * 2§Cache Optimization
The library uses Julia’s blocking strategy for cache efficiency:
- Dimensions are sorted by stride magnitude for optimal memory access
- Operations are blocked into tiles fitting L1 cache (
BLOCK_MEMORY_SIZE= 32KB) - Contiguous arrays use fast paths bypassing the blocking machinery
Modules§
- auxiliary
- Auxiliary routines ported from StridedViews.jl/src/auxiliary.jl
- execution
- Execution contract for operation-family implementations.
- view
- Julia-like dynamic-rank strided view types.
Structs§
- Adjoint
- Adjoint operation: f(x) = adjoint(x) = conj(transpose(x)) For scalar numbers, this is conj.
- Concatenate
Plan - A compiled multi-input concatenate traversal.
- Conj
- Complex conjugate operation: f(x) = conj(x)
- Copy
Plan - A compiled copy traversal for one
(dims, dst_strides, src_strides)layout pair. - Dynamic
Slice Plan - A compiled fixed-window dynamic-slice traversal.
- Dynamic
Update Slice Plan - A compiled dynamic-update-slice traversal.
- Erased
ArgReduce Plan - Dtype-erased
argmax/argminalong one axis. - Erased
Concatenate Plan - Dtype-erased concatenate wrapper.
- Erased
Copy Plan - Dtype-erased wrapper around
CopyPlan. - Erased
Norm Plan - Dtype-erased fused
layer_norm/rms_normalong one axis. - Erased
RawStrided Mut - Borrowed dtype-erased raw strided output layout.
- Erased
RawStrided Ptr - Pointer-backed dtype-erased input used by one-shot write entry points.
- Erased
RawStrided Ref - Borrowed dtype-erased raw strided input layout.
- Erased
RawStrided Uninit Mut - Borrowed dtype-erased raw strided output whose reachable elements may be uninitialized.
- Erased
Reduce Plan - Dtype-erased reduction wrapper.
- Erased
Scan Plan - Dtype-erased cumulative scan (
cumsum/cumprod) along one axis. - Exec
Context - Caller-selected execution policy for prepared kernel replay.
- Gather
Plan - A compiled gather traversal for one value layout, index layout, and output layout.
- Gather
Spec - Gather configuration shared by generic and erased replay.
- Identity
- Identity operation: f(x) = x
- Lazy
Outer Product Layout - Storage layout for a lazily ordered outer product.
- Norm
Spec - Kind,
epsand optional affine parameters of anErasedNormPlan. - PadPlan
- A compiled pad traversal.
- RawStrided
Mut - Borrowed raw strided output layout.
- RawStrided
Ref - Borrowed raw strided input layout.
- Reverse
Plan - A compiled reverse traversal over selected axes.
- Scan
Options - Direction and inclusivity of an
ErasedScanPlan. - Scatter
Plan - A compiled additive scatter traversal.
- Scatter
Spec - Scatter configuration shared by generic and erased replay.
- Slice
Plan - A compiled static-slice traversal.
- Strided
Array - Owned strided multidimensional array.
- Strided
View - Dynamic-rank immutable strided view with lazy element operations.
- Strided
View Mut - Dynamic-rank mutable strided view.
- Transpose
- Transpose operation: f(x) = transpose(x) For scalar numbers, this is identity. For matrix elements, this would transpose each element.
Enums§
- ArgReduce
Op - Runtime operation of an
ErasedArgReducePlan. - Compare
Op - Runtime comparison selected once before entering the element loop.
- Execution
Policy - Controls how a strided operation may use CPU threads.
- KernelD
Type - Dtypes supported by dtype-erased kernel entry points.
- Norm
Kind - Normalization computed by an
ErasedNormPlan. - Reduce
Op - Runtime reduction operation for dtype-erased full reductions.
- ScanOp
- Runtime operation of an
ErasedScanPlan. - Strided
Error - Errors that can occur during strided array operations.
Constants§
- BLOCK_
MEMORY_ SIZE - Block memory size for cache-optimized iteration (L1 cache target).
- CACHE_
LINE_ SIZE - Cache line size in bytes.
- RAW_
FUSED_ RANK_ LIMIT - Maximum rank fused on the stack before falling back to the view kernels.
Traits§
- Composable
Element Op - Trait for element operations that support type-level composition.
- Compose
- Helper trait for composing two ElementOp types.
- Element
Op - Trait for element-wise operations applied to strided views.
- Element
OpApply - Trait for types that support element operations (conj, transpose, adjoint).
- Gather
Index - Index scalar types accepted by
GatherPlan. - Kernel
Storage Element - A scalar type that can safely witness storage for an erased kernel view.
- Maybe
Send - Equivalent to
Sendwhenparallelis enabled; blanket-impl otherwise. - Maybe
Send Sync - Equivalent to
Send+Syncwhenparallelis enabled; blanket-impl otherwise. - Maybe
Simd Ops - Trait for types that may have SIMD-accelerated sum/dot operations.
- Maybe
Sync - Equivalent to
Syncwhenparallelis enabled; blanket-impl otherwise.
Functions§
- add
- Element-wise addition:
dest[i] += src[i]. - axpby_
accum - Update caller-owned contiguous storage:
y = alpha * x + beta * y. - axpy
- AXPY:
dest[i] = alpha * src[i] + dest[i]. - axpy_
conj_ raw dest = alpha * conj(src) + destover borrowed raw strided layouts.- axpy_
raw dest = alpha * src + destover borrowed raw strided layouts.- batched_
outer_ product_ into - Compute
dest[lhs_free..., rhs_free..., batch...] = lhs[lhs_free..., batch...] * rhs[rhs_free..., batch...]. - batched_
outer_ product_ into_ uninit - Compute a batched outer product into a fully overwritten uninitialized output.
- broadcast_
mul_ into - Broadcasted element-wise multiplication:
dest[i] = a[i] * b[i]. - broadcast_
mul_ into_ uninit - Broadcast and multiply into a fully overwritten uninitialized output.
- col_
major_ strides - Compute column-major strides (Julia default: first index varies fastest).
- compare_
into - Compare two ordered views elementwise into a Boolean destination.
- compare_
into_ uninit - Compare two views into a fully overwritten uninitialized Boolean output.
- copy_
conj - Copy with complex conjugation:
dest[i] = conj(src[i]). - copy_
into - Copy elements from source to destination:
dest[i] = src[i]. - copy_
into_ col_ major - Copy elements from
srctodst, optimized for col-major destination. - copy_
into_ uninit - Copy initialized values into a potentially uninitialized destination.
- copy_
scale - Copy with scaling:
dest[i] = scale * src[i]. - copy_
scale_ conj_ raw dest = scale * conj(src)over borrowed raw strided layouts.- copy_
scale_ raw dest = scale * srcover borrowed raw strided layouts.- copy_
transpose_ scale_ into - Copy with transpose and scaling:
dest[j,i] = scale * src[i,j]. - dot
- Dot product:
sum(OpA::apply(a[i]) * OpB::apply(b[i])). - embed_
diagonal_ into_ uninit - Embed a dense column-major input on a diagonal in a new axis.
- fma
- Fused multiply-add:
dest[i] += OpA::apply(a[i]) * OpB::apply(b[i]). - map_
into - Apply a function element-wise from source to destination.
- map_
update_ into - Update in place:
dest[i] = f(OpD(dest[i])). - mul
- Element-wise multiplication:
dest[i] *= src[i]. - mul_
into - Element-wise multiplication:
dest[i] = a[i] * b[i]. - mul_
into_ uninit - Multiply two views into a fully overwritten uninitialized output.
- plan_
lazy_ outer_ product - Plan an outer-product output whose memory order follows the inputs’ physical stride order instead of the logical output order.
- reduce
- Full reduction with map function:
reduce(init, op, map.(src)). - reduce_
axis - Reduce along a single axis, returning a new StridedArray.
- robust_
complex_ divide_ f32 - Divide two
Complex<f32>values by widening toComplex<f64>(Julia’sComplex{Float32}route), so thef32squares cannot overflow. - robust_
complex_ divide_ f64 - Divide two
Complex<f64>values without the|z|²overflow/underflow of the textbook formula. - row_
major_ strides - Compute row-major strides (C default: last index varies fastest).
- sum
- Sum all elements:
sum(src). - symmetrize_
conj_ into - Conjugate-symmetrize a square matrix:
dest = (src + conj(src^T)) / 2. - symmetrize_
into - Symmetrize a square matrix:
dest = (src + src^T) / 2. - triangular_
mask_ into_ uninit - Copy dense column-major matrices, replacing the masked triangle with
fill. - with_
execution_ policy - Execute
operationunder an explicit strided-kernel execution policy. - zip_
map2_ into - Binary element-wise operation:
dest[i] = f(a[i], b[i]). - zip_
map3_ into - Ternary element-wise operation:
dest[i] = f(a[i], b[i], c[i]). - zip_
map4_ into - Quaternary element-wise operation:
dest[i] = f(a[i], b[i], c[i], e[i]). - zip_
update2_ into - Update in place from one input:
dest[i] = f(OpD(dest[i]), OpA(a[i])). - zip_
update3_ into - Update in place from two inputs:
dest[i] = f(OpD(dest[i]), OpA(a[i]), OpB(b[i])).
Type Aliases§
- Result
- Result type for strided array operations.