Skip to main content

Crate strided_basic

Crate strided_basic 

Source
Expand description

§strided-basic

Shared typed CPU primitives and concrete copy, concatenation and reduction execution. Generic map/zip, SIMD, indexing plans and their validation stay with the traversal implementation. Concrete ordinary arithmetic/indexing dispatch lives in strided-kernel; runtime-DAG fusion lives in strided-fused. Neither of those packages is a dependency of this crate.

use strided_basic::{copy_into, StridedArray};
let src = StridedArray::<f64>::from_fn_col_major(&[2, 3], |i| (i[0] + 2*i[1]) as f64);
let mut dst = StridedArray::<f64>::col_major(&[2, 3]);
copy_into(&mut dst.view_mut(), &src.view()).unwrap();
assert_eq!(dst.get(&[1, 2]), 5.0);

copy_into_uninit copies into StridedViewMut<MaybeUninit<T>> without first initializing the destination. Success initializes every logical element, not unreachable holes. It preserves the shared bounded-thread policy.

simd is enabled by default. parallel opts into the shared execution policy; ExecContext::serial() requests the serial path. Features select implementation capabilities, not operation families.

The execution module is a documented low-level kernel-extension contract. Prevalidated entry points are unsafe, and a layout marker alone does not satisfy their requirements. Application code should use the checked root APIs.

See kernel boundaries and the packaged NOTICE / THIRD-PARTY-LICENSES for the Strided.jl lineage.

This new package has not yet been published; use the matching workspace checkout. Cache-optimized kernels for strided multidimensional array operations.

This crate is a Rust port of Julia’s Strided.jl and StridedViews.jl libraries, providing efficient operations on strided multidimensional array views.

§Core Types

§Primary API (view-based, Julia-compatible)

§Map Operations

§In-place Update Operations

§Reduce Operations

§Basic Operations

§Example

use strided_basic::{StridedView, StridedViewMut, StridedArray, Identity, map_into};

// Create a column-major array (Julia default)
let src = StridedArray::<f64>::from_fn_col_major(&[2, 3], |idx| {
    (idx[0] * 10 + idx[1]) as f64
});
let mut dest = StridedArray::<f64>::col_major(&[2, 3]);

// Map with view-based API
map_into(&mut dest.view_mut(), &src.view(), |x| x * 2.0).unwrap();
assert_eq!(dest.get(&[1, 2]), 24.0); // (1*10 + 2) * 2

§Cache Optimization

The library uses Julia’s blocking strategy for cache efficiency:

  • Dimensions are sorted by stride magnitude for optimal memory access
  • Operations are blocked into tiles fitting L1 cache (BLOCK_MEMORY_SIZE = 32KB)
  • Contiguous arrays use fast paths bypassing the blocking machinery

Modules§

auxiliary
Auxiliary routines ported from StridedViews.jl/src/auxiliary.jl
execution
Execution contract for operation-family implementations.
view
Julia-like dynamic-rank strided view types.

Structs§

Adjoint
Adjoint operation: f(x) = adjoint(x) = conj(transpose(x)) For scalar numbers, this is conj.
ConcatenatePlan
A compiled multi-input concatenate traversal.
Conj
Complex conjugate operation: f(x) = conj(x)
CopyPlan
A compiled copy traversal for one (dims, dst_strides, src_strides) layout pair.
DynamicSlicePlan
A compiled fixed-window dynamic-slice traversal.
DynamicUpdateSlicePlan
A compiled dynamic-update-slice traversal.
ErasedArgReducePlan
Dtype-erased argmax / argmin along one axis.
ErasedConcatenatePlan
Dtype-erased concatenate wrapper.
ErasedCopyPlan
Dtype-erased wrapper around CopyPlan.
ErasedNormPlan
Dtype-erased fused layer_norm / rms_norm along one axis.
ErasedRawStridedMut
Borrowed dtype-erased raw strided output layout.
ErasedRawStridedPtr
Pointer-backed dtype-erased input used by one-shot write entry points.
ErasedRawStridedRef
Borrowed dtype-erased raw strided input layout.
ErasedRawStridedUninitMut
Borrowed dtype-erased raw strided output whose reachable elements may be uninitialized.
ErasedReducePlan
Dtype-erased reduction wrapper.
ErasedScanPlan
Dtype-erased cumulative scan (cumsum / cumprod) along one axis.
ExecContext
Caller-selected execution policy for prepared kernel replay.
GatherPlan
A compiled gather traversal for one value layout, index layout, and output layout.
GatherSpec
Gather configuration shared by generic and erased replay.
Identity
Identity operation: f(x) = x
LazyOuterProductLayout
Storage layout for a lazily ordered outer product.
NormSpec
Kind, eps and optional affine parameters of an ErasedNormPlan.
PadPlan
A compiled pad traversal.
RawStridedMut
Borrowed raw strided output layout.
RawStridedRef
Borrowed raw strided input layout.
ReversePlan
A compiled reverse traversal over selected axes.
ScanOptions
Direction and inclusivity of an ErasedScanPlan.
ScatterPlan
A compiled additive scatter traversal.
ScatterSpec
Scatter configuration shared by generic and erased replay.
SlicePlan
A compiled static-slice traversal.
StridedArray
Owned strided multidimensional array.
StridedView
Dynamic-rank immutable strided view with lazy element operations.
StridedViewMut
Dynamic-rank mutable strided view.
Transpose
Transpose operation: f(x) = transpose(x) For scalar numbers, this is identity. For matrix elements, this would transpose each element.

Enums§

ArgReduceOp
Runtime operation of an ErasedArgReducePlan.
CompareOp
Runtime comparison selected once before entering the element loop.
ExecutionPolicy
Controls how a strided operation may use CPU threads.
KernelDType
Dtypes supported by dtype-erased kernel entry points.
NormKind
Normalization computed by an ErasedNormPlan.
ReduceOp
Runtime reduction operation for dtype-erased full reductions.
ScanOp
Runtime operation of an ErasedScanPlan.
StridedError
Errors that can occur during strided array operations.

Constants§

BLOCK_MEMORY_SIZE
Block memory size for cache-optimized iteration (L1 cache target).
CACHE_LINE_SIZE
Cache line size in bytes.
RAW_FUSED_RANK_LIMIT
Maximum rank fused on the stack before falling back to the view kernels.

Traits§

ComposableElementOp
Trait for element operations that support type-level composition.
Compose
Helper trait for composing two ElementOp types.
ElementOp
Trait for element-wise operations applied to strided views.
ElementOpApply
Trait for types that support element operations (conj, transpose, adjoint).
GatherIndex
Index scalar types accepted by GatherPlan.
KernelStorageElement
A scalar type that can safely witness storage for an erased kernel view.
MaybeSend
Equivalent to Send when parallel is enabled; blanket-impl otherwise.
MaybeSendSync
Equivalent to Send + Sync when parallel is enabled; blanket-impl otherwise.
MaybeSimdOps
Trait for types that may have SIMD-accelerated sum/dot operations.
MaybeSync
Equivalent to Sync when parallel is enabled; blanket-impl otherwise.

Functions§

add
Element-wise addition: dest[i] += src[i].
axpby_accum
Update caller-owned contiguous storage: y = alpha * x + beta * y.
axpy
AXPY: dest[i] = alpha * src[i] + dest[i].
axpy_conj_raw
dest = alpha * conj(src) + dest over borrowed raw strided layouts.
axpy_raw
dest = alpha * src + dest over borrowed raw strided layouts.
batched_outer_product_into
Compute dest[lhs_free..., rhs_free..., batch...] = lhs[lhs_free..., batch...] * rhs[rhs_free..., batch...].
batched_outer_product_into_uninit
Compute a batched outer product into a fully overwritten uninitialized output.
broadcast_mul_into
Broadcasted element-wise multiplication: dest[i] = a[i] * b[i].
broadcast_mul_into_uninit
Broadcast and multiply into a fully overwritten uninitialized output.
col_major_strides
Compute column-major strides (Julia default: first index varies fastest).
compare_into
Compare two ordered views elementwise into a Boolean destination.
compare_into_uninit
Compare two views into a fully overwritten uninitialized Boolean output.
copy_conj
Copy with complex conjugation: dest[i] = conj(src[i]).
copy_into
Copy elements from source to destination: dest[i] = src[i].
copy_into_col_major
Copy elements from src to dst, optimized for col-major destination.
copy_into_uninit
Copy initialized values into a potentially uninitialized destination.
copy_scale
Copy with scaling: dest[i] = scale * src[i].
copy_scale_conj_raw
dest = scale * conj(src) over borrowed raw strided layouts.
copy_scale_raw
dest = scale * src over borrowed raw strided layouts.
copy_transpose_scale_into
Copy with transpose and scaling: dest[j,i] = scale * src[i,j].
dot
Dot product: sum(OpA::apply(a[i]) * OpB::apply(b[i])).
embed_diagonal_into_uninit
Embed a dense column-major input on a diagonal in a new axis.
fma
Fused multiply-add: dest[i] += OpA::apply(a[i]) * OpB::apply(b[i]).
map_into
Apply a function element-wise from source to destination.
map_update_into
Update in place: dest[i] = f(OpD(dest[i])).
mul
Element-wise multiplication: dest[i] *= src[i].
mul_into
Element-wise multiplication: dest[i] = a[i] * b[i].
mul_into_uninit
Multiply two views into a fully overwritten uninitialized output.
plan_lazy_outer_product
Plan an outer-product output whose memory order follows the inputs’ physical stride order instead of the logical output order.
reduce
Full reduction with map function: reduce(init, op, map.(src)).
reduce_axis
Reduce along a single axis, returning a new StridedArray.
robust_complex_divide_f32
Divide two Complex<f32> values by widening to Complex<f64> (Julia’s Complex{Float32} route), so the f32 squares cannot overflow.
robust_complex_divide_f64
Divide two Complex<f64> values without the |z|² overflow/underflow of the textbook formula.
row_major_strides
Compute row-major strides (C default: last index varies fastest).
sum
Sum all elements: sum(src).
symmetrize_conj_into
Conjugate-symmetrize a square matrix: dest = (src + conj(src^T)) / 2.
symmetrize_into
Symmetrize a square matrix: dest = (src + src^T) / 2.
triangular_mask_into_uninit
Copy dense column-major matrices, replacing the masked triangle with fill.
with_execution_policy
Execute operation under an explicit strided-kernel execution policy.
zip_map2_into
Binary element-wise operation: dest[i] = f(a[i], b[i]).
zip_map3_into
Ternary element-wise operation: dest[i] = f(a[i], b[i], c[i]).
zip_map4_into
Quaternary element-wise operation: dest[i] = f(a[i], b[i], c[i], e[i]).
zip_update2_into
Update in place from one input: dest[i] = f(OpD(dest[i]), OpA(a[i])).
zip_update3_into
Update in place from two inputs: dest[i] = f(OpD(dest[i]), OpA(a[i]), OpB(b[i])).

Type Aliases§

Result
Result type for strided array operations.