Expand description
Threshold-guarded data parallelism with a thread-count-independent shape.
Two rules make every helper here safe to drop into a load path that has to stay reproducible.
The partition is a pure function of the input length. Work is split into fixed-size blocks, never into “one piece per thread”, so the same input produces the same blocks on a laptop and on a build machine. A reduction then combines block results in index order. Floating-point addition is not associative, so a thread-count-dependent partition would give a thread-count-dependent answer; a fixed one cannot.
Small inputs stay on the calling thread. Every entry point takes a minimum length and runs serially below it. A ligand must not pay for a thread pool, and a ribosome must not be denied one.
Cost is O(n / threads) above the threshold and O(n) below it, with one
Vec of block results — blocks = n / BLOCK entries — for the reducing
forms and no allocation at all for the in-place forms.
Every parallel branch schedules onto Rayon’s process-wide registry. No
scene, renderer or language binding constructs a private pool, so molgfx,
molframe and independent Python Engine objects share the same Rust workers
instead of multiplying thread counts and scratch memory.
Constants§
- BLOCK
- Elements per block.
Functions§
- all_
blocks - Whether
fholds for every element, short-circuiting per block. - for_
each_ block - Applies
fto each(offset, block)ofslice, in parallel abovemin_len. - map_
blocks_ into - Writes
dstfromsrcblock-wise, in parallel abovemin_len. - map_
zip_ blocks_ into - Writes
dstfrom two aligned sources block-wise, in parallel abovemin_len. - reduce_
blocks - Reduces
slicewith a fixed block partition and an in-order combine.