Skip to main content

Module parallel

Module parallel 

Source
Expand description

Threshold-guarded data parallelism with a thread-count-independent shape.

Two rules make every helper here safe to drop into a load path that has to stay reproducible.

The partition is a pure function of the input length. Work is split into fixed-size blocks, never into “one piece per thread”, so the same input produces the same blocks on a laptop and on a build machine. A reduction then combines block results in index order. Floating-point addition is not associative, so a thread-count-dependent partition would give a thread-count-dependent answer; a fixed one cannot.

Small inputs stay on the calling thread. Every entry point takes a minimum length and runs serially below it. A ligand must not pay for a thread pool, and a ribosome must not be denied one.

Cost is O(n / threads) above the threshold and O(n) below it, with one Vec of block results — blocks = n / BLOCK entries — for the reducing forms and no allocation at all for the in-place forms.

Every parallel branch schedules onto Rayon’s process-wide registry. No scene, renderer or language binding constructs a private pool, so molgfx, molframe and independent Python Engine objects share the same Rust workers instead of multiplying thread counts and scratch memory.

Constants§

BLOCK
Elements per block.

Functions§

all_blocks
Whether f holds for every element, short-circuiting per block.
for_each_block
Applies f to each (offset, block) of slice, in parallel above min_len.
map_blocks_into
Writes dst from src block-wise, in parallel above min_len.
map_zip_blocks_into
Writes dst from two aligned sources block-wise, in parallel above min_len.
reduce_blocks
Reduces slice with a fixed block partition and an in-order combine.