pub struct Pool { /* private fields */ }Expand description
Persistent thread pool: shared job slot, epoch dispatch, caller participation.
Implementations§
Source§impl Pool
impl Pool
pub fn new(n_workers: usize) -> Self
Sourcepub fn with_spin(n_workers: usize, spin_budget: usize) -> Self
pub fn with_spin(n_workers: usize, spin_budget: usize) -> Self
Explicit spin budget (tests pin it without touching the env).
Sourcepub fn effective_threads() -> usize
pub fn effective_threads() -> usize
The thread count from_env would use RIGHT NOW: forced (C ABI)
CMF_THREADS > big-core topology > available_parallelism−1. ≤1 means the model runs serial (no pool). Introspection (
execution_mode, status endpoints) must report THIS, not available_parallelism.
Sourcepub fn from_env() -> Option<Arc<Self>>
pub fn from_env() -> Option<Arc<Self>>
Pool sized from CMF_THREADS (see module docs). None = serial.
Without the env, heterogeneous ARM defaults to its BIG cores.
Sourcepub fn bind_numa(&self, regions: &[&[u8]])
pub fn bind_numa(&self, regions: &[&[u8]])
Keep the pool on the NUMA node that holds regions (the model’s
weight bytes). Linux with two or more nodes only; CMF_NUMA=0
turns it off, CMF_NUMA=node:<n> forces a node.
WHY: decode streams every weight once per token, and on a two-socket host the page cache holds a file on whichever node read it. Unpinned, the scheduler spreads the workers over both sockets and half the matvec rows cross the socket link. Measured on a 2×EPYC 7763 pod with the model’s pages all on node 0 (31 CPUs of cgroup quota): a STREAM-style read over a node-0 buffer gives 42 GB/s from 31 unpinned threads and 74 GB/s from 31 threads kept on node 0. The mask is the node’s physical cores (first SMT sibling) when there are enough of them for the pool, else the whole node; never narrower than the pool, so nothing oversubscribes. Threads are bound to a SET of cores, not to one core each: the scheduler still balances inside the node. The calling thread adopts the same mask on its next dispatch.
Sourcepub fn run_rows(&self, rows: usize, f: &(dyn Fn(usize, usize) + Sync))
pub fn run_rows(&self, rows: usize, f: &(dyn Fn(usize, usize) + Sync))
Run f(row_start, row_end) over 0..rows, self-balancing.
One dispatch, but workers pull row-ranges from a shared cursor instead of each taking a fixed 1/n slice. On a heterogeneous CPU (Apple Silicon: 4 P-cores + 6 E-cores here) a static split makes every matvec end at the SLOWEST core’s pace while the fast ones idle at the barrier; pulling by grain lets a P-core take several chunks for each one an E-core takes, so skew collapses to a single grain. Row ranges stay disjoint and each row’s dot is computed exactly as in the serial path → bit-identical output.
Sourcepub fn run_many(&self, parts: &[(usize, &(dyn Fn(usize, usize) + Sync))])
pub fn run_many(&self, parts: &[(usize, &(dyn Fn(usize, usize) + Sync))])
Multi-matrix job: one dispatch serves SEVERAL row spaces
(roadmap §3 P0 — «одна внешняя публикация job на слой»). Parts
are laid out back-to-back in a virtual row space and pulled by
grain from one shared cursor, so QKV or gate+up cost a single
barrier instead of one each. Each part’s f(start, end) sees its
OWN row indices — per-row math and outputs are bit-identical to
separate run_rows calls.