Expand description
§onnx-runtime-memory
Liveness-based activation memory planning for the ORT 2.0 runtime.
§Why this crate exists
The user’s north-star goal is to run any-size model even when VRAM/RAM is insufficient, as long as the whole system has enough storage. Zero-copy weight streaming already removes weight copies from the budget. The next big lever is activation memory.
Today the executor allocates one buffer per graph value for its whole
lifetime, so peak activation memory is SUM(every intermediate tensor)
rather than the concurrent peak. Most intermediates are dead long before
the run ends. A liveness-based planner shares one physical buffer among
values whose lifetimes do not overlap, cutting peak activation memory from
O(N nodes) to O(max concurrent live set) — often a multiple-× reduction
and the key to fitting big models.
This crate is the pure, deterministic planning algorithm, deliberately
decoupled from the risky executor surgery (which lands later). It depends
only on onnx_runtime_ir, contains no unsafe, and is free of PyO3 / EP /
session dependencies so it is trivially testable in isolation.
§What it computes
Given a Graph, a ViewMap of zero-copy view
aliases, and a size oracle (Fn(ValueId) -> Option<usize>), the planner
produces an ActivationPlan:
assignments: ValueId -> SlotId— which reusable slot backs each value.slots: Vec<SlotInfo>— each slot’s byte capacity.peak_bytes— the arena size the executor must allocate (sum of slot capacities) — the shared, concurrent-peak footprint.naive_bytes+savings_ratio— the one-buffer-per-value baseline and the proven reduction.
The three-step algorithm is: (1) compute each buffer owner’s live
interval [def, use_end] in topological order, folding view consumers and
graph-output liveness into the root owner; (2) size every owner via the
oracle, returning PlanStatus::Deferred if any size is symbolic; (3)
greedily walk nodes, allocating each output’s slot (best-fit reuse of a
retired slot, else a new one) and retiring inputs after the node so a
node’s own inputs and any graph output are never clobbered.
§Static vs. dynamic shapes
The same algorithm serves build-time and run-time planning through
the size oracle. plan_activations_static plans from fully-static shapes;
any symbolic-shaped activation makes the whole plan PlanStatus::Deferred
so the executor can re-plan once shapes resolve for a run. A run-time caller
passes its own oracle backed by resolved shapes.
§Zero-copy view aliasing
The executor treats layout/movement-op outputs (Slice, Reshape,
Transpose, …) as zero-copy views that own no buffer and alias a source
buffer, pinning the source so it outlives every alias. The planner mirrors
this: a view gets no slot; instead it extends its source’s live interval
(transitively — a view of a view folds to the root buffer owner). Correct
view-liveness folding is exactly what stops a reused buffer from clobbering a
still-live alias. The caller supplies the view -> source edges via
ViewMap; op names are never hardcoded here.
§Intended executor integration (deferred follow-up — out of scope now)
The follow-up PR wires this into onnx-runtime-session’s executor:
- Build the
ViewMapfrom the executor’s existing view plan (theviews/pinnedmachinery inexecutor.rs), mapping each view value to its source (root) buffer owner. - Call the planner with a size oracle backed by the run’s resolved
shapes (build-time static shapes where available; a per-run oracle
otherwise). On
PlanStatus::Deferred, re-plan once shapes resolve. - Allocate
peak_bytesas one arena (ornum_slotsDeviceBuffers), then map eachSlotIdto an offset/allocation. - Hand each value a
TensorMutwindow into its assigned slot instead of the current per-valuebuffers: HashMap<ValueId, DeviceBuffer>.
Two concerns are explicitly out of scope for both this crate and the current planning contract, and must be handled by the integration PR:
- In-place ops — an op that may safely overwrite an input (e.g. an elementwise unary) can share its input’s slot for its output. Detecting this requires per-op semantics the planner does not model; until then the planner conservatively gives each output its own (reused) slot.
- Fragmentation —
peak_bytesis the sum of slot capacities. Packing slots into a single arena with alignment/offset assignment (and any resulting internal fragmentation) is the executor’s responsibility.
§Correctness invariants (enforced by validate)
- No two values with overlapping live intervals share a slot.
- Every graph output has a slot that is never reused after its def.
- A pinned source outlives every view aliasing it (fold correctness).
Structs§
- Activation
Plan - A computed activation memory plan: which slot backs each value, and how big the whole activation arena must be.
- Interval
- The half-open-agnostic live interval of a buffer owner, in units of topological node index.
- Liveness
- Result of liveness analysis over a graph.
- Plan
Options - Options controlling what the activation planner considers part of the arena.
- SlotId
- Identifier of a reusable activation slot in an
ActivationPlan. - Slot
Info - A single reusable activation slot: an arena region the executor allocates once and hands, in turn, to every value assigned to it (their lifetimes are guaranteed disjoint).
- ViewMap
- A set of zero-copy
view → sourcealiasing relationships.
Enums§
- Plan
Error - A failure that prevents the planner from producing any plan.
- Plan
Status - The outcome of a planning attempt.
- Validate
Error - A violated correctness invariant found by
crate::validate.
Functions§
- compute_
liveness - Compute the live interval of every activation buffer owner.
- plan_
activations - Compute a liveness-based activation plan with slot reuse.
- plan_
activations_ static - Convenience: plan directly from fully-static shapes.
- static_
size - Byte size of a value from its static shape, or
Noneif any dimension is symbolic (unknown until run time) or the element count overflowsusize. - static_
size_ oracle - A size oracle closure that sizes values from their fully-static shapes.
- validate
- Validate a plan against the graph, using
size_oracleto check slot capacities. Enforces: - validate_
static - Convenience: validate a plan built from fully-static shapes.