Skip to main content

Module stack

Module stack 

Source
Expand description

Opt-in, shared autotuning above individual kernels.

Tensor operators and fused graphs use the same controller as explicit application plans. Existing LocalTuner behavior is unchanged until enable_stack_autotune is called; only sets declaring an exact workload signature and reference participate in the new path.

Structs§

Candidate
CandidateReport
Decision
Problem
RuntimeEnvironment
Supply extra tags for a deployment’s compiler flags and load/topology/power regime. Set these BEFORE constructing devices/starting inference; changing them live is unsupported.
StackPolicy
StackTuner
No state mutex is held while running a candidate, validator, profile, or disk operation. Contenders use the declared reference instead of waiting (also avoids nested-tuning deadlocks).
Stats
Tolerance
TuneFailure
TuneReport

Enums§

DecisionSource
FailureKind
Mode
Scope
Timing
Validation

Traits§

TrialRunner
Benchmarks MUST use isolated state. A completed measurement includes all device work whose cost belongs to the plan. A submit-only host timestamp is not a valid measurement.

Functions§

cache_key
enable_stack_autotune
Install once, before worker threads/model loading. No implicit environment mutation or I/O.
fields
Canonical, length-delimited encoding prevents concatenation aliases such as (ab,c)/(a,bc).
is_tuning
median
paired_score
runtime_environment
stack_autotuner
try_execute_stack
Fallible full-stack path. Actual request execution is performed ONCE after selection; an execution error invalidates future choices but is never silently replayed on another kernel.
workload_digest
This is NOT a cryptographic authenticity check. Caches must be in a trusted user directory.