Expand description
Opt-in, shared autotuning above individual kernels.
Tensor operators and fused graphs use the same controller as explicit application plans. Existing LocalTuner behavior is unchanged until enable_stack_autotune is called; only sets declaring an exact workload signature and reference participate in the new path.
Structs§
- Candidate
- Candidate
Report - Decision
- Problem
- Runtime
Environment - Supply extra tags for a deployment’s compiler flags and load/topology/power regime. Set these BEFORE constructing devices/starting inference; changing them live is unsupported.
- Stack
Policy - Stack
Tuner - No state mutex is held while running a candidate, validator, profile, or disk operation. Contenders use the declared reference instead of waiting (also avoids nested-tuning deadlocks).
- Stats
- Tolerance
- Tune
Failure - Tune
Report
Enums§
Traits§
- Trial
Runner - Benchmarks MUST use isolated state. A completed measurement includes all device work whose cost belongs to the plan. A submit-only host timestamp is not a valid measurement.
Functions§
- cache_
key - enable_
stack_ autotune - Install once, before worker threads/model loading. No implicit environment mutation or I/O.
- fields
- Canonical, length-delimited encoding prevents concatenation aliases such as (ab,c)/(a,bc).
- is_
tuning - median
- paired_
score - runtime_
environment - stack_
autotuner - try_
execute_ stack - Fallible full-stack path. Actual request execution is performed ONCE after selection; an execution error invalidates future choices but is never silently replayed on another kernel.
- workload_
digest - This is NOT a cryptographic authenticity check. Caches must be in a trusted user directory.