Expand description
Model execution interface with clear prefill/decode separation
This module provides the ModelExecutor trait that replaces the “fat” Model interface, focusing purely on tensor operations without tokenization or sampling.
Structs§
- Decode
Input - Input for decode phase (generating one token at a time)
- Decode
Output - Output from decode phase
- Executor
Admission Epochs - Stable scheduler-facing projection of one plan-runtime capacity domain.
- Executor
Attention Config - Runtime attention configuration for model executor
- Executor
Capabilities - Executor capabilities and configuration
- Executor
Capacity Wait Registration - Type-erased, single-use registration for one plan-runtime capacity wait.
- Executor
Config - Executor configuration
- Executor
Execution Capacity Deferral - Executor
Execution Capacity Evidence - Exactly one typed owner for an execution-capacity deferral.
- Executor
Execution Capacity Preemption - Request-scoped authority selected for a capacity-pressure preemption.
- Executor
Execution Capacity Preemption Receipt - Proof that the executor retired one exact request authority to a terminal state. Source-generation advancement remains independently verified by the engine before the scheduler may resume another frontier.
- Executor
Execution Maintenance Mutation - One exact physical pool mutation completed while an executor was trying to make an unsubmitted frontier runnable.
- Executor
Execution Maintenance Progress - Typed proof that a bounded executor call committed relevant backing growth before yielding the same logical frontier back to the scheduler.
- Executor
Execution Maintenance Retry - Scheduler-consumable proof binding physical maintenance to the exact logical frontiers whose unsubmitted execution attempt observed it.
- Executor
Memory Config - Memory configuration for executor
- Executor
Memory Usage - Executor memory usage
- Executor
Metrics - Executor performance metrics
- Executor
Prefill Admission - Borrowed, already-tokenized input used to probe plan-runtime prefill admission before the request can enter a device submission batch.
- Executor
Prefill Admission Receipt - Scheduler-visible proof that an executor retained request and sequence authority for future scheduler-owned prefill chunks.
- Executor
Prefill Completion - Capacity-aware result for one planned prefill frontier.
- Executor
Prefill Maintenance Deferral - Non-authoritative projection of plan-runtime maintenance work.
- Executor
Request State Deferral - Scheduler-visible proof that execution is temporarily blocked by a Request-lifetime state hazard rather than by physical capacity.
- Executor
Sequence Completion - Product-authoritative evidence for a successfully completed sequence.
- Executor
Status - Executor status information
- Greedy
Repetition Penalty - Sparse repetition-penalty metadata for model-side greedy argmax.
- KvSlot
Allocation - Per-cache outcome from a KV slot reservation attempt.
- KvSlot
Capacity Snapshot - Point-in-time model-owned paged-KV capacity snapshot.
- KvSlot
Request - One model-owned KV slot reservation request.
- KvSlot
Reservation - Model-owned paged-KV reservation evidence.
- Memory
Requirements - Memory requirements for model execution
- Optimization
Config - Optimization configuration
- Plan
Runtime Decode Input - Tensor-free decode input for an executor with plan-runtime resource authority.
- Plan
Runtime Decode Output - Tensor-free decode output from a plan-runtime executor.
- Plan
Runtime Prefill Authority - Exact executor-owned authority advanced by one prefill completion.
- Plan
Runtime Prefill Completion - Capacity-aware result for one tensor-free plan-runtime prefill frontier.
- Plan
Runtime Prefill Input - Tensor-free prefill input for an executor with plan-runtime resource authority.
- Plan
Runtime Prefill Output - Tensor-free prefill output bound to one request and committed KV extent.
- Plan
Runtime Resource Snapshot - Point-in-time memory evidence emitted by the shared plan runtime.
- Prefill
Chunk - Exact, validated prompt progress assigned to one prefill invocation.
- Prefill
Input - Input for prefill phase (processing the initial prompt)
- Prefill
Output - Output from prefill phase
- Speculative
Decode Output - Output from speculative decoding
- Token
Selection Mask - Token-validity mask for model-side greedy argmax.
- Unified
Batch - A mixed-batch forward request: any combination of in-progress prefill
chunks and decode steps. See
UnifiedBatchItemfor the per-item semantics. The producer (engine) groups all sequences active in this iter into a single batch; the consumer (model) runs one forward and returns per-item logits (only for items withis_final_chunk = true, in the order they appear initems). - Unified
Batch Item - One sequence’s contribution to a unified mixed-batch forward.
Enums§
- Attention
Type - Attention mechanism types
- Execution
Resource Authority - Declares the authoritative owner of request-lifetime accelerator resources.
- Executor
Batch Decode Outcome - Capacity-aware batch decode result.
- Executor
Batch Prefill Outcome - Transactional result of attempting one physical prefill batch.
- Executor
Execution Capacity Evidence Owner - Scheduler-visible proof that an execution attempt was not submitted and must not be retried until one of its exact capacity sources changes.
- Executor
Execution Capacity Preemption Authority - Executor
Execution Capacity Stage - Pre-submit runtime stage that could not acquire its exact dynamic capacity.
- Executor
Execution Deferral - A pre-submit execution frontier can be blocked by independently evolving sources. Capacity waits participate in scheduler pressure/yield policy; Request-state waits never do and resume only from their exact hazard coordinator.
- Executor
Prefill Admission Decision - Typed result of probing plan-runtime prefill capacity.
- Executor
Prefill Maintenance Blocker - Scheduler-visible reason that a prefill needs plan-runtime backing maintenance. These values are projections only and cannot allocate memory.
- Executor
Prefill Maintenance Outcome - Result of one bounded plan-runtime backing maintenance attempt.
- Executor
Prefill Maintenance Stage - Stage that must advance before a plan-runtime prefill can be admitted.
- Executor
Prefill Outcome - Executor
Request Origin - Typed origin for an executor-owned request lifecycle.
- Executor
Sampling Output - Typed product output returned by a plan-runtime execution wave.
- Executor
State - Executor state
- Executor
Type - Supported executor types
- Logits
Return Policy - Plan
Runtime Batch Decode Outcome - Capacity-aware, tensor-free batch decode result for a plan runtime.
- Plan
Runtime Batch Prefill Outcome - Plan
Runtime Prefill Outcome - Plan
Runtime Prefill Product - Product state emitted by one completed plan-runtime prefill chunk.
Traits§
- Batch
Model Executor - Batch model executor for processing multiple requests efficiently
- Executor
Registry - Executor registry for managing multiple executors
- Model
Executor - Core model executor trait focusing on tensor operations
- Model
Executor Factory - Model executor factory
- Speculative
Executor - Speculative execution support