Expand description
Request and response types for inference
Structs§
- ApiChat
Message - ApiChat
Request - ApiChat
Response - ApiCompletion
Request - ApiCompletion
Response - ApiFunction
- ApiFunction
Call - ApiJson
Schema - ApiResponse
Format - ApiStream
Options - ApiTool
- ApiTool
Call - ApiTool
Choice Function - Batch
Request - Batch request for processing multiple requests together
- Engine
Decode Stage Interval - A measured request-local engine decode stage in the request’s monotonic clock domain. These are explicit producer boundaries, not intervals inferred by filling gaps between other profile events.
- Engine
Token Timing Evidence - Inference
Evidence Request - Explicit request for execution evidence that is expensive or sensitive to retain.
- Inference
Execution Evidence - Evidence captured at the engine execution boundary.
- Inference
Request - Inference request
- Inference
Response - Inference response
- Scheduled
Request - Scheduled request with additional state information
- Stream
Chunk - Streaming response chunk
Enums§
- ApiFunction
Call Choice - ApiMessage
Role - ApiRequest
- ApiResponse
- ApiTool
Call Protocol - ApiTool
Choice - Engine
Decode Stage - Engine-boundary timing for one inference request.
- Request
State - Request state in the scheduler