Expand description
Request and response types for inference
Structs§
- ApiChat
Message - ApiChat
Request - ApiChat
Response - ApiCompletion
Request - ApiCompletion
Response - ApiFunction
- ApiFunction
Call - ApiJson
Schema - ApiResponse
Format - ApiStream
Options - ApiTool
- ApiTool
Call - ApiTool
Choice Function - Batch
Request - Batch request for processing multiple requests together
- Engine
Decode Stage Interval - A measured request-local engine decode stage in the request’s monotonic clock domain. These are explicit producer boundaries, not intervals inferred by filling gaps between other profile events.
- Engine
Token Timing Evidence - Inference
Evidence Request - Explicit request for execution evidence that is expensive or sensitive to retain.
- Inference
Execution Evidence - Evidence captured at the engine execution boundary.
- Inference
Request - Inference request
- Inference
Response - Inference response
- Scheduled
Request - Scheduled request with additional state information
- Stream
Chunk - Streaming response chunk
Enums§
- ApiFunction
Call Choice - ApiMessage
Role - ApiRequest
- ApiResponse
- ApiTool
Call Protocol - ApiTool
Choice - Engine
Decode Stage - Engine-boundary timing for one inference request.
- Request
State - Request state in the scheduler
- Structured
Output Branch - Branch established by the structured-output grammar for a complete result.
When both the final and tool languages accept the same bytes, the grammar
owner must select
Finalbefore invoking the classified response helper.
Constants§
Functions§
- api_
response_ from_ classified_ generated_ text - Preserve an authoritative grammar decision through product response shaping.
- api_
response_ from_ generated_ text - chat_
api_ may_ emit_ tool_ or_ function_ call - chat_
api_ response_ from_ generated_ text