Expand description
OpenAI Audio wire models.
The local Speech page documents its JSON request but not the SSE event payloads. The Transcription and Translation pages document responses and multipart examples but omit a complete body-parameter table; request-only fields not visible there are retained from the v2 public protocol model.
Structs§
- Audio
Duration Usage - Audio
Input Token Details - Audio
Token Usage - Custom
Voice - Server
VadConfig - Speech
Event - Speech
Request - Transcription
- Transcription
Diarized - Transcription
Diarized Segment - Transcription
Language - Transcription
Logprob - Transcription
Request - Transcription
Segment - Transcription
Text Delta Event - Transcription
Text Done Event - Transcription
Text Segment Event - Transcription
Verbose - Transcription
Word - Translation
- Translation
Request - Translation
Verbose - Unknown
Transcription Stream Event
Enums§
- Audio
Chunking Auto - Audio
Chunking Strategy - Audio
Duration Usage Type - Audio
Token Usage Type - Audio
Usage - Known
Speech Response Format - Known
Speech Stream Format - Known
Timestamp Granularity - Known
Transcription Include - Known
Transcription Response Format - Known
Translation Response Format - Server
VadType - Speech
Response Format - Speech
Stream Event - Speech SSE payload observed by compatible backends. The OpenAI snapshot
confirms SSE transport but does not name its events;
type,delta, andaudioare session-derived aliases and every other field remains opaque. - Speech
Stream Format - Speech
Voice - Timestamp
Granularity - Transcription
Include - Transcription
Response - Transcription
Response Format - Transcription
Segment Type - Transcription
Stream Event - Transcription
Text Delta Type - Transcription
Text Done Type - Transcription
Text Segment Type - Translation
Response - Translation
Response Format
Type Aliases§
- Speech
Response - Binary bytes returned when speech uses
stream_format: "audio".