Expand description
Transcribes audio into the input language.
Endpoint: POST /audio/transcriptions (multipart/form-data request).
Response shapes depend on response_format:
jsonandverbose_jsondeserialize into the typedTranscriptionResponseviaget_response.text,srt, andvttreturn plain text; useget_response_stringfor those.
Streaming transcriptions (stream) is not supported yet.
![warn] This module is untested! No OpenAI-compatible provider accessible to this project implements this endpoint, and no OpenAI API key was available for testing. If you encounter any issues, please report them on the repository.
Structs§
- Server
VadConfig - Manual VAD detection parameters, sent as the
server_vadchunking strategy. - Transcription
- Represents a transcription response returned by the model.
- Transcription
Language - A language detected in transcribed audio.
- Transcription
Logprob - The log probability of a token in the transcription.
- Transcription
Request - Transcribes audio into the input language.
- Transcription
Segment - A segment of the transcribed text and its corresponding details.
- Transcription
Verbose - Represents a verbose json transcription response.
- Transcription
Verbose Usage - Usage statistics for models billed by audio input duration.
- Transcription
Word - An extracted word and its corresponding timestamp.
- Usage
Tokens Input Token Details - Details about the input tokens billed for a request.
Enums§
- Chunking
Strategy - Controls how the audio is cut into chunks.
- Include
- Additional information to include in the transcription response.
- Timestamp
Granularity - The timestamp granularities to populate for a transcription.
- Transcription
Response - The typed transcription response:
jsonyields a plainTranscription,verbose_jsonyields aTranscriptionVerbose. - Transcription
Usage - Usage statistics for a transcription request. Billed either by token
usage or by audio input duration, discriminated by
type.