Skip to main content

Module transcriptions

Module transcriptions 

Source
Expand description

Transcribes audio into the input language.

Endpoint: POST /audio/transcriptions (multipart/form-data request).

Response shapes depend on response_format:

  • json and verbose_json deserialize into the typed TranscriptionResponse via get_response.
  • text, srt, and vtt return plain text; use get_response_string for those.

Streaming transcriptions (stream) is not supported yet.

![warn] This module is untested! No OpenAI-compatible provider accessible to this project implements this endpoint, and no OpenAI API key was available for testing. If you encounter any issues, please report them on the repository.

Structs§

ServerVadConfig
Manual VAD detection parameters, sent as the server_vad chunking strategy.
Transcription
Represents a transcription response returned by the model.
TranscriptionLanguage
A language detected in transcribed audio.
TranscriptionLogprob
The log probability of a token in the transcription.
TranscriptionRequest
Transcribes audio into the input language.
TranscriptionSegment
A segment of the transcribed text and its corresponding details.
TranscriptionVerbose
Represents a verbose json transcription response.
TranscriptionVerboseUsage
Usage statistics for models billed by audio input duration.
TranscriptionWord
An extracted word and its corresponding timestamp.
UsageTokensInputTokenDetails
Details about the input tokens billed for a request.

Enums§

ChunkingStrategy
Controls how the audio is cut into chunks.
Include
Additional information to include in the transcription response.
TimestampGranularity
The timestamp granularities to populate for a transcription.
TranscriptionResponse
The typed transcription response: json yields a plain Transcription, verbose_json yields a TranscriptionVerbose.
TranscriptionUsage
Usage statistics for a transcription request. Billed either by token usage or by audio input duration, discriminated by type.