Skip to main content

Module audio

Module audio 

Source
Expand description

Turn audio into text or text into audio.

Submodules: speech (text-to-speech, binary response), transcriptions (speech-to-text in the input language), and translations (speech-to-English). Streaming transcriptions (stream) and stream_format audio streaming are not supported yet.

![warn] The endpoints in this module are untested! No OpenAI-compatible provider accessible to this project (DeepSeek, Alibaba Cloud Model Studio) implements the /audio/* endpoints, and no OpenAI API key was available for testing. If you encounter any issues, please report them on the repository.

Modules§

speech
Generates audio from the input text.
transcriptions
Transcribes audio into the input language.
translations
Translates audio into English.

Enums§

AudioResponseFormat
The format of the transcription/translation output, in one of these options: json, text, srt, verbose_json, or vtt.