Expand description
Turn audio into text or text into audio.
Submodules: speech (text-to-speech, binary response),
transcriptions (speech-to-text in the input language), and
translations (speech-to-English). Streaming transcriptions (stream)
and stream_format audio streaming are not supported yet.
![warn] The endpoints in this module are untested! No OpenAI-compatible provider accessible to this project (DeepSeek, Alibaba Cloud Model Studio) implements the
/audio/*endpoints, and no OpenAI API key was available for testing. If you encounter any issues, please report them on the repository.
Modules§
- speech
- Generates audio from the input text.
- transcriptions
- Transcribes audio into the input language.
- translations
- Translates audio into English.
Enums§
- Audio
Response Format - The format of the transcription/translation output, in one of these
options:
json,text,srt,verbose_json, orvtt.