pub struct TranscriptionRequest {
pub file: PathBuf,
pub model: String,
pub language: Option<String>,
pub prompt: Option<String>,
pub response_format: Option<AudioResponseFormat>,
pub temperature: Option<f32>,
pub timestamp_granularities: Option<Vec<TimestampGranularity>>,
pub chunking_strategy: Option<ChunkingStrategy>,
pub include: Option<Vec<Include>>,
pub extra_body_map: Option<Map<String, Value>>,
}Expand description
Transcribes audio into the input language.
The keywords, languages, known_speaker_names, and
known_speaker_references parameters are not covered by this type yet.
Fields§
§file: PathBufThe audio file (as a path) to transcribe, in one of these formats: flac, mp3, mp4, mpeg, mpga, m4a, ogg, wav, or webm. The request must include enough format metadata for the file to be identified; an extension-bearing filename satisfies this.
model: StringID of the model to use. The options are gpt-transcribe,
gpt-4o-transcribe, gpt-4o-mini-transcribe, whisper-1 (which is
powered by the open source Whisper V2 model), and
gpt-4o-transcribe-diarize.
language: Option<String>The language of the input audio. Supplying the input language in
ISO-639-1
(e.g. en) format will improve accuracy and latency.
prompt: Option<String>An optional text to guide the model’s style or continue a previous audio segment. The prompt should match the audio language.
response_format: Option<AudioResponseFormat>The format of the output, in one of these options: json, text,
srt, verbose_json, or vtt. For gpt-4o-transcribe and
gpt-4o-mini-transcribe, the only supported format is json.
With json / verbose_json, use get_response; with text /
srt / vtt, use get_response_string.
temperature: Option<f32>The sampling temperature, between 0 and 1. Higher values like 0.8 will make the output more random, while lower values like 0.2 will make it more focused and deterministic.
timestamp_granularities: Option<Vec<TimestampGranularity>>The timestamp granularities to populate for this transcription.
response_format must be set to verbose_json to use timestamp
granularities. Either or both of these options are supported: word,
or segment.
chunking_strategy: Option<ChunkingStrategy>Controls how the audio is cut into chunks.
When set to auto, the server first normalizes loudness and then
uses voice activity detection (VAD) to choose boundaries. A
server_vad object can be provided to tweak VAD detection
parameters manually. If unset, the audio is transcribed as a single
block.
include: Option<Vec<Include>>Additional information to include in the transcription response.
logprobs will return the log probabilities of the tokens in the
response.
extra_body_map: Option<Map<String, Value>>Additional JSON properties, sent as extra multipart text fields (strings verbatim, other JSON values serialized).
Trait Implementations§
Source§impl Clone for TranscriptionRequest
impl Clone for TranscriptionRequest
Source§impl Debug for TranscriptionRequest
impl Debug for TranscriptionRequest
Source§impl Default for TranscriptionRequest
impl Default for TranscriptionRequest
Source§impl Post for TranscriptionRequest
impl Post for TranscriptionRequest
Source§impl PostNoStream for TranscriptionRequest
impl PostNoStream for TranscriptionRequest
Source§async fn get_response_string(
&self,
client: &Client,
base_url: &str,
options: &RequestOptions,
) -> Result<String, OapiError>
async fn get_response_string( &self, client: &Client, base_url: &str, options: &RequestOptions, ) -> Result<String, OapiError>
Sends a transcription POST request using multipart/form-data format, following the field layout of the official SDK.