openai-interface 0.14.0

A low-level Rust interface for the OpenAI API
Documentation

OpenAI Interface

A low-level Rust interface for interacting with OpenAI's API. Both streaming and non-streaming APIs are supported.

Currently, chat completions (create / retrieve / update / delete), completions, models, embeddings, moderations, file management (upload / list / retrieve / delete / download content), images (generate / edit / variation), audio (speech / transcriptions / translations), the Responses API (create / retrieve / delete / input items / cancel), batches, uploads, fine-tuning jobs, vector stores, containers, conversations, evals, and realtime session creation are supported. See the support matrix below for details.

Repository:

Codeberg: Codeberg Repo
GitCode: GitCode Repo

You are welcome to contribute to this project through any of the links above.

Features

  • Chat Completions: Full support for OpenAI's chat completion and completion API, including both streaming and non-streaming responses, and multimodal user messages (text / image / audio / file content parts).
  • Models: List, retrieve and delete models.
  • Embeddings: Create embedding vectors from text input.
  • Moderations: Classify whether text and/or image input is potentially harmful (untested).
  • Images: Generate, edit, and create variations of images (untested).
  • Audio: Text-to-speech, transcription, and translation endpoints (untested).
  • Files: Support for the OpenAI file API (create / list / retrieve / delete / download content).
  • Responses API: Create, retrieve and delete responses, list their input items, and cancel background responses, including streaming events.
  • Batches: Create, retrieve, list and cancel async batch processing jobs.
  • Uploads: Multi-part upload sessions for large files (create / add parts / complete / cancel).
  • Fine-tuning: Create, retrieve, list and cancel fine-tuning jobs, list their events and checkpoints, and restore fine-tuned models.
  • Vector Stores: Manage vector stores and their files (including file batches and semantic search) for the file_search tool.
  • Containers: Manage containers and their files for the Code Interpreter tool.
  • Conversations: Manage the stateful conversation layer of the Responses API and its items.
  • Evals: Manage evals, their runs and the runs' output items.
  • Realtime: Create ephemeral Realtime and transcription session tokens (the WebSocket transport itself is not implemented).
  • Streaming and Non-streaming: Support for both streaming and non-streaming responses, with a delta accumulator (chat::create::accumulator::ChatCompletionAccumulator) that assembles text and tool calls (joined by tool-call index) from the stream, the same way the official SDKs do.
  • Reasoning Effort: The OpenAI-compatible reasoning_effort parameter is supported out of the box for reasoning models.
  • Configurable HTTP Client: Every request method takes a reqwest::Client, so proxies, timeouts and connection pooling are under your control.
  • Strong Typing: Complete type definitions for all API requests and responses, utilizing Rust's powerful type system.
  • Error Handling: Comprehensive error handling with detailed error types defined in the [errors] module. Failed requests carry the API's error message, type and code.
  • Async/Await: Built with async/await support.
  • Musl Support: Designed to work with musl libc out-of-the-box.
  • Multiple Provider Support: Expected to work with OpenAI, DeepSeek, Qwen, vLLM, Z.ai / 智谱 GLM, and other compatible API providers. Provider-specific fields are opt-in via cargo features (see below).

Installation

[!WARNING] Versions prior to 0.3.0 have serious issues with SSE streaming responses processing: instead of a single chunk, multiple chunks may be returned in each iteration of the response stream.

Add this to your Cargo.toml:

[dependencies]
openai-interface = { version = "0.14", features = ["deepseek", "qwen"] }

Cargo Features

Fields that are proprietary to a single provider are opt-in via cargo features. Cross-vendor de-facto standards — such as reasoning_content (streamed by DeepSeek, Qwen3, ollama, vLLM and OpenRouter alike) — are always available:

  • reasoning (default): cross-vendor reasoning fields — reasoning_content on assistant messages (request and response), streamed deltas, and logprobs, plus its accumulation in ChatCompletionAccumulator.
  • deepseek: Enables DeepSeek's proprietary fields — the Beta chat prefix completion fields (prefix, and reasoning_content as the prefix-completion CoT input), the thinking and user_id request parameters, and the prompt_cache_hit_tokens / prompt_cache_miss_tokens usage statistics. Implies reasoning. See api-docs.deepseek.com.
  • qwen: Enables Qwen's proprietary request parameters (enable_thinking, thinking_budget, top_k) as direct fields of the chat request body. Implies reasoning. See the Qwen OpenAI-compatible Chat API docs.
  • vllm: Enables vLLM's proprietary fields, collected in the openai_interface::vllm module. On the request side: the extra sampling parameters (min_p, repetition_penalty, stop_token_ids, prompt_logprobs, bad_words, allowed_token_ids, ...), the chat-template controls (chat_template, chat_template_kwargs, add_generation_prompt, continue_final_message, ...), structured_outputs (vLLM's successor to the deprecated guided_json / guided_regex / guided_choice / guided_grammar keys), and the KV-transfer and scheduling parameters (kv_transfer_params, priority, cache_salt, stream_interval, ...). On the response side: stop_reason, token_ids and routed_experts per choice; prompt_logprobs, prompt_token_ids, prompt_text, kv_transfer_params and ec_transfer_params on the completion and the streamed chunks; and root / parent / max_model_len on model objects. Implies reasoning. See vLLM's OpenAI-compatible server docs.
  • zai: Enables Z.ai / 智谱 GLM (BigModel) proprietary fields, collected in the openai_interface::zai module. Generic controls GLM spells its own way (do_sample, tool_stream) are zai-gated fields of the chat request body; the keys GLM shares with another provider are unified rather than duplicated (thinking and user_id with deepseek, request_id with vllm). GLM's platform-ecosystem extensions live in zai: the watermark_enabled flag (zai::PlatformParams, flattened in through RequestBody::zai_platform), the retrieval and web_search tool types, and the web_search results GLM returns. reasoning_effort needs no gate — it is already an ungated field whose enum covers every GLM value. Implies reasoning. See the GLM chat-completions reference.
  • azure: Deprecated no-op. Streaming delta.annotations and delta.audio are now always available (the non-streaming message fields were never gated). The empty feature remains defined so existing manifests keep compiling.
  • ferritls: Unrelated to request fields — adds the pure-Rust ferritls-rustls TLS crypto backend and rest::install_crypto_provider, the helper that installs it. Off by default, so the crate never dictates your crypto backend. See Choosing the TLS Crypto Provider.

Usage

Chat Completion

This crate provides methods for both streaming and non-streaming chat completions. The following examples demonstrate how to use these features.

Non-streaming Chat Completion

use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::chat::create::response::no_streaming::ChatCompletion;
use openai_interface::rest::{
    default_client, install_crypto_provider, post::PostNoStream, RequestOptions,
};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("DEEPSEEK_API_KEY")?;
    // Needs the `ferritls` feature; skip this if you install your own
    // provider — see "Choosing the TLS Crypto Provider" below.
    install_crypto_provider().ok();
    let client = default_client();

    let request = RequestBody {
        messages: vec![
            Message::system("You are a helpful assistant."),
            Message::user("Hello, how are you?"),
        ],
        model: "deepseek-v4-flash".to_string(),
        stream: Some(false),
        ..Default::default()
    };

    // Send the request. The base URL is everything before the endpoint
    // path — for OpenAI and most compatible gateways it must include the
    // `/v1` prefix (e.g. `https://api.openai.com/v1`); DeepSeek is the
    // exception and uses the bare host.
    let options = RequestOptions::bearer(api_key);
    let chat_completion: ChatCompletion = request
        .get_response(&client, "https://api.deepseek.com", &options)
        .await?;
    let text = chat_completion.choices[0]
        .message
        .content
        .as_deref()
        .unwrap();
    println!("{:?}", text);
    Ok(())
}

Streaming Chat Completion

This example demonstrates how to handle streaming responses from the API. get_stream_response deserializes every server-sent event and stops automatically at the data: [DONE] sentinel. The ChatCompletionAccumulator assembles the fragments into a complete message — concatenating text and tool-call arguments (by tool-call index) exactly like the official SDKs.

use openai_interface::chat::create::accumulator::ChatCompletionAccumulator;
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::rest::{default_client, install_crypto_provider, post::PostStream, RequestOptions};
use futures_util::StreamExt;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("DEEPSEEK_API_KEY")?;
    // Needs the `ferritls` feature; skip this if you install your own
    // provider — see "Choosing the TLS Crypto Provider" below.
    install_crypto_provider().ok();
    let client = default_client();

    let request = RequestBody {
        messages: vec![
            Message::system("You are a helpful assistant."),
            Message::user("Who are you?"),
        ],
        model: "deepseek-v4-flash".to_string(),
        stream: Some(true),
        ..Default::default()
    };

    // Send the request. The base URL is everything before the endpoint
    // path — for OpenAI and most compatible gateways it must include the
    // `/v1` prefix (e.g. `https://api.openai.com/v1`); DeepSeek is the
    // exception and uses the bare host.
    let options = RequestOptions::bearer(api_key);
    let mut response_stream = request
        .get_stream_response(&client, "https://api.deepseek.com", &options)
        .await?;

    let mut accumulator = ChatCompletionAccumulator::new();
    while let Some(chunk_result) = response_stream.next().await {
        let chunk = chunk_result?;
        if let Some(content) = chunk.choices.first().and_then(|c| c.delta.content.as_deref()) {
            println!("message chunk: {}", content);
        }
        accumulator.push(&chunk);
        // For reasoning models such as `deepseek-reasoner`, the chain of
        // thought arrives in `choice.delta.reasoning_content`.
    }

    let message = accumulator.into_message();
    println!("complete message: {:?}", message.content);
    println!("tool calls: {:?}", message.tool_calls);
    Ok(())
}

If a provider occasionally emits chunks your code cannot deserialize, wrap the stream with openai_interface::rest::skip_deserialization_errors to drop those items instead of abandoning the stream at the first one.

Configuring the HTTP Client

Every request method takes the client as its first argument. Pass a custom client to use a proxy or a different timeout:

let client = reqwest::Client::builder()
    .proxy(reqwest::Proxy::http("http://127.0.0.1:10808")?)
    .timeout(std::time::Duration::from_secs(60))
    .build()?;

rest::default_client sets a 60s connect timeout and a 300s read timeout. The read timeout applies to each read and resets after every successful read, so a streaming response is never cut off mid-stream — but a backend that stays completely silent for over 300s (e.g. a long reasoning phase without streamed reasoning content) will be dropped; build a custom client with a larger read_timeout if you need to accommodate that.

Authenticating with Other Providers

Every request method takes a RequestOptions value, which carries the authentication scheme plus any extra headers. [RequestOptions::bearer] reproduces the classic OpenAI Authorization: Bearer behavior; Azure and Anthropic authenticate with other headers, which you can supply per request:

use openai_interface::rest::RequestOptions;

// Azure OpenAI: credentials travel in the `api-key` header.
let azure = RequestOptions::new().with_header("api-key", "azure-key")?;

// Anthropic: `x-api-key` plus a version header.
let anthropic = RequestOptions::new()
    .with_header("x-api-key", "anthropic-key")?
    .with_header("anthropic-version", "2023-06-01")?;

// OpenAI: extra headers can be layered on top of bearer auth.
let openai = RequestOptions::bearer("sk-...")
    .with_header("OpenAI-Organization", "org-...")?;

Choosing the TLS Crypto Provider

This crate depends on reqwest with its rustls-no-provider feature, so the rustls stack is compiled without a crypto backend. That keeps the build pure Rust (no C or asm toolchain, which is what makes the musl target work out-of-the-box) and leaves the backend choice to the application.

The consequence is that exactly one rustls CryptoProvider must be installed as the process default before the first reqwest::Client is built — including the client returned by default_client(). Without one, reqwest panics at client construction time.

This crate never installs a provider for you: neither default_client() nor any request method touches that global state, so the application stays in control. If you would rather not decide, the optional ferritls feature adds the pure-Rust ferritls-rustls backend plus a helper that installs it:

[dependencies]
openai-interface = { version = "0.10", features = ["ferritls"] }
// Requires the `ferritls` feature; call it once, before the first client.
openai_interface::rest::install_crypto_provider()
    .expect("a rustls crypto provider was already installed");

Without that feature, ferritls-rustls is not in your dependency tree at all and there is nothing to call — install a provider yourself instead. First install wins: whichever provider is installed when the first client is built is the one the whole process uses, so do it before any request.

// In the application crate, with `rustls = "0.23"` (feature `ring` or
// `aws-lc-rs`) as one of its own dependencies:
rustls::crypto::ring::default_provider()
    .install_default()
    .expect("a rustls crypto provider was already installed");

Note that reqwest's features are additive, so this requirement disappears altogether when your own project depends on reqwest with a crypto backend compiled in:

[dependencies]
reqwest = "0.13"          # default features: `default-tls` -> `rustls`
openai-interface = "0.10" # no provider of its own

default-tls (or rustls directly) makes reqwest fall back to the aws-lc-rs provider it ships with, and native-tls routes TLS through the system stack so the rustls path is never taken — either way you do not have to install anything, and you never call install_crypto_provider. The catch is that the backend is then decided by feature unification instead of by you, so an unrelated dependency change can move it. If you want the choice pinned, enable ferritls or install a provider yourself.

Custom Request Parameters

For provider-specific parameters, prefer enabling the matching cargo feature (deepseek, qwen, vllm or zai) so the fields are available as typed members of the request structs.

If you need a field that is not covered by the typed structs, you can inject arbitrary JSON properties. Every top-level request body carries an extra_body_map field (Option<serde_json::Map<String, serde_json::Value>>, flattened into the request body; on multipart endpoints the entries are sent as extra form fields):

use openai_interface::chat::create::request::{Message, RequestBody};
use serde_json::json;

let request = RequestBody {
    messages: vec![Message::user("Hello")],
    model: "gpt-4.1".to_string(),
    stream: Some(false),
    extra_body_map: Some(
        serde_json::from_value(json!({ "some_vendor_field": 42 })).unwrap(),
    ),
    ..Default::default()
};

Serving Against vLLM

With the vllm feature, the parameters vLLM accepts beyond the OpenAI standard are typed members of the request body instead of extra_body_map entries. They are split in two because the two text-generation endpoints do not accept the same set: vllm_sampling (decoding knobs, valid on both /v1/chat/completions and /v1/completions) and vllm_chat (chat-template rendering, structured output, KV transfer and scheduling; chat only). Both are flattened, so the JSON is exactly what the official client's extra_body would produce.

use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::vllm::{ChatParams, SamplingParams, StructuredOutputsParams};

let request = RequestBody {
    messages: vec![Message::user("Classify this sentiment: vLLM is wonderful!")],
    model: "Qwen/Qwen3-8B".to_string(),
    stream: Some(false),
    vllm_sampling: Some(SamplingParams {
        min_p: Some(0.1),
        repetition_penalty: Some(1.05),
        ..Default::default()
    }),
    vllm_chat: Some(ChatParams {
        structured_outputs: Some(StructuredOutputsParams {
            choice: Some(vec!["positive".to_string(), "negative".to_string()]),
            ..Default::default()
        }),
        // Template-defined switches, e.g. turning Qwen3's thinking off.
        chat_template_kwargs: Some(
            serde_json::from_value(serde_json::json!({ "enable_thinking": false })).unwrap(),
        ),
        ..Default::default()
    }),
    ..Default::default()
};

On the response side, vLLM's extras are plain fields of the same structs: stop_reason and token_ids per choice, prompt_logprobs, prompt_token_ids, prompt_text and kv_transfer_params on the completion and its chunks, and root / parent / max_model_len on model objects.

One difference needs calling out: vLLM streams and returns the chain of thought under reasoning, not the cross-vendor reasoning_content (it accepts the latter on input, but always emits the former). So against a vLLM backend read message.reasoning — or ChatCompletionAccumulator::reasoning() for a stream — and map it back onto reasoning_content when you feed it into a follow-up request.

Serving Against Z.ai / 智谱 GLM

With the zai feature, the parameters GLM accepts beyond the OpenAI standard are typed members of the request body. GLM is served from https://open.bigmodel.cn/api/paas/v4 (note the /api/paas/v4 suffix in place of /v1), or https://api.z.ai/api/paas/v4 for the international Z.ai brand. Generic controls (do_sample, tool_stream) are plain fields; GLM's platform-ecosystem extensions are grouped in the zai module, with zai_platform flattened into the body:

use openai_interface::chat::create::request::{Message, RequestBody, RequestTool, Thinking, ThinkingType};
use openai_interface::zai::{PlatformParams, SearchEngine, WebSearchTool};

let request = RequestBody {
    messages: vec![Message::user("最近有什么关于 Rust 的新闻?")],
    model: "glm-4.6".to_string(),
    stream: Some(false),
    do_sample: Some(true),
    thinking: Some(Thinking {
        type_: ThinkingType::Enabled,
        clear_thinking: Some(true),
    }),
    zai_platform: Some(PlatformParams {
        watermark_enabled: Some(true),
    }),
    tools: Some(vec![RequestTool::WebSearch {
        web_search: WebSearchTool {
            search_engine: Some(SearchEngine::SearchProJina),
            enable: Some(true),
            ..Default::default()
        },
    }]),
    ..Default::default()
};

GLM shares the thinking and user_id keys with DeepSeek and request_id with vLLM, so those are single fields gated on the union of features rather than per-provider copies — enabling zai alongside deepseek or vllm never emits a key twice. On the response side GLM adds a top-level request_id and a web_search results array (zai::WebSearchResult); the chain of thought arrives as the cross-vendor reasoning_content.

Modules

  • [chat]: Contains all chat completion related structs, enums, and methods.
  • [completions]: Contains all completion related structs, enums, and methods. Note that this API is getting deprecated in favour of chat and is only available for out-dated LLM models.
  • [models]: List, retrieve and delete models.
  • [embeddings]: Create embedding vectors from text input.
  • [moderations]: Classify whether text input is potentially harmful.
  • [images]: Generate, edit, and create variations of images.
  • [audio]: Turn audio into text (transcriptions / translations) or text into audio (speech).
  • [files]: Providing the capacity to upload and manage files.
  • [rest]: Providing all REST related traits and methods, plus default_client and shared status/error handling.
  • [errors]: Defines error types used throughout the crate.
  • [pagination]: Shared cursor-pagination types (Page, PaginationQuery) used by the list endpoints.
  • [vllm] (with the vllm feature): vLLM's proprietary request and response fields.
  • [zai] (with the zai feature): Z.ai / 智谱 GLM's proprietary request and response fields.

API Support Matrix

All newly added modules are untested against a live OpenAI API (no API key was available); they follow the official documentation and should work with OpenAI-compatible providers that implement the same endpoints. Please report any issues on the Codeberg issue tracker.

API group Endpoints Status
Chat Completions create / retrieve / update / delete tested
Completions (legacy) create tested
Models list / retrieve / delete tested
Embeddings create tested
Moderations create untested
Images generate / edit / variation untested
Audio speech / transcriptions / translations untested
Files create / list / retrieve / delete / content tested
Responses create / retrieve / delete / input items / cancel partially tested
Batches create / retrieve / list / cancel untested
Uploads create / add part / complete / cancel untested
Fine-tuning jobs create / retrieve / list / cancel, events, checkpoints, model restore untested
Vector Stores create / retrieve / update / delete / list, files CRUD + content, search, file batches untested
Containers create / retrieve / delete / list, files CRUD + content untested
Conversations create / retrieve / update / delete / list, items CRUD + list untested
Evals create / retrieve / update / delete / list, runs CRUD + cancel, output items untested
Realtime sessions / transcription sessions (HTTP only; WebSocket not implemented) untested

Not implemented: the Realtime WebSocket transport, the evals alpha permissions endpoints (/fine_tuning/alpha/permissions), and streaming variants of the images and audio transcription endpoints.

Error Handling

All errors are converted into [errors::OapiError]. On a failed request the response body is parsed into [errors::ApiError], which carries the API's error message, type, code and the HTTP status, so failures can be diagnosed without re-sending the request.

Musl Build

This crate is designed to work with musl libc, making it suitable for lightweight deployments in containerized environments. TLS is provided by rustls with a pure-Rust crypto backend, so OpenSSL does not need to be built from source (see "Choosing the TLS Crypto Provider" above for how the backend is selected at runtime).

To build for musl:

rustup target add x86_64-unknown-linux-musl
cargo build --target x86_64-unknown-linux-musl

Supported Providers

This crate aims to support standard OpenAI-compatible API endpoints. Unfortunately, OpenAI aggressively restricts the access from the People's Republic of China. As a result, the implementation has been tested primarily with DeepSeek and Qwen. Please open an issue if you find any mistakes or inaccuracies in the implementation.

The vllm fields were derived from vLLM's own documentation and protocol sources rather than from a live server (no vLLM deployment was available); they are covered by parsing and serialization tests against recorded payload shapes. If you run vLLM and hit a mismatch, please report it.

Note that this crate models the fields vLLM adds to the OpenAI endpoints it implements. vLLM's own endpoints — /v1/score, /rerank, /pooling, /classify, /tokenize, /detokenize, the render endpoints and POST /v1/chat/completions/batch — are not implemented.

The zai fields were derived from GLM's documentation (docs.bigmodel.cn and docs.z.ai) rather than from a live server; they are covered by parsing and serialization tests against the documented payload shapes. If you call GLM and hit a mismatch, please report it. GLM's mcp tool type and its image / video / embedding and GLM Coding Plan endpoints are not modelled.

Contributing

Contributions are welcome! Please feel free to submit pull requests or open issues for bugs and feature requests, on the Codeberg issue tracker.

  • The minimum supported Rust version (MSRV) is 1.88 (declared as rust-version in Cargo.toml); changes must keep building on it.
  • Run cargo fmt --check, cargo clippy --all-targets --all-features -- -D warnings and cargo test before submitting.
  • User-facing changes must be recorded in CHANGELOG.md.
  • Runnable sample programs live in the examples directory.

License

This project is licensed under the MIT License - see the LICENSE file for details.