openai-interface 0.9.0

A low-level Rust interface for the OpenAI API
Documentation

OpenAI Interface

A low-level Rust interface for interacting with OpenAI's API. Both streaming and non-streaming APIs are supported.

Currently, chat completions (create / retrieve / update / delete), completions, models, embeddings, moderations, file management (upload / list / retrieve / delete / download content), images (generate / edit / variation), audio (speech / transcriptions / translations), the Responses API (create / retrieve / delete / input items / cancel), batches, uploads, fine-tuning jobs, vector stores, containers, conversations, evals, and realtime session creation are supported. See the support matrix below for details.

Repository:

Codeberg: Codeberg Repo
GitCode: GitCode Repo

You are welcome to contribute to this project through any of the links above.

Features

  • Chat Completions: Full support for OpenAI's chat completion and completion API, including both streaming and non-streaming responses, and multimodal user messages (text / image / audio / file content parts).
  • Models: List, retrieve and delete models.
  • Embeddings: Create embedding vectors from text input.
  • Moderations: Classify whether text and/or image input is potentially harmful (untested).
  • Images: Generate, edit, and create variations of images (untested).
  • Audio: Text-to-speech, transcription, and translation endpoints (untested).
  • Files: Support for the OpenAI file API (create / list / retrieve / delete / download content).
  • Responses API: Create, retrieve and delete responses, list their input items, and cancel background responses, including streaming events.
  • Batches: Create, retrieve, list and cancel async batch processing jobs.
  • Uploads: Multi-part upload sessions for large files (create / add parts / complete / cancel).
  • Fine-tuning: Create, retrieve, list and cancel fine-tuning jobs, list their events and checkpoints, and restore fine-tuned models.
  • Vector Stores: Manage vector stores and their files (including file batches and semantic search) for the file_search tool.
  • Containers: Manage containers and their files for the Code Interpreter tool.
  • Conversations: Manage the stateful conversation layer of the Responses API and its items.
  • Evals: Manage evals, their runs and the runs' output items.
  • Realtime: Create ephemeral Realtime and transcription session tokens (the WebSocket transport itself is not implemented).
  • Streaming and Non-streaming: Support for both streaming and non-streaming responses.
  • Reasoning Effort: The OpenAI-compatible reasoning_effort parameter is supported out of the box for reasoning models.
  • Configurable HTTP Client: Every request method takes a reqwest::Client, so proxies, timeouts and connection pooling are under your control.
  • Strong Typing: Complete type definitions for all API requests and responses, utilizing Rust's powerful type system.
  • Error Handling: Comprehensive error handling with detailed error types defined in the [errors] module. Failed requests carry the API's error message, type and code.
  • Async/Await: Built with async/await support.
  • Musl Support: Designed to work with musl libc out-of-the-box.
  • Multiple Provider Support: Expected to work with OpenAI, DeepSeek, Qwen, and other compatible API providers. Provider-specific fields are opt-in via cargo features (see below).

Installation

[!WARNING] Versions prior to 0.3.0 have serious issues with SSE streaming responses processing: instead of a single chunk, multiple chunks may be returned in each iteration of the response stream.

Add this to your Cargo.toml:

[dependencies]
openai-interface = { version = "0.8.1", features = ["deepseek", "qwen"] }

Cargo Features

To keep the request and response types strictly OpenAI-compatible, fields that are proprietary to other providers are opt-in via cargo features:

  • deepseek: Enables DeepSeek's proprietary fields — the Beta chat prefix completion fields (prefix / reasoning_content on assistant messages), the thinking and user_id request parameters, reasoning_content in responses and logprobs, prompt_cache_hit_tokens / prompt_cache_miss_tokens usage statistics, and the insufficient_system_resource finish reason. See api-docs.deepseek.com.
  • qwen: Enables Qwen's proprietary request parameters (enable_thinking, thinking_budget, top_k) as direct fields of the chat request body. See the Qwen OpenAI-compatible Chat API docs.

Usage

Chat Completion

This crate provides methods for both streaming and non-streaming chat completions. The following examples demonstrate how to use these features.

Non-streaming Chat Completion

use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::chat::create::response::no_streaming::ChatCompletion;
use openai_interface::rest::{default_client, post::PostNoStream};

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("DEEPSEEK_API_KEY")?;
    let client = default_client();

    let request = RequestBody {
        messages: vec![
            Message::System {
                content: "You are a helpful assistant.".to_string(),
                name: None,
            },
            Message::User {
                content: "Hello, how are you?".into(),
                name: None,
            },
        ],
        model: "deepseek-v4-flash".to_string(),
        stream: Some(false),
        ..Default::default()
    };

    // Send the request
    let chat_completion: ChatCompletion = request
        .get_response(&client, "https://api.deepseek.com/chat/completions", &api_key)
        .await?;
    let text = chat_completion.choices[0]
        .message
        .content
        .as_deref()
        .unwrap();
    println!("{:?}", text);
    Ok(())
}

Streaming Chat Completion

This example demonstrates how to handle streaming responses from the API. get_stream_response deserializes every server-sent event and stops automatically at the data: [DONE] sentinel.

use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::chat::create::response::streaming::ChatCompletionChunk;
use openai_interface::rest::{default_client, post::PostStream};
use futures_util::StreamExt;

#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
    let api_key = std::env::var("DEEPSEEK_API_KEY")?;
    let client = default_client();

    let request = RequestBody {
        messages: vec![
            Message::System {
                content: "You are a helpful assistant.".to_string(),
                name: None,
            },
            Message::User {
                content: "Who are you?".into(),
                name: None,
            },
        ],
        model: "deepseek-v4-flash".to_string(),
        stream: Some(true),
        ..Default::default()
    };

    // Send the request
    let mut response_stream = request
        .get_stream_response(&client, "https://api.deepseek.com/chat/completions", &api_key)
        .await?;

    let mut message = String::new();

    while let Some(chunk_result) = response_stream.next().await {
        let chunk: ChatCompletionChunk = chunk_result?;
        if let Some(choice) = chunk.choices.first() {
            if let Some(content) = choice.delta.content.as_deref() {
                println!("message chunk: {}", content);
                message.push_str(content);
            }
            // For reasoning models such as `deepseek-reasoner`, the chain of
            // thought arrives in `choice.delta.reasoning_content` — this field
            // is only available with the `deepseek` cargo feature enabled.
        }
    }

    println!("complete message: {}", message);
    Ok(())
}

Configuring the HTTP Client

Every request method takes the client as its first argument. Pass a custom client to use a proxy or a different timeout:

let client = reqwest::Client::builder()
    .proxy(reqwest::Proxy::http("http://127.0.0.1:10808")?)
    .timeout(std::time::Duration::from_secs(60))
    .build()?;

Custom Request Parameters

For provider-specific parameters, prefer enabling the matching cargo feature (deepseek or qwen) so the fields are available as typed members of the request structs.

If you need a field that is not covered by the typed structs, you can inject arbitrary JSON properties:

use openai_interface::chat::create::request::{Message, RequestBody};
use serde_json::json;

let request = RequestBody {
    messages: vec![Message::User {
        content: "Hello".into(),
        name: None,
    }],
    model: "gpt-4.1".to_string(),
    stream: Some(false),
    extra_body_map: Some(
        serde_json::from_value(json!({ "some_vendor_field": 42 })).unwrap(),
    ),
    ..Default::default()
};

Modules

  • [chat]: Contains all chat completion related structs, enums, and methods.
  • [completions]: Contains all completion related structs, enums, and methods. Note that this API is getting deprecated in favour of chat and is only available for out-dated LLM models.
  • [models]: List, retrieve and delete models.
  • [embeddings]: Create embedding vectors from text input.
  • [moderations]: Classify whether text input is potentially harmful.
  • [images]: Generate, edit, and create variations of images.
  • [audio]: Turn audio into text (transcriptions / translations) or text into audio (speech).
  • [files]: Providing the capacity to upload and manage files.
  • [rest]: Providing all REST related traits and methods, plus default_client and shared status/error handling.
  • [errors]: Defines error types used throughout the crate.
  • [pagination]: Shared cursor-pagination types (Page, PaginationQuery) used by the list endpoints.

API Support Matrix

All newly added modules are untested against a live OpenAI API (no API key was available); they follow the official documentation and should work with OpenAI-compatible providers that implement the same endpoints. Please report any issues on the repository.

API group Endpoints Status
Chat Completions create / retrieve / update / delete tested
Completions (legacy) create tested
Models list / retrieve / delete tested
Embeddings create tested
Moderations create untested
Images generate / edit / variation untested
Audio speech / transcriptions / translations untested
Files create / list / retrieve / delete / content tested
Responses create / retrieve / delete / input items / cancel partially tested
Batches create / retrieve / list / cancel untested
Uploads create / add part / complete / cancel untested
Fine-tuning jobs create / retrieve / list / cancel, events, checkpoints, model restore untested
Vector Stores create / retrieve / update / delete / list, files CRUD + content, search, file batches untested
Containers create / retrieve / delete / list, files CRUD + content untested
Conversations create / retrieve / update / delete / list, items CRUD + list untested
Evals create / retrieve / update / delete / list, runs CRUD + cancel, output items untested
Realtime sessions / transcription sessions (HTTP only; WebSocket not implemented) untested

Not implemented: the Realtime WebSocket transport, the evals alpha permissions endpoints (/fine_tuning/alpha/permissions), and streaming variants of the images and audio transcription endpoints.

Error Handling

All errors are converted into [errors::OapiError]. On a failed request the response body is parsed into [errors::ApiError], which carries the API's error message, type, code and the HTTP status, so failures can be diagnosed without re-sending the request.

Musl Build

This crate is designed to work with musl libc, making it suitable for lightweight deployments in containerized environments. Longer compile times may be required as OpenSSL needs to be built from source.

To build for musl:

rustup target add x86_64-unknown-linux-musl
cargo build --target x86_64-unknown-linux-musl

Supported Providers

This crate aims to support standard OpenAI-compatible API endpoints. Unfortunately, OpenAI aggressively restricts the access from the People's Republic of China. As a result, the implementation has been tested primarily with DeepSeek and Qwen. Please open an issue if you find any mistakes or inaccuracies in the implementation.

Contributing

Contributions are welcome! Please feel free to submit pull requests or open issues for bugs and feature requests.

License

This project is licensed under the AGPL-3.0 License - see the LICENSE file for details.