# OpenAI Interface
A low-level Rust interface for interacting with OpenAI's API. Both streaming
and non-streaming APIs are supported.
Currently, chat completions (create / retrieve / update / delete), completions,
models, embeddings, moderations, file management (upload / list / retrieve /
delete / download content), images (generate / edit / variation), audio
(speech / transcriptions / translations), the Responses API (create /
retrieve / delete / input items / cancel), batches, uploads, fine-tuning
jobs, vector stores, containers, conversations, evals, and realtime session
creation are supported. See the support matrix below for details.
> Repository:
>
> Codeberg: [Codeberg Repo](https://codeberg.org/Hammerklavier/openai-interface)
> GitCode: [GitCode Repo](https://gitcode.com/astral-sphere/openai-interface)
>
> You are welcome to contribute to this project through any of the links above.
## Features
- **Chat Completions**: Full support for OpenAI's chat completion and completion API,
including both streaming and non-streaming responses, and multimodal user
messages (text / image / audio / file content parts).
- **Models**: List, retrieve and delete models.
- **Embeddings**: Create embedding vectors from text input.
- **Moderations**: Classify whether text and/or image input is potentially
harmful (untested).
- **Images**: Generate, edit, and create variations of images (untested).
- **Audio**: Text-to-speech, transcription, and translation endpoints (untested).
- **Files**: Support for the OpenAI file API (create / list / retrieve / delete /
download content).
- **Responses API**: Create, retrieve and delete responses, list their input
items, and cancel background responses, including streaming events.
- **Batches**: Create, retrieve, list and cancel async batch processing jobs.
- **Uploads**: Multi-part upload sessions for large files (create / add parts /
complete / cancel).
- **Fine-tuning**: Create, retrieve, list and cancel fine-tuning jobs, list
their events and checkpoints, and restore fine-tuned models.
- **Vector Stores**: Manage vector stores and their files (including file
batches and semantic search) for the `file_search` tool.
- **Containers**: Manage containers and their files for the Code Interpreter
tool.
- **Conversations**: Manage the stateful conversation layer of the Responses
API and its items.
- **Evals**: Manage evals, their runs and the runs' output items.
- **Realtime**: Create ephemeral Realtime and transcription session tokens
(the WebSocket transport itself is not implemented).
- **Streaming and Non-streaming**: Support for both streaming and non-streaming responses,
with a delta accumulator (`chat::create::accumulator::ChatCompletionAccumulator`) that
assembles text and tool calls (joined by tool-call index) from the stream, the same way
the official SDKs do.
- **Reasoning Effort**: The OpenAI-compatible `reasoning_effort` parameter is supported
out of the box for reasoning models.
- **Configurable HTTP Client**: Every request method takes a `reqwest::Client`, so
proxies, timeouts and connection pooling are under your control.
- **Strong Typing**: Complete type definitions for all API requests and responses,
utilizing Rust's powerful type system.
- **Error Handling**: Comprehensive error handling with detailed error types defined in
the [`errors`] module. Failed requests carry the API's error message, type and code.
- **Async/Await**: Built with async/await support.
- **Musl Support**: Designed to work with musl libc out-of-the-box.
- **Multiple Provider Support**: Expected to work with OpenAI, DeepSeek, Qwen, vLLM, and
other compatible API providers. Provider-specific fields are opt-in via cargo features
(see below).
## Installation
> [!WARNING] Versions prior to 0.3.0 have serious issues with SSE streaming responses
> processing: instead of a single chunk, multiple chunks may be returned in each
> iteration of the response stream.
Add this to your `Cargo.toml`:
```toml
[dependencies]
openai-interface = { version = "0.13", features = ["deepseek", "qwen"] }
```
### Cargo Features
Fields that are proprietary to a single provider are opt-in via cargo
features. Cross-vendor de-facto standards — such as `reasoning_content`
(streamed by DeepSeek, Qwen3, ollama, vLLM and OpenRouter alike) — are
always available:
- **`reasoning`** (default): cross-vendor reasoning fields —
`reasoning_content` on assistant messages (request and response), streamed
deltas, and logprobs, plus its accumulation in `ChatCompletionAccumulator`.
- **`deepseek`**: Enables DeepSeek's proprietary fields — the Beta chat
prefix completion fields (`prefix`, and `reasoning_content` as the
prefix-completion CoT input), the `thinking` and `user_id` request
parameters, and the `prompt_cache_hit_tokens` / `prompt_cache_miss_tokens`
usage statistics. Implies `reasoning`. See
[api-docs.deepseek.com](https://api-docs.deepseek.com/).
- **`qwen`**: Enables Qwen's proprietary request parameters
(`enable_thinking`, `thinking_budget`, `top_k`) as direct fields of the
chat request body. Implies `reasoning`. See
[the Qwen OpenAI-compatible Chat API docs](https://www.alibabacloud.com/help/zh/model-studio/qwen-api-via-openai-chat-completions).
- **`vllm`**: Enables [vLLM](https://docs.vllm.ai/)'s proprietary fields,
collected in the `openai_interface::vllm` module. On the request side:
the extra sampling parameters (`min_p`, `repetition_penalty`,
`stop_token_ids`, `prompt_logprobs`, `bad_words`, `allowed_token_ids`,
...), the chat-template controls (`chat_template`,
`chat_template_kwargs`, `add_generation_prompt`,
`continue_final_message`, ...), `structured_outputs` (vLLM's successor to
the deprecated `guided_json` / `guided_regex` / `guided_choice` /
`guided_grammar` keys), and the KV-transfer and scheduling parameters
(`kv_transfer_params`, `priority`, `cache_salt`, `stream_interval`, ...).
On the response side: `stop_reason`, `token_ids` and `routed_experts` per
choice; `prompt_logprobs`, `prompt_token_ids`, `prompt_text`,
`kv_transfer_params` and `ec_transfer_params` on the completion and the
streamed chunks; and `root` / `parent` / `max_model_len` on model
objects. Implies `reasoning`. See
[vLLM's OpenAI-compatible server docs](https://docs.vllm.ai/en/latest/serving/online_serving/openai_compatible_server/).
- **`azure`**: Deprecated no-op. Streaming `delta.annotations` and
`delta.audio` are now always available (the non-streaming message fields
were never gated). The empty feature remains defined so existing manifests
keep compiling.
- **`ferritls`**: Unrelated to request fields — adds the pure-Rust
`ferritls-rustls` TLS crypto backend and `rest::install_crypto_provider`,
the helper that installs it. Off by default, so the crate never dictates
your crypto backend. See
[Choosing the TLS Crypto Provider](#choosing-the-tls-crypto-provider).
## Usage
### Chat Completion
This crate provides methods for both streaming and non-streaming chat completions. The following examples demonstrate how to use these features.
#### Non-streaming Chat Completion
```rust
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::chat::create::response::no_streaming::ChatCompletion;
use openai_interface::rest::{
default_client, install_crypto_provider, post::PostNoStream, RequestOptions,
};
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let api_key = std::env::var("DEEPSEEK_API_KEY")?;
// Needs the `ferritls` feature; skip this if you install your own
// provider — see "Choosing the TLS Crypto Provider" below.
install_crypto_provider().ok();
let client = default_client();
let request = RequestBody {
messages: vec![
Message::system("You are a helpful assistant."),
Message::user("Hello, how are you?"),
],
model: "deepseek-v4-flash".to_string(),
stream: Some(false),
..Default::default()
};
// Send the request. The base URL is everything before the endpoint
// path — for OpenAI and most compatible gateways it must include the
// `/v1` prefix (e.g. `https://api.openai.com/v1`); DeepSeek is the
// exception and uses the bare host.
let options = RequestOptions::bearer(api_key);
let chat_completion: ChatCompletion = request
.get_response(&client, "https://api.deepseek.com", &options)
.await?;
let text = chat_completion.choices[0]
.message
.content
.as_deref()
.unwrap();
println!("{:?}", text);
Ok(())
}
```
#### Streaming Chat Completion
This example demonstrates how to handle streaming responses from the API.
`get_stream_response` deserializes every server-sent event and stops
automatically at the `data: [DONE]` sentinel. The
[`ChatCompletionAccumulator`] assembles the fragments into a complete
message — concatenating text and tool-call arguments (by tool-call index)
exactly like the official SDKs.
[`ChatCompletionAccumulator`]: https://docs.rs/openai-interface/latest/openai_interface/chat/create/accumulator/struct.ChatCompletionAccumulator.html
```rust
use openai_interface::chat::create::accumulator::ChatCompletionAccumulator;
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::rest::{default_client, install_crypto_provider, post::PostStream, RequestOptions};
use futures_util::StreamExt;
#[tokio::main]
async fn main() -> Result<(), Box<dyn std::error::Error>> {
let api_key = std::env::var("DEEPSEEK_API_KEY")?;
// Needs the `ferritls` feature; skip this if you install your own
// provider — see "Choosing the TLS Crypto Provider" below.
install_crypto_provider().ok();
let client = default_client();
let request = RequestBody {
messages: vec![
Message::system("You are a helpful assistant."),
Message::user("Who are you?"),
],
model: "deepseek-v4-flash".to_string(),
stream: Some(true),
..Default::default()
};
// Send the request. The base URL is everything before the endpoint
// path — for OpenAI and most compatible gateways it must include the
// `/v1` prefix (e.g. `https://api.openai.com/v1`); DeepSeek is the
// exception and uses the bare host.
let options = RequestOptions::bearer(api_key);
let mut response_stream = request
.get_stream_response(&client, "https://api.deepseek.com", &options)
.await?;
let mut accumulator = ChatCompletionAccumulator::new();
while let Some(chunk_result) = response_stream.next().await {
let chunk = chunk_result?;
if let Some(content) = chunk.choices.first().and_then(|c| c.delta.content.as_deref()) {
println!("message chunk: {}", content);
}
accumulator.push(&chunk);
// For reasoning models such as `deepseek-reasoner`, the chain of
// thought arrives in `choice.delta.reasoning_content`.
}
let message = accumulator.into_message();
println!("complete message: {:?}", message.content);
println!("tool calls: {:?}", message.tool_calls);
Ok(())
}
```
If a provider occasionally emits chunks your code cannot deserialize, wrap
the stream with `openai_interface::rest::skip_deserialization_errors` to
drop those items instead of abandoning the stream at the first one.
#### Configuring the HTTP Client
Every request method takes the client as its first argument. Pass a custom
client to use a proxy or a different timeout:
```rust
let client = reqwest::Client::builder()
.proxy(reqwest::Proxy::http("http://127.0.0.1:10808")?)
.timeout(std::time::Duration::from_secs(60))
.build()?;
```
`rest::default_client` sets a 60s connect timeout and a 300s read timeout.
The read timeout applies to each read and resets after every successful
read, so a streaming response is never cut off mid-stream — but a backend
that stays completely silent for over 300s (e.g. a long reasoning phase
without streamed reasoning content) will be dropped; build a custom client
with a larger `read_timeout` if you need to accommodate that.
#### Authenticating with Other Providers
Every request method takes a [`RequestOptions`] value, which carries the
authentication scheme plus any extra headers. [`RequestOptions::bearer`]
reproduces the classic OpenAI `Authorization: Bearer` behavior; Azure and
Anthropic authenticate with other headers, which you can supply per request:
```rust
use openai_interface::rest::RequestOptions;
// Azure OpenAI: credentials travel in the `api-key` header.
let azure = RequestOptions::new().with_header("api-key", "azure-key")?;
// Anthropic: `x-api-key` plus a version header.
let anthropic = RequestOptions::new()
.with_header("x-api-key", "anthropic-key")?
.with_header("anthropic-version", "2023-06-01")?;
// OpenAI: extra headers can be layered on top of bearer auth.
let openai = RequestOptions::bearer("sk-...")
.with_header("OpenAI-Organization", "org-...")?;
```
[`RequestOptions`]: https://docs.rs/openai-interface/latest/openai_interface/rest/struct.RequestOptions.html
#### Choosing the TLS Crypto Provider
This crate depends on `reqwest` with its `rustls-no-provider` feature, so the
rustls stack is compiled **without** a crypto backend. That keeps the build
pure Rust (no C or asm toolchain, which is what makes the musl target
work out-of-the-box) and leaves the backend choice to the application.
The consequence is that exactly one [rustls `CryptoProvider`] must be installed
as the process default before the first `reqwest::Client` is built — including
the client returned by `default_client()`. Without one, reqwest panics at
client construction time.
This crate never installs a provider for you: neither `default_client()` nor
any request method touches that global state, so the application stays in
control. If you would rather not decide, the optional **`ferritls`** feature
adds the pure-Rust [`ferritls-rustls`] backend plus a helper that installs it:
```toml
[dependencies]
openai-interface = { version = "0.10", features = ["ferritls"] }
```
```rust
// Requires the `ferritls` feature; call it once, before the first client.
openai_interface::rest::install_crypto_provider()
.expect("a rustls crypto provider was already installed");
```
Without that feature, `ferritls-rustls` is not in your dependency tree at all
and there is nothing to call — install a provider yourself instead. First
install wins: whichever provider is installed when the first client is built is
the one the whole process uses, so do it before any request.
```rust
// In the application crate, with `rustls = "0.23"` (feature `ring` or
// `aws-lc-rs`) as one of its own dependencies:
rustls::crypto::ring::default_provider()
.install_default()
.expect("a rustls crypto provider was already installed");
```
Note that `reqwest`'s features are additive, so this requirement disappears
altogether when your own project depends on `reqwest` with a crypto backend
compiled in:
```toml
[dependencies]
reqwest = "0.13" # default features: `default-tls` -> `rustls`
openai-interface = "0.10" # no provider of its own
```
`default-tls` (or `rustls` directly) makes reqwest fall back to the
`aws-lc-rs` provider it ships with, and `native-tls` routes TLS through the
system stack so the rustls path is never taken — either way you do not have to
install anything, and you never call `install_crypto_provider`. The catch is
that the backend is then decided by feature unification instead of by you, so
an unrelated dependency change can move it. If you want the choice pinned,
enable `ferritls` or install a provider yourself.
[rustls `CryptoProvider`]: https://docs.rs/rustls/latest/rustls/crypto/struct.CryptoProvider.html
[`ferritls-rustls`]: https://crates.io/crates/ferritls-rustls
#### Custom Request Parameters
For provider-specific parameters, prefer enabling the matching cargo feature
(`deepseek`, `qwen` or `vllm`) so the fields are available as typed members of
the request structs.
If you need a field that is not covered by the typed structs, you can inject
arbitrary JSON properties. Every top-level request body carries an
`extra_body_map` field
(`Option<serde_json::Map<String, serde_json::Value>>`, flattened into the
request body; on multipart endpoints the entries are sent as extra form
fields):
```rust
use openai_interface::chat::create::request::{Message, RequestBody};
use serde_json::json;
let request = RequestBody {
messages: vec![Message::user("Hello")],
model: "gpt-4.1".to_string(),
stream: Some(false),
extra_body_map: Some(
serde_json::from_value(json!({ "some_vendor_field": 42 })).unwrap(),
),
..Default::default()
};
```
#### Serving Against vLLM
With the `vllm` feature, the parameters vLLM accepts beyond the OpenAI
standard are typed members of the request body instead of `extra_body_map`
entries. They are split in two because the two text-generation endpoints do
not accept the same set: `vllm_sampling` (decoding knobs, valid on both
`/v1/chat/completions` and `/v1/completions`) and `vllm_chat` (chat-template
rendering, structured output, KV transfer and scheduling; chat only). Both
are flattened, so the JSON is exactly what the official client's `extra_body`
would produce.
```rust
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::vllm::{ChatParams, SamplingParams, StructuredOutputsParams};
let request = RequestBody {
messages: vec![Message::user("Classify this sentiment: vLLM is wonderful!")],
model: "Qwen/Qwen3-8B".to_string(),
stream: Some(false),
vllm_sampling: Some(SamplingParams {
min_p: Some(0.1),
repetition_penalty: Some(1.05),
..Default::default()
}),
vllm_chat: Some(ChatParams {
structured_outputs: Some(StructuredOutputsParams {
choice: Some(vec!["positive".to_string(), "negative".to_string()]),
..Default::default()
}),
// Template-defined switches, e.g. turning Qwen3's thinking off.
chat_template_kwargs: Some(
serde_json::from_value(serde_json::json!({ "enable_thinking": false })).unwrap(),
),
..Default::default()
}),
..Default::default()
};
```
On the response side, vLLM's extras are plain fields of the same structs:
`stop_reason` and `token_ids` per choice, `prompt_logprobs`,
`prompt_token_ids`, `prompt_text` and `kv_transfer_params` on the completion
and its chunks, and `root` / `parent` / `max_model_len` on model objects.
One difference needs calling out: vLLM streams and returns the chain of
thought under **`reasoning`**, not the cross-vendor `reasoning_content` (it
accepts the latter on input, but always emits the former). So against a vLLM
backend read `message.reasoning` — or `ChatCompletionAccumulator::reasoning()`
for a stream — and map it back onto `reasoning_content` when you feed it into
a follow-up request.
### Modules
- [`chat`]: Contains all chat completion related structs, enums, and methods.
- [`completions`]: Contains all completion related structs, enums, and methods.
Note that this API is getting deprecated in favour of `chat` and is only available
for out-dated LLM models.
- [`models`]: List, retrieve and delete models.
- [`embeddings`]: Create embedding vectors from text input.
- [`moderations`]: Classify whether text input is potentially harmful.
- [`images`]: Generate, edit, and create variations of images.
- [`audio`]: Turn audio into text (transcriptions / translations) or text into
audio (speech).
- [`files`]: Providing the capacity to upload and manage files.
- [`rest`]: Providing all REST related traits and methods, plus
`default_client` and shared status/error handling.
- [`errors`]: Defines error types used throughout the crate.
- [`pagination`]: Shared cursor-pagination types (`Page`, `PaginationQuery`)
used by the list endpoints.
- [`vllm`] (with the `vllm` feature): vLLM's proprietary request and
response fields.
### API Support Matrix
All newly added modules are **untested** against a live OpenAI API (no API
key was available); they follow the official documentation and should work
with OpenAI-compatible providers that implement the same endpoints. Please
report any issues on the [Codeberg issue tracker](https://codeberg.org/Hammerklavier/openai-interface/issues).
| Chat Completions | create / retrieve / update / delete | tested |
| Completions (legacy) | create | tested |
| Models | list / retrieve / delete | tested |
| Embeddings | create | tested |
| Moderations | create | untested |
| Images | generate / edit / variation | untested |
| Audio | speech / transcriptions / translations | untested |
| Files | create / list / retrieve / delete / content | tested |
| Responses | create / retrieve / delete / input items / cancel | partially tested |
| Batches | create / retrieve / list / cancel | untested |
| Uploads | create / add part / complete / cancel | untested |
| Fine-tuning | jobs create / retrieve / list / cancel, events, checkpoints, model restore | untested |
| Vector Stores | create / retrieve / update / delete / list, files CRUD + content, search, file batches | untested |
| Containers | create / retrieve / delete / list, files CRUD + content | untested |
| Conversations | create / retrieve / update / delete / list, items CRUD + list | untested |
| Evals | create / retrieve / update / delete / list, runs CRUD + cancel, output items | untested |
| Realtime | sessions / transcription sessions (HTTP only; WebSocket not implemented) | untested |
Not implemented: the Realtime WebSocket transport, the evals alpha
permissions endpoints (`/fine_tuning/alpha/permissions`), and streaming
variants of the images and audio transcription endpoints.
### Error Handling
All errors are converted into [`errors::OapiError`]. On a failed request the
response body is parsed into [`errors::ApiError`], which carries the API's
error `message`, `type`, `code` and the HTTP `status`, so failures can be
diagnosed without re-sending the request.
## Musl Build
This crate is designed to work with musl libc, making it suitable for
lightweight deployments in containerized environments. TLS is provided by
rustls with a pure-Rust crypto backend, so OpenSSL does not need to be built
from source (see "Choosing the TLS Crypto Provider" above for how the backend
is selected at runtime).
To build for musl:
```bash
rustup target add x86_64-unknown-linux-musl
cargo build --target x86_64-unknown-linux-musl
```
## Supported Providers
This crate aims to support standard OpenAI-compatible API endpoints. Unfortunately, OpenAI
aggressively restricts the access from the People's Republic of China. As a result, the
implementation has been tested primarily with DeepSeek and Qwen. Please open an issue if you
find any mistakes or inaccuracies in the implementation.
The `vllm` fields were derived from vLLM's own documentation and protocol
sources rather than from a live server (no vLLM deployment was available);
they are covered by parsing and serialization tests against recorded payload
shapes. If you run vLLM and hit a mismatch, please report it.
Note that this crate models the fields vLLM adds to the OpenAI endpoints it
implements. vLLM's own endpoints — `/v1/score`, `/rerank`, `/pooling`,
`/classify`, `/tokenize`, `/detokenize`, the `render` endpoints and
`POST /v1/chat/completions/batch` — are not implemented.
## Contributing
Contributions are welcome! Please feel free to submit pull requests or open
issues for bugs and feature requests, on the
[Codeberg issue tracker](https://codeberg.org/Hammerklavier/openai-interface/issues).
- The minimum supported Rust version (MSRV) is **1.88** (declared as
`rust-version` in `Cargo.toml`); changes must keep building on it.
- Run `cargo fmt --check`, `cargo clippy --all-targets --all-features -- -D warnings`
and `cargo test` before submitting.
- User-facing changes must be recorded in `CHANGELOG.md`.
- Runnable sample programs live in the [`examples`](examples/) directory.
## License
This project is licensed under the MIT License - see the LICENSE file for details.