Expand description
vLLM’s proprietary extensions to the OpenAI-compatible API.
Everything in this module is gated on the vllm cargo feature.
vLLM serves models behind OpenAI-compatible endpoints, but accepts a
large set of request parameters that OpenAI’s API does not define, and
returns extra fields OpenAI never sends. With the official Python client
those parameters have to be smuggled through extra_body; here they are
typed.
The types are split in two because the two text-generation endpoints do not accept the same set:
SamplingParams— decoding knobs, accepted by both/v1/chat/completionsand the legacy/v1/completions.ChatParams— chat-template rendering, structured outputs, KV transfer and scheduling, accepted by/v1/chat/completionsonly.
Both are #[serde(flatten)]ed into their request bodies, so the wire
format is identical to what extra_body produces: every field lands at
the top level of the JSON body.
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::vllm::SamplingParams;
let request = RequestBody {
messages: vec![Message::user("Hello")],
model: "Qwen/Qwen3-8B".to_string(),
vllm_sampling: Some(SamplingParams {
min_p: Some(0.1),
repetition_penalty: Some(1.05),
..Default::default()
}),
..Default::default()
};
let json = serde_json::to_value(&request).unwrap();
assert_eq!(json["min_p"], serde_json::json!(0.1f32));
assert_eq!(json["repetition_penalty"], serde_json::json!(1.05f32));§Note on shared keys
Two keys vLLM accepts are not defined here because another provider spells
them identically, so they live once on
RequestBody instead —
defining them in both places would emit the key twice:
top_k, gated onany(feature = "qwen", feature = "vllm").request_id(vLLM’s caller-chosen identifier replacing the generated UUID), gated onany(feature = "vllm", feature = "zai")because Z.ai / GLM uses the same key.
§Differences that need no extra field
Some vLLM divergences are behavioural, and are documented where the affected OpenAI-standard field lives:
suffixis rejected outright on/v1/completions.image_url.detailis rejected outright.response_format.json_schema.strictis parsed but has no effect: guided decoding always enforces the schema.- When both
max_tokensandmax_completion_tokensare set,max_tokenswins. - The chain of thought arrives as
reasoning, notreasoning_content— see thereasoningfields on the response message and streamed delta.
§Not implemented
vLLM also exposes endpoints OpenAI has no equivalent of: /v1/score,
/rerank, /pooling, /classify, /tokenize, /detokenize, the
POST /v1/{chat/completions,completions}/render prompt-rendering
endpoints, POST /v1/chat/completions/batch, and the admin endpoints
(/load, /sleep, /wake_up, /is_sleeping, /reset_prefix_cache).
None of them are modelled by this crate; this feature covers the fields
vLLM adds to the OpenAI endpoints it does implement.
Structs§
- Chat
Params - Extra
/v1/chat/completionsparameters of a vLLM server: chat-template rendering, structured output, KV transfer and scheduling. - Logprob
- vLLM’s log probability of one token at one position, as reported in the
prompt_logprobsresponse field. - Repetition
Detection Params - vLLM’s
repetition_detectionparameter: aborts generation once the output settles into a repeated n-gram. - Sampling
Params - Extra sampling parameters accepted by both
/v1/chat/completionsand/v1/completionson a vLLM server. - Structured
Outputs Params - vLLM’s
structured_outputsparameter: constrains decoding to a grammar so the output is guaranteed to parse.
Enums§
- Stop
Reason - Why generation stopped, as reported in vLLM’s
stop_reasonresponse field. - Truncation
Side - Which end of an over-long prompt
truncate_prompt_tokensremoves.