Expand description
vLLM’s proprietary extensions to the OpenAI-compatible API.
Everything in this module is gated on the vllm cargo feature.
vLLM serves models behind OpenAI-compatible endpoints, but accepts a
large set of request parameters that OpenAI’s API does not define, and
returns extra fields OpenAI never sends. With the official Python client
those parameters have to be smuggled through extra_body; here they are
typed.
The types are split in two because the two text-generation endpoints do not accept the same set:
SamplingParams— decoding knobs, accepted by both/v1/chat/completionsand the legacy/v1/completions.ChatParams— chat-template rendering, structured outputs, KV transfer and scheduling, accepted by/v1/chat/completionsonly.
Both are #[serde(flatten)]ed into their request bodies, so the wire
format is identical to what extra_body produces: every field lands at
the top level of the JSON body.
use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::vllm::SamplingParams;
let request = RequestBody {
messages: vec![Message::user("Hello")],
model: "Qwen/Qwen3-8B".to_string(),
vllm_sampling: Some(SamplingParams {
min_p: Some(0.1),
repetition_penalty: Some(1.05),
..Default::default()
}),
..Default::default()
};
let json = serde_json::to_value(&request).unwrap();
assert_eq!(json["min_p"], serde_json::json!(0.1f32));
assert_eq!(json["repetition_penalty"], serde_json::json!(1.05f32));§Note on top_k
top_k is not defined here: it is already a field of
RequestBody, gated on
any(feature = "qwen", feature = "vllm") because Qwen and vLLM spell it
the same way. Defining it twice would emit the key twice.
§Differences that need no extra field
Some vLLM divergences are behavioural, and are documented where the affected OpenAI-standard field lives:
suffixis rejected outright on/v1/completions.image_url.detailis rejected outright.response_format.json_schema.strictis parsed but has no effect: guided decoding always enforces the schema.- When both
max_tokensandmax_completion_tokensare set,max_tokenswins. - The chain of thought arrives as
reasoning, notreasoning_content— see thereasoningfields on the response message and streamed delta.
§Not implemented
vLLM also exposes endpoints OpenAI has no equivalent of: /v1/score,
/rerank, /pooling, /classify, /tokenize, /detokenize, the
POST /v1/{chat/completions,completions}/render prompt-rendering
endpoints, POST /v1/chat/completions/batch, and the admin endpoints
(/load, /sleep, /wake_up, /is_sleeping, /reset_prefix_cache).
None of them are modelled by this crate; this feature covers the fields
vLLM adds to the OpenAI endpoints it does implement.
Structs§
- Chat
Params - Extra
/v1/chat/completionsparameters of a vLLM server: chat-template rendering, structured output, KV transfer and scheduling. - Logprob
- vLLM’s log probability of one token at one position, as reported in the
prompt_logprobsresponse field. - Repetition
Detection Params - vLLM’s
repetition_detectionparameter: aborts generation once the output settles into a repeated n-gram. - Sampling
Params - Extra sampling parameters accepted by both
/v1/chat/completionsand/v1/completionson a vLLM server. - Structured
Outputs Params - vLLM’s
structured_outputsparameter: constrains decoding to a grammar so the output is guaranteed to parse.
Enums§
- Stop
Reason - Why generation stopped, as reported in vLLM’s
stop_reasonresponse field. - Truncation
Side - Which end of an over-long prompt
truncate_prompt_tokensremoves.