Skip to main content

Module vllm

Module vllm 

Source
Expand description

vLLM’s proprietary extensions to the OpenAI-compatible API.

Everything in this module is gated on the vllm cargo feature.

vLLM serves models behind OpenAI-compatible endpoints, but accepts a large set of request parameters that OpenAI’s API does not define, and returns extra fields OpenAI never sends. With the official Python client those parameters have to be smuggled through extra_body; here they are typed.

The types are split in two because the two text-generation endpoints do not accept the same set:

  • SamplingParams — decoding knobs, accepted by both /v1/chat/completions and the legacy /v1/completions.
  • ChatParams — chat-template rendering, structured outputs, KV transfer and scheduling, accepted by /v1/chat/completions only.

Both are #[serde(flatten)]ed into their request bodies, so the wire format is identical to what extra_body produces: every field lands at the top level of the JSON body.

use openai_interface::chat::create::request::{Message, RequestBody};
use openai_interface::vllm::SamplingParams;

let request = RequestBody {
    messages: vec![Message::user("Hello")],
    model: "Qwen/Qwen3-8B".to_string(),
    vllm_sampling: Some(SamplingParams {
        min_p: Some(0.1),
        repetition_penalty: Some(1.05),
        ..Default::default()
    }),
    ..Default::default()
};

let json = serde_json::to_value(&request).unwrap();
assert_eq!(json["min_p"], serde_json::json!(0.1f32));
assert_eq!(json["repetition_penalty"], serde_json::json!(1.05f32));

§Note on top_k

top_k is not defined here: it is already a field of RequestBody, gated on any(feature = "qwen", feature = "vllm") because Qwen and vLLM spell it the same way. Defining it twice would emit the key twice.

§Differences that need no extra field

Some vLLM divergences are behavioural, and are documented where the affected OpenAI-standard field lives:

  • suffix is rejected outright on /v1/completions.
  • image_url.detail is rejected outright.
  • response_format.json_schema.strict is parsed but has no effect: guided decoding always enforces the schema.
  • When both max_tokens and max_completion_tokens are set, max_tokens wins.
  • The chain of thought arrives as reasoning, not reasoning_content — see the reasoning fields on the response message and streamed delta.

§Not implemented

vLLM also exposes endpoints OpenAI has no equivalent of: /v1/score, /rerank, /pooling, /classify, /tokenize, /detokenize, the POST /v1/{chat/completions,completions}/render prompt-rendering endpoints, POST /v1/chat/completions/batch, and the admin endpoints (/load, /sleep, /wake_up, /is_sleeping, /reset_prefix_cache). None of them are modelled by this crate; this feature covers the fields vLLM adds to the OpenAI endpoints it does implement.

See vLLM’s OpenAI-compatible server docs.

Structs§

ChatParams
Extra /v1/chat/completions parameters of a vLLM server: chat-template rendering, structured output, KV transfer and scheduling.
Logprob
vLLM’s log probability of one token at one position, as reported in the prompt_logprobs response field.
RepetitionDetectionParams
vLLM’s repetition_detection parameter: aborts generation once the output settles into a repeated n-gram.
SamplingParams
Extra sampling parameters accepted by both /v1/chat/completions and /v1/completions on a vLLM server.
StructuredOutputsParams
vLLM’s structured_outputs parameter: constrains decoding to a grammar so the output is guaranteed to parse.

Enums§

StopReason
Why generation stopped, as reported in vLLM’s stop_reason response field.
TruncationSide
Which end of an over-long prompt truncate_prompt_tokens removes.