pub struct ChatParams {Show 23 fields
pub echo: Option<bool>,
pub add_generation_prompt: Option<bool>,
pub continue_final_message: Option<bool>,
pub add_special_tokens: Option<bool>,
pub chat_template: Option<String>,
pub chat_template_kwargs: Option<Map<String, Value>>,
pub documents: Option<Vec<HashMap<String, String>>>,
pub mm_processor_kwargs: Option<Map<String, Value>>,
pub media_io_kwargs: Option<Map<String, Value>>,
pub structured_outputs: Option<StructuredOutputsParams>,
pub priority: Option<i64>,
pub session_id: Option<String>,
pub cache_salt: Option<String>,
pub stream_interval: Option<u32>,
pub kv_transfer_params: Option<Map<String, Value>>,
pub ec_transfer_params: Option<Map<String, Value>>,
pub return_tokens_as_token_ids: Option<bool>,
pub return_token_ids: Option<bool>,
pub return_token_offsets: Option<bool>,
pub return_prompt_text: Option<bool>,
pub repetition_detection: Option<RepetitionDetectionParams>,
pub vllm_xargs: Option<Map<String, Value>>,
pub routed_experts_prompt_start: Option<u32>,
}Expand description
Extra /v1/chat/completions parameters of a vLLM server: chat-template
rendering, structured output, KV transfer and scheduling.
Flattened into the request body; see the module docs.
Rejected by the legacy /v1/completions endpoint.
Fields§
§echo: Option<bool>Echo the rendered prompt back at the start of the completion.
add_generation_prompt: Option<bool>Append the template’s generation prompt (e.g. <|assistant|>) after
the last message. vLLM’s default is true.
continue_final_message: Option<bool>Instead of appending a generation prompt, continue the final message
as if the assistant had already started it. Mutually exclusive with
add_generation_prompt.
add_special_tokens: Option<bool>Prepend the tokenizer’s BOS token to the rendered prompt. vLLM’s
default is false, because the chat template normally emits it.
chat_template: Option<String>Override the model’s Jinja chat template with this one, for this request only.
chat_template_kwargs: Option<Map<String, Value>>Extra variables handed to the chat template. This is how
template-defined switches are set — e.g. Qwen3’s enable_thinking
when the model is served by vLLM rather than by DashScope:
{"enable_thinking": false}.
documents: Option<Vec<HashMap<String, String>>>Documents made available to templates that take a documents
variable, as a list of string-to-string maps.
mm_processor_kwargs: Option<Map<String, Value>>Per-request overrides for the multimodal processor (image resizing, video frame sampling, and similar model-specific knobs).
media_io_kwargs: Option<Map<String, Value>>Per-request overrides for the multimodal I/O layer, keyed by modality.
structured_outputs: Option<StructuredOutputsParams>Constrain decoding to a grammar. vLLM’s successor to the deprecated
top-level guided_json / guided_regex / guided_choice /
guided_grammar keys.
priority: Option<i64>Scheduling priority; higher is served first. Requires the server to
run with --scheduling-policy priority. The X-Vllm-Priority
request header overrides this field.
session_id: Option<String>Session identifier, used to keep related requests on the same engine for KV-cache reuse.
cache_salt: Option<String>Prefix-cache isolation: requests with a different salt never reuse each other’s cached blocks.
stream_interval: Option<u32>Minimum number of generated tokens between streamed chunks, used to
batch output and cut per-chunk overhead. 1 streams every token.
kv_transfer_params: Option<Map<String, Value>>Parameters for disaggregated prefill (KV-cache transfer between instances). Echoed back on the response.
ec_transfer_params: Option<Map<String, Value>>Parameters for encoder-cache transfer between instances. Echoed back on the response.
return_tokens_as_token_ids: Option<bool>Report generated text as the literal token-ID strings
(token_id:1234) instead of decoded text.
return_token_ids: Option<bool>Add a token_ids list to each choice, alongside the decoded text.
return_token_offsets: Option<bool>Add per-token character offsets to the rendered prompt. Requires a fast tokenizer and text-only input.
return_prompt_text: Option<bool>Return the fully rendered prompt text in prompt_text.
repetition_detection: Option<RepetitionDetectionParams>Abort generation when the output starts repeating an n-gram pattern.
vllm_xargs: Option<Map<String, Value>>Escape hatch for parameters vLLM added after this crate was released: string-keyed scalars or lists passed straight to the engine.
routed_experts_prompt_start: Option<u32>First prompt position for which to report per-token expert routing.
Requires the server to run with --enable-return-routed-experts.