pub struct SamplingParams {Show 17 fields
pub min_p: Option<f32>,
pub repetition_penalty: Option<f32>,
pub stop_token_ids: Option<Vec<u32>>,
pub ignore_eos: Option<bool>,
pub min_tokens: Option<u32>,
pub skip_special_tokens: Option<bool>,
pub spaces_between_special_tokens: Option<bool>,
pub include_stop_str_in_output: Option<bool>,
pub truncate_prompt_tokens: Option<i64>,
pub truncation_side: Option<TruncationSide>,
pub prompt_logprobs: Option<u32>,
pub logprob_token_ids: Option<Vec<u32>>,
pub allowed_token_ids: Option<Vec<u32>>,
pub bad_words: Option<Vec<String>>,
pub length_penalty: Option<f32>,
pub use_beam_search: Option<bool>,
pub watermarking: Option<bool>,
}Expand description
Extra sampling parameters accepted by both /v1/chat/completions and
/v1/completions on a vLLM server.
Flattened into the request body; see the module docs. Every field is optional and omitted from the JSON when unset, so vLLM applies its own default.
Fields§
§min_p: Option<f32>Lower bound on the probability mass of the candidate tokens, relative
to the most likely one. vLLM’s default is 0.0 (disabled).
repetition_penalty: Option<f32>Multiplicative penalty applied to tokens that have already appeared.
vLLM’s default is 1.0 (no penalty). Not the same knob as OpenAI’s
frequency_penalty / presence_penalty; both may be combined.
stop_token_ids: Option<Vec<u32>>Additional token IDs that end generation, on top of the model’s own
EOS token and any stop strings.
ignore_eos: Option<bool>Generate exactly max_tokens tokens, ignoring EOS. Useful for
benchmarking; the output usually ends mid-sentence.
min_tokens: Option<u32>Suppress EOS until at least this many tokens have been generated, so
short outputs cannot be cut off early. vLLM’s default is 0.
skip_special_tokens: Option<bool>Strip special tokens from the decoded text. vLLM’s default is true.
spaces_between_special_tokens: Option<bool>Keep a space between adjacent special tokens when decoding. vLLM’s
default is true.
include_stop_str_in_output: Option<bool>Include the matched stop string in the returned text instead of
trimming it. vLLM’s default is false.
truncate_prompt_tokens: Option<i64>Truncate the prompt to at most this many tokens. -1 disables
truncation (vLLM’s default); any other value must be positive.
truncation_side: Option<TruncationSide>Which end of an over-long prompt truncate_prompt_tokens removes.
prompt_logprobs: Option<u32>Number of most-likely prompt tokens to report log probabilities for,
per prompt position. Returned in the response’s prompt_logprobs
field. OpenAI has no equivalent.
logprob_token_ids: Option<Vec<u32>>Report the log probability of these specific token IDs at each generated position, in addition to the sampled token.
allowed_token_ids: Option<Vec<u32>>Restrict decoding to exactly these token IDs.
bad_words: Option<Vec<String>>Substrings that must not appear in the output. Matching runs over the decoded text, so a bad word spanning a token boundary is still caught.
length_penalty: Option<f32>Exponential length penalty applied when scoring sequences. Only
meaningful for beam search; vLLM’s default is 1.0.
use_beam_search: Option<bool>Deprecated: beam search was removed from the vLLM engine. Accepted for wire compatibility only; do not rely on it.
watermarking: Option<bool>Embed an invisible watermark in the generated text. vLLM’s default is
true when the server was started with watermarking support.