pub struct Choice {
pub finish_reason: FinishReason,
pub index: u32,
pub logprobs: Option<ChoiceLogprobs>,
pub message: ChatCompletionMessage,
pub stop_reason: Option<StopReason>,
pub token_ids: Option<Vec<u32>>,
pub routed_experts: Option<String>,
}Fields§
§finish_reason: FinishReasonThe reason the model stopped generating tokens.
This will be stop if the model hit a natural stop point or a provided stop
sequence, length if the maximum number of tokens specified in the request was
reached, content_filter if content was omitted due to a flag from our content
filters, tool_calls if the model called a tool, or function_call
(deprecated) if the model called a function.
index: u32The index of the choice in the list of choices.
logprobs: Option<ChoiceLogprobs>Log probability information for the choice.
message: ChatCompletionMessageA chat completion message generated by the model.
stop_reason: Option<StopReason>vLLM: which terminator ended generation — the matched stop string, or
the matched token ID. Not part of the OpenAI schema; finish_reason
alone only reports stop for both cases.
token_ids: Option<Vec<u32>>vLLM: the generated token IDs, for tracing tokens in agent scenarios.
Only set when the request set vllm_chat.return_token_ids.
routed_experts: Option<String>vLLM: per-token expert routing decisions for mixture-of-experts
models, as base64-encoded NumPy .npy bytes of shape
(num_tokens - 1, num_layers, num_experts_per_tok).
Only set when the server runs with --enable-return-routed-experts.