pub struct ChatCompletionChunk {
pub id: String,
pub choices: Vec<CompletionChunkChoice>,
pub created: u64,
pub model: String,
pub object: Option<ChatCompletionChunkObject>,
pub service_tier: Option<ServiceTier>,
pub system_fingerprint: Option<String>,
pub usage: Option<CompletionUsage>,
pub moderation: Option<ChatModeration>,
pub prompt_token_ids: Option<Vec<u32>>,
pub prompt_text: Option<String>,
pub request_id: Option<String>,
}Fields§
§id: StringA unique identifier for the chat completion.
choices: Vec<CompletionChunkChoice>A list of chat completion choices. Can be more than one
if n is greater than 1. Empty for the final usage-only chunk
(see stream_options: {"include_usage": true}); some backends
send that chunk with "choices": null or omit the key entirely,
which deserializes as an empty list too.
created: u64The Unix timestamp (in seconds) of when the chat completion was created. Each chunk has the same timestamp.
model: StringThe model used for the chat completion.
object: Option<ChatCompletionChunkObject>The object type, which is always chat.completion.chunk.
Some only when the backend sends a recognized value; some
non-OpenAI gateways omit or repurpose the field.
service_tier: Option<ServiceTier>Specifies the processing type used for serving the request.
When the service_tier parameter is set, the response body will include the
service_tier value based on the processing mode actually used to serve the
request. This response value may be different from the value set in the
request parameter.
system_fingerprint: Option<String>This fingerprint represents the backend configuration that the model runs with.
Can be used in conjunction with the seed request parameter to understand when
backend changes have been made that might impact determinism.
usage: Option<CompletionUsage>An optional field that will only be present when you set
stream_options: {"include_usage": true} in your request. When present, it
contains a null value except for the last chunk which contains the token
usage statistics for the entire request.
NOTE: If the stream is interrupted or cancelled, you may not receive the final usage chunk which contains the total token usage for the request.
moderation: Option<ChatModeration>Moderation results for the request input and generated output.
Present on the moderation chunk when moderated completions are
requested via the moderation request parameter.
prompt_token_ids: Option<Vec<u32>>vLLM: the prompt’s token IDs after chat-template rendering. Sent on the first chunk only.
prompt_text: Option<String>vLLM: the fully rendered prompt text. Only sent on the first
chunk, and only when the request set vllm_chat.return_prompt_text.
request_id: Option<String>Z.ai / GLM: the request identifier, echoing the request’s
request_id or the one GLM generated. Not part of the OpenAI chunk
schema.