1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
// Copyright 2026 Thomas Santerre and Moderately AI Inc.
//
// SPDX-License-Identifier: MIT OR Apache-2.0
//! Language model response types.
/// Normalized reason the model stopped generating.
///
/// Maps the per-provider stop-reason strings into one shared enum so
/// downstream consumers can distinguish
/// "ran out of room" (`MaxTokens`) from "finished naturally" (`EndTurn`)
/// without knowing which provider produced the response. Without this,
/// a truncated completion might otherwise surface as a parse error with no
/// hint that the model ran out of `max_tokens`.
/// A response from a language model generation call.
/// An incremental chunk from a streaming language model response.
/// Token usage statistics for a generation call.
///
/// Under prompt caching the provider splits the prompt: `input_tokens`
/// counts only the *uncached* remainder, while the cached prefix is
/// reported separately in `cache_creation_input_tokens` (a cache write
/// on a miss) and `cache_read_input_tokens` (a cache hit). The full
/// prompt size is therefore `input_tokens + cache_creation_input_tokens
/// + cache_read_input_tokens` — summing only `input_tokens` undercounts
/// once caching engages. Providers that don't report caching leave both
/// cache fields `0`.