1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
//! Provider-agnostic streaming chat abstraction with tool-use support.
//!
//! Why: trusty-memory and trusty-search both want to support more than one
//! upstream LLM (OpenRouter for cloud, Ollama / LM Studio for local). Rather
//! than each crate re-implementing the dispatch, we expose a small
//! [`ChatProvider`] trait plus concrete implementations and an auto-detector
//! for a running local model server. The trait also surfaces OpenAI-style
//! tool/function calling so downstream agents can let the model invoke tools
//! (search, memory recall, shell, etc.).
//!
//! What: defines the [`ChatProvider`] trait, [`ToolDef`] / [`ToolCall`] /
//! [`ChatEvent`] tool-use types, an [`OpenRouterProvider`] and an
//! [`OllamaProvider`] that both speak OpenAI-compatible
//! `/v1/chat/completions` with SSE streaming (including the streamed
//! `tool_calls` shape), a [`BedrockProvider`] that uses the AWS Bedrock
//! `Converse` API (behind the `bedrock` feature flag), and
//! [`auto_detect_local_provider`] which probes `{base_url}/v1/models` with a
//! 1-second timeout.
//!
//! Test: `cargo test -p trusty-common --features unconditional-only` covers
//! default config values, the unreachable-server path of
//! `auto_detect_local_provider`, SSE delta streaming, and accumulation of
//! streamed tool-call fragments.
pub use ;
pub use ;
// Re-expose the bedrock_impl module as `bedrock_provider` so downstream
// crates can access constants (e.g. `DEFAULT_BEDROCK_MODEL`) without needing
// to depend on the bedrock feature themselves.
pub use BedrockProvider;
// Stub constant so code that references DEFAULT_BEDROCK_MODEL compiles without
// the bedrock feature. Must stay in sync with bedrock_impl::DEFAULT_BEDROCK_MODEL.
// Claude Sonnet 4.6 drops the date stamp and -v1:0 suffix (verified vs AWS docs).
pub const DEFAULT_BEDROCK_MODEL: &str = "us.anthropic.claude-sonnet-4-6";
use crateChatMessage;
use Result;
use async_trait;
use ;
use Sender;
// ── Public re-exports so callers get the full surface from `chat::*` ──────────
/// Configuration for a local OpenAI-compatible model server (Ollama, LM
/// Studio, llama.cpp's server, etc.).
///
/// Why: callers want a single struct they can deserialize from config files
/// and pass to [`auto_detect_local_provider`] without juggling defaults.
/// What: holds an enable flag, the server's base URL (no trailing slash),
/// and the default model to request. Defaults target Ollama's standard
/// localhost binding.
/// Test: `local_model_config_defaults` asserts the default values.
// ─── Tool-use types ───────────────────────────────────────────────────────────
/// JSON-Schema description of a callable tool, in OpenAI function-calling
/// shape.
///
/// Why: downstream agents (trusty-memory, trusty-search) expose tools like
/// `memory_recall` or `web_search` to the LLM. The OpenAI tool format is the
/// de-facto common denominator across OpenRouter, Ollama, LM Studio, and
/// most cloud providers.
/// What: `name` and `description` are passed verbatim; `parameters` is a
/// JSON Schema object (typically `{"type":"object","properties":{...}}`).
/// Test: `tool_def_serializes_as_function` checks the wire shape.
/// A tool invocation the model wants the host to perform.
///
/// Why: the streaming chat API emits `tool_calls` in fragments — first an
/// `id` + `function.name`, then a string of `function.arguments` deltas.
/// We accumulate fragments and surface one fully-formed [`ToolCall`] per
/// invocation to the caller.
/// What: `id` is the upstream's call id (echoed back in subsequent
/// `role:"tool"` messages); `name` is the function name; `arguments` is a
/// JSON string (NOT a parsed value — many models emit malformed JSON and
/// callers want the raw text for error reporting / repair).
/// Test: `accumulates_streamed_tool_call_fragments`.
/// Sampling knobs forwarded to an OpenAI-compatible provider.
///
/// Why (issue #3758): the streaming request wire previously sent only
/// `model`/`messages`/`tools`, so a streamed reply's style, verbosity, and
/// stopping behaviour were NOT equivalent to the blocking path's for the same
/// turn — the caller silently got provider defaults instead of its configured
/// temperature, token ceiling, and stop sequences. Carrying them in one struct
/// (rather than three provider constructor arguments) keeps `new()` stable for
/// existing callers and lets the set grow without churning call sites.
/// What: every field is optional — `None`/empty means "omit from the request
/// body and let the provider default apply", which is exactly the pre-#3758
/// behaviour, so a caller that does not opt in is unaffected. `stop` maps to
/// the OpenAI `stop` array (the direct-Anthropic dialect calls the same thing
/// `stop_sequences`).
/// Test: `sampling_params_serialize_into_request_body`,
/// `default_sampling_omits_fields`.
/// Token-usage tally surfaced at the end of a streamed response.
///
/// Why (issue #3767): Bedrock's `ConverseStream` reports usage only once, in
/// a terminal `metadata` event — unlike the per-token deltas, it has no
/// natural home in [`ChatEvent::Delta`]. Without a dedicated variant a
/// streaming provider has no way to hand token counts back to the caller, so
/// a downstream cost-reporting consumer would silently see zero usage for
/// every streamed call. Kept as a small standalone struct (rather than
/// reusing `trusty_common::inference::types::Usage`) so the `chat` module —
/// which compiles unconditionally, unlike the `inference-client`-gated
/// `types` module — never depends on a feature-gated type.
/// What: four token buckets mirroring the shape every provider in this
/// workspace already reports (prompt/completion split, plus the two
/// prompt-cache buckets). The four fields are SEPARATE and ADDITIVE — not
/// nested — matching both the Bedrock `TokenUsage` wire shape
/// (`input_tokens`/`output_tokens`/`cache_read_input_tokens`/
/// `cache_write_input_tokens` are independent fields, verified against
/// `aws-sdk-bedrockruntime` 1.132.0) and this workspace's existing
/// convention (`crate::perf::TokenUsage`,
/// `trusty-agents/src/llm/anthropic_native/mod.rs`'s usage parsing, and
/// `chat::bedrock_impl::parse_usage`'s non-streaming counterpart). A
/// provider that has none of the cache fields simply leaves them zero.
/// Test: `bedrock_stream_reports_usage_from_metadata_event` (bedrock_impl).
/// Streaming chat event.
///
/// Why: replaces the previous "string-only" channel so callers can
/// distinguish text deltas from tool invocations and from terminal
/// success/error without parsing magic markers out of the text stream.
/// What: `Delta` is a content chunk; `ToolCall` is a fully-accumulated tool
/// invocation; `Usage` carries the call's token tally (emitted zero or one
/// times, whenever the provider's wire format reports it — see
/// [`ChatUsage`]); `Done` signals the upstream stream terminated normally;
/// `Error` carries a human-readable message for stream-mid failures (the
/// provider also returns `Err` from `chat_stream`, but `Error` lets the
/// caller display partial-stream failures inline).
/// Test: `ollama_provider_streams_sse_deltas`.
///
/// # Stability
///
/// `#[non_exhaustive]`: adding a variant here is otherwise a SemVer-breaking
/// change for every downstream crate, because an exhaustive `match` over the
/// old variant set stops compiling with E0004. That is not hypothetical — the
/// `Usage` variant forced arm additions in five consumers at once
/// (`trusty-agents`, `trusty-analyze`, `trusty-memory`, `trusty-mpm`,
/// `trusty-search`) and is the direct cause of the 0.27.0 MINOR bump: a patch
/// release would have re-resolved the already-published, arm-less consumer
/// sources against the new variant and hard-failed `cargo install` (the same
/// failure that forced `trusty-analyze` 0.7.3 to be yanked).
///
/// Downstream matches must therefore carry a wildcard arm. This attribute does
/// NOT retroactively fix consumers published before it landed — it only stops
/// the next variant addition from repeating the break.
/// Streaming chat provider abstraction.
///
/// Why: downstream crates (trusty-memory, trusty-search) want to support
/// multiple LLM backends without hard-coding which one to call. Providers
/// expose a uniform streaming interface so the caller can swap them at
/// runtime based on configuration / availability.
/// What: implementors stream [`ChatEvent`]s into `tx`. Pass an empty
/// `tools` vec to disable tool use entirely (the provider MUST then omit
/// the `tools` field from the upstream request — some models error on an
/// empty array). Returning `Ok(())` means the stream completed normally;
/// the caller should also expect a final [`ChatEvent::Done`].
/// Test: implementations are covered by their own unit tests in this
/// module plus integration tests in downstream crates.