1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
use async_trait;
use crate*;
// ---------------------------------------------------------------------------
// ASR Adapter
// ---------------------------------------------------------------------------
/// Turns audio into text.
///
/// Implementors may be local (Whisper.cpp, Parakeet via ONNX) or
/// cloud-based (OpenAI Whisper API, Deepgram). The pipeline treats them
/// identically.
///
/// The whole utterance is passed at once and one transcript comes back.
/// An adapter that can also emit partial hypotheses as audio arrives
/// will implement an additional streaming trait alongside this one —
/// that capability is not part of 0.1.0, and the shape here is
/// deliberately the one every backend can satisfy.
/// Why an [`AsrAdapter`] could not produce a transcript.
///
/// The variants are the distinctions a caller can act on: a missing
/// model needs a different response than a runtime failure, and neither
/// is the same as the user simply not having spoken. Marked
/// `#[non_exhaustive]` so that finer distinctions can be added later
/// without breaking a `match`.
// ---------------------------------------------------------------------------
// Context Provider
// ---------------------------------------------------------------------------
/// Captures a snapshot of the current OS / application context.
///
/// Implementors call into platform-specific APIs: macOS Accessibility
/// (AXUIElement), Windows UI Automation, Linux AT-SPI, or a manual
/// provider for testing.
// ---------------------------------------------------------------------------
// LLM Refiner
// ---------------------------------------------------------------------------
/// Takes raw ASR text + application context and produces refined output.
///
/// Implementors may call cloud LLMs (Cerebras, Groq, OpenAI) or on-device
/// models (Apple Foundation Models, Gemini Nano, Ollama).
/// Why an [`LlmRefiner`] could not refine its input.
///
/// The pipeline degrades gracefully on any of these — it emits the
/// unrefined text rather than failing the session — so the distinction
/// exists for logging and for callers driving a refiner directly.
// ---------------------------------------------------------------------------
// Output Emitter
// ---------------------------------------------------------------------------
/// Delivers the final pipeline output to the OS / application.
///
/// Implementors handle clipboard insertion, key emulation, stdout, callbacks,
/// or any other output mechanism.