rpi-voice 0.2.2

Voice conversation extension for rpi — offline STT + auto-TTS replies
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
284
285
286
287
288
289
290
291
292
293
294
295
296
297
298
299
300
301
302
303
304
305
306
307
308
309
310
311
312
313
314
315
316
317
318
319
320
321
322
323
324
325
326
327
328
329
330
331
332
333
334
335
336
337
338
339
340
341
342
343
344
345
346
347
348
349
350
351
352
353
354
355
356
357
358
359
360
361
362
363
364
365
366
367
368
369
370
371
372
373
374
375
376
377
378
379
380
381
382
383
384
385
386
387
388
389
390
391
392
393
394
395
396
397
398
399
400
401
402
403
404
# rpi-voice

Hands-free voice conversation for the rpi Rust agent — assistant replies can be
spoken aloud (TTS), and you can talk to the agent instead of typing (STT).
Auto-TTS is opt-in and starts disabled.

Unlike a standalone voice-assistant binary, this is an **rpi extension**: it
loads into the running TUI, so the normal text conversation stays intact and
voice is layered on top.

## What it does

- **Auto-TTS** — every assistant reply is synthesized with Microsoft Edge TTS
  (free, no API key) and played through the default output device. Replies are
  queued on one worker thread, so a multi-message turn is spoken in order
  rather than dropped while audio is playing.
- **Voice input** — `/voice` records the microphone, stops automatically after a
  short trailing silence, transcribes with a Whisper-compatible API, and injects
  the text into the conversation as a user message.
- **Push-to-talk** — `/voice ptt` dedicates a key (default space) to talking:
  hold it and the footer counts down, then becomes a live mic level meter
  (`🎤 listening █████░░░ 1.1s`); release and the utterance is transcribed and
  sent. The key is only claimed while the input box is empty.
- **Markdown-aware speech** — code fences, backticks, links, headings and
  emphasis markers are stripped before synthesis, so the voice never reads
  `**bold**` or a URL.
- **Status in the TUI** — the footer shows `voice: 🎤 recording…`,
  `voice: 📝 transcribing…`, the injected text preview, and any error.

## Commands

| Command | Effect |
| --- | --- |
| `/voice` | Record → transcribe → editable draft in the input box |
| `/voice auto [on\|off]` | Continuous conversation: reply, then listen again |
| `/voice on` / `/voice off` | Enable / disable auto-TTS of replies |
| `/voice ptt [on|off]` | Push-to-talk: hold the key, release to deliver |
| `/voice stop` | Stop the reply that is being read aloud |
| `/voice output [draft|send]` | Where a transcription goes (default `draft`) |
| `/voice status` | Auto-TTS state, current voice, recording/playing, STT key |
| `/voice set <name>` | Set the TTS voice for this session |
| `/voice model` | Show STT engine + local model status |
| `/voice model download` | Pre-fetch the embedded SenseVoice model |
| `/voice help` | Usage |

## Environment

| Variable | Default | Purpose |
| --- | --- | --- |
| `RPI_VOICE` | `zh-CN-XiaoxiaoNeural` | TTS voice name |
| `RPI_VOICE_AUTO_TTS` | `off` | Set `on` to start speaking assistant replies automatically |
| `OPENAI_API_KEY` | — | STT key (only required for the hosted OpenAI endpoint) |
| `RPI_STT_API_KEY` | — | STT key; takes precedence over `OPENAI_API_KEY` |
| `RPI_STT_API_BASE` | `https://api.openai.com/v1` | OpenAI-compatible STT base URL |
| `RPI_STT_MODEL` | `whisper-1` | STT model name (API backend) |
| `RPI_STT_ENGINE` | `local` | `local` \| `auto` \| `api`; local SenseVoice is the default |
| `RPI_STT_MODEL_DIR` | `~/.rpi/agent/models/sense-voice` | Embedded model dir (`model.int8.onnx` + `tokens.txt`) |
| `RPI_VOICE_STT_LANG` | auto | ISO-639-1 language hint for STT (e.g. `zh`) |
| `RPI_VOICE_RECORD_MS` | `20000` | Hard cap on a single recording |
| `RPI_VOICE_SILENCE_MS` | `1200` | Trailing silence that ends recording (`0` disables) |
| `RPI_VOICE_MIN_SPEECH_MS` | `500` | Speech required before silence auto-stop |
| `RPI_VOICE_PTT_KEY` | `space` | Key that drives `/voice ptt` |
| `RPI_VOICE_PTT_HOLD_MS` | `600` | Hold time before the mic opens; shorter holds are ignored |
| `RPI_VOICE_PTT_MAX_MS` | `60000` | Safety cap on one push-to-talk recording |
| `RPI_VOICE_OUTPUT` | `draft` | `draft` (input box) \| `send` (straight to the model) |
| `RPI_VOICE_DRAFT_MS` | `2000` | Auto-send countdown for a draft; `0` waits for a manual Enter |
| `RPI_VOICE_NO_SPEECH_MS` | unset | Optional continuous-mode speech-start timeout; auto waits like PTT by default |
| `RPI_VOICE_WARMUP_MS` | `4000` | Continuous mode: let a Bluetooth microphone wake before the speech window starts |
| `RPI_VOICE_AUTO_EMPTY_MAX` | `3` | Continuous mode: empty turns before it stops itself |
| `RPI_VOICE_INPUT_DEVICE` | system default | Case-insensitive substring of the capture device to use |
| `RPI_VOICE_INPUT_GAIN` | `1.0` | Software input gain (e.g. `20` for a very quiet mic), capped at 100 |

### STT backends

STT uses the embedded SenseVoice model locally by default. On the first voice
input, missing model files are downloaded automatically to
`~/.rpi/agent/models/sense-voice/` (or
`RPI_CODING_AGENT_DIR/models/sense-voice/`). Use `/voice model download` to
pre-fetch the model before recording.

Select another backend with `RPI_STT_ENGINE`:

| `RPI_STT_ENGINE` | Behaviour |
| --- | --- |
| `local` (default) | Use the embedded engine and download missing model files automatically |
| `auto` | Use local SenseVoice when available; fall back to the API if unavailable |
| `api` | Force the OpenAI-compatible endpoint |

#### Embedded offline model — SenseVoice (recommended)

The default build includes the embedded engine and runs entirely on-device via
[sherpa-onnx](https://github.com/k2-fsa/sherpa-onnx) + **SenseVoice** (中文/英/日/韩/粤,
auto language detection, punctuation). No key, no server, no network, no runtime
DLL (the native libs are linked statically).

Model files live in `~/.rpi/agent/models/sense-voice/` (`$RPI_CODING_AGENT_DIR/models`
if set) and are downloaded automatically on first use — or up front with
`/voice model download`. ≈240 MB.

```bash
# build (the native libs are fetched at build time)
cargo build -p rpi-voice --release --features local-stt
```

Building needs a C++ toolchain with **cmake** and **libclang** (bindgen) — build-time
only, nothing at runtime. Cheapest way to get libclang without installing full LLVM:

```bash
pip install libclang
# then point LIBCLANG_PATH at clang/native in site-packages
```

The `sherpa-rs` cdylib also needs the Windows system lib at link time:

```bash
# Windows
export LIBCLANG_PATH="$(python -c 'import clang,os;print(os.path.dirname(clang.__file__))')/native"
export RUSTFLAGS="-C link-arg=advapi32.lib"
cargo build -p rpi-voice --release --features local-stt

# Linux (static sherpa also needs this)
export RUSTFLAGS="-C relocation-model=dynamic-no-pic"
cargo build -p rpi-voice --release --features local-stt
```

`RPI_STT_MODEL_DIR` overrides the model directory (must contain `model.int8.onnx`
and `tokens.txt`).

#### OpenAI-compatible endpoint

The client speaks the OpenAI `/audio/transcriptions` shape, so any compatible
server works. A key is **only** required for the hosted OpenAI endpoint — a local
server needs none:

```bash
# Free, fully local (faster-whisper): docker one-liner
docker run -p 8000:8000 ghcr.io/speaches-ai/speaches:latest-cpu
export RPI_STT_ENGINE=api
export RPI_STT_API_BASE=http://localhost:8000/v1
export RPI_STT_MODEL=Systran/faster-whisper-small
```

Groq's free tier also works: `RPI_STT_API_BASE=https://api.groq.com/openai/v1`,
`RPI_STT_MODEL=whisper-large-v3`, plus a key.

Common Chinese voices: `zh-CN-XiaoxiaoNeural` (female, default),
`zh-CN-YunxiNeural` (male), `zh-CN-XiaoyiNeural` (lively).

```bash
# Windows
set OPENAI_API_KEY=sk-xxx
# macOS / Linux
export OPENAI_API_KEY=sk-xxx
```

## Build

```bash
# API-only (default; no native STT deps)
cargo build -p rpi-voice --release

# with embedded offline STT (see "Embedded offline model" above for env vars)
cargo build -p rpi-voice --release --features local-stt
```

The built `rpi_voice.dll`/`.so`/`.dylib` is dropped into `.rpi/extensions/` by
`task install` (or `rpi-package-install`) and loads automatically on the next
rpi start.

## Push-to-talk

`/voice ptt` turns the configured key (default: **space**) into a hold-to-talk
button:

```
/voice ptt          # on
/voice ptt off      # off
```

- Hold the key → the footer counts down (`voice: ⏳ hold ▰▰▰▱▱▱ 0.4s to talk`),
  so you can see the press registered and know to keep holding.
- Once `RPI_VOICE_PTT_HOLD_MS` (default **0.6s**) elapses, the mic opens and the
  footer becomes a live level meter driven by your voice:
  `voice: 🎤 listening █████░░░░░░░ 1.1s — release to send`. A bar that never
  moves means the mic isn't hearing you.
- The same meter is shown for **every** recording path — one-shot `/voice` and
  hands-free turns too (they end on silence, so their hint reads `silence ends
  the turn` instead of `release to send`).
- Release → recording stops, the audio is transcribed, and the text is
  delivered according to `/voice output` (see **Where the text goes**); the
  reply is then spoken by auto-TTS.
- A hold shorter than the threshold never touches the microphone and leaves no
  trace in the footer.

## Continuous conversation (`/voice auto`)

Hands-free turn-taking, in the style of a voice assistant: you talk, it answers
aloud, then it listens again. No key, no `/voice` — just keep talking.

```
/voice auto         # on (starts listening immediately)
/voice auto off     # off
```

```
you:  why is this request failing?
   🔊 voice: ♪ ▂▅▇▅▂▁▃ playing 3.4s      ← reply spoken aloud
   voice: 🎧 listening…                   ← mic reopens by itself
you:  what if I make it synchronous?
   …
```

The loop is closed by **playback finishing**, never by a timer: the microphone
opens only after the speakers go quiet, so the assistant can never transcribe
its own voice (the classic hands-free echo bug).

### Getting out

| How | What happens |
| --- | --- |
| **Start typing** | The mode pauses itself — you've taken the keyboard. (Type to resume: `/voice auto`.) |
| `/voice auto off` | Exits now, aborting the listen in flight |
| Silence | After `RPI_VOICE_AUTO_EMPTY_MAX` (3) empty turns it stops on its own |

### When it says "nothing heard"

The status carries the **measured input peak**, so the two very different causes
are distinguishable at a glance:

```
voice: 🎧 nothing heard (peak 0.180) — still listening (1/3)   ← mic works, no speech recognized
voice: 🎧 mic is silent (peak 0.000) — check the input device (1/3)   ← nothing reached the mic
```

A peak at (or near) `0.000` means the microphone captured nothing at all — a
muted device, or the wrong input selected. A healthy peak with an empty
transcription means the audio arrived but STT found no words in it.

#### 1. Which microphone is it even using?

`/voice status` names the active capture endpoint:

```
🎤 rpi-voice
  ...
  Input device: 麦克风 (USB Audio Device)
```

A machine with a headset *and* a webcam has several capture endpoints, and the
Windows default is often **not** the one you are talking into — so you get a
recording of the room while you speak into a device nobody is listening to. Pin
it explicitly (substring match, case-insensitive):

```bash
RPI_VOICE_INPUT_DEVICE=EDIFIER rpi     # talk into the headset
RPI_VOICE_INPUT_DEVICE="USB Audio" rpi
```

#### 2. Is the microphone actually delivering audio?

```bash
cargo test -p rpi-voice --lib -- --ignored --nocapture mic_probe
```

Lists every capture endpoint with its format, records 4 seconds from the one
that would be used, and reports what arrived — through the exact code path the
extension uses. **Speak while it runs.** It prints `peak` / `rms` / `speech` and
an explicit verdict.

#### 3. Is the level simply too low?

A microphone can be perfectly healthy and still deliver almost nothing: a Windows
input level left near the bottom, or a mic sitting across the desk. Measured in
the wild — a working USB microphone whose *speech* peaked at `0.011` of full
scale (≈ -39 dBFS), which is 10–40× below a healthy capture. At that level the
voice never rises above the room, so the VAD says "nobody is talking" and STT
has almost nothing to work with.

Check the per-second level profile first (`mic_probe_compares` below). If every
second is ≈ `0.01` or less, either raise the level in Windows
(**Settings → System → Sound → Input →** your mic **→ Volume**, plus any
"Microphone boost") or apply software gain:

```bash
RPI_VOICE_INPUT_GAIN=20 rpi
```

The boost is applied where the samples enter the recorder, so the meter, the VAD
and the STT input all see the same corrected signal — boosting after the VAD
would still leave the turn ending early. `/voice status` reports the gain that
was applied, so a level reading can always be read back against the raw capture.

#### 4. Bluetooth headsets need ~4 seconds to wake up

A Bluetooth headset's microphone is a **different profile (HFP)** from its music
playback (A2DP), and Windows takes **3–4 seconds** to switch when an app starts
capturing. A short probe therefore measures a device that has not woken yet and
wrongly concludes it is dead.

Measured on an EDIFIER W820NB with a 200-second capture, per second:

```
second:  1     2     3     4     5     6   ...  11    12    13    14
peak:   0.00  0.00  0.00  0.72  0.68  0.44 ...  0.73  0.76  0.75  0.74
```

Seconds 1–3 are digital silence; the microphone comes alive at second 4. The
same speaker, captured simultaneously on a laptop's webcam microphone, never
exceeded `0.18` — a fifth of the level, because it sits across the room.

So a Bluetooth headset is usually the **best** input, provided the wait is
tolerated: the microphone first gets a wake-up grace period (`RPI_VOICE_WARMUP_MS`,
default 4s), then auto keeps the stream open like PTT until speech is detected
and trailing silence ends the turn. Set `RPI_VOICE_NO_SPEECH_MS` only when an
explicit speech-start timeout is desired.

#### 5. Speech detection adapts to your mic

The VAD threshold is **not** a fixed level. It measures the ambient floor from
quiet frames and requires speech to be ~3× that (≈ +10 dB SNR), with a low
absolute minimum. A fixed threshold is what once made a perfectly good quiet
microphone (ambient RMS ~30 against a hard floor of 260) look completely dead:
every turn reported "heard nothing" while the hardware worked fine.

Push-to-talk keeps working while continuous mode is on. So does barge-in —
speaking over a reply or typing cuts the audio.

> Continuous mode needs replies to be **spoken** (that's the turn hand-off), so
> `/voice auto` turns auto-TTS back on if it was muted, and says so.

## Interrupting speech (barge-in)

Nothing is worse than an assistant that keeps reading while you talk over it,
so any sign that you are taking the floor cuts playback off instantly:

- **Start typing** — the moment the prompt draft changes, the voice stops. This
  needs no recording and is the usual way out of a long reply.
- **Hold the PTT key** — the hold threshold stops playback *before* the mic
  opens, so the speakers can never feed back into the recording.
- **`/voice`** — starting a dictation also stops playback, as do `/voice off`
  (mute) and `/voice stop`.

While speech plays, `voice:` in the footer becomes a music-style equalizer:

```
voice: ♪ ▂▅▇▅▂▁▃ playing 1.2s
```

The bar heights follow the **real audio envelope** — playback publishes the RMS
of every ~60ms chunk it hands to the audio device — so the bars move with what
you are actually hearing, and the crest ripples left → right like a level meter.

> Barge-in on typing relies on the host emitting an `editorChange` event. On an
> older host you still get push-to-talk and `/voice stop`; typing simply won't
> interrupt.

## Where the text goes

By default a transcription is **not** sent straight to the model — it is placed
in the prompt editor as an editable *draft*, with a countdown in the footer:

```
input: 你好,帮我看看这个 bug █
footer: Auto-send in 1.4s · any key to edit
```

- Leave it alone and it is submitted automatically after `RPI_VOICE_DRAFT_MS`
  (default **2s**).
- **Type anything** to cancel the countdown — the draft stays put and the key
you pressed becomes your first edit, so you can fix a mis-recognized word
  before it ever becomes a message. (Editing at all — typing, pasting,
  Backspace — also cancels it.)
- Press `Enter` to send early.
- Set `RPI_VOICE_DRAFT_MS=0` to disable the countdown entirely and always wait
  for `Enter`.

Because the draft is submitted through the editor's normal path, it appears in
the transcript as a genuine **user message**. The old immediate path
(`SendUserMessage`) did not render a user bubble at all.

Use `/voice output send` (or `RPI_VOICE_OUTPUT=send`) when you would rather skip
the draft entirely and send the moment you release the key — faster, but with
no chance to correct a mis-recognition.

The key is only claimed while the **input box is empty**, so space still types
normally whenever you have a draft. Chords (`Ctrl+Space`, `Alt+Space`) are left
to the editor. When PTT is off the key is not claimed at all.

This needs a host that routes keys to extensions (rpi's `register_shortcut` +
`Input` events); on an older host the command still works but the key is never
delivered, so use `/voice` instead.

## Notes / limitations

- Push-to-talk relies on the host delivering key **release** events, which the
  terminal must report (true for Windows conhost and modern terminal emulators).
  If presses register but releasing never ends the recording, the
  `RPI_VOICE_PTT_MAX_MS` cap still stops it.
- Interrupting on *typing* needs a host that emits `editorChange`; on an older
  host use push-to-talk or `/voice stop` instead.
- STT needs a reachable OpenAI-compatible endpoint: the hosted API (with a key),
  a free local server, or the embedded offline engine (`--features local-stt`).
- Edge TTS is an undocumented Microsoft endpoint; if it changes, synthesis can
  break and the failure is logged to stderr.