Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
rpi-voice
Hands-free voice conversation for the rpi Rust agent — assistant replies can be spoken aloud (TTS), and you can talk to the agent instead of typing (STT). Auto-TTS is opt-in and starts disabled.
Unlike a standalone voice-assistant binary, this is an rpi extension: it loads into the running TUI, so the normal text conversation stays intact and voice is layered on top.
What it does
- Auto-TTS — every assistant reply is synthesized with Microsoft Edge TTS (free, no API key) and played through the default output device. Replies are queued on one worker thread, so a multi-message turn is spoken in order rather than dropped while audio is playing.
- Voice input —
/voicerecords the microphone, stops automatically after a short trailing silence, transcribes with a Whisper-compatible API, and injects the text into the conversation as a user message. - Push-to-talk —
/voice pttdedicates a key (default space) to talking: hold it and the footer counts down, then becomes a live mic level meter (🎤 listening █████░░░ 1.1s); release and the utterance is transcribed and sent. The key is only claimed while the input box is empty. - Markdown-aware speech — code fences, backticks, links, headings and
emphasis markers are stripped before synthesis, so the voice never reads
**bold**or a URL. - Status in the TUI — the footer shows
voice: 🎤 recording…,voice: 📝 transcribing…, the injected text preview, and any error.
Commands
| Command | Effect |
|---|---|
/voice |
Record → transcribe → editable draft in the input box |
/voice auto [on|off] |
Continuous conversation: reply, then listen again |
/voice on / /voice off |
Enable / disable auto-TTS of replies |
| `/voice ptt [on | off]` |
/voice stop |
Stop the reply that is being read aloud |
| `/voice output [draft | send]` |
/voice status |
Auto-TTS state, current voice, recording/playing, STT key |
/voice set <name> |
Set the TTS voice for this session |
/voice model |
Show STT engine + local model status |
/voice model download |
Pre-fetch the embedded SenseVoice model |
/voice help |
Usage |
Environment
| Variable | Default | Purpose |
|---|---|---|
RPI_VOICE |
zh-CN-XiaoxiaoNeural |
TTS voice name |
RPI_VOICE_AUTO_TTS |
off |
Set on to start speaking assistant replies automatically |
OPENAI_API_KEY |
— | STT key (only required for the hosted OpenAI endpoint) |
RPI_STT_API_KEY |
— | STT key; takes precedence over OPENAI_API_KEY |
RPI_STT_API_BASE |
https://api.openai.com/v1 |
OpenAI-compatible STT base URL |
RPI_STT_MODEL |
whisper-1 |
STT model name (API backend) |
RPI_STT_ENGINE |
local |
local | auto | api; local SenseVoice is the default |
RPI_STT_MODEL_DIR |
~/.rpi/agent/models/sense-voice |
Embedded model dir (model.int8.onnx + tokens.txt) |
RPI_VOICE_STT_LANG |
auto | ISO-639-1 language hint for STT (e.g. zh) |
RPI_VOICE_RECORD_MS |
20000 |
Hard cap on a single recording |
RPI_VOICE_SILENCE_MS |
1200 |
Trailing silence that ends recording (0 disables) |
RPI_VOICE_MIN_SPEECH_MS |
500 |
Speech required before silence auto-stop |
RPI_VOICE_PTT_KEY |
space |
Key that drives /voice ptt |
RPI_VOICE_PTT_HOLD_MS |
600 |
Hold time before the mic opens; shorter holds are ignored |
RPI_VOICE_PTT_MAX_MS |
60000 |
Safety cap on one push-to-talk recording |
RPI_VOICE_OUTPUT |
draft |
draft (input box) | send (straight to the model) |
RPI_VOICE_DRAFT_MS |
2000 |
Auto-send countdown for a draft; 0 waits for a manual Enter |
RPI_VOICE_NO_SPEECH_MS |
unset | Optional continuous-mode speech-start timeout; auto waits like PTT by default |
RPI_VOICE_WARMUP_MS |
4000 |
Continuous mode: let a Bluetooth microphone wake before the speech window starts |
RPI_VOICE_AUTO_EMPTY_MAX |
3 |
Continuous mode: empty turns before it stops itself |
RPI_VOICE_INPUT_DEVICE |
system default | Case-insensitive substring of the capture device to use |
RPI_VOICE_INPUT_GAIN |
1.0 |
Software input gain (e.g. 20 for a very quiet mic), capped at 100 |
STT backends
STT uses the embedded SenseVoice model locally by default. On the first voice
input, missing model files are downloaded automatically to
~/.rpi/agent/models/sense-voice/ (or
RPI_CODING_AGENT_DIR/models/sense-voice/). Use /voice model download to
pre-fetch the model before recording.
Select another backend with RPI_STT_ENGINE:
RPI_STT_ENGINE |
Behaviour |
|---|---|
local (default) |
Use the embedded engine and download missing model files automatically |
auto |
Use local SenseVoice when available; fall back to the API if unavailable |
api |
Force the OpenAI-compatible endpoint |
Embedded offline model — SenseVoice (recommended)
The default build includes the embedded engine and runs entirely on-device via sherpa-onnx + SenseVoice (中文/英/日/韩/粤, auto language detection, punctuation). No key, no server, no network, no runtime DLL (the native libs are linked statically).
Model files live in ~/.rpi/agent/models/sense-voice/ ($RPI_CODING_AGENT_DIR/models
if set) and are downloaded automatically on first use — or up front with
/voice model download. ≈240 MB.
# build (the native libs are fetched at build time)
Building needs a C++ toolchain with cmake and libclang (bindgen) — build-time only, nothing at runtime. Cheapest way to get libclang without installing full LLVM:
# then point LIBCLANG_PATH at clang/native in site-packages
The sherpa-rs cdylib also needs the Windows system lib at link time:
# Windows
# Linux (static sherpa also needs this)
RPI_STT_MODEL_DIR overrides the model directory (must contain model.int8.onnx
and tokens.txt).
OpenAI-compatible endpoint
The client speaks the OpenAI /audio/transcriptions shape, so any compatible
server works. A key is only required for the hosted OpenAI endpoint — a local
server needs none:
# Free, fully local (faster-whisper): docker one-liner
Groq's free tier also works: RPI_STT_API_BASE=https://api.groq.com/openai/v1,
RPI_STT_MODEL=whisper-large-v3, plus a key.
Common Chinese voices: zh-CN-XiaoxiaoNeural (female, default),
zh-CN-YunxiNeural (male), zh-CN-XiaoyiNeural (lively).
# Windows
# macOS / Linux
Build
# API-only (default; no native STT deps)
# with embedded offline STT (see "Embedded offline model" above for env vars)
The built rpi_voice.dll/.so/.dylib is dropped into .rpi/extensions/ by
task install (or rpi-package-install) and loads automatically on the next
rpi start.
Push-to-talk
/voice ptt turns the configured key (default: space) into a hold-to-talk
button:
/voice ptt # on
/voice ptt off # off
- Hold the key → the footer counts down (
voice: ⏳ hold ▰▰▰▱▱▱ 0.4s to talk), so you can see the press registered and know to keep holding. - Once
RPI_VOICE_PTT_HOLD_MS(default 0.6s) elapses, the mic opens and the footer becomes a live level meter driven by your voice:voice: 🎤 listening █████░░░░░░░ 1.1s — release to send. A bar that never moves means the mic isn't hearing you. - The same meter is shown for every recording path — one-shot
/voiceand hands-free turns too (they end on silence, so their hint readssilence ends the turninstead ofrelease to send). - Release → recording stops, the audio is transcribed, and the text is
delivered according to
/voice output(see Where the text goes); the reply is then spoken by auto-TTS. - A hold shorter than the threshold never touches the microphone and leaves no trace in the footer.
Continuous conversation (/voice auto)
Hands-free turn-taking, in the style of a voice assistant: you talk, it answers
aloud, then it listens again. No key, no /voice — just keep talking.
/voice auto # on (starts listening immediately)
/voice auto off # off
you: why is this request failing?
🔊 voice: ♪ ▂▅▇▅▂▁▃ playing 3.4s ← reply spoken aloud
voice: 🎧 listening… ← mic reopens by itself
you: what if I make it synchronous?
…
The loop is closed by playback finishing, never by a timer: the microphone opens only after the speakers go quiet, so the assistant can never transcribe its own voice (the classic hands-free echo bug).
Getting out
| How | What happens |
|---|---|
| Start typing | The mode pauses itself — you've taken the keyboard. (Type to resume: /voice auto.) |
/voice auto off |
Exits now, aborting the listen in flight |
| Silence | After RPI_VOICE_AUTO_EMPTY_MAX (3) empty turns it stops on its own |
When it says "nothing heard"
The status carries the measured input peak, so the two very different causes are distinguishable at a glance:
voice: 🎧 nothing heard (peak 0.180) — still listening (1/3) ← mic works, no speech recognized
voice: 🎧 mic is silent (peak 0.000) — check the input device (1/3) ← nothing reached the mic
A peak at (or near) 0.000 means the microphone captured nothing at all — a
muted device, or the wrong input selected. A healthy peak with an empty
transcription means the audio arrived but STT found no words in it.
1. Which microphone is it even using?
/voice status names the active capture endpoint:
🎤 rpi-voice
...
Input device: 麦克风 (USB Audio Device)
A machine with a headset and a webcam has several capture endpoints, and the Windows default is often not the one you are talking into — so you get a recording of the room while you speak into a device nobody is listening to. Pin it explicitly (substring match, case-insensitive):
RPI_VOICE_INPUT_DEVICE=EDIFIER RPI_VOICE_INPUT_DEVICE="USB Audio"
2. Is the microphone actually delivering audio?
Lists every capture endpoint with its format, records 4 seconds from the one
that would be used, and reports what arrived — through the exact code path the
extension uses. Speak while it runs. It prints peak / rms / speech and
an explicit verdict.
3. Is the level simply too low?
A microphone can be perfectly healthy and still deliver almost nothing: a Windows
input level left near the bottom, or a mic sitting across the desk. Measured in
the wild — a working USB microphone whose speech peaked at 0.011 of full
scale (≈ -39 dBFS), which is 10–40× below a healthy capture. At that level the
voice never rises above the room, so the VAD says "nobody is talking" and STT
has almost nothing to work with.
Check the per-second level profile first (mic_probe_compares below). If every
second is ≈ 0.01 or less, either raise the level in Windows
(Settings → System → Sound → Input → your mic → Volume, plus any
"Microphone boost") or apply software gain:
RPI_VOICE_INPUT_GAIN=20
The boost is applied where the samples enter the recorder, so the meter, the VAD
and the STT input all see the same corrected signal — boosting after the VAD
would still leave the turn ending early. /voice status reports the gain that
was applied, so a level reading can always be read back against the raw capture.
4. Bluetooth headsets need ~4 seconds to wake up
A Bluetooth headset's microphone is a different profile (HFP) from its music playback (A2DP), and Windows takes 3–4 seconds to switch when an app starts capturing. A short probe therefore measures a device that has not woken yet and wrongly concludes it is dead.
Measured on an EDIFIER W820NB with a 200-second capture, per second:
second: 1 2 3 4 5 6 ... 11 12 13 14
peak: 0.00 0.00 0.00 0.72 0.68 0.44 ... 0.73 0.76 0.75 0.74
Seconds 1–3 are digital silence; the microphone comes alive at second 4. The
same speaker, captured simultaneously on a laptop's webcam microphone, never
exceeded 0.18 — a fifth of the level, because it sits across the room.
So a Bluetooth headset is usually the best input, provided the wait is
tolerated: the microphone first gets a wake-up grace period (RPI_VOICE_WARMUP_MS,
default 4s), then auto keeps the stream open like PTT until speech is detected
and trailing silence ends the turn. Set RPI_VOICE_NO_SPEECH_MS only when an
explicit speech-start timeout is desired.
5. Speech detection adapts to your mic
The VAD threshold is not a fixed level. It measures the ambient floor from quiet frames and requires speech to be ~3× that (≈ +10 dB SNR), with a low absolute minimum. A fixed threshold is what once made a perfectly good quiet microphone (ambient RMS ~30 against a hard floor of 260) look completely dead: every turn reported "heard nothing" while the hardware worked fine.
Push-to-talk keeps working while continuous mode is on. So does barge-in — speaking over a reply or typing cuts the audio.
Continuous mode needs replies to be spoken (that's the turn hand-off), so
/voice autoturns auto-TTS back on if it was muted, and says so.
Interrupting speech (barge-in)
Nothing is worse than an assistant that keeps reading while you talk over it, so any sign that you are taking the floor cuts playback off instantly:
- Start typing — the moment the prompt draft changes, the voice stops. This needs no recording and is the usual way out of a long reply.
- Hold the PTT key — the hold threshold stops playback before the mic opens, so the speakers can never feed back into the recording.
/voice— starting a dictation also stops playback, as do/voice off(mute) and/voice stop.
While speech plays, voice: in the footer becomes a music-style equalizer:
voice: ♪ ▂▅▇▅▂▁▃ playing 1.2s
The bar heights follow the real audio envelope — playback publishes the RMS of every ~60ms chunk it hands to the audio device — so the bars move with what you are actually hearing, and the crest ripples left → right like a level meter.
Barge-in on typing relies on the host emitting an
editorChangeevent. On an older host you still get push-to-talk and/voice stop; typing simply won't interrupt.
Where the text goes
By default a transcription is not sent straight to the model — it is placed in the prompt editor as an editable draft, with a countdown in the footer:
input: 你好,帮我看看这个 bug █
footer: Auto-send in 1.4s · any key to edit
- Leave it alone and it is submitted automatically after
RPI_VOICE_DRAFT_MS(default 2s). - Type anything to cancel the countdown — the draft stays put and the key you pressed becomes your first edit, so you can fix a mis-recognized word before it ever becomes a message. (Editing at all — typing, pasting, Backspace — also cancels it.)
- Press
Enterto send early. - Set
RPI_VOICE_DRAFT_MS=0to disable the countdown entirely and always wait forEnter.
Because the draft is submitted through the editor's normal path, it appears in
the transcript as a genuine user message. The old immediate path
(SendUserMessage) did not render a user bubble at all.
Use /voice output send (or RPI_VOICE_OUTPUT=send) when you would rather skip
the draft entirely and send the moment you release the key — faster, but with
no chance to correct a mis-recognition.
The key is only claimed while the input box is empty, so space still types
normally whenever you have a draft. Chords (Ctrl+Space, Alt+Space) are left
to the editor. When PTT is off the key is not claimed at all.
This needs a host that routes keys to extensions (rpi's register_shortcut +
Input events); on an older host the command still works but the key is never
delivered, so use /voice instead.
Notes / limitations
- Push-to-talk relies on the host delivering key release events, which the
terminal must report (true for Windows conhost and modern terminal emulators).
If presses register but releasing never ends the recording, the
RPI_VOICE_PTT_MAX_MScap still stops it. - Interrupting on typing needs a host that emits
editorChange; on an older host use push-to-talk or/voice stopinstead. - STT needs a reachable OpenAI-compatible endpoint: the hosted API (with a key),
a free local server, or the embedded offline engine (
--features local-stt). - Edge TTS is an undocumented Microsoft endpoint; if it changes, synthesis can break and the failure is logged to stderr.