rpi-voice 0.2.1

Voice conversation extension for rpi — offline STT + auto-TTS replies
docs.rs failed to build rpi-voice-0.2.1
Please check the build logs for more information.
See Builds for ideas on how to fix a failed build, or Metadata for how to configure docs.rs builds.
If you believe this is docs.rs' fault, open an issue.
Visit the last successful build: rpi-voice-0.2.0

rpi-voice

Hands-free voice conversation for the rpi Rust agent — assistant replies can be spoken aloud (TTS), and you can talk to the agent instead of typing (STT). Auto-TTS is opt-in and starts disabled.

Unlike a standalone voice-assistant binary, this is an rpi extension: it loads into the running TUI, so the normal text conversation stays intact and voice is layered on top.

What it does

  • Auto-TTS — every assistant reply is synthesized with Microsoft Edge TTS (free, no API key) and played through the default output device. Replies are queued on one worker thread, so a multi-message turn is spoken in order rather than dropped while audio is playing.
  • Voice input — /voice records the microphone, stops automatically after a short trailing silence, transcribes with a Whisper-compatible API, and injects the text into the conversation as a user message.
  • Push-to-talk — /voice ptt dedicates a key (default space) to talking: hold it and the footer counts down, then becomes a live mic level meter (🎤 listening █████░░░ 1.1s); release and the utterance is transcribed and sent. The key is only claimed while the input box is empty.
  • Markdown-aware speech — code fences, backticks, links, headings and emphasis markers are stripped before synthesis, so the voice never reads **bold** or a URL.
  • Status in the TUI — the footer shows voice: 🎤 recording…, voice: 📝 transcribing…, the injected text preview, and any error.

Commands

Command Effect
/voice Record → transcribe → editable draft in the input box
/voice auto [on|off] Continuous conversation: reply, then listen again
/voice on / /voice off Enable / disable auto-TTS of replies
`/voice ptt [on off]`
/voice stop Stop the reply that is being read aloud
`/voice output [draft send]`
/voice status Auto-TTS state, current voice, recording/playing, STT key
/voice set <name> Set the TTS voice for this session
/voice model Show STT engine + local model status
/voice model download Pre-fetch the embedded SenseVoice model
/voice help Usage

Environment

Variable Default Purpose
RPI_VOICE zh-CN-XiaoxiaoNeural TTS voice name
RPI_VOICE_AUTO_TTS off Set on to start speaking assistant replies automatically
OPENAI_API_KEY — STT key (only required for the hosted OpenAI endpoint)
RPI_STT_API_KEY — STT key; takes precedence over OPENAI_API_KEY
RPI_STT_API_BASE https://api.openai.com/v1 OpenAI-compatible STT base URL
RPI_STT_MODEL whisper-1 STT model name (API backend)
RPI_STT_ENGINE local local | auto | api; local SenseVoice is the default
RPI_STT_MODEL_DIR ~/.rpi/agent/models/sense-voice Embedded model dir (model.int8.onnx + tokens.txt)
RPI_VOICE_STT_LANG auto ISO-639-1 language hint for STT (e.g. zh)
RPI_VOICE_RECORD_MS 20000 Hard cap on a single recording
RPI_VOICE_SILENCE_MS 1200 Trailing silence that ends recording (0 disables)
RPI_VOICE_MIN_SPEECH_MS 500 Speech required before silence auto-stop
RPI_VOICE_PTT_KEY space Key that drives /voice ptt
RPI_VOICE_PTT_HOLD_MS 600 Hold time before the mic opens; shorter holds are ignored
RPI_VOICE_PTT_MAX_MS 60000 Safety cap on one push-to-talk recording
RPI_VOICE_OUTPUT draft draft (input box) | send (straight to the model)
RPI_VOICE_DRAFT_MS 2000 Auto-send countdown for a draft; 0 waits for a manual Enter
RPI_VOICE_NO_SPEECH_MS unset Optional continuous-mode speech-start timeout; auto waits like PTT by default
RPI_VOICE_WARMUP_MS 4000 Continuous mode: let a Bluetooth microphone wake before the speech window starts
RPI_VOICE_AUTO_EMPTY_MAX 3 Continuous mode: empty turns before it stops itself
RPI_VOICE_INPUT_DEVICE system default Case-insensitive substring of the capture device to use
RPI_VOICE_INPUT_GAIN 1.0 Software input gain (e.g. 20 for a very quiet mic), capped at 100

STT backends

STT uses the embedded SenseVoice model locally by default. On the first voice input, missing model files are downloaded automatically to ~/.rpi/agent/models/sense-voice/ (or RPI_CODING_AGENT_DIR/models/sense-voice/). Use /voice model download to pre-fetch the model before recording.

Select another backend with RPI_STT_ENGINE:

RPI_STT_ENGINE Behaviour
local (default) Use the embedded engine and download missing model files automatically
auto Use local SenseVoice when available; fall back to the API if unavailable
api Force the OpenAI-compatible endpoint

Embedded offline model — SenseVoice (recommended)

The default build includes the embedded engine and runs entirely on-device via sherpa-onnx + SenseVoice (中文/英/日/韩/粤, auto language detection, punctuation). No key, no server, no network, no runtime DLL (the native libs are linked statically).

Model files live in ~/.rpi/agent/models/sense-voice/ ($RPI_CODING_AGENT_DIR/models if set) and are downloaded automatically on first use — or up front with /voice model download. ≈240 MB.

# build (the native libs are fetched at build time)
cargo build -p rpi-voice --release --features local-stt

Building needs a C++ toolchain with cmake and libclang (bindgen) — build-time only, nothing at runtime. Cheapest way to get libclang without installing full LLVM:

pip install libclang
# then point LIBCLANG_PATH at clang/native in site-packages

The sherpa-rs cdylib also needs the Windows system lib at link time:

# Windows
export LIBCLANG_PATH="$(python -c 'import clang,os;print(os.path.dirname(clang.__file__))')/native"
export RUSTFLAGS="-C link-arg=advapi32.lib"
cargo build -p rpi-voice --release --features local-stt

# Linux (static sherpa also needs this)
export RUSTFLAGS="-C relocation-model=dynamic-no-pic"
cargo build -p rpi-voice --release --features local-stt

RPI_STT_MODEL_DIR overrides the model directory (must contain model.int8.onnx and tokens.txt).

OpenAI-compatible endpoint

The client speaks the OpenAI /audio/transcriptions shape, so any compatible server works. A key is only required for the hosted OpenAI endpoint — a local server needs none:

# Free, fully local (faster-whisper): docker one-liner
docker run -p 8000:8000 ghcr.io/speaches-ai/speaches:latest-cpu
export RPI_STT_ENGINE=api
export RPI_STT_API_BASE=http://localhost:8000/v1
export RPI_STT_MODEL=Systran/faster-whisper-small

Groq's free tier also works: RPI_STT_API_BASE=https://api.groq.com/openai/v1, RPI_STT_MODEL=whisper-large-v3, plus a key.

Common Chinese voices: zh-CN-XiaoxiaoNeural (female, default), zh-CN-YunxiNeural (male), zh-CN-XiaoyiNeural (lively).

# Windows
set OPENAI_API_KEY=sk-xxx
# macOS / Linux
export OPENAI_API_KEY=sk-xxx

Build

# API-only (default; no native STT deps)
cargo build -p rpi-voice --release

# with embedded offline STT (see "Embedded offline model" above for env vars)
cargo build -p rpi-voice --release --features local-stt

The built rpi_voice.dll/.so/.dylib is dropped into .rpi/extensions/ by task install (or rpi-package-install) and loads automatically on the next rpi start.

Push-to-talk

/voice ptt turns the configured key (default: space) into a hold-to-talk button:

/voice ptt          # on
/voice ptt off      # off
  • Hold the key → the footer counts down (voice: ⏳ hold ▰▰▰▱▱▱ 0.4s to talk), so you can see the press registered and know to keep holding.
  • Once RPI_VOICE_PTT_HOLD_MS (default 0.6s) elapses, the mic opens and the footer becomes a live level meter driven by your voice: voice: 🎤 listening █████░░░░░░░ 1.1s — release to send. A bar that never moves means the mic isn't hearing you.
  • The same meter is shown for every recording path — one-shot /voice and hands-free turns too (they end on silence, so their hint reads silence ends the turn instead of release to send).
  • Release → recording stops, the audio is transcribed, and the text is delivered according to /voice output (see Where the text goes); the reply is then spoken by auto-TTS.
  • A hold shorter than the threshold never touches the microphone and leaves no trace in the footer.

Continuous conversation (/voice auto)

Hands-free turn-taking, in the style of a voice assistant: you talk, it answers aloud, then it listens again. No key, no /voice — just keep talking.

/voice auto         # on (starts listening immediately)
/voice auto off     # off
you:  why is this request failing?
   🔊 voice: ♪ ▂▅▇▅▂▁▃ playing 3.4s      ← reply spoken aloud
   voice: 🎧 listening…                   ← mic reopens by itself
you:  what if I make it synchronous?
   …

The loop is closed by playback finishing, never by a timer: the microphone opens only after the speakers go quiet, so the assistant can never transcribe its own voice (the classic hands-free echo bug).

Getting out

How What happens
Start typing The mode pauses itself — you've taken the keyboard. (Type to resume: /voice auto.)
/voice auto off Exits now, aborting the listen in flight
Silence After RPI_VOICE_AUTO_EMPTY_MAX (3) empty turns it stops on its own

When it says "nothing heard"

The status carries the measured input peak, so the two very different causes are distinguishable at a glance:

voice: 🎧 nothing heard (peak 0.180) — still listening (1/3)   ← mic works, no speech recognized
voice: 🎧 mic is silent (peak 0.000) — check the input device (1/3)   ← nothing reached the mic

A peak at (or near) 0.000 means the microphone captured nothing at all — a muted device, or the wrong input selected. A healthy peak with an empty transcription means the audio arrived but STT found no words in it.

1. Which microphone is it even using?

/voice status names the active capture endpoint:

🎤 rpi-voice
  ...
  Input device: 麦克风 (USB Audio Device)

A machine with a headset and a webcam has several capture endpoints, and the Windows default is often not the one you are talking into — so you get a recording of the room while you speak into a device nobody is listening to. Pin it explicitly (substring match, case-insensitive):

RPI_VOICE_INPUT_DEVICE=EDIFIER rpi     # talk into the headset
RPI_VOICE_INPUT_DEVICE="USB Audio" rpi

2. Is the microphone actually delivering audio?

cargo test -p rpi-voice --lib -- --ignored --nocapture mic_probe

Lists every capture endpoint with its format, records 4 seconds from the one that would be used, and reports what arrived — through the exact code path the extension uses. Speak while it runs. It prints peak / rms / speech and an explicit verdict.

3. Is the level simply too low?

A microphone can be perfectly healthy and still deliver almost nothing: a Windows input level left near the bottom, or a mic sitting across the desk. Measured in the wild — a working USB microphone whose speech peaked at 0.011 of full scale (≈ -39 dBFS), which is 10–40× below a healthy capture. At that level the voice never rises above the room, so the VAD says "nobody is talking" and STT has almost nothing to work with.

Check the per-second level profile first (mic_probe_compares below). If every second is ≈ 0.01 or less, either raise the level in Windows (Settings → System → Sound → Input → your mic → Volume, plus any "Microphone boost") or apply software gain:

RPI_VOICE_INPUT_GAIN=20 rpi

The boost is applied where the samples enter the recorder, so the meter, the VAD and the STT input all see the same corrected signal — boosting after the VAD would still leave the turn ending early. /voice status reports the gain that was applied, so a level reading can always be read back against the raw capture.

4. Bluetooth headsets need ~4 seconds to wake up

A Bluetooth headset's microphone is a different profile (HFP) from its music playback (A2DP), and Windows takes 3–4 seconds to switch when an app starts capturing. A short probe therefore measures a device that has not woken yet and wrongly concludes it is dead.

Measured on an EDIFIER W820NB with a 200-second capture, per second:

second:  1     2     3     4     5     6   ...  11    12    13    14
peak:   0.00  0.00  0.00  0.72  0.68  0.44 ...  0.73  0.76  0.75  0.74

Seconds 1–3 are digital silence; the microphone comes alive at second 4. The same speaker, captured simultaneously on a laptop's webcam microphone, never exceeded 0.18 — a fifth of the level, because it sits across the room.

So a Bluetooth headset is usually the best input, provided the wait is tolerated: the microphone first gets a wake-up grace period (RPI_VOICE_WARMUP_MS, default 4s), then auto keeps the stream open like PTT until speech is detected and trailing silence ends the turn. Set RPI_VOICE_NO_SPEECH_MS only when an explicit speech-start timeout is desired.

5. Speech detection adapts to your mic

The VAD threshold is not a fixed level. It measures the ambient floor from quiet frames and requires speech to be ~3× that (≈ +10 dB SNR), with a low absolute minimum. A fixed threshold is what once made a perfectly good quiet microphone (ambient RMS ~30 against a hard floor of 260) look completely dead: every turn reported "heard nothing" while the hardware worked fine.

Push-to-talk keeps working while continuous mode is on. So does barge-in — speaking over a reply or typing cuts the audio.

Continuous mode needs replies to be spoken (that's the turn hand-off), so /voice auto turns auto-TTS back on if it was muted, and says so.

Interrupting speech (barge-in)

Nothing is worse than an assistant that keeps reading while you talk over it, so any sign that you are taking the floor cuts playback off instantly:

  • Start typing — the moment the prompt draft changes, the voice stops. This needs no recording and is the usual way out of a long reply.
  • Hold the PTT key — the hold threshold stops playback before the mic opens, so the speakers can never feed back into the recording.
  • /voice — starting a dictation also stops playback, as do /voice off (mute) and /voice stop.

While speech plays, voice: in the footer becomes a music-style equalizer:

voice: ♪ ▂▅▇▅▂▁▃ playing 1.2s

The bar heights follow the real audio envelope — playback publishes the RMS of every ~60ms chunk it hands to the audio device — so the bars move with what you are actually hearing, and the crest ripples left → right like a level meter.

Barge-in on typing relies on the host emitting an editorChange event. On an older host you still get push-to-talk and /voice stop; typing simply won't interrupt.

Where the text goes

By default a transcription is not sent straight to the model — it is placed in the prompt editor as an editable draft, with a countdown in the footer:

input: 你好,帮我看看这个 bug █
footer: Auto-send in 1.4s · any key to edit
  • Leave it alone and it is submitted automatically after RPI_VOICE_DRAFT_MS (default 2s).
  • Type anything to cancel the countdown — the draft stays put and the key you pressed becomes your first edit, so you can fix a mis-recognized word before it ever becomes a message. (Editing at all — typing, pasting, Backspace — also cancels it.)
  • Press Enter to send early.
  • Set RPI_VOICE_DRAFT_MS=0 to disable the countdown entirely and always wait for Enter.

Because the draft is submitted through the editor's normal path, it appears in the transcript as a genuine user message. The old immediate path (SendUserMessage) did not render a user bubble at all.

Use /voice output send (or RPI_VOICE_OUTPUT=send) when you would rather skip the draft entirely and send the moment you release the key — faster, but with no chance to correct a mis-recognition.

The key is only claimed while the input box is empty, so space still types normally whenever you have a draft. Chords (Ctrl+Space, Alt+Space) are left to the editor. When PTT is off the key is not claimed at all.

This needs a host that routes keys to extensions (rpi's register_shortcut + Input events); on an older host the command still works but the key is never delivered, so use /voice instead.

Notes / limitations

  • Push-to-talk relies on the host delivering key release events, which the terminal must report (true for Windows conhost and modern terminal emulators). If presses register but releasing never ends the recording, the RPI_VOICE_PTT_MAX_MS cap still stops it.
  • Interrupting on typing needs a host that emits editorChange; on an older host use push-to-talk or /voice stop instead.
  • STT needs a reachable OpenAI-compatible endpoint: the hosted API (with a key), a free local server, or the embedded offline engine (--features local-stt).
  • Edge TTS is an undocumented Microsoft endpoint; if it changes, synthesis can break and the failure is logged to stderr.