# talk-rs
> **⚠️ Disclaimer: this project is entirely "vibe coded".**
>
> Every line of code, test, and documentation in this repository was
> produced by an AI coding agent under human direction. No part of it
> has been written, line-by-line audited, or formally reviewed by a
> human engineer. It works on the author's machine and ships with tests,
> but treat it accordingly: read the source before trusting it with
> anything important, and expect rough edges, dead code paths, and the
> occasional architectural oddity that a human would have caught.
Voice dictation for Linux. Record, transcribe, and paste text into any
application – all from a single keyboard shortcut.
`talk-rs` captures audio from your microphone, sends it to a
transcription API (Mistral Voxtral or OpenAI Whisper), and types the
result into the focused window. A small X11 overlay badge shows the
current state (recording / transcribing) so you always know what is
happening.
## Features
- **Dictation workflow** – press a key to start recording, press again
to stop; transcribed text is pasted automatically.
- **Multiple providers** – Mistral (Voxtral), OpenAI (Whisper / GPT-4o),
and local on-device Parakeet for batch transcription; Mistral and
OpenAI realtime streaming via WebSocket.
- **Text-to-speech** – the `speak` command synthesizes text and plays it
(or saves a WAV). Two providers mirror the transcription side:
`kokoro` (local, offline, multi-language via sherpa-onnx) and
`mistral` (remote Voxtral TTS). Language handling is config-driven and
agnostic.
- **Speaker diarization** – identify who is speaking (`--diarize`);
output is tagged with speaker labels. Currently supported with Mistral
V2 models in batch mode.
- **Multi-candidate picker** – with `--pick`, run several providers in
parallel and choose the best transcription from a GTK picker window. A
waterfall spectrogram of the recording is shown above the candidate
list; it loads asynchronously and adapts to window width. Non-default
models are transcribed on demand via a per-candidate button to avoid
unnecessary API calls.
- **Visual overlay** – non-intrusive X11 badge at the top of the screen
(works without a compositor).
- **Live transcription overlay** – a dynamic phase-coloured waterfall
replaces the static "transcribing" badge; three independent throughput
tracks (upload bytes, download bytes, paste characters) show progress
through the pipeline, with a time-grid layer over the spectrogram.
- **Dead audio detection** – overlay badge shows a red prohibit icon and
"NO SOUND" warning when no real microphone is detected (e.g. headset
unplugged); a centered full-screen overlay and repeating alert tone
reinforce the signal, and a notification also appears on the text
panel.
- **Auto-pause** – automatically pauses audio forwarding during silence,
trimming dead air from transcription input. Resumes instantly with a
300 ms lookback buffer to preserve speech onset. The badge shows
yellow pause bars and "LISTENING" during pauses. Disable with
`--no-auto-pause`.
- **Audio visualizers** – optional in-badge visualization during
recording (`--viz waterfall`, `--viz amplitude`, `--viz spectrum`);
monochrome mode with `--mono`.
- **Audio feedback** – start/stop tones plus a periodic boop while the
badge shows `LISTENING` (i.e. during auto-pause silence); the boop is
silent while you are actively speaking, and is never heard when
`--no-auto-pause` is used. Fully configurable or disabled.
- **Bluetooth headset auto-switch** – when a Bluetooth headset is
connected in A2DP mode (high-quality stereo, no microphone), `talk-rs`
automatically switches it to its Hands-Free Profile (HFP) for the
duration of the recording so the headset microphone is available, then
restores the original profile on stop. Survives unclean termination: a
state file at `$XDG_RUNTIME_DIR/talk-rs/card-profile.json` is written
before the switch, and the next invocation restores the original
profile from that file before starting a new recording. Works with any
HFP-capable headset (uses PulseAudio's standard `device.form_factor`
property, not vendor-specific identifiers). Disable per-invocation
with `--no-bt-auto-switch` or globally via the `audio.bt_auto_switch`
config key.
- **Context bias** – supply domain-specific vocabulary to improve
transcription accuracy.
- **Daemon toggle mode** – first invocation spawns a short-lived daemon
that records; second invocation signals it to stop, transcribe, paste,
and exit. Ideal for global shortcuts.
- **No idle daemon** – between dictations, zero `talk-rs` processes stay
resident. The toggle daemon spawns on press, records, and exits after
pasting. No memory footprint when idle.
- **Retry last** – re-transcribe the last cached recording without
speaking again (`--retry-last`).
- **Recordings browser** – `record --ui` opens a GTK4 window listing all
recordings (OGG from `output_dir`) and dictation cache (OGG). Imported
audio dropped into `output_dir` – `.m4a`, `.mp4`, or `.aac` (e.g.
iPhone voice memos) – is listed and playable alongside native OGG
recordings, with the same waterfall and transcription workflow. Play,
delete, open in the file manager, or transcribe on demand. A waterfall
spectrogram is shown for every recording, even those without
transcripts; sections auto-refresh via inotify when files change
externally.
- **Per-segment timing export** – `--output-yaml` includes per-segment
start/end timestamps in the metadata sidecar, suitable for subtitles
and post-processing.
- **Standalone commands** – `record` and `transcribe` can be used
independently for scripting.
- **Environment overrides** – every config value can be set via
`TALK_RS_*` environment variables.
## Prerequisites
### Build dependencies
``` bash
# Debian / Ubuntu
sudo apt install build-essential pkg-config libasound2-dev libopus-dev \
libgtk-4-dev libpipewire-0.3-dev libspa-0.2-dev libclang-dev \
libpulse-dev
# Fedora
sudo dnf install alsa-lib-devel opus-devel pkg-config \
gtk4-devel pipewire-devel spa-devel clang-devel \
pulseaudio-libs-devel
```
`libpulse-dev` / `pulseaudio-libs-devel` provides the PulseAudio client
library used to switch a Bluetooth headset's PulseAudio card profile
between A2DP (high-quality stereo) and HFP (microphone-enabled) when
recording. On PipeWire systems, `pipewire-pulse` supplies the
`libpulse.so` shared object at runtime, so no PulseAudio daemon is
needed – only the development headers for compilation.
A working Rust toolchain is required (1.87+). Install via
[rustup](https://rustup.rs) if needed.
### Runtime dependencies
`PipeWire` must be running (used for audio capture). Most modern Linux
desktops ship with PipeWire by default.
Cloud transcription providers require an API key:
- [Mistral](https://console.mistral.ai/) (default)
- [OpenAI](https://platform.openai.com/)
Parakeet is a local, on-device ASR backend using sherpa-onnx on the CPU
and needs no API key. After a consent prompt, its model is downloaded on
first use into `~/.local/share/talk-rs/models/`.
## Installation
If building from a git clone, run `./autogen.sh` first to resolve
version placeholders in `Cargo.toml`:
``` bash
./autogen.sh
cargo build --release
```
The binary is at `target/release/talk-rs`. Copy it somewhere in your
`$PATH`:
``` bash
cp target/release/talk-rs ~/.local/bin/
```
## Configuration
`talk-rs` reads `$XDG_CONFIG_HOME/talk-rs/config.yaml` (typically
`~/.config/talk-rs/config.yaml`).
Copy the example and fill in your values:
``` bash
mkdir -p ~/.config/talk-rs
cp config.example.yaml ~/.config/talk-rs/config.yaml
```
Minimal working configuration:
``` yaml
output_dir: ~/talk-rs-output
providers:
mistral:
api_key: YOUR_MISTRAL_API_KEY
```
### Required fields
| `output_dir` | Absolute path to a writable directory for recordings (a leading `~` is expanded to your home) |
| `providers.mistral.api_key` | Mistral API key (if using Mistral) |
| `providers.openai.api_key` | OpenAI API key (if using OpenAI) |
### Optional fields
| `providers.mistral.url` | `https://api.mistral.ai` | Mistral API base URL |
| `providers.mistral.model` | `voxtral-mini-2507` | Mistral transcription model |
| `providers.mistral.context_bias` | *none* | Comma-separated words for accuracy |
| `providers.mistral.tts_model` | `voxtral-mini-tts-latest` | Voxtral TTS model for `speak --provider mistral` (shares the STT api<sub>key</sub>) |
| `providers.mistral.tts_voice` | *none* | Default Mistral preset voice id for `speak` (else pass `--voice`) |
| `providers.openai.url` | `https://api.openai.com` | OpenAI API base URL |
| `providers.openai.model` | `gpt-transcribe` | OpenAI batch model |
| `providers.openai.realtime_model` | `gpt-live-transcribe` | OpenAI realtime model |
| `providers.openai.prompt` | *none* | Prompt for spelling, punctuation, and context hints |
| `providers.openai.keywords` | *none* | Expected vocabulary list for the new models |
| `providers.openai.languages` | *none* | Ordered expected-language list |
| `providers.openai.realtime_delay` | *none* | Realtime delay: `minimal`, `low`, `medium`, `high`, or `xhigh` |
| `providers.kokoro.model_dir` | *XDG data dir* | Kokoro TTS model cache dir (auto-downloaded on first `speak`) |
| `providers.kokoro.voice` | *per-language default* | Default Kokoro voice name (`af_heart`, `ff_siwis`, …) |
| `providers.kokoro.num_threads` | `4` | Kokoro inference threads |
| `providers.kokoro.lang` | `en` (model-baked) | Default Kokoro phonemization language |
| `transcription.default_provider` | `mistral` | Default transcription provider when unspecified |
| `speak.default_provider` | *kokoro-if-configured* | Default `speak` provider (`kokoro` or `mistral`) |
| `indicators.boop_interval_ms` | `5000` | Periodic boop interval in ms (`0` disables boops; also `--no-boop`) |
| `indicators.visual_overlay` | `true` | Show X11 overlay badge |
| `indicators.viz` | *none* | In-badge visualizer: `waterfall`, `amplitude`, or `spectrum` (also `--viz`; env `TALK_RS_INDICATORS_VIZ`) |
| `indicators.mono` | `false` | Monochrome visualizer (also `--mono`) |
| `paste.chunk_chars` | `150` | Max chars per paste chunk (`0` disables chunking; also `--no-chunk-paste`) |
| `audio.bt_auto_switch` | `true` | Auto-switch a connected Bluetooth headset to HFP for the duration of a recording, then restore (also `--no-bt-auto-switch`; env `TALK_RS_AUDIO_BT_AUTO_SWITCH`) |
| `recording.sample_rate` | `48000` | Sample rate (Hz) of the `record` command's `.ogg` output (env `TALK_RS_RECORDING_SAMPLE_RATE`) |
| `recording.channels` | `1` | Channels for `record` output: `1` mono, `2` stereo (env `TALK_RS_RECORDING_CHANNELS`) |
| `recording.bitrate` | `128000` | Opus bitrate (bps) of the `record` command's `.ogg` output (env `TALK_RS_RECORDING_BITRATE`) |
The `recording.*` settings control the quality of recordings meant for a
**human** to listen to or share. They are independent of transcription:
audio sent to the providers is always downsampled to 16 kHz mono
internally (both Voxtral and Whisper operate at 16 kHz), so these knobs
do not affect transcription accuracy or upload size.
### OpenAI model migration
The OpenAI batch default is exactly `gpt-transcribe`, and the realtime
default is exactly `gpt-live-transcribe`. Existing configuration files
remain valid: `prompt`, `keywords`, `languages`, and `realtime_delay`
are optional, and omitting them sends no hint fields.
- `gpt-transcribe` sends `prompt`, repeated `keywords[]`, repeated
`languages[]`, and `response_format=json` to the batch endpoint.
- `gpt-live-transcribe` nests `model`, `prompt`, `keywords`,
`languages`, and `delay` under `session.audio.input.transcription` in
the realtime `session.update`.
- `whisper-1` remains available for segment/word timestamps, subtitles,
and translation workflows. It retains `response_format=verbose_json`,
accepts `prompt` and one unambiguous singular `language`, and does not
accept `keywords` or multiple languages.
- `gpt-4o-transcribe` and `gpt-4o-mini-transcribe` remain available as
legacy batch choices. `gpt-realtime-whisper` remains available as a
legacy realtime choice, using singular `language` for one expected
language and rejecting `keywords`, multiple languages, and
`realtime_delay`.
Configured hints that the selected model cannot represent fail locally
before HTTP or WebSocket traffic; talk-rs never silently drops them.
There is no structured previous-turn configuration field in the official
schema. Earlier-turn context is service-managed where supported.
Direct-OpenAI prices checked 2026-07-31: `gpt-transcribe` costs
$0.0045/min ($0.27/hr), while `whisper-1` costs $0.006/min ($0.36/hr).
The new batch default is 25% lower, a $0.09 per hour savings.
### Environment overrides
Every config value can be overridden via environment variables:
``` bash
export TALK_RS_PROVIDERS_MISTRAL_API_KEY="sk-..."
export TALK_RS_PROVIDERS_OPENAI_API_KEY="sk-..."
export TALK_RS_PROVIDERS_OPENAI_PROMPT="Preserve punctuation and casing."
export TALK_RS_PROVIDERS_OPENAI_KEYWORDS="Kalysto, talk-rs"
export TALK_RS_PROVIDERS_OPENAI_LANGUAGES="fr, en"
export TALK_RS_PROVIDERS_OPENAI_REALTIME_DELAY="low"
```
`TALK_RS_PROVIDERS_OPENAI_KEYWORDS` and
`TALK_RS_PROVIDERS_OPENAI_LANGUAGES` are comma-separated lists. talk-rs
trims whitespace and ignores empty entries deterministically;
environment values override YAML in both existing and
environment-created OpenAI sections.
See `config.example.yaml` for the full list.
## Usage
### Global options
| `-v` | Increase logging verbosity (`-vv` debug, `-vvv` trace) |
| `--log-file <PATH>` | Write logs to a file in addition to stderr (env `TALK_RS_LOG_FILE`); propagated to the toggle daemon |
### Dictate (main workflow)
Record, transcribe, and paste into the focused application:
``` bash
talk-rs dictate
```
Toggle mode (ideal for keyboard shortcuts):
``` bash
talk-rs dictate --toggle
```
First call starts a background daemon that records. Second call stops
recording, transcribes, and pastes the result.
Options:
| `--toggle` | Daemon toggle mode |
| `--provider` | Choose `mistral`, `openai`, or `parakeet` |
| `--model` | Override model for this invocation |
| `--diarize` | Enable speaker diarization (batch mode only) |
| `--timestamp` | Include timestamps in output (`HH:MM:SS` prefix) |
| `--realtime` | Stream audio via WebSocket (incremental text) |
| `--pick` | Show multi-candidate picker (GTK window) |
| `--retry-last` | Re-transcribe the last cached recording |
| `--replace-last-paste` | Delete previous paste before inserting new text |
| `--save <PATH>` | Save audio recording to a file |
| `--output-yaml <FILE>` | Write transcription metadata YAML |
| `--input-audio-file <FILE>` | Feed a pre-recorded audio file instead of live mic |
| `--monitor` | Mix system audio (monitor) with mic input |
| `--no-sounds` | Disable audio indicators |
| `--no-boop` | Disable periodic boop sounds (keep start/stop) |
| `--no-chunk-paste` | Paste all text in one shot (disable chunking) |
| `--no-overlay` | Disable visual overlay |
| `--no-auto-pause` | Disable auto-pause during silence (forward all audio) |
| `--no-paste` | Skip pasting transcription into the focused application |
| `--upload-format <FORMAT>` | Audio format for batch uploads: `wav` (default) or `ogg` |
| `--viz <MODE>` | In-badge visualizer: `waterfall`, `amplitude`, or `spectrum` |
| `--mono` | Monochrome visualizer (theme-aware) |
| `--no-bt-auto-switch` | Disable Bluetooth headset HFP auto-switch (overrides `audio.bt_auto_switch`) |
### Record
Capture audio to an OGG/Opus file:
``` bash
talk-rs record # auto-named <output_dir>/YYYY/MM/YYYY-MM-DDTHH-MM-SS±ZZZZ.ogg
talk-rs record meeting-notes.ogg # custom filename
talk-rs record --toggle # first call starts, second call stops
talk-rs record --toggle meeting.ogg # toggle with an explicit output path
```
Toggle mode starts a background recorder on the first call. The second
call sends SIGINT; the recorder stops capture, finalizes and syncs the
audio file, then exits. Its PID remains published during finalization,
so repeated stop calls cannot start a second writer for the same toggle
slot. Toggle mode conflicts with `--ui`.
Foreground and toggle recording use the same feedback as dictation: the
start tone finishes before capture begins, an optional X11 badge
visualizes the captured PCM, and the periodic boop sounds only while the
badge detects silence. On stop, the boop and badge are torn down before
capture closes. The stop tone sounds only after the encoder has
finalized the file and `sync_all()` has made it durable. If X11 is
unavailable, recording and sound feedback continue without the badge.
Options:
| `--toggle` | Toggle background recording on the first/second call |
| `--monitor` | Mix system audio (monitor) with microphone input |
| `--no-sounds` | Disable start, stop, and boop sounds |
| `--no-boop` | Disable silence-gated boops (keep start/stop) |
| `--no-overlay` | Disable the recording badge |
| `--viz <MODE>` | In-badge visualizer: `waterfall`, `amplitude`, or `spectrum` |
| `--mono` | Monochrome visualizer (theme-aware) |
| `--ui` | Open GTK4 recordings browser (play, delete, open folder) |
| `--no-bt-auto-switch` | Disable Bluetooth headset HFP auto-switch (overrides `audio.bt_auto_switch`) |
Toggle state is stored under `$XDG_CACHE_HOME/talk-rs/` (normally
`~/.cache/talk-rs/`). Dictation keeps its existing `daemon.pid`,
`daemon.lock`, and `daemon.log` files. Standalone recording uses the
separate `record.pid`, `record.lock`, and `record.log` files, so the two
commands cannot stop or overwrite each other's toggle state.
### Transcribe
Transcribe an existing audio file:
``` bash
talk-rs transcribe recording.ogg # print to stdout
talk-rs transcribe recording.ogg output.txt # write to file
talk-rs transcribe recording.ogg --provider openai
talk-rs transcribe voice-memo.m4a # imported m4a also works
```
Options:
| `--provider` | Choose `mistral`, `openai`, or `parakeet` |
| `--model` | Override model for this invocation |
| `--diarize` | Enable speaker diarization (tag by speaker) |
| `--timestamp` | Include timestamps in output (`HH:MM:SS` prefix) |
1. Provider overload
Cloud providers periodically answer `503 high load` or
`429 backend_out_of_capacity`, especially on long recordings.
`talk-rs` treats every 5xx and 429 as "busy, try again later" and
waits between attempts: 5 s, 15 s, 30 s, 60 s, 120 s, 120 s (about
5.5 minutes in total, seven attempts). A `Retry-After` header from
the provider overrides the schedule slot (capped at 120 s). Each
wait is logged at `-v` as `server retry N/6`. Other 4xx answers are
permanent and fail immediately.
Connection failures (DNS, unreachable host, TLS timeout) use a
separate, shorter schedule and never consume the server-retry
budget. Large uploads are given the whole request time budget on
every attempt, so an 80-minute recording is not cut off by the early
connection slots.
If all seven attempts fail the recording is kept in the cache;
re-run `transcribe` on the `.ogg` (or `dictate --retry-last`) once
the provider has recovered.
### Speak (text-to-speech)
The architectural mirror of `transcribe`: text in, speech out.
Synthesize text and play it through the speakers, or save it to a WAV
file. Two providers – `kokoro` (local, offline, on-device via
sherpa-onnx) and `mistral` (remote Voxtral TTS).
``` bash
talk-rs speak "hello world" # synthesize + play
talk-rs speak --provider mistral --voice <id> "hi" # remote Voxtral TTS
talk-rs speak -o out.wav "save me to a file" # write WAV, don't play
talk-rs speak --lang fr --voice ff_siwis "Bonjour" # French (Kokoro)
```
Text is resolved in priority order: the positional argument, then
`--file`, then stdin (when stdin is not a TTY).
The default provider is resolved as `--provider` \>
`speak.default_provider` (config) \> the local `kokoro` backend when a
`providers.kokoro` section exists, else `mistral`. On first use the
Kokoro model (~350 MB) is downloaded after an interactive consent prompt
(or a non-interactive proceed-with-log, exactly like the Parakeet ASR
model) into `~/.local/share/talk-rs/models/kokoro-multi-lang-v1_0/`.
Options:
| `--provider` | Choose `kokoro` (local) or `mistral` (remote Voxtral) |
| `--voice` | Kokoro voice name (`af_heart`, `am_michael`, `ff_siwis`, …) or a Mistral preset voice id |
| `--lang` | Language for Kokoro phonemization; for Mistral, selects/validates the preset voice (auto-detected from the text when omitted) |
| `--speed` | Speech rate multiplier for Kokoro (`1.0` = normal) |
| `-f`, `--file` | Read the text to speak from a file |
| `-o`, `--output` | Save synthesized audio to a WAV file instead of playing it |
| `--force` | Bypass the voice/language mismatch guard |
Language handling is config-driven and agnostic: the requested language
selects the phonemizer. The stock Kokoro model ships one baked language;
other languages are derived on demand by patching the model's ONNX
`voice` metadata, cached as `model-<lang>.onnx` next to the stock model.
### Supported input audio formats
Commands that read audio from disk – `transcribe`,
`dictate --input-audio-file`, `dictate --retry-last`, and the
`record --ui` recordings browser – accept the following formats:
| `.ogg` | Ogg Opus | Native format written by `talk-rs record` |
| `.m4a` / `.mp4` | MP4 / AAC | Decoded via [symphonia](https://github.com/pdeljanov/Symphonia); ideal for iPhone voice memos and similar imports |
| `.aac` | Raw AAC stream | Decoded via symphonia |
| `.wav` | 16-bit PCM | Decoded directly; primarily used for legacy cache entries |
Imported files keep their original extension on upload, so the
`transcribe` command sends `.m4a` to the provider as-is – both Mistral
Voxtral and OpenAI Whisper accept these natively. Playback and waterfall
spectrograms in the records browser work for every listed format.
## GNOME keyboard shortcut
Bind `talk-rs dictate --toggle` to a key (e.g. `Super+/`):
``` bash
BASE="org.gnome.settings-daemon.plugins.media-keys"
BPATH="/org/gnome/settings-daemon/plugins/media-keys/custom-keybindings"
CURRENT=$(gsettings get "$BASE" custom-keybindings)
N=0
done
SLOT="${BPATH}/custom${N}/"
SCHEMA="${BASE}.custom-keybinding:${SLOT}"
if [ "$CURRENT" = "@as []" ]; then
gsettings set "$BASE" custom-keybindings "['${SLOT}']"
else
gsettings set "$SCHEMA" name 'talk-rs dictate'
gsettings set "$SCHEMA" command 'talk-rs dictate --toggle --viz waterfall'
gsettings set "$SCHEMA" binding '<Super>slash'
```
First press starts recording, second press stops, transcribes, and
pastes into the focused application.
To change the key, replace `<Super>slash` with the desired binding (e.g.
`<Super>semicolon`, `<Super>d`). Add `--realtime` to use streaming
transcription instead of batch mode.
## Development
``` bash
cargo fmt # format
cargo clippy --all-targets # lint
cargo test # test
cargo build # build
```
### Cargo feature flags
The default build enables everything (full dictation CLI). Library
consumers can opt out of the desktop stack with
`default-features = false`:
| (core) | always | reqwest, tokio-tungstenite, opus, rubato, symphonia, … | Transcription providers (Mistral/OpenAI batch + realtime WS), Mistral Voxtral TTS, resampler, Opus/OGG encode, file decode, config, cache |
| `parakeet` | yes | sherpa-onnx (static C++), tar, bzip2 | Local ASR backend (Parakeet TDT, CPU) |
| `kokoro` | yes | sherpa-onnx (static C++), tar, bzip2 | Local TTS backend (Kokoro multi-lang, CPU) for `speak` |
| `playback` | yes | cpal | Local audio playback (the `speak` command's speakers output; shared player) |
| `capture` | yes | pipewire, cpal (via `playback`), libpulse-binding | Live mic/monitor capture, BT headset HFP switch, indicator tones |
| `ui` | yes | gtk4, x11rb, png, fontdue, dark-light, … (+ `playback`) | X11 overlay, visualizers, clipboard/paste, picker, recordings browser |
Example headless consumer (cloud transcription only):
``` toml
talk-rs = { path = "../talk-rs", default-features = false }
```
## License
MIT – see [LICENSE](LICENSE).
# Changelog
## 0.7.1 (2026-09-17)
### Fix
* [dictate] reject a bad API key or model up front in ``--realtime`` mode. [Valentin Lab]
Batch dictation runs a pre-flight check (API key, model) and fails
immediately on a misconfiguration. Realtime dictation skipped it:
``RealtimeTranscriber::validate`` was fully implemented for Mistral
and OpenAI but never called, hidden behind three
``#[allow(dead_code)]``, so a bad key or model only surfaced from
inside the streaming loop after the user had started talking.
``dictate_realtime`` now calls ``validate`` right after creating the
transcriber and before opening the streaming session, so the error
is immediate and enriched with the available models; reconnects keep
skipping it. Capture itself has already started at that point and
buffers meanwhile: the check gates the provider session, not the
microphone. The three pragmas are gone, so the compiler would flag
any future regression. A unit test drives ``validate`` against a
mock that rejects the WebSocket upgrade with 401 and checks that
exactly one handshake happens.
## 0.7.0 (2026-09-17)
### New
* [overlay, picker] show a draining countdown while the provider is busy. [Valentin Lab]
During a server backoff the badge used to sit on a static dim-blue
stripe with ``TRANSCRIBING`` for up to two minutes — indistinguishable
from a hang, which is exactly how users reported it.
The badge now enters an amber "waiting on the provider" state: the
phase-line row becomes a solid amber band that drains right-to-left
across the wait (refilling when the next retry fires), and the centre
text reads ``BUSY · RETRY N/M`` so the user knows it is the provider,
not their setup, and how far through the retry budget they are. The
scrolling phase history records the wait as dim amber so the timeline
still shows how long the provider stayed busy after the retry
succeeds. Connection retries keep their previous rendering.
The picker row label gains the wait as well: ``server retry 1/6 in
30s…``.
``TranscriptionEvent::RetryScheduled`` carries a new ``delay`` field
(zero for connection retries) so UI consumers can render the countdown
from the event's own timestamp.
* [record] add feedback parity with ``dictate`` [Valentin Lab]
Make standalone recordings provide the same start and stop tones,
silence-gated boop, and optional X11 badge as dictation.
Share one feedback lifecycle so ``--no-sounds``, ``--no-boop``,
``--no-overlay``, ``--viz``, and ``--mono`` behave consistently in foreground
and ``--toggle`` modes.
* [record] support ``record --toggle`` lifecycle. [Valentin Lab]
Give standalone recording the same toggle workflow as dictation while keeping
each command's ``daemon`` state isolated.
Register ``SIGINT`` before startup work and retain ownership through encoder
finalization and ``sync_all()`` so rapid stop requests cannot race a second
recorder.
### Fix
* [transcription] cache model suggestions per API key, not only per endpoint. [Valentin Lab]
Two credentials pointing at the same ``api_base`` shared one cached
model list, so a second key could be served the entitlements of the
first for up to an hour. The in-process cache is now keyed by
``(api_base, api_key)``.
This also isolates the ``model_suggestions`` unit tests, which
previously failed under ``--test-threads=1`` because consecutive
``MockServer`` instances reuse the same port and therefore the same
cache entry.
* [transport] wait out provider overload instead of failing within two minutes. [Valentin Lab]
Long recordings sent to Mistral regularly failed with ``503 high load,
please retry`` or ``429 backend_out_of_capacity``: the transport retried
a 5xx at most twice with no delay, never retried a 429 at all, and capped
every attempt at the connection-phase budget (``2, 5, 8, 11, 15, 30,
120`` s), so a 26 MB upload was structurally doomed on the first six
slots. An 83-minute call thus needed seven uploads to succeed when the
API was healthy, and gave up in about two minutes when it was not.
Data-phase (server busy) retries now form their own loop, independent of
the connection-phase budget: every 5xx and 429 is retried after a wait
of ``5, 15, 30, 60, 120, 120`` s (seven attempts, about 5.5 minutes),
a server-supplied ``Retry-After`` overrides the schedule slot (capped at
120 s), and the wait is cancellable. The per-attempt cap is lifted to
the request's own ``wall_clock`` when that is larger, so large uploads
get their full proportional budget on every attempt. Each backoff is
logged at ``info`` as ``server retry N/6``.
``PipelineFailure`` annotates an exhausted 429 as "server-retry budget
exhausted" rather than "4xx permanent, no retry".
## 0.6.0 (2026-08-02)
### New
* [transcription] expose ``OrderedItemTranscript`` to downstream consumers. [Valentin Lab]
Item-aware realtime providers report transcript text as ordered
conversation items rather than as the generic ``TextDelta`` /
``SegmentDelta`` pair, and reconciling those items requires
conversation ordering, provisional-to-authoritative replacement, and a
terminal drain.
``OrderedItemTranscript`` already implements exactly that, but it was
crate-private, so an external consumer of ``TranscriptionEvent`` had no
way to handle the item-aware variants without reimplementing the same
ordering algorithm.
Publish the type and the six methods such a consumer needs, and keep
the internal ordering state and the replay helpers restricted.
* [openai] add ``gpt-transcribe`` and item-aware realtime support. [Valentin Lab]
Integrate the new batch and live models with configurable context, keyword, language, and delay hints while retaining legacy model compatibility.
Reconcile provisional realtime deltas by item so authoritative completions replace rather than duplicate picker and dictation text.
* [speak, synthesis] add TTS ``speak`` command with Kokoro and Mistral providers. [Valentin Lab]
Architectural mirror of the ``transcription`` surface on the
synthesis side: a ``OneShotSynthesizer`` trait with a cfg-gated
factory, a local Kokoro backend (``kokoro`` feature, same
``sherpa-onnx`` static lib as ``parakeet``, consent-gated ~350 MB
model download via the shared ``model_fetch`` helper) and a remote
Mistral Voxtral backend (``POST /v1/audio/speech``, sharing the
existing ``providers.mistral`` API key via new ``tts_model`` and
``tts_voice`` fields).
``talk-rs speak`` reads text from the positional argument,
``--file`` or stdin, then plays through the ``AudioPlayer`` or saves
a WAV with ``-o``.
Language handling is automatic by default: ``whichlang`` detects the
text language, selecting the phonemizer for Kokoro (non-default
languages derive a metadata-patched ``model-<lang>.onnx``, cached
next to the stock model) and the per-language voice for Mistral
(``tts_voices`` config map or built-in presets). ``--lang``
overrides detection. An explicitly pinned voice whose language
contradicts the resolved language is a hard error naming the voice,
both languages and the remedies; ``--force`` bypasses the guard.
Voices with unknowable language (custom Mistral UUIDs) skip the
guard.
Config surface: ``providers.kokoro`` (voice, ``num_threads``,
``lang``, ``model_dir``), ``providers.mistral.tts_voices``,
``speak.default_provider``, with matching ``TALK_RS_*`` env
overrides.
* [paste] composable paste-node tree configurable per target window. [Valentin Lab]
Refactor the paste subsystem into a tree of composable nodes
(router / decorator / leaf), each implementing a common ``paste()``
interface, selectable via a recursive ``paste:`` tree in the config.
Node families:
- routers: ``detect-display-server`` and ``match-wm-class`` (first-match
glob routing on ``WM_CLASS``);
- decorator: ``chunk``;
- leaves: ``clipboard`` and ``xtest-type``.
This lets the user pick chunking, paste shortcut (``ctrl_v`` vs
``ctrl_shift_v``), or XTest typing on a per-target-window basis.
Fully backward compatible: absent ``paste:`` reproduces the previous
behaviour; the old flat ``paste`` keys are still accepted and mapped to
an equivalent tree; ``--no-chunk-paste`` strips the chunk nodes.
Runtime paste behaviour is unchanged in this commit (the existing
clipboard sync is moved verbatim into the ``clipboard`` node).
Also add an ``x11_get_wm_class`` helper to drive the ``match-wm-class``
router.
* [transcription, config] add local Parakeet backend with consent-gated download. [Valentin Lab]
Add NVIDIA Parakeet TDT 0.6b v3 as a local, on-device, CPU
transcription provider selectable via ``provider: parakeet`` (config
``transcription.default_provider`` or the ``--provider`` flag), running
fully offline through a statically-linked ``sherpa-onnx``.
The ~640 MB model is NEVER downloaded silently. The transcribe pipeline
only checks for the model (``model::ensure_present``) and errors if it
is absent; the actual fetch (``model::download_model``) is an explicit,
consented step driven per entry surface:
- ``transcribe`` CLI: ``[y/N]`` prompt on a TTY; a clear ``stderr``
message then proceed when piped (selecting parakeet is the consent).
- ``dictate --toggle``: a "DOWNLOADING MODEL" overlay badge while it
fetches, before recording starts.
- ``--pick`` picker: a GTK ``AlertDialog`` on click; Cancel is a clean
no-op. Parakeet is only listed when ``providers.parakeet`` is
configured, and an absent default model is demoted to a click-to-
download row rather than auto-erroring.
The download is atomic (extract + verify into a staging dir, then a
single ``rename`` into the cache) with an ``fs2`` lock for concurrent
invocations, so an interrupted run never leaves a half-populated model
dir. Inference runs in ``spawn_blocking`` and reuses the existing
16 kHz mono decode pipeline. The backend sits behind a ``parakeet``
cargo feature (on by default; ``--no-default-features`` drops the
C++ dependency for a lean cloud-only build).
* [record, config] add human-quality recording profile for ``.ogg`` output. [Valentin Lab]
The ``record`` command reused ``AudioConfig::new()`` — the 16 kHz mono
32 kbps profile required by the transcription providers — to write
recordings meant for a human to listen to. At 16 kHz the Opus encoder
is capped to an 8 kHz audio bandwidth (Nyquist), so shared recordings
sounded muffled and dull.
Separate the two concerns: the transcription profile stays hardcoded
and hidden (an API requirement, not a user preference), while a new
user-facing ``recording:`` config section controls the ``record``
output quality. It exposes ``sample_rate`` (default 48000),
``channels`` (default ``1`` mono) and ``bitrate`` (default 128000),
each also settable via ``TALK_RS_RECORDING_*`` env vars.
The ``record`` path now:
- resolves its ``AudioConfig`` from the ``recording:`` section;
- captures via ``PipeWireCapture`` (negotiates the requested rate /
channels with the device, robust on mono-only mics) instead of the
exact-match ``CpalCapture``;
- encodes with ``OggOpusWriter::new_for_recording`` (Opus
``Application::Audio`` for fuller fidelity vs the speech-optimised
``Voip`` mode kept by the transcription path).
Verified empirically: a fresh recording now carries energy up to
~20 kHz (was a hard wall at 8 kHz). Documents the new section in
``config.example.yaml`` and ``README.org``; adds config tests for
defaults, YAML parsing, partial overrides and env overrides.
* [config] expand a leading ``~`` in ``output_dir`` to ``$HOME`` [Valentin Lab]
The README minimal example and many users naturally write
``output_dir: ~/talk-rs-output``, but there was no tilde expansion:
the value was taken literally as a relative directory named ``~``.
Combined with the new absolute-path requirement, the documented
example would now fail outright.
Expand a leading ``~`` (bare) or ``~/…`` to the user's home directory
(via ``directories::UserDirs``) at load time, after the
``TALK_RS_OUTPUT_DIR`` override and before validation, so both the
file value and the env override benefit. A tilde that is not the
first path component (e.g. ``/tmp/~/x``) and the ``~user`` form are
left untouched. Document the behavior in ``config.example.yaml`` and
``README.org``; add tests for the bare-``~``, ``~/…``, env-override,
and non-leading-tilde cases.
* Add --timestamp flag for diarized transcription output. [Boris Gallet]
Adds a CLI flag --timestamp (also -t) that prefixes each
speaker-attributed line with [HH:MM:SS] in the transcript output.
- CLI: --timestamp added to both transcribe and dictate commands
- Config: env var TALK_RS_TIMESTAMP also supported
- Output format: [HH:00:00] speaker_1 text...
- Tests: 5 new tests covering timestamp formatting, edge cases,
and integration with diarization
Motivation: when transcribing long meetings (>30min), knowing
when each speaker talked is essential for navigation and summary.
* [paste, config] configurable paste shortcut (Ctrl+V vs Ctrl+Shift+V) [Boris Gallet]
Add `PasteShortcut` enum to `PasteConfig` with two variants:
- `ctrl_shift_v` (default) — primary selection paste
- `ctrl_v` — regular clipboard paste for terminals/Emacs
Refactor `simulate_paste()` to take `PasteShortcut` and resolve
the correct X11 keysyms via a pure `paste_keysyms()` function.
Propagate the shortcut through `paste_text_to_target()`, the
realtime per-segment paste task, and the picker path.
* [picker, realtime] surface WS session phases as picker row status. [Valentin Lab]
* [picker, transcription] picker Stop button via SIGUSR1 cross-process cancel. [Valentin Lab]
* [transcription] cross-process job registry via lock-file YAML + SIGUSR1. [Valentin Lab]
Step 11 of the transport-consolidation plan at
``.sisyphus/plans/transport-consolidation.md``.
New module ``transcription::jobs`` extends the existing per-model
lock files in ``recording_cache`` with a YAML payload describing
the owner process and adds SIGUSR1-based cross-process
cancellation.
* [transcription, picker] wire telemetry sink into realtime path. [Valentin Lab]
Step 10 of the transport-consolidation plan. Realtime transcribers now
accept a telemetry sink so WS upgrade phase and retry events reach the
picker UI, fixing the silent-blocking symptom.
* [transcription, transport] introduce unified ``http_request`` / ``ws_upgrade`` API. [Valentin Lab]
Adds the public transport-consolidation API surface
(``talk_rs::transcription::transport::{http_request, ws_upgrade}``)
plus the truthful-counter, growing-budget, cancellable
``http_request`` implementation.
This is Steps 0-2 of the consolidation plan at
``.sisyphus/plans/transport-consolidation.md``. The plan resolves
three long-standing bugs surfaced by the user: (a) ``send_once``
hardcodes ``attempts=1, max_attempts=1`` in every
``PipelineFailure``, (b) the connect budget never grows across
attempts, (c) requests cannot be cancelled. After this commit:
- ``http_request`` reports truthful ``attempts``/``max_attempts``
in ``PipelineFailure`` (was ``1/1`` everywhere).
- Connection-phase retries run on the growing budget
``[2, 5, 8, 11, 15]`` seconds, wrapped in an outer
``tokio::time::timeout`` so blackhole destinations don't hang
for the OS-default ~125s TCP SYN timeout.
- Data-phase retries (5xx responses) run up to 3 attempts.
- ``CancellationToken`` aborts an in-flight request via
``tokio::select!`` within ms of the trigger.
The legacy ``with_retry`` / ``send_once`` paths are untouched —
they will be migrated in Steps 3-5 of the plan. Multipart bodies
are intentionally rejected by ``http_request`` for now; Step 4
adds a body-factory shape so retries can rebuild a fresh
``reqwest::multipart::Form`` per attempt.
Tests: ``tests/transport_integration.rs`` pins the spec for the
unified transport. Five of the ten tests pass after this commit
(connection budget, data retries, both attempts-counter tests,
cancellation). The remaining five remain ``unimplemented!()`` /
``panic!`` placeholders for Step 6 (``ws_upgrade``) and Step 11
(``jobs`` cross-process registry); they ship RED on purpose and
will flip green as those steps land.
The ``transcription::transport`` module changes from ``pub(crate)``
to ``pub`` so integration tests can pin the public surface.
* [record] m4a / mp4 / aac read support across the recordings browser. [Valentin Lab]
The recordings browser (``record --ui``) silently ignored any file that
was not ``.ogg``: the directory walker filtered on a hardcoded
extension, the waterfall decoder errored on unknown formats, the
playback path fell through to the WAV parser (producing either an
error or garbage), and the ``inotify`` live-refresh filter listed only
``.wav`` / ``.ogg`` events. Imported audio — typically ``.m4a`` voice
memos copied into ``output_dir`` — was therefore invisible in the UI
even though every downstream consumer (the transcription providers,
the GTK list, the file manager) handles those files natively.
Add full read support for ``.m4a``, ``.mp4``, and raw ``.aac`` across
every audio-consuming path of the browser:
* New ``read_m4a_as_f32`` in ``src/record/audio.rs`` uses
[[https://github.com/pdeljanov/Symphonia][``symphonia``]] for ISO/MP4
demuxing and AAC-LC decoding. Multichannel input is averaged to
mono and resampled with the existing ``resample_linear`` helper to
match the calling site's rate, mirroring the
``read_ogg_as_f32`` / ``read_wav_as_f32`` contract. The
authoritative sample rate and channel count are read from the
first decoded ``AudioBuffer``'s ``SignalSpec``, never from
``codec_params``: symphonia's ISO/MP4 demuxer is known to
underreport channels for AAC tracks (a stereo AAC-LC track
surfaces as ``codec_params.channels = 1`` even though the decoder
produces a 2-channel buffer), which would otherwise make the
stereo-to-mono mixdown walk the interleaved buffer with the wrong
stride, doubling the output sample count and halving playback
speed / pitch. Only the ``aac`` + ``isomp4`` symphonia feature
flags are enabled, so the release binary picks up an AAC decoder
+ MP4 demuxer (~200 KB) and no other codecs.
* New ``m4a_duration_secs`` walks the MP4 box hierarchy to read
``mvhd.duration`` / ``mvhd.timescale`` directly — O(box-count), no
full decode required. Mirrors the O(1) ``ogg_duration_secs`` probe
so the listing stays snappy on large files. Supports both v0 and
v1 ``mvhd`` layouts and tolerates ``size = 0`` ("to end of file")
and ``size = 1`` (64-bit extended) box headers.
* The central dispatcher ``read_audio_as_i16`` gains the
``"m4a" | "mp4" | "aac"`` arm — this automatically extends m4a
support to every caller, including the picker waterfall
(``src/dictate/picker/ui.rs``) and the shared
``audio_player_bar`` widget.
* ``src/record/entries.rs`` is generalised: ``collect_oggs_recursive``
→ ``collect_audio_recursive``, with a shared ``AUDIO_EXTENSIONS``
constant and a case-insensitive ``has_audio_extension`` helper so
``.M4A`` files from a camera are picked up alongside lowercased
ones. Listing-time duration is dispatched through a new
``audio_duration_secs`` that routes to the per-format probe.
Both ``list_ogg_recordings`` (now lists all audio under
``output_dir``) and ``list_cache_recordings`` (dictation cache)
use the new collector.
* ``WavPlayer::play`` in ``src/record/player.rs`` learns the same
``m4a | mp4 | aac`` arm. The fallthrough to ``read_wav_as_f32``
is preserved — backwards-compatible with the legacy
``.wav`` cache entries — but is no longer hit for AAC content,
which previously produced corrupted playback.
* The ``GTK4`` ``FileMonitor`` filter in ``src/record/ui.rs:908``
treats every listable extension as audio. The comparison is
lowercased so case-mixed imports trigger the live-refresh.
* ``delete_recording`` extends the YAML companion + ``.wf``
waterfall cache cleanup so deleting an ``.m4a`` row sweeps the
same sidecars an ``.ogg`` row would.
Test coverage:
* Two shipped fixtures under ``tests/fixtures/``: the original
``sine_440_0.5s_mono.m4a`` (5 345 B, 440 Hz sine, AAC-LC, mono,
0.5 s) and the new ``sine_440_0.5s_stereo.m4a`` (5 026 B, same
signal, stereo) that mirrors the format produced by the
AudioRecorder app the user imports voice memos from — both are
documented in ``tests/fixtures/README.md`` with the ``ffmpeg``
regeneration command.
* 5 new ``record::audio::tests`` cover ``m4a_duration_secs`` against
the mono fixture (±0.1 s tolerance to absorb encoder rounding),
``read_m4a_as_f32`` for mono (sample-count window + non-silent
peak), ``read_m4a_as_f32`` for stereo at both 16 kHz and 48 kHz
(the regression test that locks in the ``SignalSpec``-vs-
``codec_params`` correctness — empirically verified to fail with a
2× sample count when the bug is reintroduced),
``read_audio_as_i16`` dispatch through to the AAC arm, and the
unsupported-format error path.
* 4 new ``record::entries::tests`` cover the broadened collector
(``.m4a`` / ``.MP4`` / ``.aac`` mixed with ``.ogg``,
case-insensitive extension matching, mixed flat + nested
layouts) and the ``audio_duration_secs`` dispatcher's
``None``-on-unknown semantics.
* Full ``cargo fmt``, ``cargo clippy --all-targets --all-features``,
``cargo test`` (438 unit tests pass, 0 fail), and
``cargo build --release`` clean.
``README.org`` gains a "Supported input audio formats" subsection
under Usage and updates the recordings-browser feature bullet to
mention ``.m4a`` imports; ``README.md`` is regenerated by
``autogen.sh``.
* [overlay] surface validate phase + retry counter on the spectrogram. [Valentin Lab]
Add a dedicated ``Phase::Validating`` state to the overlay's HTTP
state machine, driven by the ``PreflightStarted`` /
``PreflightCompleted`` telemetry events emitted by the
validate-cache miss path. The phase reuses the same green as
``Phase::Done`` but at half opacity so the user can tell at a
glance that the network roundtrip in flight is the cheap
``/v1/models`` validation rather than the actual transcription
upload. ``Phase::color`` now returns ``([u8; 4], f32)`` so each
phase carries its own opacity; throughput tracks drop the opacity
component since they always render at full strength.
Also draws a small attempt counter at the top-left of the SPEC
area whenever a ``RetryScheduled`` event fires (validate retries
or transcription retries), dimmed to follow the current phase's
opacity so a validate retry shows a half-opacity "2" while a
transcription retry shows a full-opacity "2". The counter is
cleared at the start of every fresh recording cycle and at every
fresh ``PreflightStarted`` / ``RequestStarted`` so a stale digit
never bleeds into the next attempt.
To support the dimmed stripe and counter, ``PixelBuffer`` gains a
``blend_pixel(x, y, color, opacity)`` helper that linear-blends
``BGRA`` over the existing background and leaves the alpha channel
untouched (the X11 overlay window owns transparency, we never
want a hole punched through the badge layer).
Tests cover:
- the new phase transitions (``PreflightStarted`` →
``Validating``, success → ``Idle``, failure → ``Error``,
``RetryScheduled`` inside vs. outside ``Validating``);
- ``Phase::color`` returning green at 0.5 for ``Validating`` and
green at 1.0 for ``Done``, plus a regression check that every
non-validating phase renders at full opacity;
- ``render_retry_counter`` painting nothing for attempt 0,
visible pixels for attempt >= 1, and being measurably dimmer
at half opacity (skip-clean if no system font is installed);
- ``PixelBuffer::blend_pixel`` at 0.0 / 0.5 / 1.0 / clamp / OOB.
* [picker, transcription] reliable picker transcription with live status, validation cache, and timeout attribution. [Valentin Lab]
Three coupled changes that together make the picker UI usable on
flaky networks (notably VPNs that drop connection attempts mid-flight):
1. ``RequestTimeoutPolicy`` enum + ``TranscribeOptions`` struct.
Picker rows run under ``UserAttended``: no per-request wall-clock
cap, only the client-wide ``connect_timeout`` plus TCP-level
defences (``tcp_user_timeout``, TCP keepalive). Slow servers no
longer trip a 3-second wall-clock that fires after a successful
connect. Autonomous callers (``dictate`` end-of-recording,
``transcribe`` CLI, ``produce_transcript``) keep the previous
``Proportional`` policy so an unattended pipeline cannot hang.
Threaded through ``transcribe_audio``, ``create_batch_transcriber``,
and the ``with_policy`` constructors on ``MistralBatchTranscriber``
/ ``OpenAIBatchTranscriber``.
2. Disk-backed validation cache for ``/v1/models``
(``\$XDG_CACHE_HOME/talk-rs/validate-cache.yaml``, 24-hour TTL,
atomic tempfile + rename). Successful preflights memoize, so
subsequent transcriptions for the same ``(provider, model,
api_base)`` skip the network entirely within the TTL. Writes
are atomic and last-writer-wins; corrupt cache files fall
through to network rather than panicking. On cache miss the
preflight runs a growing-budget retry of [2, 5, 8, 11, 15]
seconds across 5 attempts so a transient connect blip cannot
doom the whole transcription on the first attempt; permanent
errors (HTTP 4xx, model-not-found) bail without retry. The
user-facing error now leads with ``"<Provider> model
validation failed (preflight to /v1/models)"`` so the failed
concern is unambiguous (the prior lead phrase ``"Failed to
connect to <Provider> API"`` falsely implied the transcription
itself failed).
3. Live status display in picker candidate rows. New
``PickerStatusSink`` translates HTTP-pipeline telemetry into
short italic strings rendered at 0.55 opacity in place of the
transcript area: ``pre-validating model…`` (cache-miss only),
``connecting…``, ``uploading…``, ``waiting for server…``,
``transcribing…``, ``retry N/M…``. Final candidate (success
or error) replaces the status with the actual transcript or
error. Action button tooltip mirrors the status while
in-flight and clears on terminal events. Realtime rows are
left alone — their partial transcript already serves as
status. New ``TranscriptionEvent::PreflightStarted`` /
``PreflightCompleted`` variants surface the validate phase to
any sink (only emitted on cache miss); ``PreflightCompleted``
is silent on the picker so the next event replaces the line
without flicker.
Also includes the diagnostic logging that enabled this fix:
``TimerSpec`` slice + ``format_reqwest_error_with_timers`` make
reqwest error messages name *which* timer fired
(``connect_timeout``, ``request_wall_clock``, ``validate_request``,
or ``kernel_tcp_unspecified``) and quote its budget, so log
readers no longer have to correlate elapsed time against client
state to identify a timeout source.
New module: ``src/transcription/transport/validate_cache.rs``.
New tests: 8 cache tests, 4 ``validate_model`` tests covering
schedule + cache-hit + retry semantics, 7 ``PickerStatusSink``
tests covering phase transitions / dedup / cache-hit / preflight
silence, plus the ``user_attended_policy_omits_request_wall_clock_attribution``
regression detector. 389 → 404 lib tests passing, 0 failures.
The stale comment at ``picker/backend.rs:152-155`` claiming
"no artificial outer timeout is needed" is updated to describe
the actual ``UserAttended`` semantics.
* [audio, dictate, record] Bluetooth headset HFP auto-switch. [Valentin Lab]
When a Bluetooth headset is connected in A2DP mode (high-quality stereo
output, no microphone), ``talk-rs`` now automatically switches it to
its Hands-Free Profile (HFP) for the duration of a recording so the
headset microphone is enabled, then restores the original profile on
stop. This replaces the equivalent feature of the legacy ``memo``
shell script and removes the need to manually toggle profiles when
dictating with a Bluetooth headset.
Detection uses PulseAudio's standard ``device.form_factor = "headset"``
property (not vendor-specific identifiers), with a fallback for stacks
that omit ``form_factor`` (matches ``bluez_card.*`` cards exposing a
``headset-head-unit*`` profile). Profile selection prefers
``headset-head-unit-msbc`` (16 kHz wideband) → ``headset-head-unit-cvsd``
(8 kHz narrowband) → ``headset-head-unit`` (generic), all of which are
PulseAudio-standard names.
Crash resilience: the saved profile is persisted to
``$XDG_RUNTIME_DIR/talk-rs/card-profile.json`` BEFORE the switch, so an
unclean termination (``SIGKILL``, power loss) leaves a recoverable
state file behind. A new ``recover_stale_profile()`` runs at the
start of every live-capture command and restores the profile from that
file before activating HFP for the new recording — so even after a
crash mid-recording the user gets A2DP back on next launch. This is a
real improvement over ``memo``, which overwrote its state file on
start and could permanently lose the original profile after a crash.
Implementation uses ``libpulse-binding`` v2.30 — pure Rust, no
``pactl`` subprocess and no D-Bus. On PipeWire systems,
``pipewire-pulse`` provides ``libpulse.so`` as a compatibility shim, so
``pa_context_set_card_profile_by_name`` reaches the right device
through ``WirePlumber``. This avoids stdout parsing and keeps error
handling fully typed. Build dependency added: ``libpulse-dev``
(Debian/Ubuntu) / ``pulseaudio-libs-devel`` (Fedora).
A new ``HeadsetGuard`` RAII type holds the saved profile and restores
it on ``Drop``, so the original profile is restored on normal return,
``?``-propagated errors, panics, and ``SIGINT``-driven daemon exits —
no need to thread explicit restore calls through every exit path.
Configurable via:
- ``--no-bt-auto-switch`` CLI flag on both ``dictate`` and ``record``
(also forwarded to the spawned daemon by ``--toggle``).
- ``audio.bt_auto_switch`` config key (default ``true``).
- ``TALK_RS_AUDIO_BT_AUTO_SWITCH`` env var
(true/false/yes/no/1/0/on/off).
- File-input dictation paths (``--input-audio-file``,
``--retry-last``) skip the switch since they don't touch the mic.
All ``bt_profile`` failures are non-fatal: they log a warning and let
the recording proceed on whatever input device is currently active.
Verified end-to-end:
- 14 new unit tests in ``src/audio/bt_profile.rs`` (profile picker,
headset detection, JSON round-trip, state-file path resolution,
guard drop semantics).
- ``tests/bt_profile_smoke.rs`` integration test (gated ``#[ignore]``
+ ``TALK_RS_BT_INTEGRATION=1``) runs a real A2DP→HFP→A2DP round-trip
against the running PulseAudio/``pipewire-pulse`` server.
- Full ``cargo fmt``, ``cargo clippy --all-targets``, ``cargo test``,
``cargo build``, ``cargo build --release`` clean.
* [transcription, record] surface cached transcripts in the recordings browser. [Valentin Lab]
Adds ``read_cached_transcript()`` in ``src/transcription/mod.rs`` -- a
synchronous, network-free waterfall that returns a transcript for an
audio file by walking, in priority order:
1. The pick file (``<stem>.pick.yml``) -- the authoritative,
possibly user-edited transcript.
2. The default-provider / default-model batch sidecar.
3. ``None`` -- nothing cached.
The recordings browser (``record --ui``) in ``src/record/entries.rs``
now uses this waterfall in both ``list_ogg_recordings`` and
``list_cache_recordings`` so an entry shows its transcript whenever
one is cheaply available -- not only when it has been pick-finalised.
The ``InProgress`` state from the pick lock still wins so the UI can
keep displaying "transcription ongoing" for in-flight runs.
* [transcription] persist per-segment timing in YAML sidecar. [Valentin Lab]
Providers already return timed transcript segments (Voxtral batch,
Whisper ``verbose_json``, Voxtral realtime), but talk-rs was
discarding everything except a ``segment_count`` integer. This
commit preserves the ``(start, end, text)`` tuples all the way from
the API response to the YAML sidecar so downstream consumers such
as ``activity-memo`` can reconstruct sub-minute timelines without
re-transcribing.
Data model (``src/transcription/mod.rs``):
New ``TranscriptSegment`` struct and an optional ``segments`` field
on ``TranscriptionResult``. A shared ``parse_transcript_segments``
helper extracts ``start``/``end``/``text`` from raw
``serde_json::Value`` slices, used by both Mistral and OpenAI
parsers.
Mistral batch (``src/transcription/mistral.rs``):
Both ``transcribe_file`` and ``transcribe_stream`` now call
``parse_transcript_segments`` on the response ``segments`` array.
Voxtral returns segments by default with ``speaker_id: null`` when
diarization is not requested — those were previously skipped by
``parse_diarization_segments`` (which requires a non-null
``speaker_id``). Both extractions now coexist: diarization
segments for ``--diarize`` users, transcript segments for everyone.
OpenAI batch (``src/transcription/openai.rs``):
``response_format`` is now ``verbose_json`` for whisper models
(which support it) and stays ``json`` for GPT-4o transcribe models
(which do not). ``verbose_json`` causes Whisper to return a
``segments`` array with ``start``/``end``/``text`` — parsed the
same way as Voxtral.
Realtime (``src/dictate/realtime.rs``, ``src/dictate/picker/backend.rs``):
``SegmentDelta`` events are now destructured as
``{ text, start, end }`` instead of ``{ text, .. }``. A parallel
``Vec<TranscriptSegment>`` accumulator captures timed segments
alongside the existing text accumulator. The picker UI state
(``picker/ui.rs``) carries segments through candidate selection.
YAML sidecar (``src/recording_cache.rs``):
New ``CommonSegment`` struct with ``#[derive(Serialize)]``.
``RecordingMetadata`` gains an optional ``segments`` field
(``skip_serializing_if = Option::is_none``). ``write_metadata``
accepts an extra ``segments: Option<&[TranscriptSegment]>``
parameter; call sites in ``dictate/mod.rs`` and
``dictate/picker/mod.rs`` pass ``result.segments.as_deref()``.
Resulting YAML when segments are present:
segments:
- start: 0.0
end: 1.5
text: Hello world.
- start: 2.0
end: 3.8
text: This is a test.
Absent when the provider returned no segments (old files, GPT-4o,
realtime sessions where the server omits timing).
Backward compatibility: existing YAML files without ``segments:``
continue to work. The field is purely additive.
Tests: unit tests for ``parse_transcript_segments`` (valid, malformed,
empty, mixed); YAML serialization with/without segments; wiremock
integration tests in both ``mistral.rs`` and ``openai.rs`` using
a new ``tests/fixtures/voxtral-response.json`` fixture.
* [overlay, indicator] add centered no-sound overlay and alert tones. [Valentin Lab]
When the overlay detects a dead audio device (no-sound), a large
semi-transparent overlay now appears at the centre of the screen with a
prohibit icon, "NO SOUND" title, and "No audio detected" subtitle. The
overlay uses per-pixel ARGB alpha (80% opacity) when a compositor is
available, falling back to solid black otherwise.
A periodic alert tone (triple-pulse at 550 Hz) plays every 2 seconds
while no-sound is active, and the regular boop heartbeat is suppressed
to avoid colliding sounds. An ``AlertPlayer`` handle allows the
silence-notifier thread to play sounds without owning the full
``SoundPlayer``.
Also fixes the badge prohibit icon regression introduced in commit
31d99f0: that commit enlarged ``DOT_RADIUS_MAX`` from 10 to 21 for the
volume-reactive dot, but the prohibit icon shared that constant,
making its stroke look disproportionately thin. A dedicated
``PROHIBIT_ICON_RADIUS`` (10.0) now restores the original proportions.
The rendering logic is factored into ``draw_prohibit_icon_with_stroke``
(parameterised stroke) with ``draw_prohibit_icon`` as a thin wrapper.
* [overlay] add time-grid layer over the waterfall spectrogram. [Valentin Lab]
Draw vertical dotted yellow marks every wall-clock second, alpha-blended
at 60% over the existing waterfall pixels. Each mark is a column of
alternating 1 px dot / 1 px gap in pure yellow (BGR [0, 255, 255]).
Destination alpha is left untouched to preserve window opacity.
Grid alignment uses ``columns_pushed_total``, an absolute column counter
that resets together with ``spectrogram_history``. Marks are placed
where ``abs_idx % COLUMNS_PER_GRID_MARK == 0`` and right-aligned the
same way the waterfall itself is. The counter increments even during
auto-pause empty-column pushes so the grid reflects true elapsed time
including silent gaps.
``render_time_grid`` is called in both the normal recording path and
the auto-pause path, after the waterfall and badges but before the
pulsing dot / LISTENING indicator.
Alpha was initially set to 0.3 but proved too subtle on real spectrograms;
bumped to 0.6 after visual testing.
This establishes the rendering pattern — absolute-index layer drawn over
the scrolling waterfall — that future event overlays (network phase
colors, byte-count bars) will follow.
* [transcription] add ``ProgressBody`` stream wrapper for upload telemetry. [Valentin Lab]
Add a ``futures::Stream`` wrapper that sits between a ``Vec<u8>``
audio payload and ``reqwest::Body::wrap_stream``, emitting telemetry
events as bytes flow through:
- ``ConnectionEstablished`` on the first ``poll_next`` (proxy for
TCP+TLS handshake completion)
- ``UploadProgress { bytes_sent, total, t }`` on every yielded chunk
- ``UploadComplete { total, t }`` exactly once at stream exhaustion
Boolean guards (``emitted_connection``, ``emitted_complete``) prevent
duplicate events if the runtime re-polls a terminal position.
Chunk size is controlled by ``PROGRESS_BODY_CHUNK_BYTES`` (8 KiB).
A ``len()`` accessor exposes the total byte count so callers can set
``Content-Length`` on multipart parts.
Six unit tests cover empty buffers, single-chunk and multi-chunk
bodies, cumulative byte accounting, and the exactly-once guarantees
for connection and completion events.
``#[allow(dead_code)]`` annotations carry inline justification
comments; the struct is wired into ``mistral.rs`` / ``openai.rs``
in sub-phase 1C and the pragmas will be removed at that point.
* [telemetry] add display-agnostic transcription event module. [Valentin Lab]
Introduce ``src/telemetry/mod.rs`` with the foundational types for
observing the HTTP transcription lifecycle without coupling to any
display or pipeline module.
Contents:
- ``TranscriptionEvent`` enum (13 variants covering request start
through paste completion, each carrying a monotonic ``Instant``
timestamp).
- ``TelemetrySink`` trait (``Send + Sync``, single ``emit`` method)
so producers can hold ``Arc<dyn TelemetrySink>`` without locking at
the API boundary.
- ``NoOpSink`` — zero-cost default that drops every event; useful for
tests and headless code paths.
- ``BroadcastSink`` — wraps ``tokio::sync::broadcast`` to fan events
to any number of subscribers, dropping silently on lag or when no
receivers are attached.
- 7 unit tests covering sink behaviors, ``Arc<dyn>`` usage, receiver
lifecycle, and a Clone smoke test over all event variants.
The module enforces a strict dependency rule: it imports nothing from
``crate::x11``, ``crate::audio``, ``crate::dictate``, or any other
display/pipeline module. Display adapters will depend on telemetry,
never the reverse.
This is sub-phase 1A of the telemetry system — pure skeleton with no
consumers or producers wired yet. Later sub-phases will integrate
event emission into the HTTP transcription path (1B/1C) and connect
the X11 overlay as a consumer (1D).
* [record] show waterfall spectrogram for recordings without transcripts. [Valentin Lab]
Recordings without a transcript preview now display an inline waterfall
spectrogram instead of blank space. A single background worker thread
computes the FFT data sequentially with yields between items to keep CPU
usage low. Both WAV and OGG files are supported via a new ``read_audio_as_i16``
helper.
* [dictate] add OGG format support with cache optimization. [Valentin Lab]
Add ``--upload-format <wav|ogg>`` CLI flag to enable OGG Opus encoding
during dictation. When OGG is selected, the encoder tees bytes to a cache
file, allowing retries to use the smaller OGG file (164 KB vs 1.2 MB WAV).
Also extend ``--input-audio-file`` to accept OGG Opus files, auto-detected
by extension and decoded via ``ogg`` + ``opus`` crates into PCM i16 chunks
through the ``AudioCapture`` trait.
Changes:
- ``src/transcription/mod.rs``: add ``UploadFormat`` enum (shared, ``clap::ValueEnum``)
- ``src/cli/def.rs``: add ``--upload-format`` argument
- ``src/cli/action/mod.rs``: plumb ``upload_format`` through action dispatch
- ``src/dictate/mod.rs``: compute OGG cache path, pass to streaming, use in retry
- ``src/dictate/streaming.rs``: spawn encode pipeline that tees OGG to cache
- ``src/dictate/toggle.rs``: forward ``--upload-format`` to daemon
- ``src/audio/file_source.rs``: add ``OggFileSource`` for OGG Opus decoding
* [dictate] add ``--no-paste`` flag and pipeline timing instrumentation. [Valentin Lab]
``--no-paste`` skips pasting the transcription into the focused
application (useful for benchmarking or headless use).
Add ``t_stop`` timing from recording stop through capture, WAV
flush, transcription, and first paste, logged at ``info`` level
for end-to-end latency analysis.
* [picker] write companion recording metadata YAML on text changes. [Valentin Lab]
Save a metadata YAML alongside the WAV so the record UI can
display transcript previews via its inotify watcher. Keyboard
edits are debounced (1 s); programmatic changes (row selection,
API result arrival) and window close are immediate. Also updates
the copy-to-clipboard icon from U+2398 to U+29C9 for consistency
with the record UI.
* [record] add dictate button and update copy icon in recordings UI. [Valentin Lab]
Add a "Transcribe recording" button on WAV files that have no
transcription yet — launches ``dictate --pick`` with the audio
file. Change the copy-to-clipboard icon from U+2398 to U+29C9.
Wire up ``.dictate-btn`` CSS class for consistent button sizing.
* [record] add copy-to-clipboard button for dictation transcripts. [Valentin Lab]
* [transcription] add configurable base URL for Mistral and OpenAI providers. [Valentin Lab]
Allow overriding the API base URL for both Mistral and OpenAI via
the ``url`` config field or ``TALK_RS_PROVIDERS_MISTRAL_URL`` /
``TALK_RS_PROVIDERS_OPENAI_URL`` environment variables. This lets
users point at self-hosted or API-compatible endpoints.
Batch transcription endpoints append ``/v1/audio/transcriptions``
to the base URL; realtime WebSocket endpoints convert the HTTP(S)
scheme to WS(S). Trailing slashes are trimmed automatically.
* [record] add resizable window, close button, and edge resize to ``record --ui`` [Valentin Lab]
Factorize undecorated-window helpers (``build_title_bar``,
``install_edge_resize``, close-btn CSS) into ``gtk_theme`` so both the
picker and the recordings browser share a single implementation.
The picker is refactored to call the shared helpers; the recordings
browser gains the same resize + close-button UX.
* [picker] add character-level diff highlighting, editable text area, and copy button. [Valentin Lab]
Integrate ``dissimilar`` crate for live character-level diff between
the original transcription and each candidate label: red strikethrough
for deletions, green background for insertions (both at 50% opacity).
Add ``escape_pango()`` and ``diff_markup()`` helpers for safe Pango
markup generation. Include an editable text area pre-populated on row
selection and a copy-to-clipboard button in the picker popup; edited
text is used on confirm.
* [picker] add waterfall spectrogram with playback controls and drag-to-seek. [Valentin Lab]
Render a frequency-vs-time waterfall strip above the candidate list
using ``generate_waterfall_columns()`` computed on a background
thread. A translucent cursor overlay tracks playback position with
sub-buffer interpolation for smooth movement.
Player gains ``pause()``/``resume()``/``seek()``/``progress()``
methods. The play button switches to Adwaita media icons and a
shared ``playing_flag`` replaces fragile label-text checks.
A rewind button and ``GestureDrag`` on the waterfall provide
drag-to-seek (auto-pauses during drag, resumes on release).
``README.org`` updated to mention the waterfall in the picker
feature list.
### Changes
* [transcription] always normalize file uploads to 16 kHz mono ``OGG`` [Valentin Lab]
Batch ``File`` uploads were sent to the providers verbatim. With the
new high-quality ``record`` output (48 kHz, up to 128 kbps), that meant
uploading far more data than either provider can use: Mistral Voxtral
and ``OpenAI`` Whisper / ``gpt-4o-transcribe`` both downsample to
16 kHz mono internally and ignore everything above 8 kHz. Sending
richer audio only wastes bandwidth, adds latency, and risks
``OpenAI``'s 25 MB cap — for zero accuracy gain.
Add ``normalize_file_for_upload`` (a single chokepoint shared by both
the Mistral and ``OpenAI`` batch transcribers): it decodes any
supported file via ``read_audio_as_i16`` (resampling to 16 kHz and
downmixing to mono) and re-encodes to ``OGG``/``Opus`` before upload,
advertising an ``.ogg`` file name. Decode failures fall back to the
original bytes so an upload that worked before keeps working. The
live streaming path is untouched — it already emits 16 kHz mono.
Adds tests covering stereo-44.1kHz to mono downmix, mono passthrough,
the missing-file error, and the undecodable-file raw-bytes fallback.
* [transcription] route realtime WS through unified transport. [Valentin Lab]
Steps 6-9 of the transport-consolidation plan at
``.sisyphus/plans/transport-consolidation.md``.
Step 6 (transport ``ws_upgrade``): the WS upgrade handshake now
shares the same growing-budget connection-retry schedule
(``CONNECTION_BUDGETS_SECS = [2, 5, 8, 11, 15]``) as
``http_request``, wrapping each
``tokio_tungstenite::connect_async`` attempt in a
``tokio::time::timeout`` so blackholes don't hang for the
OS-default ~125s TCP SYN timeout. Cancellation is wired through
the same ``tokio::select`` pattern. Tungstenite errors classify
into connect-retryable (``Io``, ``Tls``, ``ConnectionClosed``,
``AlreadyClosed``) vs permanent (``Http``, protocol violations).
Step 7 (realtime migration):
``MistralRealtimeTranscriber::transcribe_realtime`` and
``OpenAIRealtimeTranscriber::transcribe_realtime`` swap their
``with_retry`` + ``connect_async`` blocks for a single
``ws_upgrade`` call each. Same for the
``validate_realtime_session`` helpers. Provider-specific headers
(``OpenAI-Beta: realtime=v1``) are passed verbatim to the
transport.
The picker's "click T on realtime, nothing happens" symptom
diagnosed in plan section 1 dot 12 and Step 8 is structurally
fixed by this commit: the upgrade now uses the growing-budget
schedule (2+5+8+11+15 = 41s total instead of 5x15 = 75s) and
emits ``ConnectionEvent::RetryScheduled`` events. Step 10 wires
those events to the picker's UI sink to make the retries visible
to the user.
Step 9 (delete legacy primitives): delete
``transcription/transport/retry.rs`` entirely (``with_retry``,
``MAX_RETRIES``, and 5 unit tests). Delete ``CONNECT_TIMEOUT``
and ``build_client`` from ``transport/http.rs``. Delete
``WS_CONNECT_TIMEOUT`` from both realtime modules.
Outside the transport module, the words "retry" and "attempt" do
not appear anywhere related to network calls;
``paste.rs::FOCUS_MAX_RETRIES`` is the only remaining
``MAX_RETRIES`` and it governs X11 window focus, an unrelated
concern.
Test status: 433 lib tests green (down from 438 -- the 5 retry
tests are gone with the file). 6/10
``transport_integration`` tests green; the remaining 4 are
jobs-registry placeholders explicitly deferred to Step 11.
Net diff: -521 / +419 (transport surface keeps growing while
realtime and ``http.rs`` shrink).
* [transcription] migrate batch + validate + model-suggestions to ``http_request`` [Valentin Lab]
Steps 3-5 of the transport-consolidation plan at
``.sisyphus/plans/transport-consolidation.md``.
After this commit, **every batch HTTP call** to a transcription
provider funnels through ``transport::http_request``:
- ``transport::http::validate_model`` (validate preflight) — its
in-line growing-budget loop is deleted; the new transport's
``CONNECTION_BUDGETS_SECS`` covers the same ``[2, 5, 8, 11, 15]``
schedule.
- ``MistralBatchTranscriber::send_request`` — ``send_once`` is gone,
the multipart form is built by a factory closure that the
transport invokes per retry (``RequestBody::Multipart`` accepts a
``Box<dyn Fn() -> reqwest::multipart::Form + Send + Sync>`` because
``reqwest::multipart::Form`` is not ``Clone``).
- ``OpenAIBatchTranscriber::send_request`` — same treatment as Mistral.
- ``model_suggestions::fetch_transcription_models`` — its 5-attempt
in-line retry loop is deleted; the transport handles retries.
Stale-cache fallback preserved.
The notorious ``attempts=1, max_attempts=1`` lie at
``mistral.rs:359`` and ``openai.rs:316`` is removed: the providers
no longer build ``PipelineFailure`` themselves, so the
``http_request`` truthful counter (already in place from Step 2)
now surfaces correctly to every batch call site.
``transport/http.rs::build_client`` and ``CONNECT_TIMEOUT`` are
marked ``#[allow(dead_code)]`` with a TODO pointing at Step 9 (which
deletes them entirely along with the legacy ``with_retry`` path).
They're retained transiently because the realtime modules
(``realtime.rs`` / ``openai_realtime.rs``) still use the legacy
``with_retry`` upgrade path — Steps 6-7 migrate those to
``ws_upgrade``.
``RequestBody::Bytes`` now wraps ``Arc<Vec<u8>>`` so the transport
can ``.clone()`` the audio buffer cheaply across retries.
Test changes:
- ``validate_model_emits_one_retry_event_per_retry`` becomes
``validate_model_does_not_retry_on_malformed_response_body``.
Behaviour change rationale documented in the test: decode is a
*content* failure, not a transport failure; mid-body truncation
was already covered by the connection-retry layer. The pin on
the schedule ``[2, 5, 8, 11, 15]`` moves to
``transport_connection_phase_retries_with_growing_budget`` in
``tests/transport_integration.rs``.
Test status: 438/438 lib tests green, 21/22 integration suites
green; ``transport_integration`` keeps the 5 still-RED tests for
Step 6 (``ws_upgrade``) and Step 11 (jobs registry).
Net diff: -284 lines (697 → 413) with strictly better correctness.
* [transcription, error] consolidate HTTP pipeline failures into structured ``TalkError::Pipeline`` [Valentin Lab]
Replaces the older pattern of stuffing ``reqwest`` failures into
``TalkError::Config(String)`` / ``TalkError::Transcription(String)``
with a single structured variant
``TalkError::Pipeline(Box<PipelineFailure>)`` that carries the
provider, phase, attempts/max counts, URL, and a typed
``PipelineFailureKind`` (``Network`` / ``HttpStatus`` /
``ModelRejected`` / ``Decode``).
Motivation: the prior string-format helper
``format_reqwest_error_with_timers`` produced messages like:
Configuration error: OpenAI model validation failed
(preflight to /v1/models): error sending request for url
(https://...) [kind=timeout, name=connect_timeout, budget=2s,
url=https://...] -> client error (Connect) -> operation timed out
Three problems with that output:
1. ``Configuration error:`` prefix was misleading — a network
timeout against the preflight endpoint is not a config issue.
2. The URL was printed twice (once in the ``reqwest::Error``
preamble, once in the ``[url=...]`` tag).
3. The trailing chain layers ``client error (Connect)`` and
``operation timed out`` just restated what
``name=connect_timeout`` already conveyed; the verbosity drowned
out useful layers like DNS lookup failures or
``ECONNREFUSED (os error 111)``.
The new ``PipelineFailure::Display`` impl produces a clean
one-liner with structural source-chain dedup:
Mistral model validation failed [name=connect_timeout,
budget=2s, url=https://mistral.vps-03.0k.io/v1/models]
(after 5/5 attempts)
DNS / ECONNREFUSED / TLS layers are kept (they add information not
in the structured fields); ``client error (Connect)``,
``operation timed out``, etc. are dropped.
Architectural consolidation (per
``architectural-consolidation.md``):
- **Outcome A on ``TalkError``** — extended the existing error
vocabulary with one new variant rather than introducing a
parallel ``PreflightError`` / ``RequestError`` type. Boxed to
keep ``Result<T, TalkError>`` stack size sane.
- **Outcome B on the formatter** — the responsibility "render a
reqwest failure as text" moved out of
``format_reqwest_error_with_timers`` and into
``PipelineFailure::Display``. The transport layer keeps the
reqwest-classification mechanics
(``classify_reqwest_error``, ``build_pipeline_failure_kind``,
``TimerSpec`` → ``TimerLabel`` conversion); the error layer
knows how to display the structured value. No reqwest types
leak into ``error.rs``.
- **Outcome A on ``TranscriptionEvent``** — no new events
invented; the existing ``RetryScheduled`` is now emitted
between attempts of the validate-cache-miss retry loop, with
``attempt = 1..=4`` / ``max = 4`` for the 5-attempt budget.
Same vocabulary as the existing ``with_retry`` for the
transcription request, so the picker UI's ``retry N/M…``
rendering covers both phases without picker code changes.
- **Structural-first model-error detection** — the
bail-on-permanent logic in ``with_retry`` now matches
``PipelineFailureKind::ModelRejected`` structurally before
falling back to the legacy string match. Pre-migration call
sites that still produce string-stuffed errors continue to
work; structural producers always win.
Scope: in scope for this change is the HTTP preflight (validate)
and batch HTTP transcription request paths. Realtime/WebSocket
transport (``MistralRealtimeTranscriber``, ``OpenAIRealtimeTranscriber``)
stays on the legacy ``TalkError::Transcription(String)`` shape —
it has different transport semantics (WS upgrade, session events)
and is a separate consolidation when the user wants it.
User-visible effect: picker rows now show ``pre-validating
model…`` followed by ``retry 1/4…``, ``retry 2/4…``, ``retry
3/4…``, ``retry 4/4…`` during the validate-phase retries (this
was the missing-progress complaint). When validation or
transcription ultimately fails, the error message lacks the
``Configuration error:`` prefix, prints the URL once, and dedups
restatement layers so only novel information remains in the
chain.
Files touched:
- ``src/error.rs``: new ``Pipeline`` variant, ``PipelineFailure``
struct, ``PipelinePhase`` / ``PipelineFailureKind`` /
``NetworkKind`` / ``TimerLabel`` types, ``Display`` impl with
source-chain dedup heuristic. 7 unit tests for Display
variants, dedup, and ``From<PipelineFailure>`` transparency.
- ``src/transcription/transport/http.rs``: removed
``format_reqwest_error_with_timers`` and ``attribute_timer``;
added ``build_pipeline_failure_kind`` and inner
``classify_reqwest_error`` helper. Reworked
``validate_model_uncached`` to take a sink, emit
``RetryScheduled`` between attempts, and return structured
``Pipeline`` errors. Replaced 5 legacy string-format tests
with 5 structural classifier tests; added 3 new tests covering
retry-event emission and structural ``is_model_error``.
- ``src/transcription/mistral.rs``,
``src/transcription/openai.rs``: ``send_once`` produces
``PipelineFailure { phase: Request, .. }`` instead of
string-stuffed ``Transcription(String)``. Updated the
``user_attended_policy_omits_request_wall_clock_attribution``
regression test to assert on the structured shape.
- ``src/transcription/mod.rs``: ``is_model_error`` structural
fast path on ``PipelineFailureKind::ModelRejected``; legacy
string match preserved as fallback.
Verification: ``cargo fmt`` clean, ``cargo clippy --all-targets``
clean, ``cargo test`` 428 passed / 0 failed, ``cargo build`` and
``cargo build --release`` both succeed. No new
``unwrap()``/``expect()`` outside ``#[cfg(test)]``; no trailing
whitespace; no dead-code pragmas.
* [audio, dictate] gate periodic boop on auto-pause ``LISTENING`` state. [Valentin Lab]
The periodic boop heartbeat used to play unconditionally for the entire
recording, including while the user was actively speaking. The user's
mental model is that boops belong to the yellow ``LISTENING`` badge —
i.e. the auto-pause state where audio forwarding is held back during
silence. Make the implementation match that model.
* Extend ``SoundPlayer::start_boop_loop`` with a positive ``play_when``
gate (in addition to the existing negative ``suppress`` gate). Boops
fire iff ``play_when_ok && !suppressed``.
* Wire the overlay's existing ``pause_flag`` atomic into ``play_when``
at the call site in ``dictate``. No new shared state — the same
flag the audio tee already consumes for sample gating.
* Anchor the ``interval`` clock to the rising edge of ``play_when``,
not to loop creation. The first boop after entering ``LISTENING``
always lands a full ``interval`` later — never sooner — and brief
silences shorter than ``interval`` produce zero boops. A falling
edge mid-period cancels the in-flight wait so the next listening
period starts a fresh clock.
* Extract the loop body into a free ``run_boop_loop`` async function
so the state machine is unit-testable without an audio device.
Side effects: with ``--no-auto-pause`` the boop is silent for the
entire recording (the gate never opens), which is consistent with the
user's mental model. The dead-signal alert path still suppresses
boops in-place via ``suppress`` and does NOT reset the phase, so the
alert tone continues to stand alone.
Tests cover: ungated periodic firing, gated-off silence,
suppress-overrides-play-when, the phase-reset invariant on the rising
edge, and the no-emit guarantee for sub-interval listening bursts.
* [picker] unify row order, add T/↻ action button, preserve text area. [Valentin Lab]
Three connected picker UX fixes that all live in
``src/dictate/picker/ui.rs``:
1. **Deterministic row order.** Previously rows were appended in four
buckets (cached, batch-pending, realtime-pending, deferred), so a
given ``(provider, model, streaming)`` triple's screen position
depended on whether it had a cached transcription -- the order
shifted as results arrived. Now all rows go through one unified
sort on ``(provider_rank, model, streaming)`` where the config's
default provider gets rank 0 and others rank 1 (alphabetical
among them), and a single loop renders them.
2. **Unified ``T``/``↻`` action button.** Replaces the previous
``▶`` deferred-transcribe button (only on never-run rows) and
``↻`` retry button (only on errored rows) with a single action
button present on every row in the leftmost column. Icon is
``T`` when the row has no transcription (first run / spinner /
deferred / error / "no speech detected") and ``↻`` once a
non-empty transcript exists -- so the user can always re-run a
model, including successful ones. Button is sized 28x28 with
``padding: 0`` to match the play/copy icon footprint.
3. **Preserve text area on non-transcribed rows.** A new per-row
``has_transcription: Rc<RefCell<Vec<bool>>>`` tracks whether the
row currently holds a non-empty transcript. In
``connect_row_selected`` the ``buf_sel.set_text(text)`` call is
now guarded by this flag, so selecting a spinner / deferred /
error / empty-result row no longer wipes whatever the user is
viewing or editing. The flag flips true on first non-empty
``Candidate``/``StreamUpdate`` arrival, false on user retry-click,
error, and ``InitialBatchDone`` "no response".
Verified locally with ``./autogen.sh``, ``cargo fmt``, ``cargo clippy
--all-targets``, ``cargo test``, ``cargo build``, ``cargo build
--release``, plus a live UI smoke test confirming row order is
content-independent, the action button is leftmost and square, and
selecting a non-transcribed row leaves the text area intact.
* [transcription, dictate, record] pick-file waterfall + shared retry. [Valentin Lab]
Introduce a layered cache architecture for transcription:
- Layer 1 (``recording_cache``): ``<stem>.pick.yml`` is the
authoritative transcript for a recording. ``get_transcript``
returns ``Available(text)`` / ``InProgress`` / ``NotAvailable``.
Locks (``<stem>.pick-lock.yml``) coordinate producers.
- Layer 2 (``produce_transcript``): check pick, acquire lock,
delegate to Layer 3, write pick, release lock. Used by
``transcribe`` command (default options) and by ``dictate``
Mode C (file input, default options).
- Layer 3 (``transcribe_audio``): per-model sidecar cache
(``<stem>_<provider>_<model>_<mode>.yml``) + per-model lock.
New ``allow_api: bool`` parameter lets the picker probe the
cache without triggering API calls; returns ``CacheOnly`` on
miss.
- Layer 4 (``src/transcription/transport/``): single retry
primitive ``with_retry`` used by **both** batch HTTP POST and
realtime WebSocket upgrade. The 85-line retry loop in
``src/dictate/mod.rs`` is removed.
Concept consolidation:
- ``BatchTranscriber::transcribe_file`` and ``transcribe_stream``
collapse into one ``fetch_transcription(body: TranscriptionBody)``
method. The trait is now ``pub(crate)`` — outside callers go
through ``transcribe_audio`` or ``produce_transcript``.
- ``RealtimeTranscriber`` is also ``pub(crate)``. Its WebSocket
upgrade handshake is now wrapped in the shared ``with_retry``.
- ``src/transcription/http.rs`` moves to
``src/transcription/transport/http.rs``.
Consumer changes:
- ``dictate`` writes a pick on every successful transcription
unless a specific provider/model/diarize option is given.
Mode C (``--retry-last`` / ``--input-audio-file`` with default
options) short-circuits: if a pick exists, paste it directly.
- ``transcribe`` command's default branch now polls
``get_transcript`` on ``TranscriptInProgress``.
- Record UI reads the pick file only — no sidecar fallback.
Displays ``(no text)`` for empty picks, ``(transcription
ongoing)`` for in-progress, audio player when no pick exists.
- Picker probes the sidecar cache via
``transcribe_audio(allow_api=false)`` instead of enumerating
sidecar files directly; only the default model auto-transcribes,
other models show a transcribe button.
- Old picker JSON cache (``~/.cache/talk-rs/picker-results/``) is
deleted — ``src/dictate/picker/cache.rs`` removed.
Error variants added: ``TranscriptInProgress``, ``CacheOnly``,
``ModelInProgress``.
Telemetry: ``Arc<dyn TelemetrySink>`` threaded from CLI through
``produce_transcript`` / ``transcribe_audio`` /
``BatchTranscriber::fetch_transcription`` / ``with_retry`` /
``ProgressBody`` — retry events observable by the UI visualizer
without any new propagation mechanism.
Tests: 350 pass, up from 330. New unit tests cover the retry
primitive (initial-success, transient-then-success, model-error
bail, exhaustion), ``get_transcript`` state machine, pick-lock
acquire/release idempotence, per-model lock independence, and
pick round-trip.
* [transcription] unify all batch transcription behind one function and one cache. [Valentin Lab]
Every batch-from-file transcription now goes through a single entry
point ``transcription::transcribe_audio`` which checks a sidecar
cache before calling the provider and stores the result after.
Callers never see the cache.
``recording_cache::TranscriptionCache`` provides the abstract cache
API (``get`` / ``store``). Storage is currently YAML files next to
the source audio; the implementation detail is hidden. All sidecar
types gain ``Deserialize`` so cached results can be read back into a
``TranscriptionResult``.
``src/transcribe.rs`` (CLI ``talk-rs transcribe``) collapses to
config resolution + one ``transcribe_audio`` call + output. The
manual transcriber creation, sidecar writing, and model-name
resolution are gone.
``src/dictate/picker/backend.rs`` now takes ``Arc<Config>`` instead
of pre-built ``Box<dyn BatchTranscriber>`` and calls
``transcribe_audio``. The picker no longer creates transcribers
itself (``picker/mod.rs``, ``picker/ui.rs``); re-opening the picker
on the same audio file hits the cache and returns instantly.
``src/dictate/mod.rs`` retry loop uses ``transcribe_audio`` (cache
may short-circuit repeated retries on the same OGG). The
streaming and realtime paths — which cannot use ``transcribe_audio``
because the audio file does not exist at transcription start — now
call ``TranscriptionCache::store`` after completion, followed by a
separate ``write_last_pointers`` for dictate-specific symlinks.
``write_recording_metadata`` (picker/mod.rs) is deleted. The
picker no longer overwrites provider sidecars on text-edit or
selection — that is a future ``<stem>.yml`` user-selection sidecar
concern.
``timestamp_granularities`` is now sent unconditionally on all
Mistral API requests (not just ``--diarize``), restoring per-segment
timing from ``voxtral-mini-2602+`` which requires the explicit
parameter.
* [telemetry, transcription, overlay, paste] add three independent-scale throughput tracks with download byte streaming and paste character tracking. [Valentin Lab]
Transcription backends (``mistral``, ``openai``) now stream the HTTP
response body via ``bytes_stream()`` instead of buffering the whole
JSON payload, emitting per-chunk ``DownloadProgress`` events as bytes
arrive.
``paste_text_to_target`` accepts a ``TelemetrySink`` and emits
``PasteProgress`` after each chunk so the overlay can visualise paste
throughput alongside upload and download.
The overlay replaces the single upload throughput bar with three
independent tracks (upload, download, paste), each owning its own
16 px budget, peak tracking, and history ring buffer. Independent
scales ensure a ~1 KB JSON download is visually comparable to a
~100 KB audio upload — shared-budget scaling made the smaller track
invisible at a 500:1 byte ratio.
* [dictate] keep overlay visible during paste and emit paste telemetry events. [Valentin Lab]
Move ``o.hide()`` after ``paste_text_to_target()`` so the overlay
(dimmed waterfall, phase colours, throughput bars) stays on screen
throughout the paste phase.
Emit ``TranscriptionEvent::PasteStarted`` before paste and
``TranscriptionEvent::PasteCompleted`` after it, so the overlay
phase state machine shows a green band during paste
(``PasteStarted`` → Done, ``PasteCompleted`` → Idle).
* [overlay, transcription] add upload throughput bars and fix dead-signal gate during transcription. [Valentin Lab]
Layer 3b throughput bars: each waterfall column now draws a
phase-coloured vertical bar growing downward from the phase line,
proportional to the upload bytes transferred in that time slice
(peak-normalised, max 48 px). New state tracks
``current_upload_bytes``, ``prev_upload_bytes``, ``upload_peak_delta``,
and a parallel ``throughput_history`` vec. Bars render in all three
display branches (transcribing, auto-paused, normal).
Dead-signal gate fix: column-push was gated on ``!no_sound_active``,
which blocked all pushes during transcription because
``capture.stop()`` makes the ring buffer go stale and the dead-signal
detector fires. Changed to ``(!no_sound_active || is_transcribing)``
so columns keep advancing. Also reordered render branches so
``is_transcribing`` takes priority over ``no_sound_active``, preventing
the "NO SOUND" icon from hiding the transcribing waterfall.
``ProgressBody`` wrapping for retries: ``transcribe_file`` in both
``mistral.rs`` and ``openai.rs`` now reads the audio into a ``Vec<u8>``
and wraps it with ``ProgressBody`` + ``Body::wrap_stream``, so retry
attempts emit upload telemetry and throughput bars are visible during
file-based retries.
* [overlay, dictate] replace static transcribing badge with dynamic phase-colour waterfall. [Valentin Lab]
The static "transcribing" PNG badge is replaced by a live render that
keeps the 60 fps waterfall loop running through the transcription phase:
- Waterfall continues scrolling with empty columns at 30 % brightness,
making elapsed transcription time visible at a glance.
- A 2-pixel phase-colour line at the top of the spectrogram area shows
the HTTP lifecycle in real time: connecting (dim blue), uploading
(bright blue), waiting for server response (amber), receiving (teal),
done (green), error (red).
- "TRANSCRIBING" text is rendered in light blue, following the same
pattern as the auto-pause "LISTENING" overlay.
- The telemetry ``broadcast::Receiver`` is wired from the
``BroadcastSink`` broker (created in ``dictate``) into the overlay
thread; events are drained non-blockingly each frame via
``try_recv()``.
- Broker creation in ``src/dictate/mod.rs`` is moved before the overlay
so the receiver is available at construction time.
The static PNG path is kept as a fallback for the edge case where
``Transcribing`` is received outside an active recording session.
Layer 3b (byte-throughput bars) is deferred to a later commit.
* [transcription, dictate] wire telemetry events into batch transcription paths. [Valentin Lab]
Add ``set_sink`` method to the ``BatchTranscriber`` trait with a default
no-op implementation so concrete backends can accept a ``TelemetrySink``.
In ``MistralBatchTranscriber::transcribe_stream``, wrap the audio buffer
with ``ProgressBody`` so ``ConnectionEstablished`` / ``UploadProgress`` /
``UploadComplete`` events flow as bytes are sent. Both
``transcribe_stream`` and ``transcribe_file`` now emit
``RequestStarted`` / ``ResponseHeaders`` / ``RequestCompleted`` boundary
events on every exit path.
Apply the same boundary-event pattern to
``OpenAIBatchTranscriber::transcribe_file``.
In ``dictate::mod``, create an ``Arc<BroadcastSink>`` before the
transcriber and inject it via ``set_sink``. The retry loop emits
``RetryScheduled`` and re-injects the sink into each freshly-created
retry transcriber. No consumer subscribes yet — display wiring follows
in sub-phase 1D.
Remove the three ``#[allow(dead_code)]`` annotations from
``http.rs`` now that ``ProgressBody`` is used by ``mistral.rs``.
* [overlay] keep waterfall scrolling during auto-pause and dim history. [Valentin Lab]
Three visual improvements to the X11 recording/listening badge:
1. Waterfall continues scrolling during auto-pause by pushing empty
columns (``vec![0.0; SPEC_H]``) instead of freezing. The x-axis now
represents wall-clock time: silence appears as a growing empty hole
beside the dimmed audio history, which is useful for diagnosing
pause behavior.
2. During auto-pause the spectrogram/amplitude/spectrum renderers
accept a ``dim`` parameter. Normal mode passes 1.0; auto-pause
passes ``DIM_FACTOR_PAUSED`` (0.3) before overlaying the LISTENING
text and pause bars at full opacity.
3. Column advance rate decoupled from render frame rate via
``COLUMN_PERIOD_FRAMES = 2``: the waterfall advances at 30 cols/sec
while ``FPS`` stays at 60 so the red-dot pulse remains smooth. The
visible time window roughly doubles from ~4.4 s to ~8.8 s. Peak
tracking still runs every frame for stable normalization.
New unit tests cover the dim path and the zero-column ("hole")
rendering.
* [record] match ``memo`` audio filename scheme. [Valentin Lab]
Audio filenames produced by ``talk-rs`` now follow the exact scheme
used by the ``memo`` tool: ``YYYY-MM-DDTHH-MM-SS+ZZZZ.ogg`` -- an
ISO 8601 local timestamp with a numeric timezone offset, colons in
the time portion replaced by dashes to stay filesystem-safe.
Recordings from both tools can now coexist in the same directory
and sort chronologically.
``src/record/mod.rs`` -- ``default_filename()`` drops the legacy
``memo-`` prefix and switches the chrono format to
``%Y-%m-%dT%H-%M-%S%z.ogg``. The user-facing ``record`` command
therefore writes ``<output_dir>/YYYY/MM/2026-04-11T13-15-52+0200.ogg``
for auto-named recordings. ``src/cli/def.rs`` help text is updated
to reflect the new default path.
``src/recording_cache.rs`` -- ``generate_recording_path()`` now
includes ``%z`` in the timestamp used for the dictate cache at
``~/.cache/talk-rs/recordings/``. The paired metadata YAML sidecar
inherits the new stem automatically, so dictate now writes e.g.
``2026-04-11T13-15-52+0200_mistral_voxtral-mini-2507_batch.yml``.
``write_last_paste_state()`` picks up the same format for the
``timestamp`` field in ``last_paste.yml`` for internal consistency.
The existing unit test ``test_resolve_output_path_no_args_...`` and
the integration test ``test_record_default_filename_format`` are
rewritten to parse the generated filename through
``chrono::DateTime::parse_from_str`` and assert a round-trip instead
of string-matching a prefix, because the real specification is
"the stem must parse back as a local datetime with timezone offset".
Backward compatibility: existing cached recordings and legacy
``memo-*`` record files still list, play, delete, and rotate
correctly because the entries reader and cache rotation sort by
basename rather than by format. Old files gradually age out of the
dictate cache via ``rotate_cache()``.
* [record] namespace auto-generated recordings by ``YYYY/MM`` [Valentin Lab]
Auto-generated recording filenames are now placed under ``YYYY/MM``
subdirectories of the configured ``output_dir``, mirroring the layout
used by the ``memo`` tool so long-running users do not end up with
thousands of files in a single directory.
``src/record/entries.rs`` now walks nested ``YYYY/MM`` directories when
listing recordings for the ``record --ui`` browser. Mixed flat plus
nested layouts are supported by sorting entries on basename so existing
and newly created recordings interleave chronologically.
``src/record/ui.rs`` now installs its ``inotify`` watch recursively so
newly created ``YYYY/MM`` subdirectories are tracked as soon as they
appear. The GTK browser refreshes rows under the correct subtree when
files land in freshly created month directories, fixing a regression
where the UI stopped auto-updating after the first recording of a new
month.
* Harden HTTP timeouts and add ``--log-file`` support. [Valentin Lab]
Replace per-request wall-clock timeouts with TCP-level dead-connection
detection so that legitimate long uploads and slow server processing
are no longer killed prematurely:
- Upgrade ``reqwest`` 0.11 → 0.13 (``hyper-rustls`` TLS backend).
- Configure ``tcp_user_timeout`` (Linux), ``tcp_keepalive``,
keepalive interval/retries in the shared ``build_client()``.
- Remove ``BATCH_FILE_TIMEOUT``, ``RETRY_TIMEOUT``, and all
``tokio::time::timeout`` / ``tokio::select!`` wrappers around
transcription calls — stalls are now caught at the TCP layer
within a few seconds.
- Centralise ``model_suggestions`` to use the shared HTTP client.
- Switch error formatting to ``{:#}`` for reqwest error chains.
Add persistent file logging behind ``--log-file`` / ``$TALK_RS_LOG_FILE``:
- ``log::setup()`` accepts an optional file path; a second ``fern``
dispatch appends plain-text Info+ logs with timestamps.
- Auto-truncate at 2 MiB to prevent unbounded growth.
- Propagate the flag via environment so child processes inherit it.
- Add ``chrono`` timestamps to stderr output as well.
- Enable clap ``env`` feature for ``$TALK_RS_LOG_FILE`` support.
* [picker] defer non-default model transcription with on-demand button. [Valentin Lab]
Only the default transcription model is transcribed immediately when opening
a recording in picker mode. Other configured models are deferred and presented
with a transcribe button (▶) for on-demand transcription, reducing unnecessary
API calls.
* [record] show transcribe button on all recordings. [Valentin Lab]
Allow users to re-transcribe recordings that already have a transcript.
Previously the button was only shown for recordings without transcripts.
* [dictate] migrate dictation cache from WAV to OGG/Opus format. [Valentin Lab]
Unified cache format across realtime and batch dictation paths. Both now
write ``.ogg`` cache files using ``OggOpusWriter``. Removed redundant
batch-mode OGG tee. Updated cache management (rotation, last-recording
pointer, listing, retry) to use ``.ogg`` paths. Recordings browser UI
now shows OGG cache entries. Delete operation is backward-compatible for
legacy ``.wav`` files. Removed ``prune_ogg_cache()`` since ``rotate_cache()``
handles all cache maintenance.
Files changed:
- ``src/dictate/realtime.rs`` — ``ogg_recording_task`` replaces ``wav_recording_task``
- ``src/dictate/streaming.rs`` — uses ``ogg_recording_task``, removed OGG tee
- ``src/dictate/mod.rs`` — unified cache path, removed redundant prune call
- ``src/recording_cache.rs`` — ``.ogg`` paths, pointers, rotation
- ``src/record/entries.rs`` — ``list_cache_recordings()`` scanning ``.ogg``
- ``src/record/ui.rs`` — updated imports/references
- ``src/record/audio.rs`` — minor update
- ``README.org``, ``sample-metadata*.yml``, ``src/cli/def.rs`` — doc updates
* [record] extract ``audio_player_bar`` widget and improve recordings UI. [Valentin Lab]
Move waterfall spectrogram, playback cursor, and play/pause/rewind
controls into a shared ``audio_player_bar`` widget under ``src/widgets/``.
Replace the inline waterfall worker pool in ``record --ui`` with the new
widget. Restore a simple play button for entries with transcripts,
read OGG transcript previews from companion YAML metadata, and add
incremental FileMonitor row updates to avoid full section rebuilds.
* [dictate] always cache OGG format and prune old files. [Valentin Lab]
OGG cache files are now saved alongside WAV during dictation,
regardless of ``--upload-format`` setting. Old OGG files are
automatically pruned to keep only the 10 most recent, preventing
unbounded cache growth.
* [record] defer recording list loading to after window paint. [Valentin Lab]
Show the recording browser window immediately with a "Loading
recordings…" indicator; populate data and set up file watchers in
an idle callback triggered on the first ``map`` signal so the user
never stares at a blank wait.
Also add ``log::debug!`` timing traces to ``list_ogg_recordings``,
``list_wav_recordings``, and the window lifecycle.
* [overlay] make badge background opaque black. [Valentin Lab]
- Change ``BG_COLOR`` alpha from 0x00 to 0xFF for solid black background
- Force spectrogram pixels to full opacity (``color[3] = 0xFF``) so they
don't create transparent holes over the black background
- Apply rounded shape mask in ARGB visual path so corners stay rounded
with opaque background
- Update two tests to assert against ``BG_COLOR`` instead of alpha=0
* [record] optimize OGG duration to O(1) seek-from-end. [Valentin Lab]
Replace sequential packet iteration with a tail-read approach:
seek to the last ~64 KB, scan backward for the final ``OggS``
page header, and read the granule position directly. Constant
time regardless of file size.
Add tests for valid, too-small, and non-OGG files.
### Fix
* [audio] consume PCM buffer with a cursor in ``write_pcm`` to avoid quadratic drain. [Valentin Lab]
``OggOpusWriter::write_pcm`` drained ``samples_per_frame`` from the front
of ``pcm_buffer`` inside the encode loop. ``Vec::drain`` from the front
memmoves the whole remaining tail down on every iteration, making the
function quadratic in the number of buffered samples.
This was harmless for the streaming call sites (``src/dictate`` and
``src/record``), which feed small chunks as they arrive so the buffer
never holds more than about one frame. It was severe for
``encode_16k_mono_ogg`` in ``src/transcription/mod.rs``, reached via
``normalize_file_for_upload``: that path decodes a whole file to PCM and
passes it in a single call, so the buffer starts at full file length and
is drained 320 samples at a time. A long recording could spend hours
spinning on memmove before contacting the provider at all.
Consume the buffer with an index cursor and drain once at the end of the
call. Behaviour-preserving by construction: the leftover tail after the
single ``drain(..pos)`` is exactly the samples after the last full frame,
which is what ``finalize()`` pads and flushes.
Measured, release build, one call with 20 minutes of 16 kHz mono audio:
455.6 s before, 15.6 s after. The residual is the Opus encode itself,
which is linear and unchanged.
Tests keep a reference copy of the old draining implementation and assert
the new one is byte-identical to it, for both the bulk and the
irregularly-chunked call shapes. The regression guard asserts on a
test-only counter of samples relocated by front-drains rather than on
wall-clock time: at test-affordable sizes the linear Opus encode dominates
and the measured ratio between the two was only 1.2x, too close to noise
to discriminate, whereas the counter separates them by three orders of
magnitude deterministically.
* [paste, clipboard] gate chunk advance on target client re-fetch with retry. [Valentin Lab]
Chunked clipboard paste split long text and advanced to the next chunk
as soon as the ``served_count`` signal went positive. That counter is
incremented by ANY X11 client fetching the ``CLIPBOARD`` selection —
including clipboard managers — so the selection was routinely
overwritten before the target application had actually pasted a chunk,
dropping that chunk and (via a lingering serve thread) duplicating the
last one.
Gate each chunk on the TARGET window's X11 client re-fetching it,
identified by client-base (``requestor & !resource_id_mask``) rather
than the ephemeral requestor window id:
- The serve thread now records ``UTF8_STRING`` fetches per client-base
(our own read-back client is excluded), and the ``clipboard`` node
waits for the target client specifically. Chunk 1 LEARNS the
target's per-paste fetch count through a quiescence window; later
chunks CONFIRM the same count before the selection is overwritten.
- Transient focus failures (the paste keystroke sent before keyboard
focus was effective) are recovered by re-focusing the target window
and re-sending the keystroke up to ``target_fetch_retries`` times
(default 2) before giving up.
- When the target never fetches within ``chunk_fetch_timeout_ms``
(default raised 300 -> 500, now a per-attempt deadline) after all
retries, the paste ABORTS rather than silently corrupting the
document, surfacing a red overlay and an alert sound; the user's
original clipboard is still restored.
- ``settle_before_restore`` is removed — the target-confirmation makes
the restore race impossible by construction.
Blind pastes with no known target window (the realtime per-segment
path) fall back to the previous ``served_count`` gate unchanged.
New knobs on the ``clipboard`` node: ``target_fetch_retries`` and
``target_quiescence_ms``; ``restore_settle_ms`` is now a no-op kept for
backward compatibility.
* [paste, clipboard] wait for paste target to fetch before overwriting clipboard. [Valentin Lab]
Long pastes split text into chunks and, after pasting, restore the
user's original clipboard. Both the inter-chunk overwrite and the
final restore happened after a FIXED sleep rather than waiting for the
target window to actually fetch the offered content. A slow target
could then pull the WRONG clipboard generation — typically grabbing
the restored value instead of the last chunk — leaking the old
clipboard into the document and dropping the final chunk.
Block on the existing ``served_count`` signal instead of guessing with
a sleep:
- ``X11Clipboard::wait_until_served`` polls until the offered content
is fetched (or a timeout), so ``paste_one`` no longer overwrites a
chunk the target has not pulled yet.
- ``settle_before_restore`` waits for the last chunk's fetch activity
to stay stable for ``restore_settle_ms`` before restoring, closing
the restore race. Same treatment for the realtime per-segment path.
- On timeout a ``WARN`` is emitted so the corruption-prone case is no
longer silent.
Requestor-identity matching is intentionally avoided: the target
window id never appears as the ``SelectionRequest`` requestor, so the
served-count signal is the reliable discriminator.
The two timings are config-tunable via ``paste.restore_settle_ms``
(default 200) and ``paste.chunk_fetch_timeout_ms`` (default 400).
* [log] let ``--log-file`` honor ``-vvv`` so trace reaches the file. [Valentin Lab]
Previously the ``--log-file`` sink was pinned to Info, so the
paste-diagnostic traces (emitted at Trace) never reached the file even
with ``-vvv`` — they only went to stderr. In daemon mode that made
them effectively unreachable from the user-chosen log file.
Move the per-sink level filters onto the child dispatches and leave the
root dispatch at its pass-all default. In ``fern`` a parent
``Dispatch::level()`` is a hard gate children cannot exceed in
verbosity, so the old structure (root pinned to the stderr level, file
child nominally at Info) silently dropped everything below the root
level before the file child ever saw it — the "file always captures at
least Info" promise was in fact never kept at the default Warn
verbosity.
The file now captures ``max(base_level, Info)``: at least Info so it
stays useful at the default verbosity, and as deep as Trace under
``-vvv`` so the paste diagnostics land in the file. Extract
``base_level_for`` and ``file_level_for`` as pure helpers with unit
tests covering the floor and the ``-vvv`` case.
* [config] reject relative ``output_dir`` with a clear error. [Valentin Lab]
The generated ``config.example.yaml`` and the README state that
``output_dir`` "Must be an absolute path", but ``validate_config`` only
checked that it was non-empty. A relative value was silently accepted
and then resolved against the process working directory at runtime,
which is unpredictable — the toggle daemon spawns with an inherited,
effectively arbitrary CWD, so recordings could land anywhere.
Enforce the documented contract: ``validate_config`` now returns a
``output_dir must be an absolute path`` error for any non-absolute
value (whether from the file or the ``TALK_RS_OUTPUT_DIR`` env
override). Document the requirement on the struct field, and add
regression tests covering both the YAML and env paths.
Fixes #7
* [transport] extend connection retry budget for large uploads. [Boris Gallet]
The connect_budget timeout wraps the entire HTTP send request
(TCP + TLS + body upload) via tokio::time::timeout, not just the
connection establishment. For large audio files (>10 MB), the
upload to the Mistral API takes ~36s at ~530 KB/s, exceeding even
the 15s maximum budget on the 5th attempt.
Add two generous retry slots (30s, 120s) at the end of the
growing-budget schedule [2, 5, 8, 11, 15, 30, 120] so realistic
upload latencies (500-2000 KB/s) are accommodated.
Tested: 19.5 MB / 1h audio file transcribed successfully (~71s
total, succeeds on the 30s or 120s attempt depending on network
conditions).
* [recording_cache, dictate] persist diarization segments in YAML sidecar. [Boris Gallet]
`TranscriptionCache::store()` was silently discarding diarization
segments when writing the YAML metadata sidecar. The
`RecordingMetadata` struct had no `diarization` field, and
`dictate/mod.rs` explicitly set `diarization: None` when building
the cache entry.
This meant that `talk-rs dictate --diarize --timestamp --no-paste`
would print the correct timestamped speaker output to stdout, but
the YAML sidecar next to the OGG file would lose all speaker
attribution — making it impossible to reconstruct who spoke when
from the cache alone.
Changes:
- Add `CommonDiarizationSegment` with `speaker, start, end, text`
- Add `diarization` field to `RecordingMetadata`
- Add `common_diarization_from_result()` converter
- Wire `diarization` through `write_metadata_to_dir`,
`write_metadata`, `TranscriptionCache::store`, and
`into_transcription_result`
- Fix `dictate/mod.rs` to preserve `result.diarization` instead of
overwriting with `None`
- Update `transcribe.rs` test to pass diarization
- Add `diarization: None` to all existing test `RecordingMetadata`
initializers
TDD: `test_transcription_cache_round_trip_with_diarization` verifies
that SPEAKER_00 / SPEAKER_01 survive store → YAML → read.
* [dictate, toggle] forward all missing flags to daemon. [Boris Gallet]
`toggle_spawn() was silently dropping several dictate flags when
spawning the daemon process:
- --timestamp (new in PR)
- --no-paste
- --pick
- --retry-last
- --replace-last-paste
- --input-audio-file
- --output-yaml
This caused `talk-rs dictate --toggle --no-paste` to paste anyway,
and `talk-rs dictate --toggle --timestamp` to produce untimestamped
output.
Refactor: extract `build_daemon_args()` as a pure function, pass
`&DictateOpts` instead of individual parameters, add 10 unit tests
covering every forwarded flag.
* [overlay] flag NO SOUND only on stuck-at-rail, not on silence. [Valentin Lab]
Replace the variance-based dead-signal heuristic with a stuck-at-rail
detector. A disconnected device pins every sample at the ``i16`` rail
(constant ~-1.0); only that flat-and-near-rail signature now triggers
the NO SOUND warning.
The old variance ceiling mis-flagged the exact-zero digital silence that
Bluetooth HFP mics emit between speech (variance == 0) as a dead device,
producing spurious NO SOUND warnings whenever the user paused talking.
Add ``is_stuck_at_rail()`` helper with 7 unit tests covering a
disconnected device (both rails), zero silence, a quiet noise floor,
loud speech, a constant mid-level DC offset, and empty input.
* [openai] migrate realtime to GA API. [Valentin Lab]
short body.
* [picker, error, telemetry] retry labels, row off-by-one, 4xx-permanent annotation. [Valentin Lab]
Three user-reported bugs from the post-Step-12 review.
1. **Picker T button triggers wrong row when a primary cached
entry shares the same (provider, model) with a deferred
candidate.**
The button click handler searched ``local_candidates`` for the
first row matching ``(provider, model, streaming)``. Because
the primary entry (the cached pre-selected row) is inserted
first and shares those fields with a deferred candidate for
the same model, ``Vec::position`` returned the primary row's
index — the spinner appeared on the row ABOVE the one the
user clicked.
Fix: include the ``is_primary`` flag in the button's captured
identity tuple and match on all four fields in the click
handler.
2. **Retry counter ambiguity (``retry N/4`` vs ``N/5``,
unlabeled which phase).**
``TranscriptionEvent::RetryScheduled`` now carries a
``kind: RetryKind`` field (``Connection`` or ``Data``). The
picker renders ``connect retry N/M…`` for connection-phase
retries and ``server retry N/M…`` for data-phase retries.
3. **HTTP 4xx errors show ``(after 1/5 attempts)`` with no
indication that retrying was deliberately skipped.**
``PipelineFailureKind::HttpStatus`` rendering now annotates
4xx as ``— 4xx permanent, no retry`` and 5xx with exhausted
budget as ``— server-retry budget exhausted``, so users
understand the counter is correct rather than premature.
The known-flaky SIGUSR1 lib test
``cancel_remote_via_sigusr1_triggers_owner_token`` is marked
``#[ignore]`` to match its integration-test cousin (SIGUSR1 is
process-wide; the test passes in isolation but races against
other registered jobs in the parallel test runner).
Tests: 440 lib green (1 ignored), 9 transport_integration green
(1 ignored), 11 other integration suites green. Clippy clean.
Release build green.
* [error] dedup reqwest URL restatement from rendered failure chain. [Valentin Lab]
* [transcription] guard validate-cache writes with POSIX flock. [Valentin Lab]
The validate-cache write path was a read-modify-write race: a
process recorded its first entry, serialised only its own
in-process map, and atomically renamed onto the shared file —
clobbering any sibling-process entries it had never read.
Observed in production as the real cache being repeatedly
truncated to a single entry whenever a fresh ``talk-rs`` process
(picker spawn, ``dictate`` daemon, ``record --ui``) recorded a
validation.
``persist_to_disk()`` now:
1. Acquires an exclusive ``fs2::FileExt::lock_exclusive`` on a
sibling ``validate-cache.yaml.lock`` file.
2. Re-reads the on-disk YAML unconditionally and merges entries
into the in-process map (newer ``validated_at`` wins).
3. Writes the merged set via the existing tempfile + atomic rename.
4. Releases the lock via an ``RAII`` ``LockGuard``.
Readers stay lock-free; the atomic rename gives them either the
previous-complete or new-complete file.
Also adds the ``TALK_RS_VALIDATE_CACHE_PATH`` env override so
tests can redirect the cache at a tempfile, plus a process-wide
``__TEST_LOCK`` and ``__test_reset()`` exposed to sibling test
modules so cache-touching tests in ``http::tests`` don't race
with the ones here on the shared ``OnceLock`` statics.
Coverage:
- ``record_merges_with_sibling_disk_entries`` reproduces the
production failure (sibling entry on disk, fresh process records
a different key, both must survive).
- The four ``validate_model_*`` tests in ``http::tests`` now hold
a ``CacheTestGuard`` so they cannot pollute the dev's real cache.
Refs: production cache truncation observed after recent picker /
``dictate`` parallelisation work.
* [audio, dictate] restore BT headset profile immediately on stop. [Valentin Lab]
Previously the ``HeadsetGuard`` was dropped at the end of ``dictate()``,
which meant the Bluetooth headset stayed in HFP (lower-quality voice
profile) for the entire transcription + paste pipeline — sometimes
tens of seconds — before flipping back to A2DP. The user-visible
effect was that music / system audio came back in degraded HFP quality
for a noticeable window after they pressed the toggle to stop
recording.
The fix: drop the guard the moment the microphone capture stops, in
parallel with the rest of the dictation pipeline.
A new ``HeadsetGuard::restore_now_async()`` method takes the saved
profile out of the guard and dispatches the restore on
``tokio::task::spawn_blocking``. This off-loads the libpulse mainloop
driver (which can block on BlueZ profile renegotiation for ~1 s) from
the async runtime, so transcription and paste continue in parallel
without waiting.
The guard is now moved by value into ``dictate_streaming`` and
``dictate_realtime`` and ``restore_now_async()`` is called the moment
``capture.stop()`` returns — right next to the stop sound and the
"Transcribing" overlay swap. After the restore is dispatched the
guard is empty so its eventual ``Drop`` at function exit is a no-op.
The guard's RAII ``Drop`` remains the safety net for code paths
between activation and dispatch (panic, early ``?``-return, etc.):
those will still trigger a synchronous restore.
The standalone ``record`` command was already correct — its
``capture.stop()`` is on the last line of ``record()`` so the guard
drop already happened immediately. Untouched here.
Two new unit tests in ``bt_profile.rs`` cover the new method:
- ``restore_now_async`` on a guard with no saved profile is a cheap
no-op (does not spawn any blocking task).
- ``restore_now_async`` takes the saved profile out of the guard so
the eventual ``Drop`` does not double-restore.
Verified end-to-end:
- ``cargo fmt``, ``cargo clippy --all-targets``, ``cargo test``
(370 tests pass), ``cargo build``, ``cargo build --release`` clean.
- Live ``cargo test --test bt_profile_smoke`` against the running
PulseAudio / ``pipewire-pulse`` server still detects the headset
correctly.
* [overlay, dictate] pause recording pipeline on dead signal and skip transcription. [Valentin Lab]
Dead-signal detection (``no_sound_active``) explicitly excluded
auto-pause, so a dead microphone kept recording useless constant-value
frames into the OGG cache. After the user stopped, those silent
seconds were sent to the transcription API — wasting time, tokens, and
showing a misleading "transcribing" badge.
Now the overlay sets ``pause_flag = true`` when dead signal triggers,
stopping the audio tee from forwarding frames to the encoder. A new
``had_live_audio`` flag on ``OverlayHandle`` tracks whether any frame
with real variance was ever seen during the recording session. After
recording stops, ``dictate_streaming`` checks both ``buffer.is_empty()``
and ``had_live_audio()`` — if no usable audio exists, the transcription
pipeline is aborted immediately with an empty result.
* [overlay] enlarge red dot and fix its centering. [Valentin Lab]
Scale ``DOT_RADIUS_MAX`` from 10 to 21 and ``DOT_RADIUS_MIN`` from 3
to 6 so the volume indicator fills the spec area and is readable at a
glance. Shift ``DOT_CX`` from 20 to 26 so the larger dot plus its
gap clears the rounded corners.
Side-effect: the pause icon (derived from ``DOT_RADIUS_MAX``) also
grows, and its vertical centering improves because ``bar_h`` is now
odd (29 px), eliminating the 0.5 px asymmetry of the old even value.
* [transcription] add per-request wall-clock timeout to HTTP calls. [Valentin Lab]
Mistral transcription was observed hanging for 168s when the server
accepted the TCP connection but stalled at the application layer.
The existing ``build_client()`` defences (``tcp_user_timeout``,
``tcp_keepalive``, ``connect_timeout``) only cover TCP-level failures
and cannot detect a slow-but-alive server.
Add ``proportional_timeout(audio_bytes)`` in ``http.rs``:
``max(3s, kb / 10)`` — scales with payload size so large files are
not killed prematurely.
Apply ``.timeout()`` at each call site:
- ``mistral::transcribe_file`` — uses file metadata length
- ``mistral::transcribe_stream`` — uses drained audio buffer length
- ``openai::transcribe_file`` — adds a metadata read for file length
``openai::transcribe_stream`` is deliberately left unchanged: it uses
``Body::wrap_stream`` where total size is unknown at request-build
time; a wall-clock cap would kill long recordings.
``CONNECT_TIMEOUT`` (2s) is unchanged — log analysis across 81
sessions showed 100% cumulative success within the existing 5 retries.
* [record] copy full transcript to clipboard instead of truncated preview. [Valentin Lab]
The copy-to-clipboard button in the recordings browser was cloning the
``transcript_preview`` field (200 chars with trailing ellipsis) instead of
the full transcript. This caused users to lose data when copying recordings
to paste elsewhere.
The fix adds a ``transcript_full`` field to ``RecordingEntry`` holding the
complete single-line transcript (newlines collapsed to spaces, never
truncated). The copy button now clones ``transcript_full`` while the display
label continues using ``transcript_preview`` for brevity.
Also introduces a ``TRANSCRIPT_PREVIEW_CHARS`` constant and a
``transcript_variants()`` helper to extract both values from a single YAML
read, eliminating duplication between ``list_ogg_recordings`` and
``list_cache_recordings``.
Includes 8 unit tests covering edge cases: empty input, short text, newline
collapse, exact-200-char boundary, 201-char truncation, very long text,
multibyte CJK characters, and long text with embedded newlines.
* [record] remove waterfall spectrogram cache when deleting recordings. [Valentin Lab]
The ``delete_recording()`` function now also removes the ``.wf`` waterfall
spectrogram cache file when deleting a recording. Previously only ``.yml``
companion files were cleaned up, leaving orphaned ``.wf`` files behind.
* [record] make file monitor events incremental instead of rebuilding. [Valentin Lab]
Previously, any audio file event (Created/Deleted/ChangesDoneHint) triggered
a full ``populate_section()`` rebuild, which cleared all rows, re-read all
entries from disk, and always selected row 0. This caused scroll position
reset and visible flashing.
Now each event type is handled incrementally:
- Deleted: finds matching row by widget_name, selects adjacent row, removes
just that row, updates expander count
- Created/ChangesDoneHint: builds one new row or refreshes existing one
in-place at correct sorted position
- YAML events: extracted into shared ``update_row_for_yml`` helper
- Unhandled events: logged at debug level for future debugging
- ``populate_section()`` no longer called from inotify path (only at init)
Fixes scroll reset and flashing on file deletion.
* [transcription] buffer audio before upload to provide explicit Content-Length. [Valentin Lab]
Mistral's API rejects chunked Transfer-Encoding with 411 Length Required.
Collect all audio bytes from the mpsc channel into a buffer first, then
send with known length via ``Part::stream_with_length()``. This eliminates
the 411 failure + retry overhead, reducing transcription time from ~4.4s
to ~2.8s (-36%).
Also removed now-unused imports: ``futures::StreamExt`` and
``tokio_stream::wrappers::ReceiverStream``.
* [overlay] freeze visualization data during auto-pause and no-sound. [Valentin Lab]
Prevent silence gaps from appearing in the waterfall spectrogram when
auto-pause is active or no sound is detected. Wrap all per-viz-mode data
updates (``spectrogram_history``, ``amp_history``, ``spectrum_peak``) in
a condition that skips accumulation when ``!auto_paused && !no_sound_active``
is false.
* [record] extend ``FileMonitor`` lifetime past ``main_loop.run()`` [Valentin Lab]
Move ``_keep_monitors`` binding out of the inner block so the
inotify file watches remain alive for the entire duration of the
GTK main loop. Previously they were dropped before
``main_loop.run()`` was called, silently losing notifications.
* [gtk-theme] recompute edge in click handler to fix intermittent resize. [Valentin Lab]
The ``install_edge_resize`` click controller was reading a cached edge
from the motion controller's shared ``RefCell`` rather than computing
it from its own ``(x, y)`` coordinates. Hand jitter between the last
motion event and the press event caused the cached edge to be stale,
producing two failure modes:
- Window moves instead of resizing (cached edge was ``None``,
gesture denied, ``WindowHandle`` drag took over).
- Nothing happens (cached edge was set, gesture claimed, but
``begin_resize`` silently failed).
Extract ``detect_edge()`` and ``edge_cursor()`` helpers, call
``detect_edge`` from the pressed handler with a slightly larger
threshold (10 px vs 6 px cursor zone) to forgive press-time jitter.
* [transcription] provide explicit ``Content-Length`` in Mistral batch uploads. [Valentin Lab]
The Mistral API now rejects multipart requests without an explicit Content-Length header (HTTP 411). ``transcribe_file()`` was using ``Part::stream()`` which sends chunked TE without Content-Length. Switched to ``Part::stream_with_length()`` with the file size from metadata.
* [picker] defer ``WavPlayer`` initialization to avoid blocking UI on cpal device probing. [Valentin Lab]
Move ``WavPlayer::new()`` from synchronous construction (which
blocks ~1-2 s on PipeWire while probing audio devices) to a
``glib::idle_add_local_once`` callback fired after the window is
presented. The picker and recordings windows now appear instantly;
the play button is enabled once the player is ready.
* [picker] cap message loop and reuse labels to prevent cursor stutter. [Valentin Lab]
The transcription poll timer drained all queued ``StreamUpdate``
messages in an unbounded ``loop``, rebuilding the label widget
(destroy + create) on every streaming delta. When a burst of
results arrived the GTK main loop was blocked, starving the 16 ms
cursor timer and causing visible jumps.
Two fixes:
- Limit the loop to ``MAX_MSGS_PER_TICK`` (5) iterations per 50 ms
tick so other main-loop sources stay responsive.
- Reuse the existing ``gtk4::Label`` via ``set_text()`` instead of
tearing down and recreating the widget on every delta. The full
widget swap now only happens once (first update replacing the
spinner).
## 0.5.0 (2026-03-10)
### New
* [overlay] auto-pause recording during silence to trim dead air. [Valentin Lab]
When the user stops speaking, recording pauses after 0.5s of silence
(RMS below 0.003). Silent segments are not sent to the transcription
API, reducing noise and improving accuracy.
Recording resumes instantly when speech is detected. A 500ms lookback
buffer in the audio tee preserves the speech onset so the beginning
of words is not clipped.
Visual indicator: yellow pause bars replace the red pulsing dot, with
"LISTENING" text in the viz area.
Dead-signal detection (NO SOUND) takes priority over auto-pause. The
dead-signal trigger is increased from 1 frame to 0.5s to avoid false
flashes from momentary PipeWire glitches.
The pause mechanism uses a shared ``Arc<AtomicBool>`` between the
overlay thread and the audio tee task.
* [overlay] detect dead audio device and display NO SOUND warning. [Valentin Lab]
Unified audio capture: the overlay now reads from the same PipeWire
stream as the recording pipeline via a new audio tee task
(``src/audio/tee.rs``), replacing the independent CPAL capture that
could pick a different device or miss the dead-device condition.
Dead-signal detection uses sample variance rather than an RMS
threshold. A dead or missing PipeWire device sends constant
``i16::MIN`` (-32768) samples, which yields RMS ~1.0 (indistinguishable
from quiet speech) but variance = 0. A real microphone in a silent
room still produces random quantisation noise with variance > 0, so
the check reliably separates "no device" from "quiet room".
Visual warning: when a dead signal is detected the recording badge
switches from the red dot to a prohibit icon (circle + diagonal bar)
with "NO SOUND" text, and the text panel shows a descriptive message
explaining the likely cause.
Font and glyph rendering utilities (``draw_text``, ``measure_text``,
glyph helpers) moved from ``src/x11/visualizer.rs`` to the shared
``src/x11/render_util.rs`` so both the visualizer and the overlay can
use them without duplication.
* [x11] add ``--bw`` monochrome mode for audio visualizers with theme detection. [Valentin Lab]
Add a ``--bw`` flag that switches the amplitude and spectrum visualizer
panels to monochrome rendering. The foreground colour is chosen
automatically based on the desktop theme (``dark-light`` crate via
freedesktop D-Bus portal, with ``GTK_THEME`` env-var fallback):
white-on-black for dark themes, black-on-white for light themes.
Also fixes visualizer panel positioning to use the exported
``BADGE_W`` constant instead of a hardcoded 182, which caused overlap
after the badge was widened to 273 px.
* [transcription] add 2-second TCP connect timeout to HTTP clients. [Valentin Lab]
Add a CONNECT_TIMEOUT of 2 seconds to all reqwest::Client builders in Mistral, OpenAI, and model-suggestion modules. This fails fast when the server is unreachable without shortening the overall transcription timeout (reverted to 5 seconds). Constructors for MistralBatchTranscriber and OpenAIBatchTranscriber now return Result to propagate client-build errors.
* [x11] replace static recording badge with live spectrogram waterfall. [Valentin Lab]
Overlay now renders a real-time FFT spectrogram with a pulsing red
``rec`` dot, dynamic frequency scaling, and all-time volume peak
tracking. Extract shared ``PixelBuffer``, ``RingBuffer``, FFT, and
shape helpers into ``render_util`` module from ``visualizer``.
### Changes
* [overlay] improve FFT resolution and fix visual gaps in visualizers. [Valentin Lab]
Double ``FFT_SIZE`` from 1024 to 2048, raising frequency resolution
from 46.9 Hz/bin to 23.4 Hz/bin.
Rewrite ``map_spectrum_to_column()`` to use midpoint-boundary bin
ranges with peak (max) aggregation instead of truncating to a single
bin index with 3-neighbor averaging. This eliminates visual holes in
the waterfall caused by integer truncation skipping bins.
In ``render_spectrum_badge()``, remove the hard-coded 1 px gap between
bars and distribute bins proportionally across the full badge width so
no frequency bins are dropped.
* Rename ``--bw`` flag to ``--mono`` [Valentin Lab]
* Consolidate audio visualizers into recording badge with ``--viz`` flag. [Valentin Lab]
Replace ``--amplitude`` and ``--spectrum`` side-panel flags with a single
``--viz <waterfall|amplitude|spectrum>`` option that renders the selected
visualization inside the recording badge. No visualizer by default.
- Add ``VizMode`` enum to ``config.rs`` with config file + env var support
- Overlay badge switches renderer per viz mode (waterfall/amplitude/spectrum)
- Strip dead audio panel code from ``visualizer.rs`` (text-only now, -756 lines)
- Audio capture in overlay only when viz mode is active
* [x11] add transparent gap around pulsating red dot on recording badge. [Valentin Lab]
Clear a circle slightly larger than the red dot to transparent
before drawing the dot itself, creating a 2-pixel gap that visually
separates the dot from the spectrogram waterfall underneath. The gap
scales with the dot radius (which varies from 3 to 10 px based on
volume). Edge pixels are anti-aliased by blending towards transparent.
* [x11] widen recording badge and make spectrogram fill full width. [Valentin Lab]
Badge width increased from 182 to 273 pixels (~50% wider). Spectrogram waterfall now spans the full badge width (margin 4px each side). Red dot moved left (DOT_CX 28→20) and draws on top of the spectrogram instead of beside it.
* [x11] use 32-bit ARGB visual for compositor alpha transparency on recording badge. [Valentin Lab]
Replace opaque background + Shape extension approach with true
ARGB transparency. Add ``find_argb_visual()``,
``create_argb_overlay_window()``, and ``draw_rounded_border()``
with SDF-based premultiplied alpha rendering. The badge now shows
a visible rounded border and is see-through to the desktop.
Falls back to the old opaque+shape approach when no 32-bit visual
is available.
* [dictate] reduce transcription timeouts from 5s to 2s. [Valentin Lab]
Lower ``TRANSCRIPTION_TIMEOUT`` and ``RETRY_TIMEOUT`` from 5 seconds
to 2 seconds to shorten the wait during dictation.
* [dictate, x11] decouple text panel lifecycle from audio visualizer panels. [Valentin Lab]
The text panel is now created once at visualizer thread start and
persists for the entire thread lifetime, independently of the audio
panels (amplitude/spectrum). ``Hide`` now only hides audio panels;
the text panel self-manages its visibility based on pending messages.
This removes the need for callers to explicitly ``hide()`` the
visualizer after showing error/status messages — the text panel
stays visible until its TTL messages expire, then unmaps itself.
- Remove ``HideAudio`` command; ``Hide`` takes its semantics
- Remove ``destroy_windows()``; only ``destroy_audio_windows()`` needed
- Remove ``enable_text`` parameter from ``new()`` / ``visualizer_thread()``
- Remove explicit ``viz.hide()`` calls from error paths in ``dictate``
- Event loop now has three states: audio-active, text-only, and idle
* [dictate, x11] replace ``set_text`` with TTL-based ``push_message`` for status messages. [Valentin Lab]
Error and retry messages now stack in the visualizer text panel and
fade out after a configurable TTL instead of overwriting the live
transcription line. ``hide_audio`` dismisses only the amplitude and
spectrum panels while keeping status messages visible.
* [paste] reduce inter-paste sleep delays. [Valentin Lab]
The native ``x11rb``/``XTest`` paste path is synchronous and no
longer needs the generous delays that accommodated ``xdotool``
process-spawn overhead.
- Clipboard-set → paste delay: 50 ms → 5 ms
- Inter-paste settle: 100 ms → 15 ms
- Final settle before clipboard restore: 100 ms → 50 ms
* [dictate, x11] always create visualizer for error/retry feedback. [Valentin Lab]
The visualizer is now initialised even when ``--amplitude`` and
``--spectrum`` are not passed. When neither audio panel is enabled
the visualizer thread skips CPAL audio-device initialisation and only
manages the text window — very lightweight.
The text panel is created unconditionally but mapped on demand: it
appears when a message is set and hides when cleared. This lets the
dictation pipeline display transcription errors, reconnection
attempts, and retry status directly below the overlay badge so the
user is never left wondering why nothing is happening.
Realtime mode shows: transcription errors, "reconnecting", and
"reconnect failed". Batch (streaming) mode shows: retry count,
pipeline failure reason, and "will retry after recording".
* [dictate, audio] decouple WAV recording from transcription pipeline. [Valentin Lab]
WAV recording now writes to a shared ``AudioBuffer`` that is
completely independent of the transcription pipeline. If
transcription fails mid-recording, the audio is never lost — the
``buffer_feeder`` replays all chunks from cursor 0 into a fresh
transcriber.
Key changes:
- Replace ``audio_tee_to_wav`` with ``wav_recording_task`` +
``buffer_feeder``: the WAV task writes every chunk to disk and to
the shared buffer unconditionally; feeder tasks read from the
buffer and can be killed/restarted without affecting the recording.
- Realtime mode: on ``TranscriptionEvent::Error`` or unexpected
channel closure, abort the feeder, create a new transcriber, and
replay all buffered audio from the beginning.
- Batch (streaming) mode: same decoupling; retry the encode +
transcribe pipeline up to 3 times during live recording. If all
retries fail, recording continues and the WAV is intact for
post-recording retry.
- PipeWire capture: flush residual samples after the mainloop exits
so the last partial chunk is not silently dropped.
- Play the start-sound before starting capture so the tone is never
recorded in the audio.
### Fix
* [overlay] implement color rendering for visualizers. [Valentin Lab]
Add ``heat_map_color`` (blue→cyan→green→yellow→red) for waterfall and
``level_color`` (green→yellow→red) for bar/wave. All three visualizers
now render in color by default; ``--bw`` correctly switches to
monochrome instead of being a no-op.
## 0.4.0 (2026-02-26)
### New
* [picker] add play button to preview recorded audio. [Valentin Lab]
Add a ▶/■ toggle button at the top of the ``--pick`` window so users
can listen to their recorded audio before selecting a transcription.
Reuses the existing ``WavPlayer`` from the recordings browser by
widening its visibility to ``pub(crate)``. Playback stops automatically
on window close or escape.
### Fix
* [dictate, transcription] stop ``toggle_validate`` from killing daemon on network errors. [Valentin Lab]
Proactive ``toggle_validate`` was sending ``SIGINT`` to the daemon on
unrelated network errors, causing unexpected recording stops. Remove
proactive validation entirely from the toggle/recording path. Model
suggestions are now lazy: only fetched (with cache + retries) when
transcription fails with a ``model-not-found`` error.
Move provider-specific intelligence (error detection patterns, API base
URLs, model filters) from ``model_suggestions.rs`` into ``mistral.rs``
and ``openai.rs``. ``model_suggestions.rs`` is now a pure cache utility
with no provider knowledge. Factory dispatchers in ``transcription/mod.rs``
route calls to the correct provider module.
## 0.3.0 (2026-02-24)
### Changes
* [clipboard, x11] replace ``xclip`` with native ``x11rb`` clipboard. [Valentin Lab]
Use ``x11rb`` directly for CLIPBOARD selection get/set instead of
shelling out to ``xclip``. ``set_text`` spawns a background thread
that serves ``SelectionRequest`` events until the next write or drop.
No external runtime tools are required for clipboard operations.
* [paste, x11] replace ``xdotool`` with native ``x11rb``/``XTest`` calls. [Valentin Lab]
Remove the ``xdotool`` runtime dependency by reimplementing all
window-management and key-simulation helpers with ``x11rb`` and the
XTest extension:
- ``get_active_window`` and ``focus_window`` now use ``x11rb`` directly
- ``simulate_paste`` and ``simulate_backspace`` use XTest key events
- ``xdotool`` removed from README prerequisites
### Fix
* [record] save recordings in configured ``output_dir`` by default. [Valentin Lab]
Previously ``talk-rs record`` (with no explicit path) wrote
``memo-<timestamp>.ogg`` into the current working directory.
Now it loads the user config and uses ``output_dir`` as the parent
directory for default recordings, matching the behavior of ``dictate``.
Extracts a testable ``resolve_output_path()`` helper and updates CLI
help text and config docs to reflect the new default.
## 0.2.0 (2026-02-23)
### New
* [dictate] add ``--no-chunk-paste`` flag and ``paste.chunk_chars`` config. [Valentin Lab]
Allow users to control the 150-character paste chunking behaviour.
``--no-chunk-paste`` disables chunking entirely (pastes in one shot).
``paste.chunk_chars`` in the config file sets a custom chunk size;
``0`` also disables chunking. The CLI flag overrides the config value.
* [dictate] add ``--no-boop`` flag and honor ``boop_interval_ms`` config. [Valentin Lab]
Adds a ``--no-boop`` CLI flag that disables only the periodic boop
heartbeat during recording while keeping start/stop sounds. Also
wires up the existing ``indicators.boop_interval_ms`` config field
which was defined but never read — the boop interval was hardcoded
to 5 seconds. Setting ``boop_interval_ms`` to ``0`` in config now
disables boops permanently.
* [dictate] add 5 s transcription timeout with 5-attempt retry. [Valentin Lab]
``dictate_streaming`` now aborts the transcription task if it does not
complete within 5 seconds — this prevents zombie daemon processes
caused by the API hanging indefinitely.
On failure (timeout, API error, network issue through VPN), the batch
branch retries up to 5 times using ``transcribe_file`` against the
saved WAV. Each retry creates a fresh ``BatchTranscriber`` and applies
the same 5 s timeout. Progress is shown in the visualizer overlay:
"Transcription failed: {reason}. Retrying ({N}/5)..."
If all retries are exhausted the error is displayed for 3 seconds, then
the daemon exits cleanly — YAML metadata is skipped, paste is skipped,
and the WAV is preserved for the picker.
* [record-ui] add GTK4 recordings browser with ``--ui`` flag. [Valentin Lab]
``talk-rs record --ui`` opens a two-section window:
- **Dictation cache** — WAV files from ``~/.cache/talk-rs/recordings/``
- **Recordings** — OGG files from ``config.output_dir``
Each row shows date, duration, size, and transcript preview with
play (native ``cpal``), open-in-file-manager (``FileLauncher``), and
delete buttons. Sections auto-refresh via ``gio::FileMonitor``
(inotify) when files are added or removed externally, and the
expander counter updates on deletion.
Enables ``gtk4`` feature ``v4_10`` for ``FileLauncher`` support.
* [transcription] add ``--diarize`` speaker diarization. [Valentin Lab]
Add modular speaker diarization support. Each provider can optionally
populate a ``diarization`` field on ``TranscriptionResult`` with
per-segment speaker labels (``DiarizationSegment`` struct).
Mistral V2 batch mode: sends ``diarize=true`` and
``timestamp_granularities=segment`` form fields, parses
``segments[].speaker_id`` from the API response. OpenAI returns
``None`` (interface ready for future implementation).
CLI: ``--diarize`` flag on ``transcribe`` and ``dictate`` commands.
Client-side error when combined with ``--realtime`` (Mistral WebSocket
has no diarize parameter). Output formatting merges adjacent
same-speaker segments into ``[SPEAKER_XX]``-tagged lines via
``format_transcription_output()``.
Also fixes stale docs: default model name in ``README.org`` and
``config.example.yaml``, ``--pick`` help text (rofi to GTK), and
adds missing OpenAI provider section to config example.
### Changes
* [dictate] move stop-sound and overlay feedback into ``dictate_streaming`` [Valentin Lab]
Immediate audible + visual feedback (stop sound, "Transcribing" badge)
now fires inside ``dictate_streaming`` right after ``capture.stop()``,
so the user gets feedback the instant they toggle off — before the
API call finishes. Batch-mode overlay is kept visible until paste
or empty-transcription exit. Also applies ``cargo fmt``.
* [dictate] register early ``SIGINT`` handler and add debug tracing. [Valentin Lab]
Move the ``CancellationToken`` + ``ctrl_c`` handler above the capture
and sound-indicator setup. Without this there is a ~1 s race window
where SIGINT has no handler and the daemon becomes an unkillable
orphan. Also append ``[DBG]`` traces to ``daemon.log`` in
``toggle_dispatch`` and ``toggle_stop`` for signal-flow observability.
## 0.1.0 (2026-02-21)
### New
* [audio] add ``--monitor`` flag for mic+system audio mixing. [Valentin Lab]
Add ``MonitorCapture`` that combines microphone input with system
audio (PipeWire monitor source) into a single stream. The
``--monitor`` flag is available on both ``record`` and ``dictate``
commands, and is forwarded through toggle-mode daemon spawning.
* [monitor] add GDK4-based monitor geometry module. [Valentin Lab]
Query the GDK display for the largest monitor (by physical pixel
area) and return ``(x, y, width, height)`` scaled by the monitor's
``scale_factor``. GDK4 dropped the "primary" flag, so largest-area
is used as heuristic.
This replaces the ``xrandr`` subprocess calls previously embedded
in ``overlay.rs`` and ``visualizer.rs`` with a shared, type-safe
module that returns physical pixel coordinates suitable for direct
X11 window placement.
* [dictate] integrate realtime transcription into picker. [Valentin Lab]
The ``--pick`` flag now accepts realtime (WebSocket) transcribers
alongside batch ones. Each realtime candidate streams incremental
text updates into the GTK picker via ``PickerMessage::StreamUpdate``
and displays a ⚡ indicator in the provider column.
Key additions:
- ``OPENAI_REALTIME_MODELS`` / ``MISTRAL_REALTIME_MODELS`` constants
and ``add_known_realtime_models`` to populate retry candidates.
- ``read_wav_pcm_samples`` helper to feed recorded WAV data into
realtime transcribers as chunked PCM.
- ``split_into_word_chunks`` splits paste text into word-bounded
chunks so large transcriptions are pasted incrementally, avoiding
overwhelming the target application.
- Retry button correctly re-spawns either batch or realtime
transcription depending on the candidate type.
- Removes the ``--pick is currently supported only in batch mode``
guard.
* [daemon] add ownership-safe ``signal_daemon`` and ``stop_if_owner`` helpers. [Valentin Lab]
``signal_daemon`` sends SIGINT to a daemon process group and removes
the PID file immediately so the next toggle-on sees ``NotRunning``.
The exiting daemon uses ``remove_pid_file_if_owner`` to avoid
clobbering a PID file written by a newly spawned replacement.
``stop_if_owner`` re-acquires the lock and only performs a full
graceful stop when the PID file still belongs to the expected process.
* [audio] add native PipeWire capture and ``rubato`` resampler. [Valentin Lab]
Add two new audio modules:
- ``pipewire_capture``: captures audio directly from PipeWire using
the Rust ``pipewire`` bindings, matching ``pw-cat --record`` routing
(including Bluetooth devices). Runs on a dedicated thread with
fixed-size chunk emission.
- ``resample``: sinc-interpolated downsampler (48 kHz → 16 kHz) using
``rubato`` with Blackman2 windowing and 128-tap anti-aliasing
filter. Includes ``spawn_resample_task()`` that bridges two
``mpsc`` channels, passing through unchanged when rates match.
Also adds ``preferred_capture_rate()`` to ``CpalCapture`` for
querying the default input device's best mono rate (capped at 48 kHz).
Dependencies: ``pipewire``, ``rubato``, ``audioadapter-buffers``,
``gdk4-x11``.
* [dictate] add transcription metadata capture and GTK multi-provider picker. [Valentin Lab]
Enrich transcription responses with structured ``TranscriptionMetadata``
capturing latency, token usage, detected language, and provider-specific
diagnostics (rate-limit headers, realtime session/event counters).
The ``BatchTranscriber`` trait now returns ``TranscriptionResult`` (text +
metadata) instead of a bare ``String``, and the recording cache YAML
includes the full metadata payload for post-hoc analysis.
Add a GTK4 picker window (``--pick``) that fires parallel transcription
requests across all known provider/model combinations and displays
candidates progressively as they complete. Results are cached per audio
file in ``~/.cache/talk-rs/picker-results/`` so reopening the picker
skips API calls entirely.
Supporting changes:
- ``--retry-last``: reuse the last cached recording as input audio
- ``--replace-last-paste``: delete previously pasted text before
inserting the new selection (tracked via ``last_paste.yml``)
- ``last_recording.wav`` / ``last_metadata.yml`` symlink pointers in
the recording cache for quick access to the most recent entry
- ``simulate_backspace`` helper and ``paste_text_to_target`` refactoring
to share paste logic between batch mode and picker
- ``x11_centre_and_raise`` for single-instance picker detection and
monitor-aware centring via RandR
- New ``TranscriptionEvent`` variants (``SessionInfo``,
``RateLimitsUpdated``, ``TransportMetadata``) for realtime metadata
collection
- Mistral realtime validation now skips the REST ``/v1/models`` check
(realtime-only models are not listed there)
* [dictate] add ``--output-yaml`` flag for metadata export. [Valentin Lab]
Copy the recording cache metadata YAML to a user-specified path
after transcription completes. Reuses the existing cache metadata
file written by ``write_metadata`` rather than generating a
separate one.
* [dictate] add ``--input-audio-file`` flag for file-based transcription. [Valentin Lab]
Feed a pre-recorded WAV file through the transcription pipeline
instead of live microphone capture. Works with both batch and
realtime modes, enabling reproducible benchmarks across providers
and models.
The ``WavFileSource`` implements ``AudioCapture`` and validates the
file is 16 kHz / mono / 16-bit PCM (with ``ffmpeg`` conversion
hints on mismatch). Batch mode races Ctrl+C against natural file
completion via a oneshot signal from the encode task. Realtime
mode handles file exhaustion through the existing channel cascade.
Also removes the now-redundant ``save_file`` parameter from
``dictate_streaming()`` since ``--save`` copies from the recording
cache.
* [dictate] add ``--save`` flag and recording cache. [Valentin Lab]
Replace the positional ``FILE`` argument with ``--save <PATH>`` on the
``dictate`` command. The last 10 recordings are now always cached in
``~/.cache/talk-rs/recordings/`` as timestamped WAV files with companion
YAML metadata (provider, model, realtime flag, transcript).
Both batch and realtime paths tee raw PCM to the cache. When ``--save``
is specified, the cache WAV is copied to the user path after recording.
Oldest entries beyond 10 are automatically rotated out.
* [transcription] add multi-provider support with OpenAI backend. [Valentin Lab]
Add ``OpenAI`` as a second transcription provider alongside ``Mistral``,
supporting both batch (REST) and realtime (WebSocket) modes.
Infrastructure:
- ``Provider`` enum (``Mistral`` | ``OpenAI``) with ``Deserialize``/``FromStr``
- ``BatchTranscriber`` and ``RealtimeTranscriber`` traits with factory
functions and lazy API-key validation
- ``--provider`` and ``--model`` CLI flags on ``dictate`` and ``transcribe``
- ``TranscriptionConfig`` with ``default_provider`` in config file
- ``OpenAIConfig`` with ``model`` (batch) and ``realtime_model`` fields
- Environment variable overrides for all provider settings
OpenAI batch (``openai.rs``):
- ``OpenAIBatchTranscriber`` with streaming upload via ``reqwest::Body::wrap_stream``
- Wiremock-based unit tests for success, error, streaming, and edge cases
OpenAI realtime (``openai_realtime.rs``):
- ``OpenAIRealtimeTranscriber`` using WebSocket with ``session.type: "transcription"``
- 16 kHz → 24 kHz PCM resampling (linear interpolation, 3:2 ratio)
- ``server_vad`` turn detection, ``input_audio_buffer.append`` streaming
- Post-commit timeout for completion detection (no explicit "done" event)
- Default model: ``gpt-4o-realtime-preview``
Preflight validation:
- ``validate()`` method on both traits, called before audio capture
- REST ``GET /v1/models`` check: verifies API key and model existence,
lists available transcription models on bad model name
- WebSocket session check (realtime only): catches "model not supported
in realtime mode" before daemon spawn in ``--toggle`` mode
- Validation in ``toggle_start()``, ``dictate()``, and ``transcribe()``
* [visualizer] add live transcription text overlay. [Valentin Lab]
Render live transcription text below the recording badge using
``fontdue`` for glyph rasterisation. Features include:
- Async font loading from system paths (``/usr/share/fonts``) to
avoid blocking the render loop at startup
- CJK-aware font fallback (Noto Sans CJK)
- Centered, single-line text with automatic left-clip on overflow
- Pulsing "..." dots while waiting for speech
- Rounded-corner text background via ``XShape``
- ``dictate_realtime()`` pushes ``TextDelta``/``SegmentDelta`` text
to the overlay in real time
* [cli] add ``--amplitude`` and ``--spectrum`` flags to ``dictate`` [Valentin Lab]
Integrate the visualizer module into the dictate command lifecycle:
init on start, show alongside the recording badge, hide on stop.
Both flags are forwarded through ``--toggle`` daemon mode.
* [visualizer] add real-time audio visualizer module. [Valentin Lab]
Amplitude history (RMS-based bar chart) and ``FFT`` spectrum panels
rendered via X11 ``put_image`` at 60 fps. Each panel is independently
toggleable, positioned on either side of the recording badge, and runs
its own ``CPAL`` capture stream decoupled from the recording pipeline.
Includes ring buffer, radix-2 Cooley-Tukey ``FFT``, pixel buffer
helpers, multi-monitor geometry detection, and comprehensive unit tests.
* [cli] add debug WAV capture for realtime transcription. [Valentin Lab]
Every realtime ``dictate`` session now saves a copy of the raw PCM audio
to ``$XDG_CACHE_HOME/talk-rs/debug-capture.wav`` (or to the user-specified
file path). This tees the audio stream so exactly what is sent to Voxtral
is also written to disk, allowing the user to verify that:
- recording starts when expected
- recording stops when expected
- the full audio content is captured correctly
The WAV header is patched with the final data size on completion. The
tee task continues writing even if the transcriber channel closes early.
* [cli] add structured logging with ``-v``/``-vv``/``-vvv`` verbosity. [Valentin Lab]
Replace unused ``tracing``/``tracing-subscriber`` with ``log``+``fern``+``colored``
following the ``fyl`` project pattern. All ``eprintln!`` diagnostic messages are
now routed through ``log::`` macros at appropriate levels:
- ``log::error!`` for failures (WebSocket errors, encode errors)
- ``log::warn!`` for degraded operation (missing overlay, stream drops)
- ``log::info!`` for user-visible events (start/stop, transcription result)
- ``log::debug!`` for pipeline steps (connect, session, clipboard, overlay)
- ``log::trace!`` for high-frequency data (audio chunks, WS frames)
The ``-v`` flag is forwarded to daemon subprocesses so ``--toggle`` mode
inherits verbosity. Refactors ``dictate()`` arguments into ``DictateOpts``
struct to satisfy ``clippy::too_many_arguments``.
* [overlay] add X11 visual overlay indicator for ``dictate`` command. [Valentin Lab]
Pure Rust X11 overlay using ``x11rb`` with Shape extension for binary
transparency (works without compositor). Displays embedded PNG badges
(recording/transcribing) centered on primary monitor.
- ``src/core/overlay.rs``: background thread with command channel,
``OverlayHandle`` for show/hide/quit, Shape mask from alpha channel,
pixel drawing grouped by color via ``poly_point``
- ``--no-overlay`` flag on ``dictate`` (passed through ``--toggle``)
- Overlay shows "Recording" badge on start, hides on stop
- PNG assets embedded via ``include_bytes!``, decoded with ``png`` crate
- Screen position from ``xrandr --query`` with multi-monitor support
- 8 unit tests covering PNG decode, geometry parsing, type invariants
* [audio] add sound indicators to ``dictate`` command. [Valentin Lab]
Synthesize short tones via ``cpal`` output to give audible feedback
during recording: ascending major-third on start, periodic soft boop
as heartbeat, descending major-third on stop. A single-channel
``SoundPlayer`` with preemption ensures sounds never overlap.
- ``indicator.rs``: tone synthesis, ``SoundPlayer``, boop loop via
``tokio`` task with ``CancellationToken``
- ``--no-sounds`` CLI flag to disable indicators
- Flag forwarded through ``--toggle`` daemon spawn
* [cli] add ``--toggle`` daemon mode to ``dictate`` command. [Valentin Lab]
First invocation spawns a background daemon (``--daemon``) that records
and transcribes via WebSocket. Second invocation sends ``SIGINT`` to
stop recording, complete transcription, and paste the result.
Uses kernel-level ``flock`` on a lock file to prevent races between
concurrent toggle calls. PID file at ``$XDG_CACHE_HOME/talk-rs/daemon.pid``
with stale-PID detection via ``kill(pid, 0)``. Graceful shutdown waits
up to 10 s before escalating to ``SIGTERM``.
Active window is captured by the toggle caller and forwarded to the
daemon via the hidden ``--target-window`` argument so paste targets the
correct window.
* [transcription] add ``realtime`` WebSocket module for Voxtral Realtime API. [Valentin Lab]
Implement ``MistralRealtimeTranscriber`` that connects to the Voxtral
Realtime API via WebSocket, streams base64-encoded PCM audio, and
receives incremental ``TranscriptionEvent`` variants (text deltas,
segment boundaries, language detection, errors).
The module includes event parsing, sender/receiver loops, and
comprehensive unit tests for all event types and PCM encoding.
Adds dependencies: ``tokio-tungstenite``, ``base64``, ``url``;
moves ``serde_json`` from dev to main dependencies.
* [config] add ``model`` and ``context_bias`` to ``MistralConfig`` [Valentin Lab]
Add configurable model name (defaulting to ``voxtral-mini-latest``)
and optional ``context_bias`` field for improved transcription accuracy
of proper nouns and technical terms.
Both fields support environment variable overrides via
``TALK_RS_PROVIDERS_MISTRAL_MODEL`` and
``TALK_RS_PROVIDERS_MISTRAL_CONTEXT_BIAS``.
The ``MistralTranscriber`` now passes these fields through to the
Mistral API in both file-based and streaming transcription requests.
* [audio] add ``AudioWriter`` trait with ``OggOpusWriter`` and ``WavWriter`` [Valentin Lab]
Replace raw ``OpusEncoder`` usage with container-aware writers that
produce valid ``OGG``/Opus (RFC 7845) and WAV output.
- ``OggOpusWriter``: encodes PCM → Opus, wraps in OGG pages with proper
``OpusHead``/``OpusTags`` headers and granule positions
- ``WavWriter``: wraps raw PCM in a 44-byte WAV header, with finalize
returning a corrected header for seekable files
- ``record`` command dispatches writer by file extension (``.wav`` vs
default ``.ogg``)
- ``dictate`` streaming and chunked modes now produce self-contained OGG
payloads per chunk (each with its own header), fixing transcription
API compatibility
- Chunked mode buffers raw PCM instead of pre-encoded Opus, enabling
per-chunk OGG encapsulation
- New dependencies: ``ogg`` 0.9, ``byteorder`` 1
* [cli] add ``--chunked`` mode to ``dictate`` command. [Valentin Lab]
Add ``--chunked`` flag and ``-n`` / ``--chunk-seconds`` option to split
recording into time-based chunks, each transcribed separately via
``MistralTranscriber``. Results are accumulated and pasted at the end.
Chunk duration is resolved from: CLI ``-n`` flag > ``dictate.chunk_seconds``
in config > error. Adds ``Clone`` to ``MistralConfig`` for per-chunk
transcriber instantiation.
* [cli] add ``dictate`` command with streaming transcription and clipboard paste. [Valentin Lab]
Record audio, stream it to Mistral API for transcription via
``transcribe_stream()``, then paste the result into the focused
application using ``xclip`` + ``xdotool``:
- Capture active window before recording (``xdotool getactivewindow``)
- Encode PCM → Opus and stream to API during recording
- On stop: refocus window, save clipboard, set text, paste via
``xdotool key ctrl+shift+v``, restore clipboard
- Optional ``--file`` arg to save audio alongside transcription
* [transcription] add ``transcribe_stream()`` for streaming audio upload. [Valentin Lab]
Extend ``Transcriber`` trait with ``transcribe_stream()`` method that accepts
a ``tokio::sync::mpsc::Receiver<Vec<u8>>`` for incremental audio upload.
``MistralTranscriber`` converts the receiver to a ``futures::Stream`` via
``tokio-stream`` and uses ``reqwest::Body::wrap_stream()`` for chunked
HTTP transfer encoding — audio is uploaded as it's recorded, so
transcription completes near-instantly after recording stops.
Adds ``tokio-stream`` dependency for ``ReceiverStream`` wrapper.
* [clipboard] add ``Clipboard`` trait with ``X11Clipboard`` and ``MockClipboard`` [Valentin Lab]
Introduce clipboard module for Phase 2 dictate functionality:
- ``Clipboard`` trait with ``get_text()`` and ``set_text()`` async methods
- ``X11Clipboard`` implementation using ``xclip`` command-line tool
(matches 0k-memo reference implementation)
- ``MockClipboard`` with ``Arc<Mutex<String>>`` for thread-safe testing
- ``TalkError::Clipboard`` error variant for clipboard operations
- Unit tests for ``MockClipboard`` (4 tests)
- Integration tests: ``xclip`` binary check, X11 roundtrip,
save/restore pattern (2 ignored, require X11 display)
* [cli] add ``transcribe`` command with ``MistralTranscriber`` backend. [Valentin Lab]
Transcribes audio files via Mistral API (``voxtral-mini-latest`` model).
Outputs to stdout or file. Config refactored with ``providers`` namespace
for future backend extensibility.
Includes:
- ``Transcriber`` trait with ``MockTranscriber`` and ``MistralTranscriber``
- ``transcribe`` CLI command with arg parsing
- Integration tests: mock pipeline, error handling, real Mistral API call
* [cli] add ``record`` command with ``CpalCapture`` and ``OpusEncoder`` pipeline. [Valentin Lab]
Captures audio from system microphone, encodes with Opus, writes to file.
Supports optional output path (defaults to ``memo-YYYY-MM-DD-HH-MM-SS.ogg``).
Graceful shutdown via SIGINT (Ctrl+C) with encoder flush.
Includes integration tests verifying:
- Mock capture pipeline creates valid output files
- Real hardware capture creates non-empty Opus files
- Default filename format generation
* [audio] add ``AudioEncoder`` trait with ``MockEncoder`` and ``OpusEncoder`` implementations. [Valentin Lab]
Implement audio encoding trait with two implementations:
- MockEncoder: Pass-through encoder for testing (converts i16 to little-endian bytes)
- OpusEncoder: Real Opus codec encoder using the opus crate with configurable bitrate
Features:
- AudioEncoder trait with encode() and flush() methods
- Stateful encoding with internal buffering for frame-based codecs
- Configuration from AudioConfig (sample_rate, channels, bitrate)
- Support for mono and stereo channels
- Comprehensive unit tests including encode/decode roundtrip verification
- Proper error handling with TalkError::Audio variant
Changes:
- Created src/core/audio/encoder.rs with trait and implementations
- Added Clone derive to AudioConfig for test flexibility
- Exported AudioEncoder, MockEncoder, OpusEncoder in src/core/audio/mod.rs
- All tests pass (10 audio tests)
* [audio] add ``AudioCapture`` with ``CpalCapture`` and full sample format support. [Valentin Lab]
Includes:
- ``AudioCapture`` trait for swappable backends
- ``MockAudioCapture`` for testing
- ``CpalCapture`` with support for all CPAL sample formats
(I8, U8, I16, U16, I32, U32, I64, U64, F32, F64)
- Integration tests verifying real audio device compatibility
All unit and integration tests pass on real hardware.
* [error] add ``TalkError`` enum with ``thiserror`` derives. [Valentin Lab]
Implements comprehensive error handling with variants for:
- ``Config`` - Configuration errors
- ``Audio`` - Audio capture/encoding errors
- ``Transcription`` - API transcription errors
- ``Io`` - IO operations (with ``#[from] std::io::Error``)
- ``Session`` - Session management errors
Includes unit tests for error conversions and documentation.
Refs #1.3
* [config] add ``Config`` struct with ``YAML`` loading. [Valentin Lab]
Implements configuration loading with:
- ``Config`` struct with nested ``MistralConfig``, ``AudioConfig``
- ``Config::load()`` with optional custom path parameter
- ``XDG`` directory support via ``directories`` crate
- Environment variable overrides (``TALK_RS_*`` prefix)
- Fail-fast on missing required fields
- Unit tests for loading and validation
Uses patterns from ``insight-cli`` for directory handling.
Refs #1.4
### Changes
* [visualizer] render amplitude as symmetrical waveform. [Valentin Lab]
Replace bottom-anchored vertical bars with a mirror-image waveform
centred on the vertical midline. Each column extends equally
upward and downward, producing a classic audio waveform look.
Pre-fill the amplitude history buffer with zeros so the waveform
starts at the right edge and scrolls leftward — instead of
stretching a few early samples across the whole panel width.
* [visualizer] skip text panel when not in realtime mode. [Valentin Lab]
The text bar with pulsing dots below the recording badge is only
useful in realtime mode where live transcription text streams in.
In batch mode it just showed empty dots.
Add an ``enable_text`` flag to ``VisualizerHandle::new()`` and
``create_windows()``, gated on ``opts.realtime``. The error
overlay path keeps text enabled so validation failures remain
visible.
* [dictate] switch paste chunking to character-based and init GTK4 in daemon path. [Valentin Lab]
Replace word-based ``split_into_word_chunks`` with character-based
``split_into_char_chunks`` (limit ``PASTE_CHUNK_CHARS`` = 150).
Splits on word boundaries to avoid cutting words. The 150-char
threshold keeps each paste below the point where terminal
applications collapse it into an opaque summary block.
Also call ``gtk4::init()`` before creating the overlay and
visualizer in the daemon code path. Without this, GDK4 monitor
queries silently fail and the overlay/visualizer never appear when
``talk-rs`` runs as a toggle-mode daemon.
* [overlay,visualizer] replace ``xrandr`` with GDK4 monitor geometry. [Valentin Lab]
Remove the per-module ``xrandr --query`` subprocess calls and their
associated parsing helpers (``parse_geometry``, ``parse_primary``,
``parse_geom_word``) along with their tests.
Both modules now call ``monitor::primary_monitor_geometry()`` from
the main thread before spawning their X11 worker threads, passing
the ``MonitorGeometry`` tuple in. This is required because GDK
must be queried from the thread that called ``gtk4::init()``.
* [picker-cache] add ``streaming`` flag to cache entries. [Valentin Lab]
``SelectedEntry`` and ``CachedResult`` now carry a ``streaming`` bool
(defaulting to ``false`` for backwards compatibility) so the picker
cache can distinguish batch from realtime transcription results.
``write_selected`` takes the flag as a new parameter.
* [dictate] improve picker UX and integrate PipeWire capture. [Valentin Lab]
Picker improvements:
- Sort entries by (provider, model) for stable display order.
- Add retry button (↻) on error rows via ``PickerMessage`` enum,
allowing re-transcription without reopening the picker.
- Centre window using ``gdk4_x11::X11Surface::xid()`` in the
``map`` signal, eliminating the ``_NET_CLIENT_LIST`` polling race.
- Refactor ``x11_centre_and_raise`` into ``x11_centre_and_raise_xid``
(direct XID) and a title-search wrapper.
- Record selected (provider, model) via ``picker_cache::write_selected``
so the choice persists across picker reopens.
Focus and paste fixes:
- Add ``ensure_focus()`` with exponential-backoff retries to confirm
the target window is active before pasting.
- Prefer ``read_last_paste_state()`` over recording metadata for
replacement character count (correct across successive picker
selections).
- Pass ``--delay 0`` to ``xdotool key`` in ``simulate_backspace``.
Audio capture:
- Replace ``CpalCapture`` with ``PipeWireCapture`` for live recording
at 48 kHz, piped through ``resample::spawn_resample_task`` to
downsample to 16 kHz before encoding.
* [picker-cache] add selection tracking and legacy format migration. [Valentin Lab]
Introduce ``PickerCache`` struct wrapping results + an optional
``SelectedEntry`` so the picker remembers which (provider, model)
the user last chose.
- ``read()`` now returns ``PickerCache`` and transparently migrates
the old flat-array JSON format.
- Split ``write()`` into ``write_results()`` (preserves selection)
and ``write_selected()`` (preserves results).
- ``selected`` field is omitted from JSON when ``None``
(``skip_serializing_if``).
- Updated tests for new format, legacy migration, and serialisation.
* [config] update default Mistral model to ``voxtral-mini-2507`` [Valentin Lab]
``voxtral-mini-latest`` now aliases to ``voxtral-mini-2602``; pin to
the explicit ``voxtral-mini-2507`` version so users can compare
results between model generations.
* [audio] hardcode ``AudioConfig`` parameters. [Valentin Lab]
The ``audio`` config section (``sample_rate``, ``channels``,
``bitrate``) exposed Opus encoder internals as user-facing
configuration. The only sensible values for voice dictation are
16 kHz / 1 channel / 32 kbps — any other combination either fails
or wastes bandwidth with no transcription quality gain.
``AudioConfig`` now has a ``new()`` constructor returning hardcoded
defaults. The ``audio`` field is removed from ``Config``, along
with ``TALK_RS_AUDIO_*`` env var overrides and the ``env_var_u32``
/ ``env_var_u8`` helpers that were only used for audio.
* [cli] make batch mode the default for ``dictate`` [Valentin Lab]
Replace ``--batch`` flag with ``--realtime``. Streaming upload
(batch) is now the default behavior; pass ``--realtime`` to get
incremental WebSocket transcription instead.
* [cli] paste transcription per segment in ``dictate`` realtime mode. [Valentin Lab]
Instead of accumulating all segments and pasting once after recording
stops, each segment is now pasted into the focused application as it
arrives. This provides real-time feedback while dictating.
``dictate_realtime()`` accepts an optional segment channel. A spawned
paste consumer reads from it and does clipboard-set + ``ctrl+shift+v``
per segment (~150 ms cadence). Clipboard is saved before recording and
restored after the paste task drains.
Batch mode (``--batch``) retains the previous single-paste behavior.
* [cli] replace chunked mode with realtime default and ``--batch`` flag. [Valentin Lab]
The ``dictate`` command now uses WebSocket-based realtime transcription
by default, streaming audio to the Voxtral Realtime API for incremental
results. The previous ``--chunked`` / ``--chunk-seconds`` flags and
``DictateConfig`` are removed in favor of a simpler ``--batch`` flag
that falls back to the original upload-after-stop workflow.
Add client-side ``flush_sentences()`` that splits the live transcription
buffer on sentence-ending punctuation (``.`` ``!`` ``?`` and CJK
equivalents ``。`` ``!`` ``?``), printing completed sentences to stdout
as they arrive. Includes unit tests for Latin and CJK punctuation,
trailing partials, and no-punctuation edge cases.
### Fix
* [picker] switch retry channel to ``tokio::sync::mpsc`` [Valentin Lab]
Replace ``std::sync::mpsc`` with ``tokio::sync::mpsc::unbounded_channel``
for the picker retry channel. The std channel's blocking ``.recv()``
holds a tokio worker thread and prevents the runtime from shutting
down cleanly when the picker window closes.
* [picker] use theme foreground color for selected rows. [Valentin Lab]
Pin ``color`` on ``row:selected`` to ``theme_fg_color`` so text
stays readable on light GTK themes. Without this, some themes
switch selected text to white which is invisible against our
translucent accent selection background.
* [transcription] use ``transcription_session.update`` for OpenAI realtime. [Valentin Lab]
The OpenAI Realtime API requires a different endpoint and event
format for transcription-only sessions:
- URL uses ``?intent=transcription`` instead of ``?model=``
- ``OpenAI-Beta: realtime=v1`` header is required
- Client event is ``transcription_session.update`` (not ``session.update``)
- Session schema uses flat ``input_audio_format``/``input_audio_transcription``
fields instead of nested GA ``audio.input`` structure
- Server responds with ``transcription_session.created``/``.updated``
Also fix the default ``realtime_model`` from ``gpt-4o-realtime-preview``
(a session model) to ``gpt-4o-mini-transcribe`` (a transcription model).
* [audio] fire stop sound immediately on toggle. [Valentin Lab]
Pass ``SoundPlayer`` and ``CancellationToken`` into
``dictate_realtime()`` so the stop sound plays the instant SIGINT
is caught — before ``capture.stop()`` and before the WebSocket
finishes collecting transcription results.
Previously the stop sound played only after ``dictate_realtime()``
returned, adding 100-500ms of perceived delay on toggle.
Batch mode still plays the stop sound in ``dictate()`` after the
recording completes.
* [audio] use deterministic flush signal on capture stop. [Valentin Lab]
Replace the heuristic 50ms sleep in ``CpalCapture::stop()`` with a
deterministic ``std::sync::mpsc`` signal from the callback thread.
The callback stores its final partial buffer in a shared
``final_buffer`` and signals completion via a one-shot channel.
``stop()`` waits for this signal (with a 2-second safety timeout for
unresponsive hardware), then injects the final samples into the
``tokio::sync::mpsc`` channel before closing it.
This guarantees the last audio chunk is never silently dropped,
regardless of callback timing or system load.
* [audio] wait for output device warmup before playing start sound. [Valentin Lab]
``SoundPlayer::new()`` now blocks until the CPAL output callback has
actually fired at least once, proving the audio pipeline is live. A
``Condvar`` bridges the callback thread and the caller — no arbitrary
sleep.
Previously, ``stream.play()`` returned before the device was ready,
so the first few milliseconds of the start sound were lost to device
startup latency, making the sound appear clipped.
* [audio] flush partial buffer on capture stop. [Valentin Lab]
The CPAL callback was discarding all incoming audio once ``running``
became false — any samples accumulated in its internal buffer that
had not yet reached ``samples_per_chunk`` were silently lost. This
caused the end of every recording to be chopped off.
Now the callback flushes its partial buffer exactly once when it
detects the stop signal. A 50ms grace period before dropping the
stream ensures the callback has time to see the flag and flush.
* [transcription] harden networking against hangs and silent drops. [Valentin Lab]
The realtime WebSocket and batch HTTP paths had several unhandled
failure modes that could hang the app indefinitely:
- ``connect_async`` and ``wait_for_session_created`` had no timeout,
hanging forever on unreachable servers or unresponsive sessions
- ``transcribe_file`` (batch) had no request timeout at all
- No WebSocket ping keepalive, making silent network drops undetectable
- Sender and receiver tasks ran independently with no coordination;
one dying left the other hanging
- Cleanup task silently swallowed panics from spawned tasks
- Silent WS stream end (TCP RST) was not logged
Fixes:
- 15s timeout on WS connect and ``session.created`` handshake
- 300s timeout on batch ``transcribe_file`` request
- 30s periodic Ping frames in sender loop for keepalive
- Shared ``CancellationToken`` between sender/receiver for coordinated
shutdown on either side's failure
- Cleanup task now logs panics from ``JoinHandle`` results
- Unexpected stream end logged and cancellation propagated
### Other
* Test: [audio] add closed-loop loopback integration tests. [Valentin Lab]
5 tests using a PipeWire null sink as a virtual audio loopback:
- ``loopback_start_sound_is_captured``: start sound produces non-silent
audio through the loopback pipeline.
- ``loopback_stop_sound_is_captured``: same for the stop sound.
- ``loopback_start_sound_has_two_notes``: RMS energy envelope confirms
two distinct tone bursts (sound completeness, not clipped).
- ``loopback_preemption_replaces_sound``: at most 3 tone onsets when
stop preempts start (4 would mean both played fully).
- ``loopback_capture_receives_played_audio_i16``: f32-to-i16 conversion
path verified through loopback.
Uses the ``"pulse"`` ALSA device (not ``"default"``) so cpal routes
through PulseAudio where ``$PULSE_SINK`` / ``$PULSE_SOURCE`` are
respected — tests are completely silent (no sound on real speakers).
Infrastructure: shared null sink via ``OnceLock`` with stale cleanup,
serial lock with poison recovery, ``SoundPlayer::from_device()``
targeting.
Run with: ``cargo test --test audio_loopback -- --ignored --test-threads=1``