gavinbowden.me — home

Inkblot Dictation

inkblot-dictation is a standalone Rust crate that turns a live microphone stream into cleaned-up text with an on-device Whisper model. Audio is chunked on speech and silence, transcribed with no network calls, and exposed as a plain callback API any Rust app can use.

Microphone in, text out, nothing else

I'm building a Tauri + React app that wants live dictation, and I didn't want that app's codebase clogged up with an entire speech-to-text pipeline. So the pipeline lives somewhere else.

Microphone in, cleaned transcript out. No cloud calls, no Tauri dependency. It loads a local Whisper model once, listens on the default input device, decides for itself where one spoken phrase ends and the next begins, and hands back both a raw transcript and a cleaned one over a callback, whether the caller is a Tauri command or a bare CLI loop.

Two stages, one boundary that matters

The pipeline is two async tasks connected by bounded channels: a chunk builder that owns the audio, and a transcription worker that owns Whisper. They are kept apart because of one rule.

Never block the audio thread. cpal calls the input callback on a real-time thread, and holding onto it for even a few milliseconds drops audio. So that callback only converts samples and does a try_send, and it always throws a packet away rather than wait for room.

Whisper inference is the opposite kind of work: CPU-bound, and slow on purpose. It runs inside tokio::task::spawn_blocking rather than on the async runtime, where it would stall every other task sharing the thread.

  • The audio callback is one generic build_typed_input_stream<T> that handles whatever format the default input device actually hands over (F32, I16, or U16), rather than assuming F32 and failing on devices that don't offer it.
  • Every packet gets downmixed to mono and linearly resampled to the 16kHz Whisper expects, regardless of the mic's native rate. It's plain linear interpolation, not a proper sinc resampler. Dictation isn't where resampling artifacts show up, so I didn't pull in a DSP crate to avoid them.
  • The Whisper model loads once and is shared behind an Arc, with a fresh WhisperState per chunk, so concurrent inference never shares mutable state.

Chunking a stream that never stops

Whisper doesn't take a live stream, it takes a buffer, so something has to decide where one buffer ends and the next begins. Voice activity detection here is a single RMS-threshold gate, not a model: root-mean-square against a floor (0.012 by default). It's only used for that boundary decision, not for cleaning the audio. Once speech crosses the threshold, a chunk opens. It closes on whichever fires first: 900ms of continuous silence (once the chunk is at least 1.2 seconds old, so a quick breath doesn't split "um... okay" into two chunks), or a 12-second hard cap, so one long run-on sentence can't grow the buffer without bound.

  • While a chunk is still open, it gets re-transcribed every 1500ms and emitted as a Partial event, which gives a UI something to show mid-sentence instead of a blank field until the pause. Each partial re-transcribes the whole growing buffer from scratch, and Whisper runs with no_context(true), so nothing carries between calls. Simple, at the cost of doing some of the same work twice.
  • That statelessness is what makes a failed chunk harmless: if one chunk errors, there's no cross-chunk state for it to corrupt. It is not, however, the same thing as ordering. More on that below.

Cleanup, and the bugs I found writing this

Whisper already returns cased, punctuated text. My cleanup pass throws that away on purpose: it lowercases everything, turns spoken punctuation ("comma", "period", "new paragraph") into the real thing, fixes spacing before punctuation, then re-capitalizes sentence starts and a standalone "i" (checking the characters on either side, so "is" and "it" survive). The upside is one predictable output format whether you said "comma" or just paused. The downside is that proper nouns and acronyms get flattened, so "NASA" comes out as "nasa". Both the raw and cleaned text ship in every event, so a caller can always fall back to what Whisper actually heard.

Re-reading that module for this write-up turned up three bugs, all of which are, honestly, pretty funny. The spoken-punctuation rules are substring matches with a leading space and no trailing boundary, which means " comma" also matches inside " command".

So "run the command" comes out as "run the,nd".

  • Periodic, colonel, and colony all break the same way, for the same reason.
  • The " quote " rule runs before " end quote ", so "end quote" becomes "end “" and the end-quote rule can never fire at all.
  • " new paragraph " needs a trailing space, but Whisper usually hands back "new paragraph." with a period attached. The most useful command is the one that most often doesn't work.

The fix for all three is the same: split into words first and match whole tokens, longest phrase first, instead of patching the string in place.

Staying out of the app's way

The crate never imports tauri, and that's on purpose. The public surface is DictationService (load a model, start/stop a session, ask its status) plus an Arc<dyn Fn(DictationEvent)> callback. DictationError converts into a String, so a Tauri command can return it with a plain .map_err(Into::into) without the crate knowing Tauri commands exist, and events are serde-tagged (kind plus payload, camelCase) so they cross Tauri's IPC as-is. The included live_dictation example drives the exact same service and callback from a bare CLI loop, which both proves the boundary holds and doubles as a usage demo. GPU backends work the same way: forwarded to whisper-rs as Cargo features (metal, vulkan, cuda), so the host app opts in per device instead of the crate guessing.

Where this goes next

This is the working prototype, not the hardened version, and the honest list is longer than I'd like.

The biggest item is ordering. The worker fires off each spawn_blocking inference and never awaits the handle, so several can run at once and finish in any order. A slow partial can land after its own Final, and Finals can arrive after Stopped. A per-session sequence number, or simply awaiting each job in turn, fixes it.

There's also no automated test suite, which is precisely how three string-matching bugs survived long enough to get written up on a portfolio. The cleanup module is pure string in, string out, so it's the easiest possible place to start.

After that, the fixed RMS threshold needs noise-floor adaptation (a threshold tuned for a quiet room clips soft speech, or hangs open in a noisy one), and re-transcribing the whole buffer should become incremental, so partials stop getting more expensive as a chunk grows.