Alfred, part 2 — getting voice to work
Teaching a computer to hear me, understand me and reply out loud — all on one $300 mini PC, all offline, all free.
Problem
The constitution says local, private, $0 — so no cloud speech-to-text, no cloud text-to-speech, no API credits. Every voice capability had to run on the PC itself with free, open-source models, and the first voice engine sounded mechanical.
Approach
My instinct was to blame the engine and shop for a better one — that was wrong. Most of the problem was the text: lists joined with semicolons gave it no breath, ALL-CAPS place names got shouted. A text-normalisation layer fixed that, and I swapped engines anyway to a more natural-sounding one — after measuring that the ‘fast’ quantised model was actually three and a half times slower on my CPU than the full-quality one. For hearing, whisper handled speech-to-text, with a cleanup layer to strip hallucinated phrases and voice-activity detection to stop waiting out a fixed five-second window.
Outcome
The full loop works end to end: press the mic or say the wake word, whisper transcribes, an intent router decides what was asked — keeping factual answers 100% deterministic so Alfred can never invent data — and the voice engine replies. The end-to-end test that sealed it: asking what's in my calendar and hearing it back naturally, from real data, all offline, all $0.
Lessons learnt
Measure before you build, especially the thing you're excited about — the quantised model, a proposed warm server, and a GPU path I skipped were all killed or inverted by a ten-minute measurement. And diagnose, don't replace: the voice sounded bad, and the fix was the text, not the engine.