My Voice Assistant Spent 9.9 Seconds Loading Whisper. Every Single Turn.
I built a thing that answers when I talk to it. It takes the microphone, turns what I said into text, hands that to a model, and speaks the reply back out loud. I never touched a cloud speech API. The brain in the middle is a different story, and I deal with it head on further down.
The first version that worked end to end wasn’t a conversation. It was a form submission. I would finish talking, and then nothing happened for long enough that I started checking whether the thing had crashed.
All of that delay turned out to be sitting in one place. Every single turn was launching the whisper CLI from scratch. Reading a 1.5GB model off disk and initialising Metal again cost 9.9 seconds per turn, so I spent ten seconds before anything in the pipeline knew what I had said.
The fix was neither new nor clever. Leave the model loaded. That one change took the same stage to 1.0 second a turn.
This post is about those nine seconds, and about the delays that were still standing afterwards.

What I actually built
Three stages, and that’s the whole pipeline.
mic → STT (whisper.cpp, local) → LLM → TTS → speaker → (auto re-listen) → mic
- STT is whisper.cpp running
large-v3-turbowith beam size 5. It stays up as a resident server, and I throw audio at it over HTTP. - The LLM is a CLI model I already pay a subscription for, on the small fast tier. That part does not run on my machine, and it gets its own section below.
- TTS defaults to the device’s built-in speech synthesis, with a local XTTS-v2 server as the option. The reason that order ended up backwards is a section of its own too.
- The loop is hands-free. I press a button once and never touch it again. A stretch of silence counts as the end of my turn and sends automatically, and once the answer finishes it starts listening again.
Three processes stay resident: the backend, the TTS server, the STT server. All three are registered as OS services, so a reboot brings them back without me.
Push-to-talk was a deliberate choice. A wake word assumes the microphone is recording all the time, so I pushed that down the list at the very start. I didn’t want to build a thing that listens continuously.
The expensive part: 9.9 seconds, every turn
The original STT path shelled out to the CLI. Convert the audio to 16kHz mono wav, spawn whisper-cli as a child process, read standard output, tidy it up. It worked fine.
The problem is that the whisper CLI starts from nothing on every call. It reads the model file, initialises the GPU backend, and only then looks at my audio. The actual transcription was a small slice buried inside all of that.
So I moved to whisper-server. The process stays up, the model stays in memory, and all I do is POST audio to /inference.
| CLI spawn | Resident server | |
|---|---|---|
| STT per turn | ~9.9s | ~1.0s |
Close to nine seconds a turn simply disappeared. The only thing that changed in the code was how the call gets made.
I kept one piece of the old path. If the resident server dies, it falls back to spawning the CLI. Slow and answering beats fast and occasionally silent.
The model changed as well. I started on base at 141MB, and it kept missing short words and proper nouns, writing the same word down a different way each time. Moving up to large-v3-turbo with beam size 5 cut that down noticeably. That judgement was made with my ears, and I never measured WER, so I have no percentage to hand anybody.
Getting the first sentence out early
Once STT was down to a second, the next wall was visible. Nothing made a sound until the LLM had finished writing the entire answer.
That was the obvious consequence of the structure I had. Take the whole response, hand it to TTS, play it. The longer the answer, the longer the silence in front of it.
Two things fixed it.
1. Stream the tokens and speak sentence by sentence. I ran the CLI in streaming output mode and put a sentence-boundary streamer and a speech queue in front of it. The moment the first sentence is complete, speech starts, and the rest of the answer keeps generating behind it.
2. Turn thinking off. A conversational reply has no use for reasoning tokens. I pinned it to zero through an environment variable and moved on.
The result was first sound at ~3s instead of ~5s. The full response takes exactly as long as it always did. The only thing that changed is how long I spend listening to nothing.
The screen stopped sitting idle during the wait, too. A short acknowledgement sound fires immediately, listening and thinking and speaking each animate as separate states, and Android gets a vibration. None of that removes a single millisecond of delay. It removes the stretch where I stared at the screen wondering whether the thing was still alive.
I know where the floor is. A cold first token out of the CLI takes ~3 seconds, and this structure has no way to get underneath that.
TTS: I built the better one, then took it out of the default
This is the least flattering part of the project.
I attached a local neural TTS server — XTTS-v2 running on Apple Silicon MPS as a resident Python service. The voice is clearly more natural than the alternative.
And then I put the default back to the device’s built-in speech synthesis.
Two reasons.
- XTTS adds ~2.5 seconds before it starts speaking. Cutting two seconds out of the LLM stage and handing 2.5 back at the TTS stage doesn’t add up.
- It hangs intermittently on MPS. The server stops responding and hangs there. So a failure falls back to device TTS, and needing that fallback is the same as admitting it can’t be the default.
I had tried other engines before that. Every one of these problems exists because I need Korean.
| Engine | What I ran into |
|---|---|
| Kokoro | Korean is not among the supported languages |
| MeloTTS | I couldn’t get it to build. I never reached the end |
| MMS-TTS (kor) | Works. Needs uroman romanisation in front of it, and the output is clear but flat |
| XTTS-v2 | The most natural of them. ~2.5s, plus intermittent hangs |
| Device built-in | Arrives instantly. The voice differs from device to device |
My real criterion turned out to be latency rather than audio quality. I took an ordinary voice that arrives now over a lovely voice that arrives ~2.5 seconds late. That’s something I only admitted after building the whole local TTS server.
All the pre-processing in front of TTS stayed exactly where it was, because it’s needed whichever engine wins.
- Strip emoji and symbols. Skip this and the model’s punctuation gets read out loud, character by character.
- Split the Korean and English spans. Handing a mixed sentence over in one piece wrecks the pronunciation on one side of it.
- Transcribe English words into Korean phonetics.
- Chunk long sentences.
The numbers
| Stage | Before | After | What changed |
|---|---|---|---|
| STT (per turn) | ~9.9s | ~1.0s | CLI spawn to resident server |
| First sound (LLM stage) | ~5s | ~3s | Streaming, sentence-level speech, thinking off |
| TTS start of speech | XTTS ~2.5s | Instant | Device TTS as the default |
| Remaining floor | — | ~3s | Cold first token from the CLI |
What the table doesn’t have: the real end-to-end time from tapping the button to hearing the first sound. I measured the stages one at a time and never once put a stopwatch across the whole thing. Adding the stages together would be arithmetic rather than measurement, so it isn’t in here.
Silence detection sets the rhythm of the conversation
The hands-free loop contains one value that has nothing to do with any model. How many milliseconds of quiet counts as me being finished. That single number decides when a turn ends at all.
Set it long and I wait that much longer after I’ve already stopped talking. That wait is longer than the entire second that STT now takes.
Mine sits at 1100ms today. 700ms is written down as the next thing to try, and I haven’t tried it. A value I have never run is a value I won’t publish. This one belongs to the same family as a hardcoded constant that goes quietly wrong without ever throwing an error.
Don’t do these
Every one of these is something I did.
- Don’t spawn a heavy model binary fresh on every turn. I was spending 9.9 seconds a turn on model loading, and the transcription inside that was tiny. Switching to a resident server alone brought it to ~1.0 second.
- Don’t run a resident server without a fallback. The moment that process dies, the entire product dies with it. Keep the slow path around.
- Don’t wait for the whole response before speaking. Cutting at sentence boundaries and speaking the first one took time-to-first-sound from ~5s to ~3s. The total doesn’t shrink. The silence does, and the silence is what I was feeling.
- Don’t measure conversational latency with reasoning tokens still on. I turned them off, and the numbers meant something afterwards.
- Don’t pick a TTS default on audio quality. I built a more natural local neural TTS and then removed it from the default. No voice is pretty enough to beat ~2.5 seconds and intermittent hangs.
- Don’t test a phone microphone without HTTPS.
getUserMediarefuses to open at all outside a secure context. I could not test the microphone path until HTTPS was in place. - Don’t judge STT accuracy on the smallest model.
basewrote short proper nouns down differently every time. I never measured WER either, so this claim stops at “it got noticeably better.” - Don’t blame the model for the rhythm of a conversation before looking at the silence threshold. I run 1100ms, and 700ms is still untested.
Honest limits
- The brain is not on my machine. STT and TTS run entirely on hardware I own. The LLM in the middle is a CLI model on a subscription. The metered API bill is zero, and none of this is offline. Cut the internet and this thing can only take dictation.
- I could have gone all the way and didn’t. The local model I run is coding-specialised and wrong for conversation. Standing up a separate conversational model is a path I designed and never implemented or measured, which is why there’s no local LLM latency figure in this post.
- I never measured WER. The whisper model comparison was done entirely by ear.
- I never measured tap-to-first-sound end to end. See the note under the table.
- I never measured how often XTTS hangs. “Intermittently” is as far as I can go, and I don’t know whether that’s one turn in five or one in fifty.
- 700ms silence detection is untried.
- Device built-in TTS gives me no control. The voice differs per device, so the same sentence sounds different depending on what plays it, and unifying that means going back to local neural TTS. The trade-off above is not settled.
- These are Apple Silicon numbers. I have never run this on other hardware, and the MPS hang in particular is something I have only seen in this one combination.
- That engine table exists because I need Korean. Anybody who only needs English will find this part far less painful.
So is this worth building?
- Not worth it if asking a question out loud and hearing an answer is the whole job. The assistant already on the phone finishes that, and there is no reason to build any of this.
- Worth it when the requirement is that recorded audio never leaves hardware I own. What this design protects stops right there. The transcribed sentence still goes out to the brain. For anyone who treats the audio and the sentence as the same thing, this structure does nothing at all.
- The hardware was already running. I put three services on a Mac I leave on anyway. Extra spend: zero.
Closing
The thing that held me up longest here was never model selection. My latency was piled up in one place, and it was setup work rather than computation. It was whisper, booting again on every single turn. That cost 9.9 seconds a turn, and leaving the model loaded turned the same turn into 1.0 second.