This is Part 3 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, and Part 2 where rules beat models. This part is about latency: where the seconds went, and why the fixes that mattered were almost never “use a smaller model.”
Before our first round of latency work, an average response took 4.5 seconds end to end. On a phone line that doesn’t sound slow; it sounds broken. Four and a half seconds of silence is long enough for a caller to wonder whether the line dropped, say “hello?”, and start talking over the agent just as it begins to answer.
My first instinct was to ask whether a smaller LLM would do. Wrong first question, and the answer turned out to be backwards: a 35B model ended up answering faster than a 4B. The mental model I use now:
Per-turn latency ≈ LLM calls per turn × (prompt processing + tokens you choose to say × time per token) + whatever your serving stack wastes.
Model size is one input to “time per token,” and rarely the biggest one. Most of my wins came from everywhere else: the inference engine, numeric precision, CUDA graphs, the ONNX execution provider, and how many times a single turn called the model.
Profile Per Stage Before You Shrink Anything
The first useful thing I did was boring: I timed each stage separately over about 100 responses. Here is the before/after from the first round of serving work, reported in early January, next to the team’s ~820 ms streaming budget (which also set aside ~120 ms for the telephony legs):
| Stage | Before | After (Jan) | Change | Team budget |
|---|---|---|---|---|
| STT | 351 ms | 157 ms | −55% | 250 ms |
| LLM | 584 ms | 324 ms | −45% | 250 ms |
| TTS | 3,555 ms | 713 ms | −80% | 200 ms |
| Total | 4,492 ms | 1,196 ms | −73% | ~820 ms |
The lesson is in the “Before” column. TTS was 79% of the wait; the LLM was 13%. If I had started by shrinking the LLM, the best possible outcome was saving part of that 13%.
What moved each row:
- STT: decoding with faster-whisper on CTranslate2 (training is covered in my Whisper series).
- LLM: moving serving from
transformersto vLLM, which alone took a single call from 0.6–1.5 s to 0.1–0.3 s. - TTS: synthesizing sentence by sentence and streaming, so the first sentence plays while the rest renders.
Even after a 73% cut we were over budget, with TTS still at 3.5× its share. And the table reports means. That was a mistake, and the section on LLM calls shows what the means were hiding.
Why the 35B Beat the 4B
In March I swapped the dense 4B for Qwen3.5-35B-A3B, a mixture-of-experts model that activates roughly 3B parameters per token. I expected to buy quality with latency. In April I benchmarked it against the alternatives with one harness: the team’s long production prompt, 407 turns per model, BF16 weights, CUDA graphs on, wall-clock time per complete answer.
| Model | Total / active params | Median turn | p90 | Mean answer length |
|---|---|---|---|---|
| Qwen3-4B (dense) | 4B / 4B | 279 ms | 474 ms | 86 chars |
| Qwen3.5-27B (dense) | 27B / 27B | 1,225 ms | 1,769 ms | 71 chars |
| Gemma-4-26B-A4B (MoE) | 26B / ~4B | 251 ms | 414 ms | 62 chars |
| Qwen3.5-35B-A3B (MoE) | 35B / ~3B | 289 ms | 414 ms | 65 chars |
| GPT-5 (API, default settings) | — | 19,840 ms | 41,443 ms | 78 chars |
Three things stand out:
- Active parameters set the pace, not total parameters. The dense 27B was ~4× slower than the 35B MoE; the two MoEs and the 4B landed in the same band.
- The 4B talked more. Its answers ran a third longer than the 35B’s, and half again as long with retrieved documents in the prompt (106 vs 72 characters). Going back over the data for this post, a linear fit gives about 2.8 ms per output character for the 4B and 2.3 ms for the 35B. Per character they were close; the 4B had the lower fixed cost and spent it on extra words. Net result: a tie at the median, and a tighter tail for the 35B.
- GPT-5 was out regardless of quality. A ~20 s median is not a phone conversation.
Precision turned the tie into a win. In BF16 the 35B’s weights alone are ~70 GB, nearly the whole 80 GB card before any KV cache. The FP8 checkpoint halves that, along with the bytes moved per generated token. In a separate model sweep with CUDA graphs on, the FP8 35B finished a turn faster than a dense 4B running in BF16. That is roughly how I explained it to the business team: the same speed per character, longer answers from the small model, a lighter number format for the big one.
On a phone line, extra tokens cost twice. Once to generate, and again when TTS synthesizes them and the caller sits through them. Answer length is a latency setting.
CUDA Graphs and the Cost of a “Temporary” Eager Flag
A couple of hours before the swap, I had added an environment switch, VLLM_ENFORCE_EAGER, as an escape hatch for GPU or driver compatibility problems. Shortly after the swap, the new MoE crashed at startup inside torch.compile with a FakeTensorMode error. The one-line fix was right there: set the flag, skip compilation, move on.
I didn’t, because in vLLM eager mode also turns off CUDA graph capture. At batch size 1, a decode step is hundreds of small kernel launches whose overhead can rival the math; a CUDA graph records the step once and replays it as a single launch. MoE layers, with routing and many small per-expert matmuls, have the most to lose. So I disabled only torch.compile and kept graph capture. Four minutes later I reverted that and hand-patched vLLM on the server instead, a fragile fix of its own, but graphs stayed on.
A later model sweep showed what that decision was worth. Turning CUDA graphs off slowed every model, and the MoE models I was betting on lost the most, by multiples per turn. In eager mode the 35B lost to the dense 4B; with graphs it won. Had I flipped that flag in March, the swap would have looked like a latency regression, and I would probably have rolled back to the 4B for the wrong reason.
An escape hatch is a performance setting. Any flag that disables graph capture, compilation or a fused kernel should log loudly at startup and appear next to the latency numbers. “Temporary” flags outlive the bugs they were added for.
Count LLM Calls Per Turn
By mid-March I had moved several regex-based extractions into LLM calls (Part 2), because the model seemed to handle messy STT text better than my patterns did. Each change was reasonable on its own. Together, a single intake turn could fire intent classification, a mid-flow relevance check, location extraction, time extraction, a coherence check and a query rewrite, in sequence, each reading its own prompt first. Six calls at 100–300 ms apiece blow the budget before TTS even starts.
Over March 24–25 I collapsed most of that into one structured call that returns intent, request category, location, time and a flow status (continue, end, or new topic). Downstream steps read those slots instead of asking again. The two-call chains for query rewriting and contact lookup became single calls, one yes/no question went to a regex, and the coherence check became a confidence threshold. By my estimate at the time, model calls per request dropped by about 48%.
What I didn’t measure was the shape of the result. The April benchmark did: same 35B, 407 turns, hard-coded orchestration vs the single long prompt.
| Path (Qwen3.5-35B-A3B) | Mean | Median | p90 | p99 | Turns over 600 ms |
|---|---|---|---|---|---|
| Orchestrated flow | 295 ms | 161 ms | 858 ms | 1,237 ms | 84 / 407 |
| Single prompt | 302 ms | 289 ms | 414 ms | 502 ms | 2 / 407 |
Pick your statistic, get your conclusion. The mean says it’s a tie. The median says orchestration is nearly twice as fast, because 41% of orchestrated turns came back in under 5 ms: canned templates that never touched the model. The p90 says it’s twice as slow, because the turns that did reach the model often chained calls. On the dense 27B the orchestrated p90 was 3.75 s against 1.77 s, and in the model sweep the largest models came out slower orchestrated, since every extra call is paid at big-model speed. The tail is what a caller hears.
Report latency as a distribution. Median, p90 and p99 per stage, plus how many turns cross the budget. A mean can make a bimodal system look fine.
TTS: GPU Is a Hypothesis, Not an Optimization
The TTS target was to match a commercial TTS API on MOS and on ASR round-trip CER (synthesize, transcribe, compare): better, or at least not significantly worse. The path there:
| When | Choice | What happened |
|---|---|---|
| Dec | Qwen3-Omni (speech LLM) | Rejected: emotion and speaking rate only steerable through the prompt; too new for reliable vLLM serving |
| Dec–Jan | Zonos, sentence-level synthesis + streaming | TTS stage 3,555 → 713 ms average |
| Mar | Supertonic, a ~305 MB ONNX model, on CPU | Small, CPU-friendly; ONNX thread counts tuned |
| Mar | Same model, plain CUDA execution provider | On the GPU, but (as I found in May) slower than CPU |
| May | Supertonic 3 + TensorRT EP, cached engines | ~2 s → ~100 ms per synthesis |
Getting onto the GPU at all took three commits in 13 minutes. Supertonic builds four ONNX sessions internally, and overriding the library’s default provider constant did nothing, first in one module and then in a second, presumably because the value had been read before my override ran. What worked was rebuilding all four InferenceSessions with explicit providers and logging get_providers() for each. Ask each session which provider it actually got; don’t infer it from the config you passed.
But “on the GPU” is a placement, not a speedup. When I dug in during May, plain CUDA EP was inserting host↔device Memcpy nodes around ops it couldn’t keep on the device, which on this model made it slower than CPU. The TensorRT execution provider compiles the graph into GPU-resident engines. With engine and timing caches on disk, synthesis went from ~2 s to ~100 ms, and a benchmark of 218 agent turns went from 12.1 minutes to 22 seconds, about 30×. (The January and May TTS numbers were measured differently, so don’t chain them.)
The costs: the first engine build takes 1–5 minutes per model; the CPU and GPU onnxruntime wheels provide the same module and conflict; and the TensorRT → CUDA → CPU fallback chain degrades silently. I wrote that last one down as a warning in the docs. A warning in the docs is not an assertion.
One GPU, Three Models
All of this shared one 80 GB H100:
| Service | Engine | VRAM |
|---|---|---|
| LLM | vLLM, 35B-A3B MoE in FP8 (gpu_memory_utilization=0.75) |
~60 GB |
| STT | faster-whisper, fine-tuned Whisper large-v3-turbo (fp16) | ~8 GB |
| TTS | Supertonic 3, TensorRT engines | ~1.7 GB |
vLLM pre-allocates its fraction of the card for weights plus KV cache, so gpu_memory_utilization is a budget you negotiate with your neighbors, not a ceiling. Mine went 0.4 → 0.9 → 0.8 → 0.75 within a week: up to fit the BF16 MoE on swap day, then down as FP8 freed memory and STT and TTS needed room.
Tooling fought for the same card. In March I loaded a QA judge as a second vLLM engine beside the server. It needed 5.62 GiB for its context; 5.52 GiB was free. Four commits in 35 minutes followed: a context cap, 5% utilization with a 2,048-token context, and a two-commit switch to transformers. The real fix came four days later and was embarrassingly simple: a /generate endpoint on the running server, so tooling reuses the model that is already loaded. I had moved the judge to a 30B-class model earlier that afternoon, so that saved roughly 60 GB.
The sizing rule I now apply before any multi-model run: bytes per parameter × total parameters. The “A3B” in an MoE’s name describes compute, not memory; every expert is loaded. 35B is ~70 GB in BF16 and ~35 GB in FP8. A 397B MoE in BF16 is ~794 GB of weights before a single token of KV cache.
What I’d Change
- Serve asynchronously from day one. In the code as of late April,
/chatwas anasynchandler calling the synchronous orchestrator directly, on one uvicorn worker, over vLLM’s offlineLLMclass, so simultaneous callers would have queued instead of being batched. Nothing in my notes says it bit us in testing, but phone calls arrive concurrently by nature. Next time: vLLM’s async engine or its OpenAI-compatible server, plus a concurrency test in every latency report. - Time what the caller hears. My benchmarks timed complete answers; the team’s budget assumed streaming, which is closer to time to first audio.
- Percentiles from the first report. The team’s design sheet asked for p50/p95 logging. My January report gave means.
Summary
- Profile per stage, by percentile, before touching the model. TTS was 79% of the original latency and the LLM 13%, and means hid the orchestration tail entirely.
- Speed is active parameters × tokens you choose to say, on a stack that isn’t wasting time. An FP8 MoE with ~3B active parameters, CUDA graphs on and concise answers beat a dense 4B; a dense 27B was ~4× slower. Count calls per turn as carefully as parameters.
- Verify the accelerator instead of assuming it. Ask each ONNX session which provider it got, benchmark “on the GPU” rather than trusting it, keep CUDA graphs on, and budget a shared card in bytes × total parameters.
What’s Next
One of the speedups in this post later vanished after a routine deploy, with no error, no crash and no alert. Part 4 is about failures like that: silent fallbacks, defaults standing in for missing data, truncation you never see, and an evaluation pipeline that happily scored empty answers.