This is Part 4 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, and Part 3 latency. This part is about the failures that never threw an exception.


When a web app breaks, someone at least sees a 500 page. When a voice agent breaks, the caller hears nothing. They say “여보세요?” into the silence, wait a few seconds, and hang up. No stack trace reaches them, and often none reaches you either.

Almost every failure in this post looked like that from the outside. Nothing crashed, and each one traced back to a choice that looked reasonable on its own. Together they left me with four rules:

  • “Failed” must never share a code path with “decided”.
  • Every automatic fallback needs an assertion, not a log line.
  • Everything that shapes a cached output belongs in its key.
  • Evaluation infrastructure must treat empty output as an error, not as a bad answer.

A Taxonomy of Quiet Failures

The first five rows below happened on this project; I found the last one by reading the code, not through an incident.

Failure class What it looked like Guard I’d add first
Failure mapped to a decision An LLM error read as “caller is done” → dead line Distinct error value + spoken fallback
A default hiding missing data Renamed JSON keys → location_type always "none" Schema-constrained decoding; contract tests
Truncation Department list over the context; Whisper prompt over budget Count tokens first; make overflow loud
A fallback that degrades TensorRT → CPU, ~100 ms → 2.7 s Assert the provider; post-deploy latency check
Eval infra scoring errors 399 of 407 GPT-5 answers empty Empty is an error; count failures per model
A stale cache Voice tuning reached only uncached sentences Key on every input that shapes the output

In every row, the system kept producing plausible output. A silent hang-up looks like a caller who said goodbye, and a default "none" looks like a caller who never mentioned a place.

When “Failed” Means “Hang Up”

After answering a question, the agent asks whether there is anything else. Callers often reply with something short and ambiguous, like “아, 네…”, which might be a goodbye or the run-up to a second question. A small LLM call judged it: return what to handle next, or None for “the caller is finished”.

The judge also returned None when the LLM call itself threw, and the executor routed None to the termination node, which ends the call without a word. So a timeout or a malformed response meant silence, then a dial tone, while the logs showed a warning followed by what looked like an ordinary goodbye. The fix (April 9) was small: a dedicated sentinel for “the judge failed”, a short spoken “please try again in a moment” before ending, and the log level raised to ERROR with the traceback.

Once I saw the shape, I found it in three more places: one I had patched weeks earlier without naming the pattern, one I fixed the same afternoon, and one a week later. Each time, uncertainty had been wired to the most destructive action available:

Uncertain situation Old action New action
Classifier can’t place the request (maybe just a garbled transcript) “That isn’t something I can help with” One clarifying question, then re-classify
Location still missing after the allowed re-asks “Sorry, I can’t process this”, discarding what was collected Forward the partial details to the department
After a detour, the caller answers the earlier question (“next to the convenience store”) Treated as a new query; only a short “yes” resumed the old flow Resume unless the caller declines

The QA harness was the one place that already got this right: its semantic judge treats its own request failure as “NO”, so the row fails. Failure counts against the system, never for it.

Uncertainty should buy the least destructive action. A clarifying question costs a few seconds; a wrong hang-up costs the call. When code can’t tell “the user decided” from “something broke”, it will eventually do the expensive thing for the cheap reason.

Model Swaps Break Silent Contracts

On March 24 I moved the LLM from Qwen3-4B to the Qwen3.5-35B-A3B MoE. The loud part of that afternoon was almost comic: three wrong Hugging Face model IDs in 31 minutes — a size that doesn’t exist, a real model from the previous family, then an -Instruct suffix the release doesn’t use. I was moving fast with an AI coding assistant writing most of the diffs, and neither of us checked the hub first. Then vLLM hit the torch.compile error on load from Part 3. I committed a workaround, reverted it four minutes later, and hand-patched vLLM on the server instead — outside version control, on a deploy that reinstalls vLLM whenever torch changes. A fix with a fuse on it.

Those were the good kind of failure: the model wouldn’t load, so I knew. The quiet one came that night. I had just merged several LLM calls into one classifier prompt that returned JSON with location_type and a new flow_status field. I asked for the JSON in the prompt only, with no schema-constrained decoding. The model started returning locationtype and flowstatus, and the parser was tolerant in the worst way: result.get("location_type", "none"). So location_type was always "none", flow_status was never detected, and nothing raised. From the outside, the classifier had simply stopped understanding places and endings.

My fix (March 25) stripped underscores from every key, which is a patch, not a guard. Going back over the code for this post, I found that a different prompt in the same codebase spelled that field locationtype all along. So I still can’t say whether the new model, the merged prompt or the inconsistent spelling caused the drift. Two things had changed about ten hours apart, a third had been wrong from the start, and nothing in the system could tell me which one broke the contract.

Run the structured-output contract tests first after a swap; better, use schema-constrained decoding. vLLM can constrain generation to a JSON schema. A missing required key should be a parse error you count, not a default you silently accept. Prompt-only JSON is the first thing I’d change.

Truncation You Never See

The department list. When the documents couldn’t answer a question, the agent asked the LLM to pick a department from the client site’s list. For the largest site, that meant hundreds of names in one prompt, overflowing a 4,096-token limit. The lookup caught the error and returned None — the same value that meant “no matching department”. So the caller heard “Sorry, I couldn’t find that information”, a perfectly reasonable sentence that no one listening would flag. What caught it was failing QA rows, not the warning buried in the logs.

My fix (March 25) prefiltered departments by keyword overlap to the top 60 candidates, raised to 80 later that day. Reading it again, the fix is itself a silent truncation: when no keyword overlaps, every candidate scores zero, and the “top” candidates are just the first names in alphabetical order. It should log whenever the cap cuts anything and have an explicit “no confident candidate” path.

The Whisper prompt. On March 31 I added per-site keyword prompts to our fine-tuned Whisper (from my earlier series) through faster-whisper’s initial_prompt, shared terms first and site terms after. The longest site list was 124 words, 137 with the shared terms in front. Whisper’s decoder context is 448 tokens, and with max_new_tokens fixed at 225, that left the prompt 223 tokens. By the rule of thumb I later wrote into the code (about 1.8 tokens per Korean word), the longest prompt came to roughly 250 tokens, and faster-whisper refuses a request like that.

Ten days later (April 10) I put the arithmetic into code: count the prompt’s tokens first, then cap output at min(225, max(50, 448 - prompt - 1)). Researching this post, I found a residual risk. As I read faster-whisper, it keeps only the last ~223 prompt tokens and silently drops the rest. Since the shared must-keep terms come first, a long site list pushes out exactly the terms every site needs.

Count tokens before the call. Every context window is a truncation point. If your code doesn’t decide what gets cut, a library will, and it won’t tell you.

Fallbacks Need Assertions

Part 3 described moving TTS onto ONNX Runtime’s TensorRT execution provider on May 15, which brought synthesis down to about 100 ms. The provider chain was TensorRT → CUDA → CPU, and every step down is a regression. I knew that, and my only response was one line in the service docs: if TensorRT fails to load, ONNX Runtime falls back without a warning, so check the startup log.

So I documented the failure mode and built nothing to catch it. About two weeks later, a routine deploy’s dependency install put the CPU onnxruntime wheel back over the GPU build. Synthesis went from ~100 ms to 2.7 s, roughly 27× slower, and every request still succeeded. A teammate tracked it down. Even the cleanup had a trap: the CPU and GPU wheels share one package directory, so removing only the CPU wheel half-removes the GPU install too, and the service crash-looped for about two minutes until the cleanup script was made idempotent.

The diagnosis and fixes were the team’s work, not mine. My lesson is about the sentence I wrote: a warning in the docs is not a guard. What I’d add on day one:

  • Assert at startup. Ask each ONNX session which providers it actually got (get_providers()), and refuse to start, or fail the health check, if TensorRT isn’t first.
  • Check latency after every deploy. A ten-sentence synthesis smoke test against a stored baseline catches a 27× regression in seconds.
  • Prefer loud to slow. A crash loop announces itself; a 27× slowdown at a 100% success rate doesn’t.

Your Eval Pipeline Fails Silently Too

In late April I ran a benchmark: several models replaying the same 407 test turns, scored afterwards by an LLM judge. GPT-5’s first run came back with 399 of 407 responses as empty strings, stored as ordinary answers rather than marked [ERROR]. The runner’s content or "" had turned “the API returned no content” into “the model said nothing”. Because conversations are replayed turn by turn, each empty answer also went into the history for the next turn. Scored as-is, GPT-5 would have looked like the worst model on the board.

The rerun came back with 405 answers and two explicit errors. The runner behind it omits temperature and any output-token cap for reasoning models, and the judge gets a 16,384-token completion budget. My guess is that reasoning tokens ate the output budget before any visible text came out, but I wasn’t logging finish_reason, so I can’t prove it.

The rest held up: judge failures are excluded, not averaged in as fake low scores, checkpoint and resume removes the temptation to “just use what finished”, and test data frozen to a JSON cache keeps runs comparable. I’d add one thing: the excluded count next to every mean, since a model whose hardest rows failed to score can look better than it is.

Count empties per model before reading any leaderboard. Count empty strings, [ERROR] rows and judge failures for each model. If any count isn’t zero, the leaderboard is still a draft.

Caches Remember Your Old Settings

This one comes from reading the code, not from an incident. The TTS service caches synthesized audio in SQLite, keyed by the SHA-256 of the text and nothing else. For an agent that repeats the same greeting all day, that’s a big win. But the key ignores everything else that shapes the audio. On March 16 I changed the default voice (F1 → F2), the speed (1.05 → 1.15) and the synthesis step count (5 → 10); on May 15 I moved from Supertonic 2 to Supertonic 3. Each change reached only sentences that weren’t cached yet.

Eviction removes the least-used entries first, so the phrases callers hear most are exactly the ones that never refresh, and a tuned voice and a stale one could alternate within a single call. Worse, the WAV header is written at read time with the current model’s sample rate, so a sample-rate change would silently pitch every cached clip wrong. I have no record of either problem being reported, and I don’t know whether anyone cleared the file by hand. Not knowing is the problem.

(The same day as the voice tuning, I fixed a cousin of this bug: the cache’s database path was relative, so starting the process from another directory silently gave you a different cache.)

What I’d do now: put the model, voice, speed, steps, sample rate and normalizer version into the key, or version the cache and bump it on every change that affects audio. A cache key is a claim that two outputs are interchangeable; make sure the claim is true.

Summary

  1. Give failure its own value. None, an empty string or a default must never double as a decision. Map uncertainty to the least destructive action, and let failure count against the system in evaluation.
  2. Assert what you assume. Provider chains, JSON contracts and context budgets degrade without a sound: check the active provider at startup, constrain structured output, count tokens before a library truncates for you.
  3. Version everything that persists. A cache key must cover every input that shapes the output, and a benchmark row must tell “no answer” apart from “a bad answer”.

What’s Next

Most of the fixes in this post were small patches: a sentinel here, a prefilter there, one more normalization step. Part 5 covers what happened when patches like these became the main way the dialog system evolved, with every failing QA row turning into a rule.