This is Part 2 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me. This part is about the text around the LLM: numbers going in and out, names that STT gets almost right, and where an LLM call is worth its latency.
At 2:17 one Friday morning in March, I fixed a bug in my TTS normalizer that made it read an amount like 27,000 won as “two-seven-zero-zero-zero won”. Half an hour later I gave up hand-writing rules for how Korean reads numbers and handed the job to the LLM. It speaks Korean better than my regexes do, after all.
The commit log from the next 34 minutes:
- 02:49: TTS normalization goes through the LLM with a short prompt.
- 02:57, 03:00: the prompt grows tables: all twelve native-Korean hours, irregular months, how to space phone digits, ‘다시’ for address hyphens.
- 03:12: a hybrid. Regex handles the patterns it already knew; the LLM handles parentheses, flow and leftover digits.
- 03:17: a new line asks the model not to change a single character of what the regex already converted.
- 03:23: “Ditch llm for tts norm.” Parentheses become a one-line regex.
Six commits, and I ended up back where I started. This post is about that half hour, and the other places I had to choose between a rule and a model — sometimes twice.
A Test for Rules vs LLM
The conclusion first, as three questions. Does the input have exactly one right output? Can you check it cheaply, with a unit test or a set lookup? Is it on the hot path? If the first two answers are yes, it’s code. The LLM earns its place only where language is genuinely ambiguous, and even there it sits behind a gate, gets handed the facts, and has a static fallback.
| Task | Output space | Checkable? | Hot path? | Choice |
|---|---|---|---|---|
| Reading digits aloud for TTS | One right reading | Yes, unit tests | Every turn | Rules |
| Spoken numerals → digits | One answer, except homographs | Mostly | Every turn | Rules + gated LLM |
| Is this road or contact name real? | Closed list | Yes, membership | Intake turns | Dictionary; LLM may propose |
| Ending the call after “감사합니다” (“thank you”) | Closed set | Yes | Closing turns | Rules |
| “Oh, then I’ll just call them myself” | Open | No | Rare | LLM, behind a regex gate |
| Re-asking for missing information | Fixed facts, open wording | Facts only | Retries | Canned first, LLM paraphrase on retry |
This is a pro-rules post. The rules that later became a ceiling (Parts 5 and 7) lived in dialog control; this table marks the line.
Output Side: 34 Minutes of LLM Text Normalization
TTS reads whatever string it gets, so ‘08:30’ or a phone number has to become spoken Korean. That’s fiddly: native-Korean hours but Sino-Korean minutes (한 시 삼십 분), irregular months (유월, 시월), money grouped by 만 and 억, phone numbers digit by digit.
The embarrassing part is that the rules already existed: the previous morning I had given the rule-based normalizer a rework and a 335-line test file. At 2:49 AM I demoted it to a fallback, and by 03:00 I was rebuilding it inside the prompt. Those two prompt commits are my rule set rewritten as prose and run through a sampler. The 03:17 plea only makes sense if the model was rewriting correct readings.
None of the reasons it failed were about intelligence:
- The output space is closed. ‘10시 30분’ has one correct reading; any variation is an error.
- The only check is the rules. If the rules validate the LLM, the rules are the normalizer.
- It runs on every turn. Every TTS call gets a generation’s worth of extra dead air, in exchange for a chance to mangle a phone number.
If your prompt contains a lookup table, it’s code. A prompt that accumulates “1 → 한, 2 → 두” tables is a function written in a language with no tests, no determinism and a per-call latency. Write the function.
Parentheses were the one piece I thought needed “understanding”. They became \(([^)]+)\) → , \1,, which reads the bracketed text as an aside between two pauses.
Rule Bugs Are Real, but Reproducible
Going back to rules didn’t make the bugs go away. My normalizer ran the specific patterns first and, last, a catch-all that spelled leftover digits one by one — right for phone numbers and little else. The year rule only matched four-digit years, so ‘10년’ (ten years) fell through to the catch-all and came out as ‘일공년’, “one-zero years”. The 2:17 AM currency bug had the same cause.
My first fix is the part I’m least proud of. When QA flagged ‘일공년’, I rewrote the knowledge-base document with the numbers spelled out, so the normalizer never saw ‘10년’. I was patching the data to hide a code bug. A month later QA flagged it again, now along with page counts (‘N면’): other client sites’ documents still had digits, and nothing stopped the LLM from writing its own. The real fix was two specific rules ahead of the catch-all. Even then I only checked that the 75 existing normalization tests still passed; I never added a test for ‘10년’.
A currency bug from the same family was worse. A teammate’s 60-case normalization matrix later found that my code read 120,000,000 won as ‘억 이천만 원’: I had dropped the ‘일’ before 억 the way Korean drops it before 만. And my own unit test asserted that 100,000,000 reads ‘억’, so it passed while being wrong.
What these bugs share is that they’re reproducible. The same string gives the same wrong output, and once fixed and tested, it stays fixed. Three days after that night, the non-LLM code had 261 unit tests, which no prompt can have. But tests only cover the cases you thought of. It took someone else’s matrix to find what mine missed.
Two Texts Per Turn
The first requirements sheet I mentioned in Part 1 already had the line: ‘119’ for display, ‘일일구’ for the voice. A phone agent needs both — a canonical form for display, logic and the LLM, and a spoken form for TTS — so the /chat response carries response and normalized_for_tts.
QA needed the same split. Reviewers read the LLM’s text, and ‘10년’ is fine as written Korean, so a bug like that exists only in what the caller hears. The afternoon after that night, I made the QA harness write a tts_normalized column next to every answer. QA that reads only what the LLM typed is reviewing the wrong artifact.
Input Side: A Gated LLM for Real Ambiguity
The other direction is messier. My fine-tuned Whisper (see the Whisper series) spells numbers out in Hangul like its training transcripts, the same convention that skewed my CER in Part 1. So a phone number arrives as ‘공일공 일이삼사’, while the LLM and the slot validators need digits.
The rules run first:
- Digit-word runs → digits (‘공일공’ → ‘010’).
- Sino-Korean compounds with 십/백/천/만/억 (‘이백 구십 오’ → 295), and mixed output merged (‘300 이십 칠’ → 327).
- ‘다시’ and ‘에’ as hyphens, and a blacklist of lookalikes: 만일 (“if”), 천사 (“angel”), 오만 (“arrogance”).
Rules run out at homographs. 공 is zero or “ball”, 일 is one or “work”, and 이 is two or the subject particle. “끝자리가 공이에요” (the last digit is zero) and “공이 날아왔어요” (a ball came flying) share a syllable and a particle.
So the LLM gets a narrow job:
- Gate. A heuristic calls it only when something is left over after the rules: a digit syllable before a particle, digit words glued to Arabic digits, or a correction like “X이 아니라” (“not X, but…”).
- Context. It sees the chat history, because “공이요” right after “what’s the last digit?” is a zero.
- Bounds. If the output is more than twice the input’s length, or the call fails, the rule result stands. The model only ever proposes.
Let the Dictionary Decide
Names are the other thing STT gets almost right. One swapped consonant or blurred vowel turns a real road or office name into one that doesn’t exist. Asking the LLM to fix it sounds natural, but the valid answers form a closed list I already have. So the LLM may propose, and the dictionary decides:
- Roads and neighbourhoods. Fifty-six minutes after deleting the LLM normalizer, I committed jamo-level fuzzy matching. It splits syllables into letters, runs Levenshtein, and takes the closest known name for that site within distance 2.
- Contacts. An office name from the LLM is snapped to the closed contact list with jamo matching; anything not on the list is dropped, not read out.
For Korean, match on jamo, not syllables. A Hangul syllable packs two or three letters, so syllable-level edit distance can’t tell a slip of the tongue from a different word.
| STT heard | Known name (invented) | Syllable edits | Jamo edits | Should match? |
|---|---|---|---|---|
| 세봄로 | 새봄로 | 1 | 1 | Yes, an ㅐ/ㅔ blur |
| 세본로 | 새봄로 | 2 | 2 | Yes, two small slips |
| 국봄로 | 새봄로 | 1 | 3 | No, a different syllable |
A syllable threshold of 1 rejects the second row and accepts the third, which is exactly backwards. A jamo threshold of 2 gets both right.
The counterexample came at the end of March, when I asked the query-rewrite LLM to also fix STT errors by sound. It “corrected” proper nouns that were already right and filled in dates nobody had said, and three and a half weeks later I had to forbid both explicitly. Without a list of valid answers, a model asked to fix errors also fixes things that aren’t broken.
The Pendulum: Regex → LLM → Hybrid
On March 18 I replaced a batch of brittle regexes with small single-purpose LLM calls — topic extraction for retrieval, “when will this be handled?” checks in two intake flows, pulling one more field out of a request — each adding an LLM call to the turns it touched. A week later, under latency pressure, one “when” check went back to a regex, and an LLM coherence check gave way to the classifier’s own confidence.
One afternoon, acknowledgement handling, which decides whether the caller is done, swung twice in ten minutes:
- 14:25: a regex split is replaced by a single LLM judge.
- 14:35: a hybrid. Clear closers (“알겠습니다”, “감사합니다”) end the call deterministically. Only reflective ones like “아 그럼 제가 그쪽에 직접 전화해 볼게요” go to the LLM, which chooses between closing and a warm wrap-up.
The final design is right: the closed set gets a rule, the open set gets the model, and a regex decides which is which. Even so, the rule needed context. A month later, a plain ‘알겠습니다’ after a short detour ended a QA call while an intake was still pending, so the suspended-flow check had to move in front of it. (The judge’s failure path hid another surprise, which comes back in Part 4.)
What I’d do differently: none of these commits records a latency or accuracy number. Every swing should come with two measurements on a fixed set: milliseconds added per turn and errors per hundred turns. Without them, the pendulum swings toward whichever failure you saw last.
Grounded Paraphrase: Canned First, LLM on Retry
Intake flows, which collect what, where and when from a caller, had the opposite problem. The canned re-asks were exact, but a caller who couldn’t give a usable location heard the same sentence on every retry, like a broken machine.
On April 6 at 13:27 I had the LLM generate every re-ask from the request category, the missing field and the caller’s last utterance. At 13:41 I reverted it. The questions were softer but had lost the canned hints, such as asking for an address or a nearby building: I had let the model decide what to say, not just how. The version that stuck came at 13:43. The first ask stays canned, and on retries the LLM gets the canned sentence itself to paraphrase.
The next commit taught the same lesson. A “polish” call for the final confirmation opened with “안녕하세요” (“hello”) mid-call, although its prompt already said not to greet. It had seen one canned sentence and nothing else, so as far as it knew, the call hadn’t started. The fix told it the call was in progress and banned ‘안녕하세요’ by name.
By April 22 this was a small generator. Facts such as the location, the action taken and where the request goes are inputs the model only phrases. Re-asks escalate with the retry count, from neutral to a format hint to “any nearby landmark is fine”. Endings must be formal-polite (합쇼체), and every function falls back to its canned string. On the 127 turns of the QA conversations that contain a re-ask (Qwen3.6-35B-A3B-FP8), 116 passed the QA sheet (91.3%), and re-asks that repeated the canned sentence verbatim fell from 22 to 8. What that pass rate is really worth is Part 6’s story. The point here is the split: the LLM phrases, the code decides.
Summary
- If it has exactly one right answer, it’s code. Reading digits aloud, parsing spoken numerals, name lookups and explicit closers belong in deterministic, tested code. When a prompt grows a lookup table, move the table into a function. Rule bugs reproduce, and a test keeps them fixed — if you write it.
- Where language is genuinely ambiguous, gate and bound the LLM. A cheap heuristic decides when it runs, the output is bounded or snapped into a closed list, and the rule result is the fallback. Measure Korean name distance in jamo, and measure latency and errors on a fixed set every time a task crosses the line.
- Let the model phrase, not decide. Hand it the facts and the canned sentence, keep the exact version for the first ask and the fallback, and QA what the caller hears.
What’s Next
Every LLM call in this post cost dead air. Part 3 is about latency: profiling each stage before shrinking anything, counting LLM calls per turn, and how a 35B model ended up answering faster than a 4B.