This is Part 5 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, Part 3 latency, and Part 4 failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time.


On March 25 I made 32 commits to the dialog service. The first landed at 00:56, the last at 16:59, and 26 of the titles start with “Fix”. At 11:00 I replaced one request category’s opening line — three sentences of caveats before the actual question — with one short, natural question. At 11:07 I reverted it, back to “the original detailed greeting for QA compatibility,” as the commit message puts it. The short line sounded better on the phone. It lost because the answer key had been written against the old one.

Nobody designed the system that was running by the end of April. It accreted: a regex here, a node there, one more prompt line marked important, each justified by a failing test row. This post takes that accretion apart, along with the warning signs that sat in my commit log for weeks before I read them.

The Decision That Made Sense in December

The rules didn’t start as a mistake. In November I tried LoRA fine-tuning. By mid-December I’d put it on hold and hard-coded the dialog logic as an orchestration around the small model we could serve at the time. My read was that a model that size couldn’t reliably follow one long prompt full of per-category procedures, so I moved the decisions into code. It was done on December 19, and every happy-path case passed.

In late January I refactored it into a graph. Each decision step became a class that behaved like a node, so any client site’s request flows could be assembled from parts: regex fast paths → an LLM intent classifier → slot filling → per-site intake nodes, plus RAG nodes for information questions. “Intake” meant collecting what happened, where and when, then confirming and submitting.

For a small model this was the right call, and I’d make it again. The mistake came later and was quieter: I never built anything to pay rules down. My sprint plan for two weeks in January lists exactly one KPI: “pass QA.”

Anatomy of Accretion

The service moved into its own repository on March 9, running Qwen3-4B. Here it is at three points:

Snapshot Python lines (core + API) re.compile calls Request-routing patterns Answer-prompt rules (chars) “Important” markers
Split (Mar 9) 5,888 12 11 6 (348) 0
First version tag (Mar 22) 12,783 82 63 9 (620) 2
Peak (Apr 27) 16,405 126 121 25 (2,572) 10

In seven weeks the code nearly tripled, compiled regexes grew tenfold, and the keyword patterns that routed a sentence into a flow went from 11 to 121. On March 24 the model underneath changed from Qwen3-4B to a 35B-A3B mixture-of-experts Qwen, and the curve didn’t bend. Some rules had probably stopped being necessary, but nothing in my process checked.

Much of the growth was per-site logic written as code:

  • One site got 9 hand-written intake node classes in a single 1,136-line file on March 12. A week later I wrote a generic, data-driven intake node — the right instinct — but the per-site files kept growing next to it.
  • Another site’s config was 113 lines when I first committed flows to it at 10:46 on March 24, 514 lines at 12:57, and 618 by the next afternoon. At the peak it held nearly 60 keyword regexes, dozens of intake configs and over a dozen hard-coded redirect messages. One commit from that first day says, accurately, that it refined patterns “for QA pass rate.”

Each new site brought its own scripted conversations, and each one that failed became a pattern, a node or a config field.

The Loop That Fed It

The QA harness replayed scripted multi-turn conversations against the live service and compared each turn with a reference answer (Part 6 is about that metric). My loop was run → read a failing row’s trace → patch → rerun. Here is what it does to a single constant.

At 09:12 on March 25 I added a first-turn rule: if the caller’s opening sentence is longer than 15 characters, treat it as the description and don’t ask for one. At 09:17 a short opener triggered a premature submit, so I raised the threshold to 20. At 09:25 another category’s row needed the old value, so I lowered it back to 15, in a commit named after that category. Eight minutes, two commits, and the number was back where it started. It encoded nothing about callers and everything about two test rows. Five minutes later I deleted the rule outright.

Then I made the loop faster. vLLM took about five minutes to restart, so on March 30 I wrote a 318-line hot-reloader that swapped prompts and constants in place. Three days later I deleted it inside an unrelated fix. The reloader wasn’t the problem; it just shows I was optimizing how fast the loop turned, not what it produced.

Tooling sped it up further. Of the 32 commits on March 25, 27 carry an AI coding assistant’s tag. Paste the failing turn and its trace, and a plausible patch appears within a minute. That made one more special case almost free to write. It did nothing to the cost of owning one.

Review the shape of the system, not just the patch. Every diff from that period reads fine on its own: small, commented, tied to a real failure. What I never reviewed was the weekly trend in lines, regexes, prompt rules and “important” markers. A chart of those four numbers would have shown the problem a month before I saw it.

When Fixes Start Undoing Fixes

The clearest case is mid-flow interruption: halfway through an intake, the caller says something that sounds like a new request. Suspend the intake or not?

  1. Mar 20. Any information question suspended the intake. That broke callers who were simply answering the bot’s own question, so I put an LLM relevance check in front.
  2. Mar 26, 08:21. A time-only answer, “오늘 오후 세 시쯤이요” (“around three this afternoon”), was classified as a new intent and derailed the flow. I added that node to a non-interruptible list.
  3. Mar 26, 16:39. An ID number given as an answer, and nothing else, looked like a new query. I made every specialized intake node non-interruptible until it had submitted.
  4. Apr 23. In testing, callers who asked a side question or tried to cancel mid-intake were now silently stuck. I removed the blanket guard, trusted the classifier’s continue/new/terminate decision, and added one explicit fast path for cancel phrases.

Step 3 defeated the purpose of step 1, and four weeks later step 4 undid step 3. Each step was a rational response to a real row. Together they were whack-a-mole.

The same pattern showed up elsewhere:

  • Termination vs. resume. After a short detour, a plain “알겠습니다” (“okay”) ended the call while an intake was still suspended, because the deterministic closer rule from Part 2 ran before the resume check. Each rule was right on its own; the bug lived in their order.
  • Order as hidden logic. One category’s keywords had to be checked before another’s, or the broader pattern swallowed them. On one site, checking global patterns first sent at least 17 failing rows down the wrong path. When reordering a dictionary fixes 17 rows, the order is the program, and I hadn’t written it down anywhere.
  • One phrase, twelve owners. The closing “anything else I can help with?” was appended independently in at least 12 files, sometimes twice in one reply, until I centralized it on April 9.

Importance Inflation: When Every Rule Is “Important”

The RAG answer prompt went from 6 rules and 348 characters at the split to 25 rules and 2,572 characters at the peak, 10 of them flagged “important” or “very important.” The classification prompt roughly tripled, from 861 to 2,667 characters.

Read together, the rules had started to argue:

  • Brevity vs. completeness. A rule to answer only what was asked, plus a two-to-three-sentence cap, sat a few lines from a rule to cover every sub-item when a question is broad.
  • Three phone-number rules. March 26: give contact details only when asked. April 23: give the number along with the department. April 24, the next morning: answers now led with a phone number instead of the content, so a third rule said content first, number after.
  • One-row rules. Several rules read like answers to a single scripted question: a value that held only on one date, a place with two names, internal details to leave out.

On April 6 at 10:19 I removed the sentence cap and the “no filler” rule and added guidance on empathy, to make answers sound less curt. At 10:56 I restored the cap, banned opening greetings, and added a regex that stripped greetings after generation. Thirty-seven minutes, from one position to its opposite, plus a post-processor.

Meta-talk shows where this road ends. The model kept saying some version of “the provided documents don’t include that,” which means nothing to a caller. A prompt ban didn’t hold, so a regex non-answer detector went in behind it, falling back to a referral line. Weeks later the detector needed another pattern.

If everything is important, nothing is. The “important” count is a cheap, honest measure of prompt health, because each marker admits that an earlier rule didn’t hold. When a prompt rule needs a regex behind it to be obeyed, the prompt has stopped being a specification and become a list of grievances.

Redrawing the Same Line

Retrieval went around the same loop, over one question: when should conversation history go into the retrieval query?

  • Mar 18, 11:10. I added history-augmented retrieval: prepend recent turns to the query so a vague follow-up still finds the right document. 11:15: reverted. It came back on April 1, gated to turns the classifier marked as same-topic continuations (flow_status == "CONTINUE").
  • Apr 24, 08:46. A QA-round fix widened the gate: any short follow-up (under 30 characters) pulled in recent history, even when classified as a new topic. In one scripted conversation the caller filed a report, then asked when an unrelated public facility was open. The query dragged the report along, and the reranker scored the report’s document 0.98 and the right one 0.00. 09:13: narrowed again, to one turn of history, and none when the question names its own subject.
  • Apr 27, 11:01. A TopicState tracker to keep topics from bleeding into each other: 858 lines across 12 files in one commit. Twenty-three minutes later the codebase hit its peak.

Five minutes, eight, 27, 37: the gaps between a fix and its correction, which at the time I read as agility. Each retrieval fix moved the line between “use history” and “don’t,” then needed more state to defend it. I never asked whether that line had to be drawn by hand at all.

The next day I started the clean-slate experiment that Part 7 is about.

Warning Signs I Should Have Read Earlier

None of these needed hindsight. Every one was in the commit log at the time:

Warning sign Where it showed up What it meant
A constant tuned to one row 15 → 20 → 15 in eight minutes Fitting the test set, not describing callers
A revert “for compatibility” with the answer key The 11:07 greeting rollback The test set had become the spec
Order-dependent rules 17 rows fixed by reordering a dictionary Undocumented control flow
Fix → correction within the hour Gaps of 5, 8, 27 and 37 minutes Once is normal; as a habit, I wasn’t thinking past the triggering row
A rising “important” count 0 → 10 in seven weeks Earlier rules aren’t holding
Guards that reference other guards Non-interruptible lists, skip-the-interrupt flags, resume-before-terminate ordering The rules, not the model, had become the system
Tenant logic in code 9 intake classes in one site’s file Code grows with clients × test rows

The compatibility revert is the LLM-product version of tuning on your test set: the score goes up, and generalization pays for it.

What I’d Do Differently

Most of it is bookkeeping:

  1. Give every rule a receipt. Tag it with the QA row that motivated it, the model weakness it compensates for, and a date by which it has to be re-justified.
  2. Split the scripted conversations into a set you may patch against and a holdout you only score.
  3. Chart the four shape numbers next to the pass rate, every week.
  4. Name the class of failure before writing the fix. If the honest answer is “this row,” don’t write it.

Summary

  1. Nobody designs guardrails; they accrete. From a defensible start, the dialog layer went from 5,888 to 16,405 lines, 12 to 126 compiled regexes and 0 to 10 “important” markers in seven weeks, one failing row at a time.
  2. A fast loop makes special cases nearly free, and that’s the danger. Scripted QA, a hot-reloader and an AI coding assistant made each patch cheap to write. None of them made a patch cheap to own, and fixes started undoing fixes.
  3. Review the shape, not the patch. One-row thresholds, compatibility reverts, order-dependent rules and a rising “important” count were all in the log. Track them like the pass rate.

What’s Next

Every patch in this post made a QA number go up. Part 6 asks why that number was so easy to please: a metric that compares each turn with a gold answer written in the system’s own voice rewards exactly this kind of hard-coding, and a judge tuned alongside the system drifts with it.