This is Part 7, the last of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 showed how the dialog layer’s rules piled up, and Part 6 why the evaluation kept rewarding them. This part is about taking them out.


The model sweep that spring had a 4B model at one end and a 397B mixture-of-experts at the other. Behind the team’s long production prompt, the score mostly climbed with size, the way scaling curves are supposed to. Behind my orchestration, it barely moved. A dense 27B, a 122B and a 397B (the last two in both FP8 and Int4) all landed on the same score: not close, identical to the decimal. A second version of the orchestration scored lower across the board and was just as flat.

When a hundredfold increase in parameters changes nothing, the parameters aren’t what’s producing the answers. This post is about deleting the layer that was, and about what I had to put back.

The Flat Line

Going back over the raw outputs of the April benchmark from Part 6 (407 user turns, four open models, both architectures) for this post, I found the same pattern without a judge, just by counting how often different models said exactly the same thing.

April benchmark, 407 turns Orchestration Long prompt
Replies byte-identical across all four open models 253 (62%) 78 (19%)
Replies byte-identical between Qwen3-4B and Qwen3.5-35B-A3B 293 (72%) 89 (22%)
Reference-free score, 4B → 35B (1–5 scale) 3.70 → 3.86 3.14 → 4.10

Under orchestration, a 4B and a 35B MoE produced the same bytes in nearly three turns out of four. That isn’t two models agreeing; it’s the classifier sending both down the same branch and a template doing the talking. On the judge that never sees the answer key, going from 4B to 35B bought the long prompt almost a full point and my orchestration a sixth of one.

The same day, runs with Part 6’s reactive simulated callers and the 397B behind the orchestrator showed what that sounds like:

  • Told that their question would go to the relevant department, a caller asked whether they could just call it themselves. The bot said goodbye and ended the call.
  • After a referral to 119, a caller asked for another number to try. Goodbye again.
  • After confirming a report as filed, the bot offered to resume the caller’s earlier report and, when the caller said yes, asked for the address again.

None of those decisions was the 397B’s. The graph had classified a question as a goodbye, and a location check had judged the address too vague. The largest model I had was reading lines from my state machine.

Scaffolding is a bias knob. In bias–variance terms, every hard-coded rule narrows what the system can say. For a weak model, that removes more error than it adds. For a strong one it’s mostly bias, and a bigger model can’t reduce bias that lives outside the model.

Every Guardrail Encodes a Model Weakness

On March 18 I added a helper that flattened pipe tables in retrieved documents into bullet lists. Its docstring says why: the 4B typically read one row of a table and ignored the shared headers. On March 24, in the commit that moved us to the 35B MoE, I deleted it with a one-line justification: the new model handles tables natively.

That was the only deletion in the migration that cited the new model. The rest of the 4B-era machinery (regex fast paths, the intent classifier, slot-filling nodes, templates) existed because in December I’d decided a model that size wouldn’t reliably follow one long prompt (Part 5). All of it stayed for five more weeks, and I never asked whether the 35B still needed it.

Part 6’s reference-free scores answer that. Orchestration minus long prompt: +0.56 for Qwen3-4B, +0.18 for Gemma-4-26B-A4B, 0.00 for Qwen3.5-27B, −0.24 for Qwen3.5-35B-A3B. The scaffolding’s value tracked the weakness it covered and crossed zero right around the model we had already moved to. The clean-slate commit the day after the benchmark named the target: remove all small-model scaffolding built around Qwen3-4B.

Tag every rule with the weakness it covers. “The 4B reads one row of a pipe table” is a testable claim: when the model changes, re-run the test and delete the rule if it passes. Untagged, every rule looks load-bearing.

Deleting the Graph, Keeping the API

I did it twice. The day after the benchmark, a clean-slate experiment on a branch stripped the service down to retrieval plus one LLM call: −9,566 lines, with chat history deliberately kept. On May 12 I made the real cut, also on a branch: 40-plus graph nodes, the intent classifier, slot filling and the per-site overrides, 13,528 lines including tests. The commit message is blunt: the graph was too restrictive, and plain LLM with raw RAG had outperformed it in evaluation.

What I kept on purpose was the API. session_id was still accepted, DELETE /session/{id} still returned 200, and the schema didn’t change; intent just came back null. Telephony and the QA harness kept calling it as before, and a rollback would have been a deploy, not a migration.

Service code Peak (Apr 27) Graph removed (May 12) Final branch (May 13)
Python lines (core + API) 16,405 3,707 4,020
re.compile calls 126 35 46
Largest per-site config 629 lines 12 lines 12 lines

That’s −75% from the peak. The 46 remaining regexes live in text normalization, location checks and the guards below; apart from the end-of-call gate, none decides where a conversation goes. That’s where Part 2 argued rules belong: places with exactly one right answer.

Tenants became data: a per-site config that once held keyword regexes, intake definitions and redirect messages shrank to four fields (key, names, document folders, closing sentence), with site-specific wording in a prompt file. I also condensed the team’s long production prompt, plus my own additions from the day before, from 430 lines to 96. Part 5 counted “important” markers; here the word was “absolute,” and lines containing it went from 22 to 3.

What the Deleted Layer Was Doing for Free

The graph-removal commit describes the new /chat as stateless, and I wrote that as if it were a feature. The orchestrator had held each session’s conversation context. Without it, the model saw only the current utterance and the retrieved documents, so every turn was a first turn. The clean-slate branch had kept history on purpose; the real cut dropped it.

Less than two hours later, the guard commit added a buffer of the bot’s last three replies, built for the repetition guard and also passed to the model as history; its own comment calls it “approximate.” The next morning the window went to 20. Going back over the code for this post, I found that the caller’s earlier lines never came back. The model could see everything it had said and nothing the caller had.

Memory wasn’t the only duty nobody had written down:

Implicit duty Owner in the graph Owner after the cut
Conversation memory Session context The bot’s last 3, then 20, replies; caller turns dropped
Context for vague follow-ups in retrieval History-aware retrieval node None; retrieval used the bare utterance
Refusing a request for a private person’s number Privacy check in the contact node Nothing I could find in the code or either prompt

The privacy row bothers me most. “I can’t give out a neighbor’s number” is behavior, not data, and after the cut I can’t find anything that owned it.

Before deleting a layer, inventory what it does that no interface names. Give each duty an owner, a test, or an explicit “we accept losing this.” I didn’t make that list before the cut; the table above is me making it after the fact.

Three Guards, Each With a Receipt

The guard commit describes itself in one line: no state machine, no intent classifier, only the hooks the benchmark data actually justifies. Three hooks, 19 unit tests:

Guard Failure it targets Context gate Action Tests
End-of-call short-circuit Re-summarizing before goodbye; rewording a closing line that must be exact Previous bot turn ended with “anything else?” and the reply is a short closer (≤ 30 chars, fixed list) Skip the LLM; return the fixed closing sentence 5
Repetition guard Repeating the same fallback (non-repetition 4.0/10 vs GPT-5’s 9.38 in a later benchmark round) Reply body ≥ 0.7 similar to an earlier reply in the session One reroll with an anti-repetition directive, then a fixed graceful exit 6 + 2
Form-character stripping Markdown, parentheses, slashes and arrows reaching TTS Always, before TTS Strip them, keep the content 6

The gates are the point. The old termination logic had ended a call on a bare “okay” while a report was still suspended (Part 5). This one fires only when “okay” answers “anything else?”, and a test checks that it stays quiet otherwise. Each guard has a measured failure behind it, a gate narrower than “the caller said X,” a deterministic fallback and tests. None of them routes; they act at the edges of a turn (the closer, the repeat, the formatting) and leave the middle to the model.

Where the Facts Went Instead

Without rules, every fact has to come from a document written for a reranker and for the ear. Most of that work came before the teardown, and it’s what made the teardown possible:

  • Self-contained items (March 18). Tables with shared headers became lists where each item repeats its own hours and services, so any chunk answers alone.
  • Whole documents (March 19). Short documents stayed whole instead of being cut into 450-token windows, and retrieval returned 10 candidates, not 3.
  • Spoken style (March 23). Written the way an agent would say it. Spelling out numbers that TTS had misread was the exception: as Part 2 explains, that hid a code bug.
  • Top 3 after reranking (April 9). One document failed questions that touched two topics.

Fine-tuning was the road not taken for facts; my original plan already argued for retrieval (less hallucination, easier QC, no retraining when a fact changes). I did try it for style and sub-tasks: 4-bit QLoRA of the 35B on a single H100, not the full fine-tune of my Whisper series, on a style-only set with phone numbers, departments and addresses replaced by placeholders, plus task-aligned sets. It reached an eval loss of 0.27 and 93.6% token accuracy. The toolkit around it grew ten different training scripts; settle the stack before you write the second one.

Now the self-critique. In that run’s training file, 487 of 835 examples were the orchestrator’s own sub-tasks: intent JSON and location JSON. I was training the model to be a better component of the graph, and I committed the toolkit 96 minutes after deleting the graph. Token accuracy measures fit to your formats, not whether a caller is better served; that takes a blind comparison like the one in Part 6. The evidence said to remove the scaffolding, and I was fitting the model to it.

What Actually Happened, and What I’d Do Differently

The teardown was validated on a branch and handed over. The next sprint’s plan paired “simplify the prompt” with “implement orchestration,” with the goal of passing QA without excessive guardrails. What was delivered was a compromise: a much lighter flow and a simplified prompt. That’s a reasonable landing, because the other side had real advantages:

  • Templates are fast. In the April benchmark the orchestrated 35B replied in a median 161 ms, against 289 ms behind the long prompt.
  • Templates are predictable. The business team could read exactly what the bot would say. A prompt hands them a distribution.
  • The 4B really needed it. The +0.56 was real. The scaffolding was right for the model it was built for; carrying it over to the next one was the mistake.

What I’d do differently is hygiene, starting from the first rule:

  1. Tag every rule with its model weakness and the QA row behind it.
  2. Re-run scaffolded vs. minimal after every model swap, on a metric that doesn’t reward the scaffolding.
  3. Cap the rule budget. A new rule retires or merges an old one.
  4. Add back only narrow, gated, tested guards that a measured failure justifies.

The change that stuck was my default. I used to add a rule whenever a row failed. Now a rule has to justify itself, and the bar rises every time the model underneath improves.

Closing the Loop

Seven parts, one pipeline, and mostly one mistake in different costumes: trusting a number or a rule more than what it stood for.

  1. STT evaluation: measure in the channel you ship, and QA the eval set; my references started as my model’s output.
  2. Rules vs. models: if it has exactly one right answer, it’s code. Let the model phrase, not decide.
  3. Latency: profile per stage; speed is active parameters × the tokens you choose to say.
  4. Silent failures: give failure its own value, and assert what you assume.
  5. Guardrails: rules accrete one failing row at a time; review the shape, not the patch.
  6. Metrics: a gold set in one system’s voice crowns that system; keep a judge that never sees the answer key.
  7. Simplification: scaffolding built for a weak model caps a strong one. Re-justify every rule after a model swap, and keep facts in retrieval.

Measure the failure, justify the rule, delete the rest.