Too Many Guardrails Made the Bigger Model Worse: Deleting 13,528 Lines of Orchestration

This is Part 7, the last of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 showed how the dialog layer’s rules piled up, and Part 6 why the evaluation kept rewarding them. This part is about taking them out. ...

September 22, 2026 · 11 min

The Metric Chose the Architecture: How Gold-Answer QA Rewarded Our Guardrails

This is Part 6 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 covered how the dialog layer’s rules piled up one failing QA row at a time. This part is about the evaluation that kept rewarding them. ...

September 1, 2026 · 12 min

Every Failing QA Row Became a Rule: How Guardrails Pile Up in an LLM Dialog System

This is Part 5 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, Part 3 latency, and Part 4 failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time. ...

August 11, 2026 · 11 min

Nothing Crashed: Silent Failures in a Production Voice AI Stack

This is Part 4 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, and Part 3 latency. This part is about the failures that never threw an exception. When a web app breaks, someone at least sees a 500 page. When a voice agent breaks, the caller hears nothing. They say “여보세요?” into the silence, wait a few seconds, and hang up. No stack trace reaches them, and often none reaches you either. ...

July 21, 2026 · 12 min

Don't Let the LLM Read Numbers Aloud: Where Rules Beat Models in a Voice Pipeline

This is Part 2 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me. This part is about the text around the LLM: numbers going in and out, names that STT gets almost right, and where an LLM call is worth its latency. ...

June 9, 2026 · 12 min