Too Many Guardrails Made the Bigger Model Worse: Deleting 13,528 Lines of Orchestration

This is Part 7, the last of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 showed how the dialog layer’s rules piled up, and Part 6 why the evaluation kept rewarding them. This part is about taking them out. ...

September 22, 2026 · 11 min

Every Failing QA Row Became a Rule: How Guardrails Pile Up in an LLM Dialog System

This is Part 5 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, Part 3 latency, and Part 4 failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time. ...

August 11, 2026 · 11 min