Too Many Guardrails Made the Bigger Model Worse: Deleting 13,528 Lines of Orchestration

This is Part 7, the last of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 showed how the dialog layer’s rules piled up, and Part 6 why the evaluation kept rewarding them. This part is about taking them out. ...

September 22, 2026 · 11 min

The Metric Chose the Architecture: How Gold-Answer QA Rewarded Our Guardrails

This is Part 6 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 covered how the dialog layer’s rules piled up one failing QA row at a time. This part is about the evaluation that kept rewarding them. ...

September 1, 2026 · 12 min

Every Failing QA Row Became a Rule: How Guardrails Pile Up in an LLM Dialog System

This is Part 5 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, Part 3 latency, and Part 4 failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time. ...

August 11, 2026 · 11 min

Nothing Crashed: Silent Failures in a Production Voice AI Stack

This is Part 4 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, and Part 3 latency. This part is about the failures that never threw an exception. When a web app breaks, someone at least sees a 500 page. When a voice agent breaks, the caller hears nothing. They say “여보세요?” into the silence, wait a few seconds, and hang up. No stack trace reaches them, and often none reaches you either. ...

July 21, 2026 · 12 min

The Bigger Model Answered Faster: Latency Lessons from a Real-Time Voice Agent

This is Part 3 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, and Part 2 where rules beat models. This part is about latency: where the seconds went, and why the fixes that mattered were almost never “use a smaller model.” ...

June 30, 2026 · 12 min

Don't Let the LLM Read Numbers Aloud: Where Rules Beat Models in a Voice Pipeline

This is Part 2 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me. This part is about the text around the LLM: numbers going in and out, names that STT gets almost right, and where an LLM call is worth its latency. ...

June 9, 2026 · 12 min

Four Ways My STT Evaluation Lied to Me: Whisper Meets Real Phone Calls

This is Part 1 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. The Whisper fine-tuning series covered how I trained the model; this one starts with what happened when it met real calls. By the end of last year my Whisper fine-tuning curve looked about as good as curves get. On the held-out split of a public Korean 8 kHz telephone corpus (AI Hub’s low-quality telephone-network speech data), every round beat the last: 9.0% CER, then 6.3%, then 5.5%, then 3.8% for a fine-tuned large-v3-turbo. Off-the-shelf large-v3-turbo sat at 6.5%. I had set myself a target of under 4%, and 3.8% cleared it. ...

May 19, 2026 · 12 min

Bias–Variance and How to Actually Diagnose Overfitting

Everyone can recite “high bias is underfitting, high variance is overfitting.” Far fewer can look at a training run and say which one they have and what to do about it. This post does both: first the bias–variance decomposition tightly enough to be useful, then a practical playbook — learning curves, the train/validation gap, cross-validation, and the traps — for diagnosing overfitting on a real model. It’s the diagnostic companion to the L1/L2 and dropout posts, which cover the fixes. ...

March 23, 2026 · 8 min

Dropout: Why It Helps Generalization, and the Train/Inference Scaling Trick

Dropout is two ideas bolted together: randomly switch off units during training, then quietly turn them all back on for inference. The interesting parts are (1) why switching units off randomly improves generalization at all, and (2) the fact that “turn them all back on” is not free — the activations come out at the wrong scale unless you correct for it. Getting the scaling wrong is one of the most common deep-learning bugs, so this post works through both, at the same engineering depth as the L1/L2 post. ...

March 16, 2026 · 9 min

L1 vs L2 Regularization: Why Does L1 Produce Sparse Solutions?

“L1 gives you sparse weights, L2 gives you small weights” What Regularization Actually Is A model with enough capacity will happily drive its training loss to zero by memorizing the data — including the noise. That is overfitting: great training numbers, bad predictions on anything new. In a linear model it shows up as coefficients that blow up to large, opposing values. Why does fitting noise require large coefficients? Because of the amplification in the OLS solution \(\hat{\mathbf{w}} = (X^\top X)^{-1}X^\top y\). If \(X^\top X\) has a small eigenvalue (a direction of almost no variance in the data — e.g. two nearly identical features), its inverse has a large eigenvalue, and noise in \(y\) projected onto that direction gets amplified into a large coefficient. Concretely, if \(x_1 \approx x_2\), explaining \(y\) needs only \(w_1=1,\,w_2=0\) — but squeezing out the last bit of noise-fit might use \(w_1=1000,\,w_2=-999\). Since \(x_1 - x_2\) is a near-zero direction, you need enormous coefficients to produce the tiny output that matches the noise. Large \(|\mathbf{w}|\) is the fingerprint of a function contorting itself to fit noise. ...

March 2, 2026 · 11 min