The Metric Chose the Architecture: How Gold-Answer QA Rewarded Our Guardrails

This is Part 6 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 covered how the dialog layer’s rules piled up one failing QA row at a time. This part is about the evaluation that kept rewarding them. ...

September 1, 2026 · 12 min

Four Ways My STT Evaluation Lied to Me: Whisper Meets Real Phone Calls

This is Part 1 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. The Whisper fine-tuning series covered how I trained the model; this one starts with what happened when it met real calls. By the end of last year my Whisper fine-tuning curve looked about as good as curves get. On the held-out split of a public Korean 8 kHz telephone corpus (AI Hub’s low-quality telephone-network speech data), every round beat the last: 9.0% CER, then 6.3%, then 5.5%, then 3.8% for a fine-tuned large-v3-turbo. Off-the-shelf large-v3-turbo sat at 6.5%. I had set myself a target of under 4%, and 3.8% cleared it. ...

May 19, 2026 · 12 min

Did Fine-Tuning Actually Help? Evaluating and Benchmarking Whisper for Korean STT

This is Part 3 of a three-part series on fine-tuning Whisper for Korean speech-to-text: Preprocess → Train → Evaluate. Here we measure whether the fine-tuned model actually improved, and by how much. Part 1 covered preprocessing; Part 2 covered training. A trained model without evaluation is just a checkpoint on disk. You can stare at the training loss curve and hope it went down, but until you run the model on held-out data and measure something concrete — CER, WER, per-category breakdowns — you don’t know if the fine-tuning worked, whether it regressed on certain domains, or how it compares to the baseline you started from. ...

February 19, 2026 · 8 min