Every Failing QA Row Became a Rule: How Guardrails Pile Up in an LLM Dialog System

This is Part 5 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, Part 3 latency, and Part 4 failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time. ...

August 11, 2026 · 11 min

Nothing Crashed: Silent Failures in a Production Voice AI Stack

This is Part 4 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, Part 2 where rules beat models, and Part 3 latency. This part is about the failures that never threw an exception. When a web app breaks, someone at least sees a 500 page. When a voice agent breaks, the caller hears nothing. They say “여보세요?” into the silence, wait a few seconds, and hang up. No stack trace reaches them, and often none reaches you either. ...

July 21, 2026 · 12 min

The Bigger Model Answered Faster: Latency Lessons from a Real-Time Voice Agent

This is Part 3 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. Part 1 covered how my STT evaluation misled me, and Part 2 where rules beat models. This part is about latency: where the seconds went, and why the fixes that mattered were almost never “use a smaller model.” ...

June 30, 2026 · 12 min

Did Fine-Tuning Actually Help? Evaluating and Benchmarking Whisper for Korean STT

This is Part 3 of a three-part series on fine-tuning Whisper for Korean speech-to-text: Preprocess → Train → Evaluate. Here we measure whether the fine-tuned model actually improved, and by how much. Part 1 covered preprocessing; Part 2 covered training. A trained model without evaluation is just a checkpoint on disk. You can stare at the training loss curve and hope it went down, but until you run the model on held-out data and measure something concrete — CER, WER, per-category breakdowns — you don’t know if the fine-tuning worked, whether it regressed on certain domains, or how it compares to the baseline you started from. ...

February 19, 2026 · 8 min

Training Whisper on Precomputed Features: Full Fine-Tune for Korean STT

This is Part 2 of a three-part series on fine-tuning Whisper for Korean speech-to-text: Preprocess → Train → Evaluate. Here we load the preprocessed dataset and run the training loop. Part 1 covered preprocessing; Part 3 will cover evaluation and benchmarking. With precomputed mel spectrograms and tokenized labels on disk, the next step is to plug them into a training loop and optimize the model. That sounds straightforward until you start making choices: full fine-tuning or LoRA? What learning rate and batch size? How do you pad variable-length sequences correctly for an encoder-decoder, and how do you avoid wasting GPU memory or blowing up training? This post walks through the training setup I use for Whisper large-v3 on Korean telephonic audio — and the engineering trade-offs behind each decision. ...

February 15, 2026 · 7 min

Stop Recomputing Mel Spectrograms: Preprocessing Your Data Before Whisper Fine-Tuning

This is Part 1 of a three-part series on fine-tuning Whisper for Korean speech-to-text: Preprocess → Train → Evaluate. In this post, we build the data preprocessing pipeline. Parts 2 and 3 will cover the training loop and evaluation/benchmarking, respectively. When I first started fine-tuning OpenAI’s Whisper for Korean speech-to-text, I noticed something frustrating. Every single time I kicked off a training run — whether I was tweaking the learning rate, adjusting the batch size, or experimenting with a new scheduler — the framework would spend hours churning through raw audio files before a single gradient was computed. The preprocessing step was identical each time: load WAV files, resample, compute mel spectrograms, tokenize transcriptions. Nothing about the data had changed, yet I was paying the full cost of data preparation on every attempt. ...

February 11, 2026 · 7 min