This is Part 1 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. The Whisper fine-tuning series covered how I trained the model; this one starts with what happened when it met real calls.
By the end of last year my Whisper fine-tuning curve looked about as good as curves get. On the held-out split of a public Korean 8 kHz telephone corpus (AI Hub’s low-quality telephone-network speech data), every round beat the last: 9.0% CER, then 6.3%, then 5.5%, then 3.8% for a fine-tuned large-v3-turbo. Off-the-shelf large-v3-turbo sat at 6.5%. I had set myself a target of under 4%, and 3.8% cleared it.
Then in January I built a 400-utterance test set from real calls to the line and ran every model on it:
| Model | Public 8 kHz test | Real calls (400 utt.) |
|---|---|---|
| Stock large-v3-turbo | 6.5% | 13.1% |
| Medium v1 (LoRA, ~0.9M utt., one domain) | 9.0% | 10.2% |
| Medium v2 (full FT, 1.7M utt., all domains, cleaned) | 6.3% | 10.8% |
| Fine-tuned large-v3-turbo (2.0M utt.) | 3.8% | 1.3% |
Three of the four models got worse — stock Whisper twice as bad — and their ranking reshuffled. The fourth got better on messy real calls than on the clean benchmark it was tuned for, by nearly a factor of three. I reported that 1.3% with a caveat (“not enough data; retest at 2,000+ utterances”) and moved on to the LLM side of the pipeline. Going back over the data for this post, I found that caveat far too gentle. The table misleads in four separate ways, and none of them is a bug in the CER code.
The evaluation post covered the metric mechanics. This one is about the eval set.
Measure in the Channel You Ship
First, the one thing I got right early. Phone audio arrives at 8 kHz, and most of what Whisper was trained on doesn’t. My first benchmark ran stock Whisper sizes on a 16 kHz Korean test set and on the 8 kHz telephone test set:
| Model | 16 kHz test set | 8 kHz telephone test set |
|---|---|---|
| tiny | 44.0% | 95.3% |
| base | 29.2% | 53.1% |
| small | 11.3% | 29.3% |
| medium | — | 29.2% |
| large-v2 | — | 25.1% |
| large-v3-turbo | — | 6.5% |
The two columns come from different test sets, so this is a rough comparison, not a controlled ablation. Still, tiny going from 44% to 95% is hard to blame on anything but the channel. Telephone bandwidth knocked out most of the model zoo before I trained anything, which set the plan: fine-tune on telephone-band data, and run the biggest model the latency budget allows.
The serving side follows the same rule. Every input path into the Whisper service — call audio, a WAV upload, a browser microphone recording — goes through ffmpeg to 8 kHz mono PCM16 first, and only then gets resampled to the 16 kHz Whisper expects. To the model, a 48 kHz demo recording sounds like a phone call, so offline numbers stay comparable to live ones. Degrade every path the same way, or your demo and your eval measure a product you don’t ship.
The channel was right. The problems were in the data I measured with.
Lie #1: Domain Shift
The benchmark said v2 — full fine-tune, nearly twice the data, cleaned transcripts — beat v1 by 2.7 points. On real calls, v1 came out ahead, 10.2% vs 10.8%. Stock Whisper, essentially tied with v2 on the benchmark, trailed both fine-tunes by 2–3 points. All three models changed rank.
The public corpus is telephone speech, but it isn’t these calls: people thinking out loud, half-sentences, local place names the corpus has rarely heard. This is the distribution-shift trap from the bias–variance post: a held-out split from your training distribution tells you how well you fit that distribution. It is a training signal, not a production estimate.
The question I didn’t ask at the time was whether the real-call numbers could rank those models at all. For this post I recomputed per-utterance CER (spaces and punctuation stripped, so the figures differ slightly from the originals) and bootstrapped the 400 utterances 2,000 times:
| Comparison | Point estimate | 95% CI |
|---|---|---|
| v2 − v1 | +0.1 pts | [−1.3, +1.6] |
| Stock − v2 | +3.0 pts | [+1.1, +5.0] |
The whole set holds only about 6,000 reference characters. The fine-tunes really were better than stock on real calls. But v1 vs v2 is a coin flip — v2 came out worse in 55% of resamples — so the v1/v2 reversal was noise.
400 utterances catch regressions; they can’t rank close models. With a paired-difference CI of about ±1.5 points, any smaller gap disappears. The interval narrows only as 1/√n, so even my 2,000-utterance retest target would resolve only about two-thirds of a point — and that assumes independent utterances. Utterances from one call share a speaker and a line, so resample whole calls; correlated errors can widen the interval.
Lie #2: My References Were My Model’s Output
Now the 1.3%. I built the 400 references the fast way: each was drafted from the newest fine-tune’s output, then corrected by hand. That same model was one of the four being scored.
I knew it was a shortcut. I didn’t appreciate how big a shortcut. In the re-audit for this post, I compared each reference with each model’s output after stripping spaces and punctuation:
| Model | Reference identical to output |
|---|---|
| Newest fine-tune (drafted the labels) | 349 / 400 (87%) |
| Stock large-v3-turbo | 162 / 400 (41%) |
| Medium v1 | 175 / 400 (44%) |
| Medium v2 | 184 / 400 (46%) |
In 105 utterances, the reference matches the newest model and none of the other three. Some of those are genuine wins. But in at least two, a place name in the reference carries the exact one-syllable misspelling the newest model produced, while stock Whisper had it right. A reviewer facing a fluent, plausible draft fixes what’s obviously wrong and leaves the rest. Every error they miss becomes a “correct” answer for the drafting model and a mistake for every model that got it right.
How much did that inflate the score? I can only estimate. On the benchmark, the newest model made roughly 40–60% fewer errors than the other three. Scale their real-call CERs by that advantage and you land somewhere around 4–8%: plausibly still the best of the four, nowhere near 1.3%. The honest answer is that this set cannot measure that model.
What I’d do instead:
- Never pre-fill labels with a model you will score. Transcribe from scratch, or pre-fill with a model outside the comparison.
- If you need drafts for speed, show several hypotheses blind. Shuffle them, hide which model wrote which, and let the annotator pick or edit.
- Run the smell test. Compute the reference-equals-hypothesis rate per model. One model at 87% while its peers sit at 41–46% isn’t a win; it’s a labeling artifact.
Lie #3: You Are Partly Scoring Formatting
The public corpus writes numbers the way they’re spoken, in Hangul: “오후 세 시 사십 분”, not “오후 3시 40분”. The fine-tuned models learned that convention. Stock Whisper prefers digits, and writes emergency numbers like “112” rather than “일일이”. My references — drafted, remember, by a model fine-tuned on that corpus — contain zero digits. Stock Whisper put digits in 34 of its 400 outputs, and every one of those was scored as character errors, even where it heard the caller perfectly.
Here is the recomputed CER with the format-dependent rows taken out:
| Rows included | Stock | Medium v1 | Medium v2 |
|---|---|---|---|
| All 400 | 13.8% | 10.6% | 10.7% |
| Drop 34 rows where any model wrote digits | 11.8% | 10.0% | 9.6% |
| Also drop rows with Latin letters (356 left) | 11.1% | 9.6% | 9.4% |
The 44 digit-or-Latin rows hold about 30% of stock Whisper’s character errors: roughly 2–3 points of “the baseline is bad” was transcription style, not recognition. The v1/v2 order flips again between the first two rows. Spacing adds a smaller layer: for stock Whisper and each medium fine-tune, 12–14 rows differed from the reference only in where the spaces went, which a scorer that keeps spaces counts as errors.
The fix is boring and should come first: write a normalization spec — numbers, spacing, punctuation, fillers, Latin-script terms — and apply it to references and hypotheses before computing CER. Otherwise your leaderboard partly ranks which model shares your annotators’ style.
The irony: my very first requirements sheet already had a line about keeping “119” for display apart from the spoken “일일구”. I just hadn’t applied it to my own scoring.
Lie #4: Staged Traffic
The 400 utterances span several months of calls, unevenly. In the early slice (the first two months, 302 utterances), three callers produced 57% of the utterances, and their calls read like scenarios: complete sentences, tidy times and places, the same few situations recited clearly. They look like internal test and demo calls. The later slice (the remaining months, 98 utterances) sounds like real callers — fragments, restarts, half-remembered names.
| Model | Early (302 utt.) | Later (98 utt.) | Ratio |
|---|---|---|---|
| Stock large-v3-turbo | 12.1% | 19.0% | 1.6× |
| Medium v1 | 9.3% | 14.8% | 1.6× |
| Medium v2 | 9.1% | 16.1% | 1.8× |
| Newest fine-tune | 1.1% | 2.1% | 1.8× |
Every model was 1.5–1.8× worse on the later slice. With only 98 utterances the later numbers carry wide error bars, but the direction holds for all four models. Three-quarters of my “real-call” set came from a period dominated by the friendliest audio the line will ever receive.
So tag every utterance with its source and period, and report them separately. Early traffic is disproportionately the team testing its own product. If you can only report one number, take it from the traffic that looks least like a demo.
Not the Model’s Fault: Where in the Utterance Errors Fall
Overall CER hides where the errors are. In the same re-audit, I compared the first two and the last two characters of each output with the reference, skipping utterances that open with a filler like “어” or “그”. For stock Whisper and both medium fine-tunes, the first two characters were wrong in roughly a third of utterances and the last two in roughly one in eight. The head was two and a half to three times as error-prone as the tail.
Head errors run in both directions, and the models lean different ways:
- Stock Whisper drops the front. In 36 utterances its output is exactly the reference with the beginning missing — a pattern like “OO초등학교 앞인데요” coming out as “초등학교 앞인데요”.
- The fine-tunes add to it. In 32–38 utterances each, their output is the reference with something extra in front: words bleeding in from the previous segment, or an outright guess.
The fine-tunes also picked up a habit stock Whisper almost never showed. Each medium fine-tune wrote a real Korean city name absent from the reference in 10–12 utterances, mostly within the first few characters; stock Whisper essentially never did. Usually the city replaced a similar-sounding local name; a few times it came out of nothing (“버스 정류장 앞인데요” → “OO시 버스 정류장 앞인데요”). My reading: a corpus full of place-name text gave the decoder a prior toward the names it had seen most, and when the audio starts ragged, a familiar city is a likely first word. So the next training round I planned would add production audio and recorded place names, to pull those priors toward places callers actually mention.
The ragged start most likely had a cause upstream. This spring I found, and shared with the team, that the front of caller audio was being clipped before it reached STT. The service runs Whisper with vad_filter=False because an upstream gateway handles segmentation and endpointing, so the model only sees what that boundary hands it. No fine-tuning or decoding setting recovers a syllable that was never sent. The fix belongs in the audio front end — pre-roll buffering and endpointing — not in the STT model.
Utterance boundaries are part of the model’s input. If errors pile up at the start, suspect segmentation before you blame the acoustic model, and fix the boundary before you spend another training run on it.
What I’d Do From Day One
| Check | Why | What I saw |
|---|---|---|
| Build an in-channel, in-domain test set before the second training run | Benchmarks are a training signal | All three comparable models changed rank |
| Label independently of every model you score | Reviewers anchor on the draft | 349/400 references matched the drafting model vs 162–184 |
| Write a normalization spec; apply it to both sides | Otherwise you score format | Digit/Latin rows held ~30% of stock errors |
| Put a paired, bootstrapped CI on every comparison | Small sets can’t rank close models | v2 − v1: [−1.3, +1.6] points |
| Tag source and period; report each slice | Early traffic is staged | Later traffic 1.5–1.8× worse |
| Plot errors by position in the utterance | Boundaries are input too | Head ~2.5–3× more error-prone than tail |
| Budget for thousands of utterances, resampled by call | CI width falls only as 1/√n | 2,000 utterances still resolve only ~0.65 points |
Each check is cheap, and each would have changed how I read that first table.
Summary
- A public benchmark is a training signal, not a production estimate. My 8 kHz curve was real, but on real calls three models got worse and reshuffled. Build an in-channel, in-domain set early, and put a confidence interval on it before ranking anything.
- Treat the eval set as a product with its own QA. Labels drafted from a scored model, mismatched number formats and staged early traffic each moved my numbers more than the differences I was trying to measure. Check per-model reference-match rates, normalize both sides, and split by source and period.
- Not every STT error is the model’s. Errors clustered at the start of utterances, right where I later found audio being clipped before Whisper saw it, and where the fine-tunes wrote city names nobody said. Look at where errors fall, and fix the boundary before another training run.
What’s Next
Even a well-measured STT model will mishear place names and write numbers in a form the next stage doesn’t expect. Part 2 moves downstream: correcting STT errors after the fact, keeping display text and spoken text apart, and why I stopped letting the LLM read numbers aloud.