Four Ways My STT Evaluation Lied to Me: Whisper Meets Real Phone Calls
This is Part 1 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. The Whisper fine-tuning series covered how I trained the model; this one starts with what happened when it met real calls. By the end of last year my Whisper fine-tuning curve looked about as good as curves get. On the held-out split of a public Korean 8 kHz telephone corpus (AI Hub’s low-quality telephone-network speech data), every round beat the last: 9.0% CER, then 6.3%, then 5.5%, then 3.8% for a fine-tuned large-v3-turbo. Off-the-shelf large-v3-turbo sat at 6.5%. I had set myself a target of under 4%, and 3.8% cleared it. ...