<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Data-Quality on Juntak Noh — AI Notes</title>
    <link>https://ai.klavierhye.cc/tags/data-quality/</link>
    <description>Recent content in Data-Quality on Juntak Noh — AI Notes</description>
    <generator>Hugo -- 0.147.7</generator>
    <language>en</language>
    <lastBuildDate>Tue, 19 May 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://ai.klavierhye.cc/tags/data-quality/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Four Ways My STT Evaluation Lied to Me: Whisper Meets Real Phone Calls</title>
      <link>https://ai.klavierhye.cc/posts/stt-eval-real-calls/</link>
      <pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/posts/stt-eval-real-calls/</guid>
      <description>&lt;p&gt;&lt;em&gt;This is &lt;strong&gt;Part 1&lt;/strong&gt; of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. The &lt;a href=&#34;https://ai.klavierhye.cc/posts/whisper-preprocessing/&#34;&gt;Whisper fine-tuning series&lt;/a&gt; covered how I trained the model; this one starts with what happened when it met real calls.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;By the end of last year my Whisper fine-tuning curve looked about as good as curves get. On the held-out split of a public Korean 8 kHz telephone corpus (AI Hub&amp;rsquo;s low-quality telephone-network speech data), every round beat the last: 9.0% CER, then 6.3%, then 5.5%, then 3.8% for a fine-tuned large-v3-turbo. Off-the-shelf large-v3-turbo sat at 6.5%. I had set myself a target of under 4%, and 3.8% cleared it.&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
