<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Data-Quality on 노준탁 — AI 노트</title>
    <link>https://ai.klavierhye.cc/ko/tags/data-quality/</link>
    <description>Recent content in Data-Quality on 노준탁 — AI 노트</description>
    <generator>Hugo -- 0.147.7</generator>
    <language>ko</language>
    <lastBuildDate>Tue, 19 May 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://ai.klavierhye.cc/ko/tags/data-quality/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>STT 평가가 나를 속인 네 가지 방식: 실제 전화 통화 앞에 선 Whisper</title>
      <link>https://ai.klavierhye.cc/ko/posts/stt-eval-real-calls/</link>
      <pubDate>Tue, 19 May 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/ko/posts/stt-eval-real-calls/</guid>
      <description>&lt;p&gt;&lt;em&gt;이 글은 한국어 전화 Voice Agent 현장 노트 7부작 중 &lt;strong&gt;1부&lt;/strong&gt;다. 내가 리드했던 프로젝트 — 공공 서비스 전화 상담용 실시간 한국어 voice agent(STT → LLM → TTS) — 를 만들면서 겪은 일을 정리한다. &lt;a href=&#34;https://ai.klavierhye.cc/ko/posts/whisper-preprocessing/&#34;&gt;Whisper fine-tuning 시리즈&lt;/a&gt;가 모델을 어떻게 학습했는지를 다뤘다면, 이번 시리즈는 그 모델이 실제 통화를 만났을 때 벌어진 일에서 시작한다.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;작년 말, 내 Whisper fine-tuning 곡선은 더 바랄 게 없을 만큼 예뻤다. 공개 한국어 8 kHz 전화망 corpus(AI Hub의 저음질 전화망 음성인식 데이터)의 held-out split에서 라운드마다 기록이 갱신됐다: CER 9.0% → 6.3% → 5.5% → fine-tuned large-v3-turbo의 3.8%. 손대지 않은 large-v3-turbo는 6.5%였다. 스스로 잡은 목표가 4% 미만이었으니 3.8%면 통과다. (그때는 꽤 뿌듯했다.)&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
