<?xml version="1.0" encoding="utf-8" standalone="yes"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/">
  <channel>
    <title>Llm on Juntak Noh — AI Notes</title>
    <link>https://ai.klavierhye.cc/tags/llm/</link>
    <description>Recent content in Llm on Juntak Noh — AI Notes</description>
    <generator>Hugo -- 0.147.7</generator>
    <language>en</language>
    <lastBuildDate>Tue, 22 Sep 2026 00:00:00 +0000</lastBuildDate>
    <atom:link href="https://ai.klavierhye.cc/tags/llm/index.xml" rel="self" type="application/rss+xml" />
    <item>
      <title>Too Many Guardrails Made the Bigger Model Worse: Deleting 13,528 Lines of Orchestration</title>
      <link>https://ai.klavierhye.cc/posts/too-many-guardrails/</link>
      <pubDate>Tue, 22 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/posts/too-many-guardrails/</guid>
      <description>&lt;p&gt;&lt;em&gt;This is &lt;strong&gt;Part 7&lt;/strong&gt;, the last of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → &lt;strong&gt;Simplification&lt;/strong&gt;. &lt;a href=&#34;https://ai.klavierhye.cc/posts/how-guardrails-pile-up/&#34;&gt;Part 5&lt;/a&gt; showed how the dialog layer&amp;rsquo;s rules piled up, and &lt;a href=&#34;https://ai.klavierhye.cc/posts/llm-eval-metric-chose-architecture/&#34;&gt;Part 6&lt;/a&gt; why the evaluation kept rewarding them. This part is about taking them out.&lt;/em&gt;&lt;/p&gt;</description>
    </item>
    <item>
      <title>The Metric Chose the Architecture: How Gold-Answer QA Rewarded Our Guardrails</title>
      <link>https://ai.klavierhye.cc/posts/llm-eval-metric-chose-architecture/</link>
      <pubDate>Tue, 01 Sep 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/posts/llm-eval-metric-chose-architecture/</guid>
      <description>&lt;p&gt;&lt;em&gt;This is &lt;strong&gt;Part 6&lt;/strong&gt; of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → &lt;strong&gt;Metrics&lt;/strong&gt; → Simplification. &lt;a href=&#34;https://ai.klavierhye.cc/posts/how-guardrails-pile-up/&#34;&gt;Part 5&lt;/a&gt; covered how the dialog layer&amp;rsquo;s rules piled up one failing QA row at a time. This part is about the evaluation that kept rewarding them.&lt;/em&gt;&lt;/p&gt;</description>
    </item>
    <item>
      <title>Every Failing QA Row Became a Rule: How Guardrails Pile Up in an LLM Dialog System</title>
      <link>https://ai.klavierhye.cc/posts/how-guardrails-pile-up/</link>
      <pubDate>Tue, 11 Aug 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/posts/how-guardrails-pile-up/</guid>
      <description>&lt;p&gt;&lt;em&gt;This is &lt;strong&gt;Part 5&lt;/strong&gt; of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. &lt;a href=&#34;https://ai.klavierhye.cc/posts/stt-eval-real-calls/&#34;&gt;Part 1&lt;/a&gt; covered how my STT evaluation misled me, &lt;a href=&#34;https://ai.klavierhye.cc/posts/dont-let-the-llm-read-numbers/&#34;&gt;Part 2&lt;/a&gt; where rules beat models, &lt;a href=&#34;https://ai.klavierhye.cc/posts/bigger-model-faster-voice-latency/&#34;&gt;Part 3&lt;/a&gt; latency, and &lt;a href=&#34;https://ai.klavierhye.cc/posts/silent-failures-voice-ai/&#34;&gt;Part 4&lt;/a&gt; failures that never threw. This part is about the dialog layer, and how its rules multiplied one failing QA row at a time.&lt;/em&gt;&lt;/p&gt;</description>
    </item>
    <item>
      <title>Nothing Crashed: Silent Failures in a Production Voice AI Stack</title>
      <link>https://ai.klavierhye.cc/posts/silent-failures-voice-ai/</link>
      <pubDate>Tue, 21 Jul 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/posts/silent-failures-voice-ai/</guid>
      <description>&lt;p&gt;&lt;em&gt;This is &lt;strong&gt;Part 4&lt;/strong&gt; of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. &lt;a href=&#34;https://ai.klavierhye.cc/posts/stt-eval-real-calls/&#34;&gt;Part 1&lt;/a&gt; covered how my STT evaluation misled me, &lt;a href=&#34;https://ai.klavierhye.cc/posts/dont-let-the-llm-read-numbers/&#34;&gt;Part 2&lt;/a&gt; where rules beat models, and &lt;a href=&#34;https://ai.klavierhye.cc/posts/bigger-model-faster-voice-latency/&#34;&gt;Part 3&lt;/a&gt; latency. This part is about the failures that never threw an exception.&lt;/em&gt;&lt;/p&gt;
&lt;hr&gt;
&lt;p&gt;When a web app breaks, someone at least sees a 500 page. When a voice agent breaks, the caller hears nothing. They say &amp;ldquo;여보세요?&amp;rdquo; into the silence, wait a few seconds, and hang up. No stack trace reaches them, and often none reaches you either.&lt;/p&gt;</description>
    </item>
    <item>
      <title>Don&#39;t Let the LLM Read Numbers Aloud: Where Rules Beat Models in a Voice Pipeline</title>
      <link>https://ai.klavierhye.cc/posts/dont-let-the-llm-read-numbers/</link>
      <pubDate>Tue, 09 Jun 2026 00:00:00 +0000</pubDate>
      <guid>https://ai.klavierhye.cc/posts/dont-let-the-llm-read-numbers/</guid>
      <description>&lt;p&gt;&lt;em&gt;This is &lt;strong&gt;Part 2&lt;/strong&gt; of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. &lt;a href=&#34;https://ai.klavierhye.cc/posts/stt-eval-real-calls/&#34;&gt;Part 1&lt;/a&gt; covered how my STT evaluation misled me. This part is about the text around the LLM: numbers going in and out, names that STT gets &lt;em&gt;almost&lt;/em&gt; right, and where an LLM call is worth its latency.&lt;/em&gt;&lt;/p&gt;</description>
    </item>
  </channel>
</rss>
