This is Part 6 of Field Notes from a Korean Phone Voice Agent, a seven-part series about a project I led: a real-time Korean phone voice agent (STT → LLM → TTS) for a public-service call line. STT evaluation → Rules vs. models → Latency → Silent failures → Guardrails → Metrics → Simplification. Part 5 covered how the dialog layer’s rules piled up one failing QA row at a time. This part is about the evaluation that kept rewarding them.
In late April I had a leaderboard I liked. Same 407 user turns, five models, two architectures: the team’s long production prompt versus the graph orchestrator I had spent months building — intent classifier, slot filling, canned templates. Every orchestrated configuration beat every prompted one, by up to a little over a point on a 5-point scale. Qwen3-4B, the smallest model in the test, scored 4.41 behind my templates. GPT-5 behind the long prompt scored 3.38 — dead last.
That table was the best argument I ever had for the architecture. It was also, mostly, a measure of how closely each system reproduced the answer key’s wording. This post is how I found out, and the eval stack I’d build so the metric can’t quietly pick the architecture again.
How the Loop Started
The first harness was simple, and I’d build it again. The business side of the team kept a shared sheet of scripted conversations, with a gold answer for every bot turn. A script replayed each conversation through the live chat API and wrote the reply and a pass mark back into the sheet. Exact string match ran first; an LLM judge ran only on mismatches, and a judge reply that came back empty or failed counted as not passed.
That gave us a same-day regression loop owned by the people who knew what a good answer sounded like. It also rested on three assumptions nobody had written down:
- The gold answer is the right answer, and close to the only one.
- The caller’s next line doesn’t depend on what the bot just said.
- The judge is a fixed instrument.
All three broke. The third broke first, and I was the one who broke it.
The Judge That Moved With Me
During the late-March QA push from Part 5 — more than forty fix commits in three days — I loosened the judge twice in 26 minutes. At 00:56 it started accepting extra information and differences in spacing or wording. At 01:22 it also accepted listing more items than the gold answer, spelling a number in Hangul where the gold used digits (or the reverse), and an added contact number. Each rule was defensible on its own; I’d still accept most of them.
The problem was how they arrived. The same person, on the same night, in the same commits that patched the system, changed the ruler. The pass rate went up, and nobody could say how much of the rise was the system and how much was the measurement.
One earlier decision saved part of the signal. A few days before, I had started reporting exact passes (O) and judge-upgraded passes (△) separately, and that split paid off a month later, when I switched the bot’s re-ask turns (the ones that ask the caller to fill in a missing detail) from fixed strings to LLM paraphrases (Part 2). A run on the 45 conversations that contained re-asks, 127 turns in all, passed 116: O = 51, △ = 65, X = 11. More than half the passes were the judge’s opinion, and the judge was the serving model grading its own paraphrases. On rows whose gold answer was a re-ask, replies matching the old fixed wording fell from 22 to 8. The system had become more natural, and the answer key could no longer see it.
The harness had a second moving part: the clock. A few intake flows behaved differently depending on the day or the hour, and the QA script never pinned the time, so the same set could score differently depending on when it ran. The later benchmark pinned one fixed date and time, but going back over the code for this post, I found the pin only reached the prompt path; the orchestrator’s time-aware flows still read the wall clock.
Freeze the judge before tuning the system, or give it to someone else. Version the judge prompt and model like code, change them only between runs, and report exact and judge-assisted passes as separate numbers.
Same Turns, Two Judges, Different Winners
The April benchmark replayed the same 407 user turns through four open models under both architectures, plus GPT-5 with the long prompt only. I scored every reply twice.
The gold-reference scorer compared each reply to the answer key. Reading its output back for this post, it was string arithmetic: an overlap of 0.85 or more with the gold answer always meant a 5; below that, a similarity value mapped onto 1–4, with side rules for keyword overlap, phone-number agreement and “both ask for the location”. All 3,663 of its scores carry a short formulaic reason like that; not one is a judge’s sentence.
The reference-free judge (gpt-4.1-mini, temperature 0) saw the last few turns and the reply but not the gold answer, and was told to ignore factual accuracy. It rated one thing: would a good call-center agent say this?
| Model | Gold-ref: Orchestration | Gold-ref: Long prompt | Ref-free: Orchestration | Ref-free: Long prompt |
|---|---|---|---|---|
| Qwen3-4B | 4.41 | 3.47 | 3.70 | 3.14 |
| Qwen3.5-27B | 4.69 | 3.68 | 4.09 | 4.09 |
| Qwen3.5-35B-A3B | 4.53 | 3.78 | 3.86 | 4.10 |
| Gemma-4-26B-A4B | 4.65 | 3.57 | 4.07 | 3.89 |
| GPT-5 | — | 3.38 | — | 3.50 |
Mean score on a 1–5 scale, 407 turns per cell.
With the answer key, orchestration led by 0.75–1.09 points for every model. With the key hidden, it led clearly only for the 4B (+0.56). For the other three, a lead of about a point shrank to a quarter point or less: the 27B tied, Gemma kept a small edge (+0.18), and the 35B MoE — the model we had switched to a month earlier — flipped toward the prompt (−0.24). A conversation-level bootstrap I ran for this post puts roughly ±0.1–0.15 on each gap (95% intervals), so the 35B and Gemma differences are real but small.
The cause wasn’t hidden. Between 52% and 59% of orchestrated replies were character-for-character identical to the gold answer; for the prompted models it was 20–24%. The gold answers were written in a fixed script voice, and months of QA pushes had taught my templates to reproduce that voice — one commit even reverted a greeting “for QA compatibility”. A template system graded against its own phrasing isn’t being measured on quality. It’s being measured on conformity.
A gold set written in one system’s voice will crown that system. If your gold answers and your templates share wording, a high reference score mostly tells you that they share wording. Run at least one judge that never sees the answer key.
I don’t read GPT-5’s last place as a verdict on GPT-5. It ran through a different API path with different defaults, and its first run came back almost entirely empty (Part 4). Its rank says more about the harness than the model.
When the Gold Answer Is Wrong
A caller says someone is following them. The gold answer and my template both told them to call 119 — fire and ambulance. Every prompted model said 112, the police, which is the number I would want them to dial. The gold-reference scorer gave the prompted models 2/5 on a similarity of 0.44–0.45, and gave the orchestrator 5/5 for copying the wrong answer word for word.
The reference-free judge missed it too. By design it doesn’t check facts, so it rated both versions 4–5. Neither metric found the wrong label. Reading the disagreements did.
For the 35B pair, 59 turns had orchestration at 5 and the prompt at 2 or less. Forty-nine were low string similarity. Reading them for this post, only some were the same answer in different words; more were genuinely different answers, such as an intake question where the gold gave a referral number, or a confident statement that wasn’t true. Four were closing lines. Six were phone-number mismatches — exactly the hard facts a reference should check, and the only part of that scorer I’d keep. The reference-free judge rated 42 of those 59 prompted replies 4 or higher, including five of the six number mismatches. Each score was blind in its own way: the string scorer couldn’t tell a paraphrase from a wrong answer, and the reference-free judge waved wrong numbers through.
I had even written the principle down. The rubric I’d drafted for an LLM reference judge said, in effect, that an appropriate answer that differs from the intended one still deserves a high score. A string scorer can’t honor that sentence. What I’d do now: route metric disagreements to a review queue, and when a gold answer is wrong, fix the answer key, not the system.
Replaying a Script Is Off-Policy
The 407 turns came from 184 conversations — about 2.2 user turns each, six at most — and the user lines were fixed regardless of what the bot said. If the bot asked for a time when the script expected a question about location, the next scripted line answered a question nobody had asked, and the bot was graded on its reply to nonsense.
In reinforcement-learning terms this is off-policy evaluation: scoring one policy on trajectories generated by another. The scripts assumed the conversation flow I had spent months tuning the orchestrator to follow, so any system that took the conversation somewhere else paid for it. My first multi-turn capture script was worse. One scripted second turn argued with a specific wrong answer from an earlier run (“you just told me X, but…”), so it only made sense if the bot repeated the mistake.
The second version replaced the scripts with 13 small caller functions that read the bot’s actual reply and branched on keywords: if it mentions an online option, ask about that; if it hedges, ask for a number to call; if it says the closing phrase, hang up. Crude, but on-policy. The simulated caller later became an LLM, and by mid-May the benchmark was also scoring non-repetition. Our open model averaged 4.0/10 there, against GPT-5’s 9.38/10 — a failure that shaped the next stage of the work, and one a turn-by-turn gold comparison has no column for.
Who Grades the Grader
The QA judge ran on the serving model itself, through a raw generate endpoint I added so tooling would stop fighting the server for GPU memory (that saga is in Part 3). For “did anything regress since yesterday?” that’s fine. For “which system is better?” it’s wrong: the model grades its own outputs, and when you swap the model, the judge swaps with it. For the April comparison I moved LLM judging to an external model and froze the test data to a file, so a mid-run sheet edit could no longer change the set.
That still left the question under every automated metric: does it agree with people? The last piece was a blind A/B survey, a metric for the metric. It paired base and fine-tuned answers that the QA harness had recorded side by side:
- Randomized order: each pair’s A/B position was a coin flip, so position bias can’t pick a winner.
- Near-duplicates dropped: pairs more than 90% similar were skipped, so raters only judged real differences.
- Four options, including “both bad”: a forced choice between two wrong answers hides the outcome you most need to see.
- Full context: each question showed the conversation so far, not just the last turn.
The form had 26 questions and went out in two consecutive sprints, the second after the simulated caller was rebuilt. Its stated purpose was to decide whether the automated scores deserved trust. I’m describing the design, not the outcome. The point is where it sits: a human check on the metric belongs before the metric gets to pick an architecture.
The Eval Stack I’d Build Again
If I started over, this is the stack, roughly in the order I’d add it:
| Layer | What it does | What taught me |
|---|---|---|
| Hard facts by reference | Extract numbers, departments, hours; compare exactly | Six phone-number mismatches among 59 low scores |
| Quality without references | Judge from context, never shown the gold answer | The verdict flipping once the key was hidden |
| On-policy callers | Reactive, then LLM-driven; score whole conversations | A repetition failure no gold column could show |
| Frozen, external judge | Versioned; owned by someone not tuning the system | Two loosenings in 26 minutes |
| Separate counts | Exact vs judge passes; errors and empties per model | O = 51 vs △ = 65; a nearly empty GPT-5 run |
| Pinned conditions | Clock, seed, frozen test data | A wall clock leaking into one path |
| Held-out rows | Cases you never patch against | A QA push tuned to its own test set |
| Human check on the metric | Blind A/B with “both bad” | Not knowing whether any of this tracked quality |
Most of this is classical ML hygiene in new clothes. Patching the system row by row against the QA sheet is tuning on your test set, and a gold set that has absorbed your system’s wording is a quiet form of leakage; both are in the traps section of the bias–variance post. I knew both rules and broke them anyway. I made the same mistake on the speech side, where my reference transcripts were drafted from my own model’s output (Part 1). And fail toward “not passed”: a judge error or an empty generation should count against the system, never for it.
Summary
- Reference scores measure similarity to the answer key, not quality. When gold answers share wording with your templates, the template system wins by construction. Here, hiding the key shrank a lead of about a point to a quarter point or less for the three larger open models.
- The judge and the harness are part of the system under test. A judge tuned by the person tuning the system drifts with it, a scripted replay rewards the bot it was written against, and an unpinned clock measures the day of the week. Freeze, externalize and pin them.
- Split the eval by what each part can actually check. Hard facts against references, quality without them, conversations on-policy, and the whole thing checked against blind human preference. Then read the disagreements; that’s where the wrong gold labels are.
What’s Next
Once the answer key stopped choosing for me, the real question was what the scaffolding was buying a 35B model, and why a much bigger one couldn’t get past it. In Part 7, I delete 13,528 lines of orchestration, keep three guards that each come with a receipt, and look at what the deleted layer had been doing for free.