Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why is evaluating an LLM so much harder than testing normal software?

Unit tests pass or fail. LLM outputs don't. The hard part isn't running the eval — it's deciding what 'correct' even means when there are a million right answers.

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 11 min read

On this page

The picture version

Five pictures for a reader who has only ever seen a green test suite. The prose below fills in the seams the pictures skip.

1 · The problem

The fix worked. Something else quietly broke.

your ticket summarizer, before two are wrong, so you tweak the prompt tweak after those two fixed — three you’d solved now broken in normal software the suite is green or it isn’t here, there is no green You cannot answer “is this change good?” — and you shipped anyway. the ticket summarizer is the running example for the whole post
Prompt edits are nonlocal: changing one bullet moves behaviour on inputs that never mentioned it. Without an eval you can’t detect the regression, let alone localise it — which is how “improved AI” updates make the thing you relied on worse.

2 · What a unit test gets for free

Three pieces. Two of them were invisible because they were trivial.

a unit test inputs you wrote down == all green / not the bottom two are free, so you never think of them as parts of a test at all an LLM eval a dataset that represents production a grader you can defend a metric over noisy outcomes you have to build all three explicitly, and each one is its own research problem Unit testing gave you all three for free. An LLM takes all three back.
There is no single expected output — two correct summaries written by two humans will not be string-equal — and the output isn’t deterministic, so one run tells you “this sample failed”, not “the model fails here”.

3 · The grader problem

Every option is flawed, and one of them is another model.

exact match works for multiple choice, numbers, code that runs most real tasks aren’t this shape overlap metrics BLEU, ROUGE, embedding similarity reward looking like the reference, not being right an LLM judge flexible, scales, grades open-ended text measured biases: position, verbosity, self-preference humans the gold standard slow, expensive, and they disagree with each other “how do I know my judge is right?” Pick any grader and you have added a second thing that needs evaluating.
Zheng et al. (2023) measured the judge biases and the useful calibration: human-to-human agreement on their pairwise comparisons is comparable to GPT-4-to-human agreement. The gold standard is itself noisy.

4 · The trap that eats eval suites

Optimise against a fixed set and you stop solving the task.

everything real users actually send the five you thought of the long tail you never imagined they aren’t a sample — they’re the inputs the developer could already think of and then Goodhart arrives tune the prompt against the set now you solve for those tickets not for summarizing tickets A measure that becomes a target stops being a good measure. held-out evals exist for exactly this — and people forget to keep them held out
The same trap has a public-benchmark cousin: if your eval inputs are anywhere on the open internet, the model may have seen them in training and ace them for reasons that don’t generalise. There is no clean way to confirm contamination from outside the lab.

5 · Keep this card

The whole thing on one index card.

LLM eval = test inputs that look like production + a grader you can defend + a metric that aggregates noisy outcomes ∴ and each of the three is itself a thing you must evaluate
Picture to keep: not a green/red test runner, but a panel of judges scoring a diving competition — where you also have to argue about whether the judges are any good, and every dive scores slightly differently if you run it again. Where it breaks: diving judges at least agree on what a dive is.

Why it exists

You’ve had this happen as a user: an app you rely on ships an “improved AI” update, and the one thing you actually used it for gets worse. Somebody at that company shipped a change believing it was an improvement. Here’s the wall they hit.

Anyone who has shipped a feature backed by a LLM knows the shape of it. Say the feature is a summarizer for incoming support tickets — that’s the running example for this whole post. You write a prompt. It looks great on the five tickets you tried. You ship it. Two weeks later support is forwarding you summaries that are wrong in ways your five tickets never hinted at. You tweak the prompt. The new version fixes those cases — and silently regresses three others you’d already considered solved.

Normal software has a clean answer to “is this change good?” — the test suite is green or it isn’t. With an LLM there is no green. Outputs are free-form text, the same input can produce different outputs, “correct” is a judgment call, and the model under test is usually a black box you didn’t train and can’t introspect. Every team building on top of these things ends up reinventing the same painful machinery: a synthetic eval set, a scoring function that mostly works, and a haunted feeling that they don’t really know if today’s model is better than yesterday’s.

This post is about why that machinery is so hard to build, not how to build it. The shape of the problem is what trips people up.

Why it matters now

Every team using agents, chatbots, summarizers, classifiers, or “AI features” of any kind faces this. Three things make it especially painful right now:

So eval is the load-bearing thing that makes the rest of LLM engineering not-a-vibes-exercise. And it is much harder than it looks.

The short answer

LLM eval = test inputs + a grader you can defend + a metric that aggregates noisy outcomes

Picture to keep: not a green/red test runner, but a panel of judges scoring a diving competition — where you also have to argue about whether the judges are any good, and every dive gets scored slightly differently if you run it again.

A regular test suite hides two pieces of that equation because they’re trivial: the grader is == and the aggregator is “all green.” An LLM eval forces you to build both pieces explicitly, and each one is its own research problem. That’s the whole reason it’s hard.

The analogy breaks in one place worth naming: diving judges at least agree on what a dive is. In LLM eval you often don’t have a fixed target output at all, so the judges are scoring against a standard you also had to invent.

How it works

Walk through what a passing test means in normal software:

  1. Run the function on a known input.
  2. Compare output to a known expected output with ==.
  3. Repeat for many inputs. Aggregate: any failure → fail.

Now try to port that, step by step, to your ticket summarizer. Every line above breaks, and each break forces a fix that opens the next problem:

  1. There is no single expected output. A good summary can use different words, different ordering, different emphasis. Two correct summaries written by two humans will not be string-equal. So == doesn’t work, and neither does string-similarity — a rewording can be near-identical to the reference and still wrong, or wildly different and still right.

  2. The output isn’t deterministic. Even at temperature 0, you can get different tokens across runs (see why temperature 0 isn’t deterministic). So a single run doesn’t tell you “the model fails on this input”; it tells you “this sample failed.” If you want a stable failure rate, you often need multiple samples per input — which multiplies cost.

  3. You need a grader. Something has to decide if an output is correct. The realistic options are all flawed:

    • Exact-match / regex. Works only for narrow tasks (multiple choice, numeric answers, code that runs). Most real tasks aren’t this shape.
    • Reference-based metrics (BLEU, ROUGE, embedding similarity). Cheap, but they reward looking like the reference more than being right. Embedding similarity captures some meaning; it routinely misses the task-relevant kind.
    • LLM judges. Flexible, scale well, and can grade open-ended outputs. They also have measured biases — Zheng et al. (2023) document position bias, verbosity bias (preferring longer answers), and self-enhancement bias (preferring outputs that look like their own writing). They are not free of the same hallucination problem they’re meant to detect.
    • Humans. The gold standard, slow and expensive — and human annotators disagree with each other on subjective judgments more than teams budget for. (Zheng et al. report human-to-human agreement on their pairwise comparisons that’s comparable to GPT-4-to-human agreement, which is the useful calibration: the gold standard is itself noisy.)

    Whichever grader you pick, you’ve added a second model-shaped thing that itself needs to be evaluated. “How do I know my judge is right?” is a real question with no clean answer.

  4. You need a dataset that represents production. This is where most eval suites quietly die. The five examples the developer tried are not a sample — they’re the inputs the developer can already think of. Real users hit edge cases the developer never imagined. Building an eval set that catches the long tail means harvesting real traffic, which raises privacy issues, requires labelling, and goes stale every time the product changes.

  5. The aggregate is a distribution, not a pass/fail. You end up with a summary like “94% of outputs are acceptable” (numbers illustrative, not measured). Whether today’s number beats yesterday’s depends on confidence intervals, on which slices regressed, and on whether the failures got worse even if fewer.

Now stack those problems together and you can see why LLM eval feels fractal. Every single step that was free in unit testing is its own project.

Where it gets especially weird

A few specific traps that hit teams over and over:

A gap worth marking, and it belongs to the field rather than to this post: there is no settled, widely accepted methodology for evaluating open-ended LLM outputs. Public benchmarks exist, judge models exist, frameworks like Evals and others exist, but the question “is your model better than mine on my actual task” mostly does not have a turnkey answer.

You started with LLM eval = test inputs + a grader you can defend + a metric that aggregates noisy outcomes. What did this post add that the line doesn’t show? — + each of those three is itself a thing you have to evaluate. Your ticket dataset can be unrepresentative, your judge can be biased, and your aggregate can be inside the noise floor. That recursion is why eval feels bottomless: unit testing gave you all three for free, and an LLM takes all three back.

Check yourself

Before you go — you swap your grader from exact-match to an LLM judge and your ticket summarizer’s score jumps from 71% to 89%. Did the summarizer get better?

Answer

You have no idea, and that’s the point: you changed the measuring instrument, not the thing measured. Exact-match was probably marking correct-but-differently-worded summaries as failures, so some of that gain is real signal it was throwing away. But LLM judges have known biases — a documented tendency to prefer longer answers, and to prefer outputs stylistically like their own — so some of the gain may be the judge rewarding verbosity. The only way to tell them apart is to human-label a sample and check how often each grader agrees with the humans. Score changes across a grader swap are uninterpretable on their own.

And one more, on a different task shape — a colleague is evaluating a coding agent that runs tests until they pass. They point out that unlike your summarizer, this one has a real ==: the tests either go green or they don’t. Have they escaped the problem in this post?

Answer

They’ve escaped one third of it. Verifiable end states are a genuine advantage: the grader problem largely dissolves, because “tests pass” is checkable without a judge. The other two pieces don’t. The dataset problem is unchanged — their task set is still only the tasks they thought to write down, and the agent will meet weirder repos in production. And the aggregation problem gets worse, not better, because an agent is a trajectory: the same task can pass on one run and fail on the next, and “60% success” over a handful of tasks has a very wide confidence interval. They should also watch for the reward-hacking version of Goodhart — an agent that makes tests pass by editing the tests has satisfied == perfectly.

Going deeper

Note what is missing from that list: a canonical explainer — a “read this one thing and you’ll get it” piece. There isn’t one. The literature is a pile of papers and vendor writeups, and the absence of a standard reference is itself part of the answer to why this is hard.