Why is evaluating an LLM so much harder than testing normal software?
Unit tests pass or fail. LLM outputs don't. The hard part isn't running the eval — it's deciding what 'correct' even means when there are a million right answers.
On this page
The picture version
Five pictures for a reader who has only ever seen a green test suite. The prose below fills in the seams the pictures skip.
1 · The problem
The fix worked. Something else quietly broke.
2 · What a unit test gets for free
Three pieces. Two of them were invisible because they were trivial.
3 · The grader problem
Every option is flawed, and one of them is another model.
4 · The trap that eats eval suites
Optimise against a fixed set and you stop solving the task.
5 · Keep this card
The whole thing on one index card.
Why it exists
You’ve had this happen as a user: an app you rely on ships an “improved AI” update, and the one thing you actually used it for gets worse. Somebody at that company shipped a change believing it was an improvement. Here’s the wall they hit.
Anyone who has shipped a feature backed by a LLM knows the shape of it. Say the feature is a summarizer for incoming support tickets — that’s the running example for this whole post. You write a prompt. It looks great on the five tickets you tried. You ship it. Two weeks later support is forwarding you summaries that are wrong in ways your five tickets never hinted at. You tweak the prompt. The new version fixes those cases — and silently regresses three others you’d already considered solved.
Normal software has a clean answer to “is this change good?” — the test suite is green or it isn’t. With an LLM there is no green. Outputs are free-form text, the same input can produce different outputs, “correct” is a judgment call, and the model under test is usually a black box you didn’t train and can’t introspect. Every team building on top of these things ends up reinventing the same painful machinery: a synthetic eval set, a scoring function that mostly works, and a haunted feeling that they don’t really know if today’s model is better than yesterday’s.
This post is about why that machinery is so hard to build, not how to build it. The shape of the problem is what trips people up.
Why it matters now
Every team using agents, chatbots, summarizers, classifiers, or “AI features” of any kind faces this. Three things make it especially painful right now:
- Models change under you. A provider rolls out a new snapshot. Your prompt that worked yesterday produces subtly different outputs today. Without an eval, you can’t detect the regression — let alone localize it.
- Prompt changes are nonlocal. Editing one bullet in a system prompt can change behaviour on inputs that didn’t mention that bullet at all. There is no equivalent of “this function is now pure, so the diff is bounded.”
- The cost of being wrong is real. Hallucinations in customer-facing outputs aren’t a “fix it next sprint” bug; depending on the domain they’re a refund, a compliance incident, or worse.
So eval is the load-bearing thing that makes the rest of LLM engineering not-a-vibes-exercise. And it is much harder than it looks.
The short answer
LLM eval = test inputs + a grader you can defend + a metric that aggregates noisy outcomes
Picture to keep: not a green/red test runner, but a panel of judges scoring a diving competition — where you also have to argue about whether the judges are any good, and every dive gets scored slightly differently if you run it again.
A regular test suite hides two pieces of that equation because they’re
trivial: the grader is == and the aggregator is “all green.” An LLM eval
forces you to build both pieces explicitly, and each one is its own
research problem. That’s the whole reason it’s hard.
The analogy breaks in one place worth naming: diving judges at least agree on what a dive is. In LLM eval you often don’t have a fixed target output at all, so the judges are scoring against a standard you also had to invent.
How it works
Walk through what a passing test means in normal software:
- Run the function on a known input.
- Compare output to a known expected output with
==. - Repeat for many inputs. Aggregate: any failure → fail.
Now try to port that, step by step, to your ticket summarizer. Every line above breaks, and each break forces a fix that opens the next problem:
-
There is no single expected output. A good summary can use different words, different ordering, different emphasis. Two correct summaries written by two humans will not be string-equal. So
==doesn’t work, and neither does string-similarity — a rewording can be near-identical to the reference and still wrong, or wildly different and still right. -
The output isn’t deterministic. Even at temperature 0, you can get different tokens across runs (see why temperature 0 isn’t deterministic). So a single run doesn’t tell you “the model fails on this input”; it tells you “this sample failed.” If you want a stable failure rate, you often need multiple samples per input — which multiplies cost.
-
You need a grader. Something has to decide if an output is correct. The realistic options are all flawed:
- Exact-match / regex. Works only for narrow tasks (multiple choice, numeric answers, code that runs). Most real tasks aren’t this shape.
- Reference-based metrics (BLEU, ROUGE, embedding similarity). Cheap, but they reward looking like the reference more than being right. Embedding similarity captures some meaning; it routinely misses the task-relevant kind.
- LLM judges. Flexible, scale well, and can grade open-ended outputs. They also have measured biases — Zheng et al. (2023) document position bias, verbosity bias (preferring longer answers), and self-enhancement bias (preferring outputs that look like their own writing). They are not free of the same hallucination problem they’re meant to detect.
- Humans. The gold standard, slow and expensive — and human annotators disagree with each other on subjective judgments more than teams budget for. (Zheng et al. report human-to-human agreement on their pairwise comparisons that’s comparable to GPT-4-to-human agreement, which is the useful calibration: the gold standard is itself noisy.)
Whichever grader you pick, you’ve added a second model-shaped thing that itself needs to be evaluated. “How do I know my judge is right?” is a real question with no clean answer.
-
You need a dataset that represents production. This is where most eval suites quietly die. The five examples the developer tried are not a sample — they’re the inputs the developer can already think of. Real users hit edge cases the developer never imagined. Building an eval set that catches the long tail means harvesting real traffic, which raises privacy issues, requires labelling, and goes stale every time the product changes.
-
The aggregate is a distribution, not a pass/fail. You end up with a summary like “94% of outputs are acceptable” (numbers illustrative, not measured). Whether today’s number beats yesterday’s depends on confidence intervals, on which slices regressed, and on whether the failures got worse even if fewer.
Now stack those problems together and you can see why LLM eval feels fractal. Every single step that was free in unit testing is its own project.
Where it gets especially weird
A few specific traps that hit teams over and over:
- Goodhart on the eval. Once you optimize your ticket prompt against a fixed set of tickets, you start solving for those tickets rather than for summarizing tickets. Held-out evals exist for exactly this reason, and people forget to keep them held out.
- Contamination. If your eval inputs are anywhere on the public internet, the model may have seen them in training and “ace” them for reasons that don’t generalize. (Your support tickets are probably safe; the public benchmark you also track is probably not.) There’s no clean way to confirm contamination from outside the lab. The usual workaround is generating fresh eval data privately; how widely that is actually practiced isn’t something anyone has surveyed.
- Tasks where the right answer is “I don’t know.” Eval sets often reward producing some answer. A ticket with genuinely insufficient information should produce “not enough detail to summarize” — but a model that confidently invents a plausible summary can outscore one that correctly refuses.
- Multi-step / agent eval. When the LLM is in a loop calling tools, there is no single output to grade. You’re grading a trajectory. Trajectory eval is its own research area and the public state of the art is rough; the common fallback is scoring only the end goal, which tells you a run failed but not where.
A gap worth marking, and it belongs to the field rather than to this post: there is no settled, widely accepted methodology for evaluating open-ended LLM outputs. Public benchmarks exist, judge models exist, frameworks like Evals and others exist, but the question “is your model better than mine on my actual task” mostly does not have a turnkey answer.
You started with LLM eval = test inputs + a grader you can defend + a metric that aggregates noisy outcomes. What did this post add that the
line doesn’t show? — + each of those three is itself a thing you have to evaluate. Your ticket dataset can be unrepresentative, your judge
can be biased, and your aggregate can be inside the noise floor. That
recursion is why eval feels bottomless: unit testing gave you all three
for free, and an LLM takes all three back.
Check yourself
Before you go — you swap your grader from exact-match to an LLM judge and your ticket summarizer’s score jumps from 71% to 89%. Did the summarizer get better?
Answer
You have no idea, and that’s the point: you changed the measuring instrument, not the thing measured. Exact-match was probably marking correct-but-differently-worded summaries as failures, so some of that gain is real signal it was throwing away. But LLM judges have known biases — a documented tendency to prefer longer answers, and to prefer outputs stylistically like their own — so some of the gain may be the judge rewarding verbosity. The only way to tell them apart is to human-label a sample and check how often each grader agrees with the humans. Score changes across a grader swap are uninterpretable on their own.
And one more, on a different task shape — a colleague is evaluating a
coding agent that runs tests until they pass. They point out that
unlike your summarizer, this one has a real ==: the tests either
go green or they don’t. Have they escaped the problem in this post?
Answer
They’ve escaped one third of it. Verifiable end states are a genuine
advantage: the grader problem largely dissolves, because “tests pass”
is checkable without a judge. The other two pieces don’t. The
dataset problem is unchanged — their task set is still only the
tasks they thought to write down, and the agent will meet weirder
repos in production. And the aggregation problem gets worse, not
better, because an agent is a trajectory: the same task can pass on
one run and fail on the next, and “60% success” over a handful of
tasks has a very wide confidence interval. They should also watch for
the reward-hacking version of Goodhart — an agent that makes tests
pass by editing the tests has satisfied == perfectly.
Famous related terms
- LLM-as-a-judge —
LLM judge = one model graded by another— fast and flexible, biased in known ways. Not a free pass. - Held-out eval —
held-out eval = test set you don't tune against— the main defence against Goodharting your own benchmark. Useful only as long as it stays held out. - Hallucination —
hallucination = confident output, no grounding— eval is partly a defence against this leaking into prod. - Temperature —
temperature = how peaked the sampling distribution is— affects how many samples per input you need to estimate a rate. - Goodhart’s law —
Goodhart = "a measure that becomes a target stops being a good measure"— every eval suite eventually has to outrun this. - Trajectory eval —
trajectory eval = grading a sequence of tool calls and observations, not one answer— open research area for agents.
Going deeper
- Zheng et al., Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena (2023) — the primary source for “how good is an LLM judge, actually,” including the specific biases (position, verbosity, self-preference) they measured.
- Stanford HELM (crfm.stanford.edu/helm) — answers “what does a serious, many-metric eval suite look like when someone builds one properly?” Useful as a shape to copy, not as a stand-in for evaluating your task.
- The model cards and system cards published by frontier labs (a category, not a single link — each lab posts its own) — the rabbit hole, for how the people with the most eval budget describe their own methodology and its limits. Read the caveats section, not the numbers.
Note what is missing from that list: a canonical explainer — a “read this one thing and you’ll get it” piece. There isn’t one. The literature is a pile of papers and vendor writeups, and the absence of a standard reference is itself part of the answer to why this is hard.