Heads up: posts on this site are drafted by Claude and fact-checked by Codex. Both can still get things wrong — read with care and verify anything load-bearing before relying on it.
why → how

Why is structured output so hard?

You ask the model for JSON. Sometimes it gives you a trailing comma. Sometimes a markdown fence. Sometimes prose. Why is this still a problem?

AI & ML intermediate Apr 29, 2026 · updated Aug 25, 2026 · 15 min read

On this page

The picture version

Six pictures for a reader who has never fought a parser. The prose below fills in the seams the pictures skip.

1 · The problem

It writes working code in a dozen languages. It cannot reliably close a brace.

the same request, sent over and over: “reply with JSON matching this schema” parsesparsesparsesparsesparses explodes Sure! Here’s the JSON: {"vendor": "Acme", "total": 91.20,} A chatty preamble. A trailing comma. Your pipeline stops. and a firmer prompt shrinks this without ever reaching zero — which is the puzzle
Turning invoices into database rows is the running example, and it mostly works — until the occasional reply is almost JSON. The failures are rare, structural, and not fixable by asking more firmly.

2 · Why asking nicely can’t win

One side is a set of preferences. The other is a yes or a no.

what the model produces a preference for every possible next piece — soft, graded, always non-zero somewhere what the schema is the whole string parses or it does not no partial credit, no “mostly” These two objects don’t combine on their own. Every fix is a way of forcing them to.
The model always assigns some weight to a token that would break the format, and a prompt can only make that weight small. A grammar is all-or-nothing about the finished string, so shrinking a probability never gets you a guarantee.

3 · The two things people try first

Ask harder, or parse and retry. Both are statistics, not guarantees.

ask more firmly the failure rate goes down and never to zero you moved a probability, not a rule parse, and retry on failure works surprisingly well and costs a whole call each time on a long request that is not free Fine if your pipeline treats a parse failure as noise. Not fine if it treats one as a bug.
Both moves are real and both are widely used, and neither changes the shape of the problem. They lower a failure rate; they do not remove the failure — which is exactly the distinction that decides whether you need the next scene.

4 · The fix

Put a turnstile in front of the model’s mouth.

what has been written so far: {"vendor": "Acme" what could come next ,}Sure!" gates still open locked — every one of these leads somewhere that can’t parse the model still picks whichever it likes best — among the ones left open Now the format isn’t requested. It is impossible to violate.
At every step the grammar looks at what has been emitted, closes every option that could not lead to a valid document, and lets the model choose freely among the rest. The choice is still the model’s; it just can’t leave the building.

5 · Why that’s a library and not a script

The rules are about characters. The model emits chunks that straddle them.

what the grammar talks about ",  " one character at a time what the model actually emits ",  " all five, in one indivisible chunk So “which chunks are legal here?” is a different, harder question than “which characters are legal here?” and the answer depends on this model’s own chunk list
The format’s rules are written over characters, but the model can only emit its own multi-character chunks, one of which may carry a quote, a comma and two spaces at once. Working out which chunks keep the document valid is the engineering — and it has to be redone per model vocabulary.

6 · Keep this card

The whole thing on one index card.

why it’s hard = a soft spread of preferences meeting a hard yes-or-no rule ∴ the only guarantee is to let the rule veto, step by step prompting and retries get you there statistically, never categorically And it buys you syntax. Only syntax. a perfectly-formed invoice record can still carry the wrong total
Picture to keep: a turnstile in front of the model’s mouth — at every token the grammar locks the gates that lead somewhere invalid and lets the model walk through whichever of the rest it likes. Where it breaks: the turnstile checks the shape of what comes out, never whether it is true.

Why it exists

Here’s the thing that should feel weird the first time you hit it.

You have a pile of invoices and you want them in a database. So you ask an LLM to “respond with valid JSON matching this schema” — {"vendor": string, "total": number} — and hand it invoice after invoice. (Keep that schema in mind; it’s the running example for the whole post.) It mostly works. Then, on maybe one in fifty calls — or one in five if your schema is awkward — it gives you back something that is almost JSON: a stray trailing comma, an unescaped quote, a markdown code fence wrapped around the JSON, a chatty preamble like “Sure! Here’s the JSON:” before the actual object, or a key name that matches the vibe of your schema but not the spelling. Your parser explodes. Your pipeline retries. Your ops dashboard lights up.

This should be surprising. The model can write working code in fifteen languages. It can summarize a legal contract. It clearly knows what JSON is — JSON is everywhere in any plausible web-scale training corpus, even if no provider tells you the exact count. Why does this one specific task — “produce a string of bytes that satisfies a formal grammar” — fail in a way that “write a haiku about kubernetes” does not?

You probably assume this is a prompting problem — that with a firm enough system prompt, or a better model, the failures go to zero. Prompting alone can’t get you there, and the reason is structural rather than a matter of model quality. (Something else can get you there; that’s the rest of the post.) The answer lives in the gap between two things that look similar from the outside but aren’t: a model that has seen a lot of JSON and a model that can only emit valid JSON. By default the model is the first kind. Making it the second kind is harder than it looks, and the awkwardness of every “JSON mode,” “structured outputs,” “function calling,” and tool-use API you’ve used is downstream of that.

Why it matters now

Structured output is the seam between LLMs and everything else.

If you don’t have a feel for why this is hard, you’ll either over-trust prompt-only JSON (“works on my eval, fails in prod”) or over-engineer around a problem that has a clean solution at the inference layer.

The short answer

structured output = next-token sampling + a hard constraint that the whole emitted string parses against a grammar

Picture to keep: a turnstile in front of the model’s mouth. At every token, the grammar looks at what has been emitted so far, locks every gate that would lead somewhere invalid, and lets the model walk through whichever gate it likes best among the ones still open. The model still chooses; it just can’t leave the building.

A language model is a probability distribution over the next token. A schema is a hard yes/no constraint on the whole sequence of tokens. Those two objects don’t compose for free. Every approach to structured output is some way of forcing them to compose — either by hoping (“please return JSON”), by fixing it after the fact (parse, retry), or by changing what the model is allowed to sample at each step (constrained decoding). Each option is a different trade between reliability, speed, and how much of the schema the model actually understood.

How it works

To see why this is hard, you have to look at what the model is actually doing when it generates text.

The mismatch: distributions vs. grammars

At each step of generation, the model produces a probability over its entire vocabulary — typically tens to hundreds of thousands of token IDs. Sampling picks one. The next step conditions on that pick. The model has no built-in notion of “I am currently inside a JSON string” or “the next character must be a closing brace.” It just has the conditional distribution that fell out of training on a giant pile of text.

A schema, on the other hand, is a hard constraint. The output either parses or it doesn’t. There’s no “70% valid JSON.” A trailing comma turns a 5,000-character valid response into a 0-character valid response.

The model has learned, statistically, that JSON-looking prefixes tend to be followed by JSON-looking continuations. That’s good enough most of the time. It is not good enough all of the time, because “JSON-looking” is a soft prior and “valid JSON” is a hard property. Soft priors fail at the tails, and the tails are where production lives.

Three approaches, in increasing order of “actually works”

1. Prompt only — “please return JSON.”

You write the schema in the system prompt. You add “respond ONLY with JSON, no preamble, no markdown fence.” The model mostly complies. This is what every developer tries first.

What goes wrong: the failure modes are exactly the ones you’d predict from training data. On invoice #37 the model returns {"vendor": "Acme", "total": 412.00} — exactly right — wrapped in a markdown code fence your parser chokes on. Markdown code fences are extremely common in JSON examples on the web, so the model has a strong prior to wrap output in ```json ... ```. Tutorial-style preambles (“Here’s the result:”) are common too. And subtle invariants — every key from the schema is present, no extra keys, enums use the right casing — aren’t things the model can verify; it can only imitate.

2. Retry on parse failure.

Run the model. Try to parse. If it fails, send the error back and ask it to fix. This is a common pattern in LLM apps, and it works surprisingly well in practice — handed the parse error, models often fix their own JSON.

What goes wrong: every retry is another full call, with full prompt and full context. Latency and cost at least double for that request, and worse if the repair loop runs more than once. There’s no upper bound on retries that’s both safe and cheap.

3. Constrained decoding — restrict what the model is allowed to sample.

This is the approach behind the strict, schema-guaranteeing structured-output features both OpenAI and Anthropic now document explicitly (OpenAI’s “Structured Outputs,” Anthropic’s strict tool use and structured-output modes). Not every JSON-ish mode is strict — older “JSON mode” features only promise parseable JSON, not schema conformance — but the strict ones are doing the same trick. At each generation step, before sampling, take the model’s distribution over the vocabulary and mask out every token that would make the output invalid under the schema. Renormalize. Sample from what’s left. Repeat.

The grammar — typically expressed as a regex or context-free grammar and compiled to a finite-state machine or pushdown automaton — tells you, given the prefix emitted so far, which tokens can still keep the output valid. On our invoice schema: after the prefix {"vendor": ", any token whose bytes fit inside a JSON string body is legal, and the token ``` is not — so the markdown fence that broke invoice #37 is now simply unreachable, no matter how much probability mass the model wanted to put there. You set the probability of every illegal token to zero.

The model still chooses which legal token, weighted by its own distribution. It still picks the vendor name, the number, the wording. It just can’t go off the rails of the grammar. That’s the turnstile from The short answer — with one place the analogy breaks worth naming: a turnstile only blocks you at the moment you try to walk through it, while a grammar has to look ahead, refusing tokens that are locally fine but would paint the output into a corner it can’t close. That look-ahead is the whole engineering problem.

This is why providers can promise “the output will parse.” They aren’t trusting the model; they’re forbidding everything else. Read the promise carefully, though — the provider docs carve out the cases where generation ends early for reasons the grammar can’t control, such as hitting a max-token limit or a refusal. Constrained decoding guarantees you never take an invalid step; it can’t guarantee you reach the closing brace.

Where it gets subtle

You started with structured output = next-token sampling + a hard constraint that the whole emitted string parses against a grammar. What did the invoice pipeline add? — + the constraint has to be enforced at every step, over tokens rather than characters, and it buys you syntax and nothing else. Structured output isn’t a prompt-engineering problem the model happens to be bad at; it’s a mismatch between a soft distribution and a hard predicate, and the only way to guarantee the format is to let the predicate veto the distribution token by token. Prompting and retries get you most of the way statistically, never categorically — which one you reach for depends on whether your pipeline treats parse failures as noise or as bugs.

Which settles the thing that felt weird at the top. The model that writes working code in fifteen languages really does fail at “emit a string satisfying a formal grammar,” and it isn’t a gap in its competence. Writing a haiku has no wrong answer to step on; every token of your invoice object does. Nothing about the model got better when the failures stopped — the sampler just stopped being allowed to make that mistake.

Check yourself

Before you go — you turn on strict structured outputs, and your invoice parse-failure rate drops to zero. A week later, finance says the totals are wrong on about 3% of invoices. Did the feature fail?

Answer

No — it did exactly what it promises, which is narrower than what finance needed. Constrained decoding guarantees the bytes satisfy the grammar: total will be a number, vendor will be a string, the braces will close. It says nothing about whether that number is the number printed on the invoice. Those 3% are a reading/extraction problem — hallucination, or genuine ambiguity in the document — and the fix lives somewhere else entirely: better prompting, a vision model that actually sees the layout, or a second pass that cross-checks line items against the total. The mistake is treating “guaranteed to parse” as a proxy for “guaranteed to be right.” They’re separate guarantees and only one of them was on sale.

And one more — someone proposes skipping the grammar engine entirely: just filter the model’s output character by character, blocking any character that would make the JSON invalid. Why doesn’t that work?

Answer

Because the model doesn’t emit characters, it emits tokens, and a single token can carry several characters that span more than one grammar state — a token might be ", " (quote, comma, two spaces, quote) all at once. You can’t block “the next character”; your only lever is to zero out entire token IDs before sampling. So the real question is which tokens keep the output on a path to a valid document, given the current grammar state and this specific model’s vocabulary — a set that has to be computed per state, per vocabulary. That computation is the actual product Outlines and XGrammar ship, and it’s why constrained decoding is a library rather than a regex.

Going deeper