How to get reliable structured output from an LLM
- Reading time
- 6 minutes
- Assumes
- You're feeding model output into code
- Updated
- Sep 6, 2026
The failure isn't malformed JSON
Modern models rarely produce syntactically broken JSON when asked for a schema. The failures are subtler and more expensive.
A field arrives in the wrong shape. You asked for a number, you got
"about 40". Valid JSON, useless to your code.
An enum picks up a new member. You listed three categories and the model returned a fourth, because your input genuinely didn't fit any of them and the model helpfully invented somewhere to put it.
Optional becomes absent, or absent becomes null. Your consumer handles one and not the other.
Values are confidently wrong but well-typed. The schema is satisfied
perfectly. The dueDate is a hallucination.
Schema enforcement fixes the first three. Nothing fixes the fourth except evaluation, which is why "we use structured outputs" is not the end of the reliability conversation.
Design the schema for the model, not for your database
The schema you'd design for a database column is often the wrong one to hand a model, and the difference costs accuracy.
Prefer enums over free strings wherever the set is genuinely bounded. A model choosing from five categories is a far more reliable operation than a model writing a category name you then have to match.
Make ambiguity representable. If your enum is approve / reject and the
real world contains "not enough information," the model will pick one of your
two rather than fail. Add the third value. This single change eliminates more
bad outputs than any amount of prompt tuning.
Use flat structures where you can. Deep nesting is harder for a model to populate consistently, and most nesting in a first draft is organizational rather than necessary.
Name fields the way a person would. customerIntent outperforms ci_val.
The field name is part of the prompt whether you meant it to be or not.
The highest-value field to add
Add a way for the model to say it doesn't know — a nullable field, an
insufficient_information enum member, or a confidence value. Without one you
have guaranteed that uncertainty gets expressed as a confident wrong answer.
Ask for reasoning before the answer, not after
If the task involves any judgment, let the model produce its reasoning as a field, and put it before the conclusion in the schema.
Order matters because generation is sequential. A schema that puts category
first and reasoning second gets a category chosen with no analysis, followed
by a justification written to fit it. Reversed, the reasoning genuinely informs
the choice.
The reasoning field also pays for itself in debugging. When a classification is wrong, the reasoning tells you whether the model misread the input or applied your criteria differently than you intended, and those need different fixes.
Strip it before the value reaches your database if you don't want to store it, but generate it.
Validate at the boundary, always
Treat model output exactly like input from an external API you don't control: parse it into your types, and handle the failure explicitly.
Never pass raw model output into business logic on the assumption the schema held. Schema enforcement is a strong constraint, not a guarantee, and the version where it silently doesn't hold is the one that corrupts data quietly for a month.
Decide what happens on a validation failure before you need to know. Retry once with the error included, then fall back to a defined default or a human queue. A silent retry loop is how a cheap operation becomes an expensive one at three in the morning.
Test the inputs you didn't design for
The test set that matters is not the well-formed examples. It is the ones that don't fit.
Empty and near-empty inputs. A blank message, a subject line with no body, a single word. These reliably produce interesting failures.
Inputs that fit no category. If your enum has five members, find inputs that belong to none of them and check what happens.
Inputs matching several categories. Genuinely ambiguous cases, where you mostly want to know that the choice is stable rather than that it is correct.
Very long inputs. Where the relevant detail is buried in the middle, which is where models are weakest at retrieval.
Adversarial content. Text that contains instructions. A model extracting fields from user-supplied content will sometimes follow instructions inside that content, and you want to discover this in testing.
Common mistake
Treating schema conformance as the pass condition. A run where every output parsed cleanly and 15% of the values were wrong is a failing run. Score the values against a rubric; conformance is the floor, not the bar.
Watch for drift after a model change
Structured output is unusually sensitive to model swaps. The same schema and prompt against a new version will typically still conform and can quietly shift its distribution — more items in one category, more nulls, longer reasoning.
Keep a frozen set of inputs and compare category distributions before and after any model change. A distribution shift with no input change is the signal, and it is invisible in a conformance check.
This is also the argument for storing the reasoning field in production for at least a sample. When the distribution moves, the reasoning tells you why in minutes.
Before you wire it to anything
0 of 6 checked