What structured outputs actually guarantee
structured outputs is a feature where you hand the model a schema, and the API promises the reply will match it. if you ask for a JSON object with a name string and an age number, you get exactly that shape back, not a chatty paragraph with some JSON buried in it.
i care about this because i run small automations that call language models, and the most boring failure in all of them used to be a parse error at 3am. structured outputs fixes that specific failure. it does not fix the failures people assume it fixes, and that gap is what this article is about.
what it is
structured outputs means the model’s output is constrained to follow a format you define, usually a JSON Schema. the schema says which fields exist, what type each one is, and which values are allowed (for example, a field that must be one of “refund”, “exchange” or “other”).
there are three things people mix up, so let me separate them:
- prompting for JSON: you write “reply only with JSON” in the prompt. the model usually complies, and sometimes it doesn’t. nothing enforces it.
- JSON mode: the API guarantees the output is syntactically valid JSON, but not that it has the fields you wanted.
- structured outputs: the API guarantees the output matches your schema, so the fields, types and enums are what you declared.
OpenAI introduced its version under that name in August 2024, and the details are in its structured outputs guide. Anthropic documents a related route through tool use, where you define a tool with an input schema, in its tool use docs. the exact feature names and which models support them change over time, so check the vendor page for your model before you build on it.
the schema language itself is JSON Schema, an open spec. vendors support a subset of it, not all of it. more on that below.
how it works
a language model produces text one token at a time. if you want a refresher on what a token is, i wrote about it in what is a token and why your bill depends on it. at each step the model scores every token in its vocabulary, and then one is picked by sampling. i covered that part in temperature, top-p and sampling explained.
structured outputs works at that sampling step. the general technique is called constrained decoding, and the idea is simple:
- the vendor converts your schema into a grammar, basically a state machine that knows what can legally come next
- at each step, the grammar works out which tokens would keep the output valid
- every other token is blocked before sampling, so the model cannot pick it
- the model still chooses among the legal tokens, using its normal judgement
so if the schema says the next thing must be a number, the model literally cannot emit a letter. that is why the guarantee is real and not just “much more likely”. it is enforced by the decoder, not requested in the prompt.
there are practical consequences of this design. OpenAI’s docs note that the first request with a new schema has extra latency while the schema is processed, and later requests with the same schema are faster. the supported schema subset is also restricted. for OpenAI, for example, all fields have to be marked required and objects have to disallow additional properties. if you want an optional field, you model it as a value that can be null. these rules differ between vendors and change, so read the current docs rather than trusting my summary.
one more mechanical point. the guarantee applies to a finished response. if the output hits your max token limit halfway through an object, you get a truncated, invalid result. the API tells you this through a stop reason or finish reason, and you need to check it.
why it matters
here is where i actually use it.
- parsing without defensive code: before this, my scripts had retry loops, regex cleanup for markdown code fences, and a fallback for when the model wrote “Sure! Here’s the JSON”. with schema enforcement, the parse step stops being a place where things break. that is less code to maintain and fewer odd edge cases.
- classification and routing: if a support message must go to one of five queues, an enum in the schema means the answer is always one of the five. no “Billing / Refunds” invented on the fly. this is the cleanest use case.
- extraction from messy text: pulling invoice numbers, dates and totals out of emails or scraped pages into a fixed shape. if you scrape pages for this, the folks at proxyscraping.org cover the collection side, and a schema is what turns the raw page text into rows you can load into a database.
- tool calls and agents: when a model calls a function, the arguments have to be well formed or the call fails. this is part of why tool protocols lean on schemas. i touched on the wider picture in model context protocol explained.
there is a cost angle too. fewer malformed replies means fewer retries, and every retry is paid tokens plus extra latency plus another hit on your quota. if you are near your limits, that interacts with what i wrote in handling API rate limits.
common misconceptions
“the output is guaranteed to be correct”
no. it is guaranteed to be well formed. a schema can force "total" to be a number, but it cannot force it to be the right number. if the model misreads an invoice, you will get a perfectly valid JSON object with a wrong total, delivered with total confidence. i would argue this is slightly more dangerous than a parse error, because a parse error is loud and a wrong value is silent.
so you still need checks on the content. validate ranges, cross-check totals against line items, and sample outputs by hand. if you have a pile of real examples, turning them into a test set is worth the effort, and building an eval set from support tickets walks through how i would do that.
“it works with any schema”
no. vendors support a subset of JSON Schema. some keywords are unsupported, there are limits on nesting depth and the number of properties, and the rules differ between providers. a schema that works on one API may be rejected by another. if you plan to switch vendors, test your schemas on the new one early, not on migration day. this is also one reason i keep an eye on what happens when a model is deprecated, since a model swap can change what schema features you get.
“the model will always return the structure”
mostly, but there are exceptions you should handle in code:
- the response was cut off by the max token limit, so the JSON is incomplete
- the model refused the request for safety reasons. OpenAI’s docs describe a separate refusal field for this case, and a refusal does not follow your schema
- the request itself failed, for example a timeout or a rate limit error
if your code assumes every response is a clean object, the first refusal will crash it. check the stop reason, check for a refusal, then parse.
“a stricter schema always gives better results”
not necessarily. a very tight schema can push the model into awkward answers. for example, if your enum has no good option for an input, the model must still pick one of the bad ones. add an “other” or “unknown” value so the model has an honest way out. also, putting a free text reasoning field before the final answer field sometimes helps on harder tasks, because the model gets to write its thinking before committing. i have seen this help on my own tasks, but i would treat it as something to test, not a rule.
where to go from here
if this made sense, these are the next things i would read, roughly in order:
- temperature, top-p and sampling explained: constrained decoding makes more sense once you understand the sampling step it sits inside.
- building an eval set from support tickets: because valid shape is not valid content, you need a way to measure the content.
- when caching answers beats calling the model: structured, repeatable outputs are exactly the kind of thing that can be cached, which saves money.
- model context protocol explained: how schemas show up when models talk to tools.
if you want to browse everything on the site, the index is at /blog/. and if you handle personal data in these extraction pipelines, theprivacywire.com is worth a look for the privacy side of sending customer text to a third party model.
the short version: structured outputs guarantees the shape of the reply, and nothing more. treat the shape as solved, and spend your attention on whether the content inside it is right.
Written by Xavier Fok
disclosure: this article may contain affiliate links. if you buy through them we may earn a commission at no extra cost to you. verdicts are independent of payouts. last reviewed by Xavier Fok on 2026-10-03.