← all articles

Getting structured output that validates

structured-output json-schema llm data-extraction

The mail classifier that sorts inbound messages across my domains has a field for company name. It is a string and it is required, because at the time I wrote the schema every message I was thinking about came from a business.

Newsletters do not come from a business in any sense that field can hold. Neither do password resets. The classifier filled it anyway, usually by taking a word out of the sender domain and capitalising it.

Every one of those responses passed validation. They were well formed, every key present, correct types throughout. The parse failure rate on that job has been zero for months.

That is the whole problem in one field. Schema enforcement gave me a guarantee about shape and I read it as a guarantee about content.

What the constraint layer buys you

Constrained decoding is real and it works. The model emits one token at a time, and a layer between the model and your response masks out every candidate token that could not lead to a document matching the grammar. If the next character has to be a quote or a closing brace, nothing else is available to sample.

The result is that the model cannot leave a bracket open, cannot wander into a paragraph of explanation before the json, and cannot omit a key you marked required. Providers name this differently and some expose more control than others, but that is the mechanism underneath.

Worth knowing which version you are paying for. Enforcing the grammar during generation is a guarantee. Generating freely and validating afterwards, retrying behind the scenes until something passes, is a probability with a cost attached. The second still fails, and when it does the evidence lands on your bill rather than in your logs.

I used to keep a regex that hunted for the first json shaped block in a wall of prose. I do not miss it.

The schema is where fabrication gets decided

A schema describes form. This key exists. Its value is a string. It must be present.

An invented company name is a string. It meets the contract exactly, and the validator has no mechanism for caring, because you never gave it one and it cannot check the world.

So one class of engineering work disappears and the other does not move. The parsing problem is genuinely solved. The correctness problem relocates into the schema, which is the part nobody audits, because it looks like a type definition rather than like a prompt.

Required with no null

The largest single source of invented values in my own systems is a required field with no way to say the value is absent.

Look at what that does. You give the model a document, ask for a purchase order number, and build a machine that physically cannot finish a response without producing one. My invoice extractor spent three weeks doing that. On invoices from smaller suppliers, which mostly carry no purchase order number at all, it returned values with a sensible length and prefix that nobody would query at a glance.

There is no output available that satisfies both the schema and the document. The schema is enforced in code. The document is not enforced by anything.

A required field with no null option is an instruction to make something up. I will defend that in strong terms, because the fix costs about a minute and I have never seen it presented as a correctness measure.

Make the field nullable, or add an explicit value meaning this document did not contain it. My fabrication rate on that field did not improve when I did that. It stopped.

The other direction then becomes your problem. Make everything optional and the model starts reaching for null on values that are present but awkward to read off a page, and recall drops without anything appearing broken. So measure both: how often it fills a field that should be empty, and how often it empties a field that should be filled. I did mine over two evenings on about eighty documents I labelled by hand, which is a small sample and was enough.

Enumerate, then leave a door

Free text is the weakest constraint available. Any sequence of characters satisfies it, so every wrong answer is also a valid one.

An enumeration changes the failure mode. With six permitted payment terms the model either picks correctly or picks a wrong member, and a wrong member is a category error you can find by counting rows. Before I enumerated that field I had four spellings of the same thirty day term inside a single month of documents, because the suppliers spell it four ways and the model was copying them faithfully.

Anything with a fixed vocabulary should be enumerated. Status. Document type. Currency. Carrier.

Then the trap, which is the required field problem again in different clothes. An enumeration with no other member and no null forces a pick, so the model picks the nearest neighbour and hands you a confident miscategorisation where a visible gap would have been more useful.

Pair them. An enumeration that includes other, plus a free text field carrying the raw value when other is selected. The enumerated column stays clean enough to group by and the odd cases stay legible instead of being flattened into whichever category sat closest.

Two levels deep, no more

Output quality degrades as the structure gets deeper. I have no clean number for that and anyone quoting you one is guessing, but the direction has held across every model I have run this on.

An object inside an object inside an array inside an object asks the model to hold a lot of bookkeeping while doing the actual work.

Two levels is my limit now. An array of records with five flat fields each comes back more accurate than one record nested five deep, and it is easier to validate at the other end. The same logic applies to enumerations by size: six members is fine, four hundred product codes is a lookup problem wearing a schema.

Field names are prompt text

The model reads your field names and your descriptions. They are instructions, and vague ones produce vague values.

A field called amount gets you whichever number the model liked. A field called total including tax, with one line of description saying which figure to take when the document shows several, gets you the figure you wanted.

I write the descriptions before anything else now, in the register I would use writing a note for somebody doing the job by hand.

Ask for the span

The most useful field I have added to any schema asks for the exact text a value came from. A short verbatim quote from the source, per extracted value.

Then a plain string search, no model involved, checking the quote appears in the document character for character. If it does not, the value was invented and the record gets dropped.

That check is deterministic, costs nothing, and would have caught the purchase order problem in the first hour instead of the third week.

It has a limit worth stating. It works for extraction, where the answer is supposed to already be in front of the model. For anything the model is working out instead of copying there is no span to point at, and you are back to trusting it.

Rules the schema cannot hold

The schema is the floor. Correctness gets established above it, in code you write.

Dates that must fall inside a window. A total that must equal the sum of the line items. A currency code restricted to the four I actually deal in. An identifier that must already exist in a table I own.

None of that is expressible in a type system, and all of it is a handful of lines you were going to write anyway. In my extractor the line item sum check finds more genuine errors than every other rule put together. It is arithmetic. It has no view on how confident the model sounded.

A retry that says why

The common failure path is a try, a catch, and a resend of the identical prompt. Same input, same instruction, same settings, different result expected.

Sampling is stochastic so you do sometimes get one. But when a document genuinely lacks a field, or the model has genuinely misread a table, the unchanged retry mostly reproduces its own previous answer and you have paid for it twice.

Put the specific error into the retry. The line total does not match the sum of the items, and here are both numbers. The currency code fell outside the permitted set. That is a new prompt, with something in it the model can act on.

Cap the attempts at two and send the rest to a queue a person reads. A loop retrying against a document that does not contain the field will never converge, and I have watched one spend a night proving that.

The threshold that did nothing

I put a confidence field on every record for a while. The model’s own score, with a threshold under it sending anything below eight out of ten to manual review.

It achieved nothing. The scores clustered in a narrow band near the top and tracked correctness almost not at all. The fabricated purchase order numbers scored high, which makes sense in hindsight, since the model had no more doubt about those than about anything else it produced that day.

Self reported confidence as a number is close to useless to me now. The binary version earns its place: did you find this, and where. That question has an answer.

The part I actually got wrong is smaller and worse. I ran that threshold for a month before I checked whether it separated anything. I had shipped a control and never verified that it controlled.

Per tool notes on which providers enforce a schema during generation rather than validating after the fact, with current pricing, are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →