← all articles

Testing an AI feature before you ship it

llm-testing regression-testing prompt-engineering shipping-ai

I edited one sentence in the prompt behind my invoice extractor on a Thursday. Eleven days later I worked out it had started reading the delivery date into the invoice date field on roughly one document in six.

Nothing failed in that time. Every record parsed, every required field was present, every type was correct, and the pipeline reported a clean run each night.

There was also nothing to run before that deploy. The repository has over a hundred tests covering the queue, the retry path, the writes and the routing. The prompt sits in a file by itself with nothing attached to it, which is how it looks in nearly every codebase I have been inside.

This is about a feature you are building and will keep changing. There is a separate piece on this site about evaluating a third party tool before you buy it, and that job ends when you sign. This one runs for as long as the feature is in production, and its question is narrower: did what I changed this morning break something that worked yesterday.

The suite you know how to write does not apply

Same input, different output. That is the property, and it removes the assertion sitting at the centre of every test any of us has written.

The obvious move is to pin temperature at zero and go back to exact matching. It half works. Batched inference on shared hardware does not guarantee identical tokens, and more to the point providers update models behind an unchanged name, so one morning your suite goes red in forty places for something you did not do.

A suite that goes red without your involvement gets switched off. Nobody announces it. Someone adds a skip, someone stops opening the output, and within a month the thing is decoration.

So exact matching is worse than absent. It teaches you to disregard a red run, which is the only behaviour a test suite exists to produce in you.

So give up on knowing the answer and assert what is true of any acceptable answer.

Shape is the floor

Did it parse. Are the keys present. Are the types right. Does the field you called a date survive a date parser.

Cheap, fast, catches the loud failures. There is a piece on this site about schema enforcement and the ways a valid document can still be false, so I will leave it there. For a test suite, shape is a precondition. Passing it is not a result you report.

The invoice run above cleared every shape check for eleven days while putting the wrong date in the right field.

Invariants, which is where the work is

An invariant is a statement that must hold whatever the model wrote. You do not need the correct answer to check one. You need one property the correct answer has.

Some of these you may already run in production as validators. A total equal to the sum of its line items, a currency inside the set you deal in, an identifier that exists in a table you own. Those belong in both places, and the copy in your test suite is the one that stops a bad deploy.

The ones that only make sense in a test suite are the interesting half.

A summary is shorter than its input. Obvious until it fires, which mine did in week one, on short dense inputs the model padded past their original length.

Every number in the output appears somewhere in the source text. Same for any url. A link that was not in the material does not exist.

A quoted span appears in the source character for character. Plain string search, no model in the loop, no judgement. Fail the record.

The answer is in the language of the input.

Then the cross run checks, available only because you own a fixed set of cases. Run the same input twice and get the same category. Run all nineteen and confirm the distribution of labels has not moved outside a tolerance you chose. A model that has developed a new favourite label raises no error anywhere, and a distribution check finds it in a single run.

Each of those is an ordinary assertion. Deterministic, free, and when one fails it points at a line.

If I had to keep one check it would be the span search. It converts a whole class of invention into a test failure with a line number.

The cases where refusing is correct

The tests almost nobody writes are the ones where the right output is a decline.

A blank file. A question whose answer is not in the material. A document that genuinely lacks the field. An instruction the feature was never built for.

I sent my extractor a scan that turned out to be a photograph of a whiteboard, and it returned four line items with prices on them. Plausible ones. Nothing errored, the shape checks passed, and the only reason I caught it was that I happened to be watching that batch.

These get skipped because your example set is assembled out of the feature working, which is what is lying around. A refusal does not look like a case until one has cost you something.

I check refusals before accuracy now. A feature that never declines is a feature that will invent as soon as production hands it something nobody planned for.

Nineteen cases and your own eyes

Pull ten to twenty real inputs out of your traffic and write down the correct output for each. Mine for the classifier is nineteen and took a Sunday morning.

The collection method is covered in the buying piece and I am not repeating it. What differs here is the lifespan. The set gets checked in, runs on every change, and outlives every model you point it at.

The part people cut is the part that pays. For the first several runs, read every output yourself. Not the pass count. The outputs.

You are not measuring yet. You are learning what wrong looks like on this particular task. Mine turned out to be two categories quietly merging into one, in about one message in fifteen. I had been braced for fabrication and got a filing error.

After four or five passes the failure shapes start repeating, and you write invariants for the ones that repeat. Reading turns into assertions. That is the whole mechanism.

Keep it small enough to actually read. Five hundred cases is a set you will never open, and a set you never open is a score with extra steps.

Nineteen calls on a cheap model is a fraction of a cent, so run it on prompt edits, which is the trigger everyone forgets. A prompt is code. It is the code that changes behaviour most and it usually has no tests within a mile of it.

The failure catalogue costs a minute an entry

Every bad output anyone reports goes in one file. The input, what came back, and one line on why it is wrong. Mine is a text file and has never needed to be anything else.

Then it runs with the rest.

What makes it better than cases you invent is the weighting. Your imagination reaches for failures you have already considered. A complaint does not.

One rule keeps it alive: never fix a reported problem without adding the case first. The regression test arrives attached to the fix.

After a year the catalogue on one job was past a hundred entries with about thirty still failing, which sounds like a mess and is the useful part. Split it in two. A must pass file, red only when you have broken something. A known bad file, which is a written statement of what the feature cannot do, run every time so you notice when one starts passing. Then you promote it across.

That second file has answered more questions from other people than any documentation I have written.

The number you cannot act on

Skip the quality score you cannot interpret.

A model grades each output out of ten, you average, the average lands on a dashboard, and now there is a metric with a line chart under it.

Then it moves from 7.4 to 7.1 and nobody in the room can say what to do. The number has no unit. It will not tell you whether one case collapsed or twenty degraded slightly, and those want opposite responses. It will not tell you whether the model changed or the judge did.

Judges have taste as well. Mine reward length and headings, so a fixed share of any score I produced was measuring formatting. Public benchmarks carry a larger version of the same problem and there is a piece here on that, so it gets a clause.

The position, and I will argue it: twenty real cases you look at beats an automated metric you do not understand. On cost, on how fast you can respond, and because only one of them ever tells you what to change.

There is one use for a judge that I keep. Score the set, sort ascending, read the bottom five. That is a way of choosing what to open, and it stays off the dashboard.

Where this stops working

My first invariant capped a summary at a quarter of its input length. I picked the number in about four seconds.

It fired on good outputs roughly one time in thirty, because some inputs arrive dense with nothing to cut. Two weeks later I was rerunning the suite rather than reading the failure, which is the habit I have spent this piece complaining about. A check that is wrong one time in thirty trains you faster than it protects you. I moved it to a half and it has caught two real regressions and cried wolf zero times since.

The larger gap I have not closed is that none of this works when correctness is a matter of taste. Tone, house style, whether a draft is any good, whether a reply reads like a person wrote it. No invariant holds, refusals do not apply, and a judge on those is a second opinion with a number attached.

For that work I put a person in the loop and pay for it, or I ship something unmeasured and say so out loud. I do not have a third answer and I distrust anyone selling one.

The invariant list I start new features from and current per token prices for running a suite like this are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →