← all articles

When to put a human in the loop

human-in-the-loop ai-review llm-operations automation-design

I went looking for the rejection rate on my outreach approval queue and found that I had never recorded one. Four months of approvals, sitting in a chat thread, with no count of how many drafts I had sent back.

I could not tell you the number today. What I can tell you is that the true figure was somewhere near zero, because I remember the batches and I remember clearing thirty of them in a couple of minutes on a Tuesday.

One of those went to a company that had shut down more than a year earlier, opening with a friendly line about following their progress.

That queue existed so a person would see every email before it left. A person did see them. Seeing is not the same as reading, and I had built no way of telling the difference.

The count that tells you whether the review is real

Count rejections. Not reviews, not throughput, not how many items a reviewer cleared per hour.

A rejection rate flat at zero has exactly two readings, and both of them condemn the setup you have. Either the model is good enough that the review is buying you nothing, so stop paying for it. Or the model is not that good and the review is catching nothing, so stop telling people it exists.

I will take an argument on this. An approval step where nobody ever rejects anything is theatre, and it is the most common form the phrase human in the loop takes in practice. It gets built so there is a person in the diagram when somebody asks who checks this.

It costs twice. You pay for the reviewer’s hours, and you walk away believing the output has been checked, which is worse than knowing plainly that it has not.

The reviewer is not the problem

The instinct when a reviewer starts waving things through is to talk to the reviewer. That will not work, and it will not work with the next person either.

Put someone on a watch where nearly every item is fine and the rare one matters, and their detection rate drops inside the first half hour. This was measured on radar operators in the 1940s and has been measured since on baggage screeners and factory inspection lines. The variable driving it is how often the bad item actually appears. Below a certain rate, people stop catching them, and they stop catching them regardless of how much they care.

A model that is right nine times out of ten produces precisely that watch. You have constructed a low prevalence vigilance task and then asked a person to be the safety layer on top of it.

So the useful question was never whether to add a person. It is what job that person is being handed.

Ask for one named error

The cheapest fix, and the one people leave out, costs nothing at all.

Do not ask anybody to check whether the output is good. Ask them to check one specific thing.

My outreach reviewer, who is me, has one question: is this company still trading, and does the opening line describe something that is genuinely on their site. Nobody is being asked to assess the writing.

The invoice extractor has one question: does that purchase order number appear on the page.

The article queue has two: does every internal link resolve, and is the claim in the second paragraph one I can defend if a reader pushes back.

A narrow question has an answer and you can see from outside whether it got answered. Be attentive has neither property. The narrow version also survives boredom far better, because looking something up is a different cognitive act from forming a judgement, and it degrades more slowly.

Shrink the pile before you show it

Second fix: stop showing the reviewer everything.

If the system can flag which outputs it is unsure about, the queue shrinks to those. That raises the share of genuinely wrong items in front of the reviewer, which is the exact variable the vigilance problem turns on. You are no longer asking for sustained attention across a stream of correct output.

Asking the model for a confidence score does not get you there. Those numbers cluster near the top and track correctness weakly, which I have written about at length in the context of schema design and structured output.

Deterministic checks do get you there. Line items that do not sum to the stated total. An identifier absent from a table you own. A currency outside the four you deal in. A response that only arrived after two failed parses. A retrieval where nothing scored above your floor.

Those are a few lines each. They are also the difference between a queue of six hundred and a queue of eleven, and eleven is a queue somebody will finish.

The interface is the policy

Whatever the written process says, the buttons decide what happens.

If approve is a single keystroke and reject means writing a paragraph of justification into a text box, rejection has been priced above approval and under load you will get approvals. Make both directions cost the same. One key for yes, one for no, and if you want a reason then offer four buttons rather than a free text field.

Then add friction to the skip. Bulk approve, select all, the approve everything control tucked in the corner. Mine had a select all and I used it constantly, which is what happens when a control that clears the whole screen is available.

The strongest version of expensive to skip I have used is asking the reviewer to type in a value that only exists inside the output. The invoice total. The last four characters of a reference number. It feels petty while you are building it. It works, because there is no way to supply it without having looked.

Arrangements worth pulling out

A single approve control over a batch, whatever the batch contains. That is a signature.

A queue arriving faster than a person can read it. Do the arithmetic first. Four hundred items a day at ninety seconds of honest reading each is a ten hour job, and it will be done in two. The queue does not care and the items ship either way.

A reviewer used as liability cover. This one gets built in meetings rather than in code. Nobody trusts the system, nobody will say so, so a person is added to the diagram and the risk becomes formally theirs. The model still writes the output. The reviewer still cannot read four hundred of them. What changed is whose name is attached afterwards.

If the honest reason for the review step is that the system is not ready to ship, say that. The review will not fix it, and it hides how bad things are, because rejections nobody counts look identical to a system with no errors in it.

Decide by what a wrong output costs

Here is the criterion I use now, and it has nothing to do with how important a feature feels in a planning conversation.

What does one wrong output cost when it gets through, and how hard is it to undo.

The tagger in my video pipeline picks footage for each paragraph of narration. When it gets one wrong, a viewer watches a shot of a server rack while I talk about a SIM card. I have never reviewed that queue and I do not intend to. A person standing over it would cost more in a week than every bad clip it has produced in its life.

That is the honest half of this that rarely gets said. For cheap, reversible errors the review costs more than the errors do, and declining to review is the correct call rather than a lapse.

The other end looks different. An email going out under my name to a real person cannot be recalled. A credit adjustment on a customer’s account moves money. An article published under my byline sits in the index and gets read for years, and its failure mode is slow and quiet enough that I would not notice on my own.

Those get a reviewer. Each one with a named error to hunt for and a queue short enough to finish.

Change the failure direction instead

Before adding anybody, check whether you can move the exposure.

My mail triage used to file uncertain messages into whichever folder scored highest. A real customer landing in the wrong folder meant a slow reply, which is reversible and irritating and not worth a person supervising. So I changed what it does when it is unsure: uncertain mail goes to me, unfiled, instead of being sorted confidently.

No reviewer, no queue, no approval step. The cost of being wrong went down because the system stopped committing when it did not know.

That option is available more often than the review option, and it is cheaper every time.

What I have not solved

Once a person sits on the end of a pipe, everyone upstream relaxes. I shipped a prompt change to the outreach tool without testing it properly because I knew I would see the output before it went anywhere. That was rational of me and it is also how a review step ends up carrying weight nobody designed it to carry.

I have no fix for that beyond naming it, which is what this is.

The other limit is that all of the above assumes you can see your rejection rate. If review happens in an inbox or a chat thread, you cannot, and moving it somewhere countable is the first job rather than the last. Mine ran in a thread for four months and I had no idea what my own numbers looked like.

The routing checks I use to decide what reaches a review queue, and how I track whether a review step is doing anything, are here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →