← all articles

Guardrails that stop the bad output without annoying users

The two ways guardrails fail

Most teams find out their guardrails are broken in one of two ways. Either something bad slips through (a jailbreak gets a refusal bypassed, the model hallucinates a refund policy that doesn’t exist, a RAG pipeline echoes a chunk of someone else’s private document) or the guardrail is so trigger happy that support tickets start piling up with “why did it refuse to answer a normal question.”

Both failures come from the same root cause: treating “guardrail” as one filter instead of a set of checks with different jobs. A single regex blocklist or a single moderation API call can’t do input validation, prompt injection defense, and output correctness checking at once. When you try to make it do all three, you get a filter tuned so loose it misses real problems, or so tight it blocks real users, and usually both at different times of day.

If you’re running this in production and paying per-token API bills, the fix isn’t a smarter single filter. It’s putting the right check at the right stage, and being honest about what each stage can and can’t catch.

Input side versus output side

Guardrails split cleanly into two families, and conflating them is where most implementations go wrong.

Input guardrails run before the model call. They’re cheap because they don’t require an LLM at all: allowlist/denylist matching on user input, length caps, a classifier trained specifically to detect prompt injection patterns, or a check for known jailbreak phrasing. Their job is narrow: decide whether to send the request to the model at all, or whether to strip/flag something in it first (like instructions embedded in a pasted document that try to override the system prompt).

Output guardrails run after the model responds, before the response reaches the user. This is where you catch hallucinated facts against a source of truth, PII leakage, formatting violations, or a response that drifts off the system prompt’s intended scope. Output checks are more expensive because the thing you’re checking is variable length and semantically fuzzy, which usually means a second model call (a smaller, cheaper model judging the first model’s output) rather than a static rule.

The mistake I see constantly: teams put all their guardrail budget on the input side because it’s cheaper and easier to reason about, then wonder why the model still says something wrong. Input guardrails can’t catch hallucination. They never see the output. If your failure mode is “the model made up a policy,” you need an output check, full stop.

Why regex loses to structured output

A lot of guardrail effort goes into pattern matching on free text: block this phrase, flag that keyword. This works for known bad strings and fails for anything novel, which is most of what actually goes wrong. A model doesn’t need to say a blocked word to give a wrong answer.

The more durable fix, where it applies, is constraining the shape of the output instead of policing its content after the fact. If you’re asking the model to return a decision, a category, or a structured field (say, a support bot returning an action like refund, escalate, or answer), use JSON schema constrained generation or function calling so the model literally cannot emit a value outside your enum. This isn’t a guardrail bolted on after generation, it’s a constraint baked into decoding. There’s no free text to slip a bad value through, because the token sampler is restricted to valid continuations at each step.

This doesn’t solve everything. It solves the “wrong shape” and “invalid category” class of problems, not the “right shape, wrong facts” class. For factual correctness you still need a separate check, usually a retrieval grounding comparison (does the claim in the output trace back to a retrieved chunk) or a second model call asking “is this claim supported by the provided context, yes or no.” That second call costs you another round trip and another set of tokens, which is the real tradeoff nobody puts in the architecture diagram: every output guardrail you add is latency and money, paid on every single request, not just the ones that would have failed.

The false positive tax

This is the part that gets skipped in guardrail writeups. A guardrail that blocks 100% of bad outputs and also blocks 5% of good ones isn’t a win, it’s a support burden you’ve shifted from “the AI said something wrong” to “the AI refused to help and now a human has to explain why.”

The instinct when a bad output gets through is to tighten the filter. The instinct when a filter blocks something reasonable is usually to leave it alone, because loosening a safety check feels riskier than a support ticket. Left unchecked, this ratchets guardrails tighter over time until the assistant is unusable for edge cases that were never actually dangerous, just unfamiliar to whoever wrote the rule.

The way out is treating false positives as a tracked metric, not an anecdote. Every time a guardrail fires, log what triggered it and whether a human later confirmed it was a correct block. If you don’t have that data, you’re tuning blind and defaulting to “tighter is safer,” which is how you end up with a bot that refuses to discuss anything remotely adjacent to a blocked topic.

Layer checks by cost, not by paranoia

A workable structure looks less like one filter and more like a funnel, cheapest checks first:

  1. Static input checks (length, format, known injection strings). Near zero cost, runs on every request.
  2. A lightweight classifier for intent or injection risk, only for requests that pass step one. This is a small model call, cheap relative to the main generation.
  3. The main model call, constrained to a schema wherever the task allows it.
  4. An output check against source material, only for claims that need grounding (not every response needs this, a greeting doesn’t need fact checking).
  5. A human review queue for anything that hits a confidence threshold in step 4, not an automatic hard block.

Step 5 matters more than it looks. Automatic hard blocks are what create the annoying refusals. A queue that holds a borderline response for a few seconds, or degrades to “let me check that and follow up” instead of an outright refusal, protects the user experience while still catching the genuinely bad case. Not every guardrail failure needs to be a wall. Some of them just need to be a pause.

Log everything the guardrail touches, not just the blocks

If you only log the requests a guardrail rejected, you can’t tell the difference between a filter that’s working and one that’s silently letting things through because the trigger condition was written too narrowly. Log the classifier’s confidence score on every request, blocked or not, along with the final action taken. Over a couple of weeks you’ll have a real distribution to look at instead of a gut feeling, and you can move the threshold based on where the false positives and false negatives actually cluster, not where you assumed they would.

This is also the only way to catch guardrail drift. Models get updated, user phrasing shifts, and a classifier tuned against last quarter’s traffic starts missing patterns it used to catch. Without a log of the full distribution, that drift is invisible until something gets through and someone notices the hard way.

The part that’s actually a UX decision

Not every guardrail question is technical. What the user sees when a block happens is a product decision as much as an engineering one. “I can’t help with that” with no explanation reads as broken. A response that names the category it flagged, or offers an alternative framing, reads as intentional even when it’s still saying no. The mechanism catching the bad output and the message shown to the user are two separate design surfaces, and treating the second one as an afterthought is usually why a technically correct guardrail still generates complaints.

Guardrails that work in production aren’t the ones with the strictest filter. They’re the ones tuned against real traffic, split across input and output stages by what each stage can actually see, and layered so the expensive checks only run on the requests that need them.

If you’re working through this kind of production tradeoff on your own stack, you can find more breakdowns like this on the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →