← all articles

Routing the easy queries to a cheaper model

If you’re running any LLM feature at real volume, you’ve probably had the moment where you open the billing dashboard and wince. The instinct after that moment is almost always the same: stop sending every request to the biggest model you have, and start sending the easy ones somewhere cheaper. That’s model routing. It sounds simple. It mostly is simple, at the concept level. Getting it right in a live system is where the actual engineering lives.

This isn’t a pitch for any particular router product or model family. It’s a walkthrough of how routing works mechanically, what it actually buys you, and where teams get burned trying to save money on inference.

The basic idea

Every LLM API call you make has a cost tied to two things: which model answered it, and how many tokens went in and out. Flagship models cost more per token than smaller models in the same family or from competing labs, because they’re bigger, slower to run, and more expensive to serve. That gap is real and it’s persistent across the industry: a small, fast model is routinely an order of magnitude cheaper per token than the frontier model sitting next to it in the same provider’s lineup.

The problem is that most production traffic doesn’t need the frontier model. A support bot answering “what’s your refund window” doesn’t need the same model that’s debugging a race condition in a Kubernetes operator. A classifier task, a summarization of a short document, a yes/no extraction from structured text: these are things a smaller model handles fine, often at comparable quality to the expensive model, because the task itself doesn’t require deep reasoning.

Routing is the practice of deciding, per request, which model tier actually needs to handle it, instead of defaulting every request to your most capable (and most expensive) option.

How a router actually decides

There are three common approaches, and they trade off complexity against accuracy.

Heuristic routing. You write rules based on things you can measure cheaply: input length, presence of code blocks, keyword matches, whether the request came from a known “simple” endpoint in your app. If a user query is under 20 tokens and doesn’t contain code fences, send it to the small model. This is fast to build and completely transparent, but it’s blunt. Query length has almost no correlation with actual difficulty. “Summarize this” and “explain why this proof of Fermat’s last theorem is wrong” can both be short.

Embedding or classifier-based routing. You train or fine-tune a small, cheap classifier (sometimes a tiny model, sometimes just an embedding similarity check against a labeled set of past queries) that predicts task difficulty or task category before the real call goes out. This is more accurate than heuristics because it’s actually looking at semantic content, not surface features. It also adds a call in the critical path, which means added latency and its own (small) cost. You’re now running two model calls for the cheap path: the router and the answer.

Cascade routing. Instead of deciding up front, you send the request to the cheap model first, then check the output against a confidence signal, refusal, or a cheap secondary check, and only escalate to the expensive model if that check fails. This avoids a separate classifier call but means every request touches the cheap model at minimum, and the failed attempts on hard queries cost you a wasted small-model call before the retry. For workloads where most traffic really is easy, this ends up cheaper in aggregate than always calling the expensive model, even accounting for the wasted cheap calls on the hard fraction.

None of these approaches is universally “the right one.” Cascade routing is the simplest to reason about and the easiest to debug, which is why a lot of production systems start there and only add a classifier once they have enough logged traffic to train one properly.

What “easy” actually means

This is where a lot of routing setups go wrong: they define easy in terms of the input, not the task. Short doesn’t mean easy. Simple vocabulary doesn’t mean easy. A one-line prompt like “is this contract enforceable in Delaware” is short and plainly worded and completely wrong for a small model.

The more reliable signal is task type, not surface features of the text. Extraction, classification, formatting, short factual lookups grounded in provided context, translation of short passages: these are places where smaller models close most of the gap with flagship models, because the task doesn’t require long chains of reasoning or synthesis across a lot of context. Multi-step reasoning, anything requiring the model to catch its own errors mid-generation, long-context synthesis, and anything where a wrong answer is expensive to have shipped: that’s where you keep paying for the bigger model.

If you’re building a router, the first thing worth doing before writing any routing logic is logging your actual traffic by task type for a couple of weeks and looking at the distribution. Teams that skip this step tend to route on gut feel and end up either sending too much to the expensive model (no savings) or sending too much to the cheap model (quality complaints that take weeks to trace back to a routing decision, because the failure mode is “answer is subtly worse,” not “answer is missing”).

The costs routing introduces

Routing isn’t free. It trades API spend for a few other kinds of cost that are easy to underweight when you’re staring at a billing chart.

Latency. A classifier-based router adds a network round trip before the real work starts. A cascade router can add a full failed generation before the retry. If your product has a latency budget, this matters more than the dollar savings in a lot of cases.

Misrouting. Every router gets some fraction of decisions wrong. A hard query sent to the cheap model either produces a worse answer that ships anyway, or triggers an escalation that costs you both calls. There’s no router that gets this to zero, and tightening the router to reduce false negatives (hard queries misrouted as easy) usually increases false positives (easy queries needlessly escalated), which eats into your savings.

Two systems to maintain instead of one. You now have prompt behavior, safety behavior, and output format expectations that need to hold across two (or more) different models. A prompt tuned against a flagship model’s instruction-following often behaves differently on a smaller model. If you’re not testing both paths on every prompt change, you’ll ship a regression that only shows up on the cheap-model traffic, which is usually the traffic nobody’s manually reviewing because it was “the easy stuff.”

Eval drift. If you have quality evals, they need to run against both tiers, not just the one you’re most worried about. A router that’s saving you money but silently degrading the small-model path is not actually a win, it’s a deferred complaint.

A reasonable place to start

If you’re adding routing to an existing system, start with cascade routing on the narrowest task category you can find: one endpoint, one clearly bounded task, not your whole product surface. Log every routing decision and every escalation. Give it a few weeks of real traffic before you touch the thresholds. Resist the urge to route everything on day one just because the billing chart is annoying you.

And if your volume is low, or the product is quality-critical and used by a small number of people who will notice a worse answer immediately, it’s worth asking whether routing is solving a problem you actually have. Saving a fraction of a cent per request across a thousand monthly calls is not worth the maintenance burden of running two model paths. Routing pays off at volume, where the aggregate savings are large enough to justify the added surface area. Below that, you’re usually better off just picking one model and moving on.

Model routing is a real lever, not a trick. It works because task difficulty and model cost are only loosely correlated with the surface features people usually guess at, and the systems that get value out of it are the ones that measure their actual traffic before deciding how to split it.

For more breakdowns like this on the tools and infrastructure behind shipped LLM products, head back to the AI Tool Gazette homepage.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →