← all articles

When caching answers beats calling the model

The bill that made me look at caching

Somewhere around the third month of running a RAG-backed support bot in production, I pulled up the API dashboard and stopped scrolling. A huge chunk of spend was going to questions we had already answered, sometimes minutes apart, sometimes with the exact same wording. The model was doing full retrieval, full context assembly, and a full generation pass for a question it had already solved that same hour. That’s the moment caching stops being a “nice to have” architecture note and starts being a line item you can point to.

Caching an LLM response is not exotic. It is the same idea as caching a database query or a rendered web page: if the inputs are the same, skip the expensive work and hand back the stored output. What makes it interesting for LLMs is that “the same inputs” is a fuzzier concept than a SQL query string, and the cost of getting it wrong (serving a stale or wrong answer with total confidence) is higher than a slightly outdated product listing.

Three different things people mean by “cache”

When people say “we added a cache” to an LLM pipeline, they usually mean one of three distinct mechanisms, and mixing them up leads to wrong expectations.

Exact-match caching hashes the literal request (prompt, model, parameters) and stores the response keyed on that hash. Next time the same hash shows up, you skip the model call entirely. This is cheap to build, cheap to reason about, and only fires when someone sends the identical prompt again. In practice this happens more than people expect: FAQ-style chatbots, coding assistants completing common boilerplate, support tickets that funnel into a handful of canonical questions. It does nothing for “how do I reset my password” versus “how can I reset my password,” because those hash differently.

Semantic caching fixes that gap by embedding the incoming query, comparing it against embeddings of past queries, and treating anything above a similarity threshold as a cache hit. This catches paraphrases. It also introduces a new failure mode: two questions can be semantically close and still need different answers. “Can I cancel my subscription” and “can I pause my subscription” sit close together in embedding space but should not share a response. Every semantic cache lives or dies on where you set that threshold, and tuning it is an ongoing job, not a one-time config value.

Prompt-prefix caching (sometimes called context caching or KV caching, depending on the vendor) is different from the other two because it doesn’t skip the model call, it skips redoing work inside the call. When a prompt starts with a long, unchanging block (a system prompt, a big document, a tool schema) the model can reuse the internal key-value state it already computed for that prefix instead of reprocessing it token by token. This matters most for RAG pipelines and agents where the same large context gets reused across many turns with only the tail end of the prompt changing. It’s a latency and compute win on the fixed part of the prompt, not a full response cache, so it doesn’t help when every request is genuinely novel end to end.

Knowing which of these three you’re actually building changes everything from the storage layer to the invalidation logic, so it’s worth being explicit about which one a given feature needs before writing code.

Where the win is real

The strongest case for caching shows up when three conditions line up: the query space is repetitive, correctness doesn’t hinge on freshness, and the underlying model call is expensive relative to a cache lookup.

Support and documentation bots are the clearest example. A meaningful share of user questions cluster around a small set of intents, and the answer to “what’s your refund policy” doesn’t change between 9am and 9:05am. Coding assistants get a similar benefit on boilerplate completions and common library usage patterns, where the same snippet of code gets requested by different developers on different days. RAG systems over static document sets benefit twice: once from caching the retrieval step (don’t re-embed and re-search for a query you’ve already resolved) and once from caching the generation step if the retrieved context and question are close enough to something already answered.

The economics also depend on where the expense sits. If your pipeline does retrieval plus reranking plus generation, and the model call is the smallest part of that chain, a cache hit saves you the whole chain, not just the generation cost. That compounds the win beyond whatever the per-token savings look like.

Where it quietly backfires

The failure mode I’ve seen bite teams isn’t “caching didn’t help.” It’s “caching served something confidently wrong and nobody noticed for a while.” A few specific traps:

Personalized or account-specific answers get cached across users if the cache key doesn’t include enough context. If your key is just the question text and not the account, plan, or role, you can serve a cache hit meant for one user’s entitlements to a different user entirely. This is a correctness bug, not a performance one, and it’s the kind that erodes trust fast.

Time-sensitive facts rot silently. Pricing, inventory, policy text, anything tied to “current state” needs either a short TTL or an explicit invalidation hook tied to the source system changing. A cache with no expiry on this kind of data doesn’t fail loudly, it just gets quietly wrong and stays that way until someone complains.

Low repetition rates make the infrastructure not worth it. If your query distribution is long-tailed (most questions are asked once and never again), you’ll pay for embedding lookups, storage, and cache-maintenance code and get a low hit rate back. Before building a semantic cache, it’s worth actually looking at your query logs and asking how many requests in the last month were near-duplicates of an earlier one. If that number is small, the engineering effort is better spent elsewhere.

Semantic drift in the threshold is the subtler version of the same problem. A threshold tuned on last quarter’s query patterns can start misfiring as your product surface changes and users start asking about features that didn’t exist when you set it. Semantic caches need periodic review of what’s actually matching, not just a one-time similarity cutoff.

A rough decision path

When I’m deciding whether a given endpoint deserves a cache layer, I look at the query logs first, not the architecture diagram. If a meaningful fraction of requests are literal or near-literal repeats, exact-match caching is close to free to add and has almost no downside beyond storage. If the repeats are paraphrases rather than exact strings, semantic caching can help, but only after you’ve looked at real examples of what would and wouldn’t have matched under a candidate threshold, because the threshold you’d guess from intuition is usually wrong in one direction or the other.

If the pipeline reuses a large, static context across many calls, prompt-prefix caching is worth checking regardless of how repetitive the user-facing questions are, since it targets the fixed part of the prompt rather than the whole exchange. And if the data behind the answer changes on any meaningful cadence, the cache needs an explicit expiry or invalidation path tied to that change, not a “we’ll clear it manually if it looks wrong” plan, because nobody remembers to do that consistently.

None of this replaces monitoring. A cache that silently serves stale or misrouted answers is worse than no cache, because it looks like the system is working. Log cache hits separately from fresh generations, sample them periodically, and treat a rising hit rate as a metric to investigate, not just celebrate.

Caching an LLM response is one of the few optimizations in this space that pays off in both cost and latency at the same time, but it only pays off where the query pattern actually supports it. The teams that get burned are usually the ones that added a cache because it sounded like the responsible thing to do, not because they looked at their own traffic first.

More breakdowns like this, on the parts of shipping with LLMs that don’t show up in the marketing copy, are on the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →