What prompt caching actually saves you
What prompt caching actually does
Every time you call a model API, you pay for every token in the prompt, not just the new part. If your system prompt is 3,000 tokens of instructions, tool definitions, and a few examples, and you send that same prompt 200 times a day, you’re paying to reprocess those 3,000 tokens 200 times. Prompt caching lets the provider skip most of that reprocessing on repeat calls, and pass some of the savings back to you as a cheaper token rate.
Mechanically, it works because inference over a given prefix produces the same internal state every time, provided nothing before it in the prompt has changed. If you send the same 3,000 token block first, followed by a different question each time, the provider can store the computed state for that block and reuse it instead of recomputing it from scratch. That’s the whole trick. Underneath, it’s a memoization layer sitting in front of inference.
This matters for the prompt caching cost question because the discount only applies to the part of the prompt that’s byte for byte identical to something cached earlier, and only within the cache’s time window. Change one word early in the prompt and the cache breaks for everything after it.
Two prices, not one
The number that gets quoted on pricing pages is the cache read discount, and it’s real: reading from cache is meaningfully cheaper than a normal input token. What gets left out of the pitch is that writing to the cache in the first place usually costs more than a normal input token, not less. The first call that populates the cache carries a premium, and only later calls that hit that cache get the discount.
So the actual arithmetic is: the first call is more expensive than normal, every call after that within the cache window is much cheaper, and if nothing ever reuses that cache, you paid the premium for nothing. Providers differ on the details. Some apply caching automatically once a prompt crosses a token threshold with no separate write charge, others require you to explicitly mark the cacheable portion, and exact rates move often enough that I won’t put a specific number in this piece beyond the shape of it: writes cost more, reads cost less, check your provider’s current docs before you budget against either one.
Where it pays off
Caching earns its keep when the same large block of tokens gets reused many times before it expires. The workloads where this works share a pattern:
- A coding assistant that sends the same repo context, file tree, or style guide on every turn of a session
- A RAG pipeline where the retrieved chunks stay the same across a batch of related queries
- An agent loop where the tool definitions and system prompt are fixed and only the latest turn changes
- Long few shot prompts reused across many requests inside the same session window
In all of these, the ratio of stable prefix to new content per call is high, and the number of calls that reuse the same prefix within the cache window is also high. That second condition matters as much as the first. A 10,000 token cached system prompt used twice is worse than not caching at all, because you paid the write premium and only collected one cheap read against it.
Where it does not
Caching quietly loses money in a few specific patterns I’ve watched teams walk into.
Low reuse. If your prompt structure changes every call, because you’re injecting different retrieved context each time or building the prompt dynamically with variable length data early in it, there’s no stable prefix to hit. You pay write costs repeatedly and never collect a read discount.
Short cache windows against slow traffic. Cache entries expire, typically in minutes rather than hours unless you pay for an extended window. If your app gets a burst of requests followed by a quiet stretch, by the time the next request lands the cache has already evicted and you’re back to paying write price.
Caching the wrong part of the prompt. The savings only apply to whatever sits in the cached prefix. Teams sometimes cache the small dynamic part and leave the large static part uncached because of how they built the request, which gets the discount backwards. The big stable block, tool schemas, style guides, long context documents, is what needs to sit at the front of the prompt and stay identical byte for byte.
Small prompts. Most providers have a minimum token count before caching is even available, often in the low thousands of tokens. A 400 token system prompt isn’t going to hit the caching path at all on most APIs, so there’s nothing to save.
The break-even math you actually need to run
Before adding caching to a pipeline, work out three numbers against your own traffic: the size of your stable prefix in tokens, how many times that exact prefix gets reused before the cache window expires, and the write-versus-read price ratio your provider publishes. If a prefix gets reused only once or twice per window, the write premium can eat the entire benefit or leave you slightly behind where you started. If it gets reused a few dozen times, the read discount dominates and the write premium becomes a rounding error.
This is worth calculating with your own logs rather than assuming caching is free money. A pricing page advertising a steep discount on cached reads is describing the read price relative to a normal input token, not describing your total bill relative to what you were paying before you turned caching on. Those are different numbers, and conflating them is how teams end up surprised when the bill doesn’t drop as much as they expected.
What this means for RAG and coding assistants specifically
RAG pipelines that re-rank or re-retrieve on every query, so the injected context genuinely changes each call, get less out of caching than teams expect going in. The part that stays fixed, retrieval instructions, output format rules, tool definitions, is usually a small fraction of the token count next to the retrieved chunks themselves. Cache that fixed part and you still pay full price for the chunks, which are often the majority of the prompt.
Coding assistants and agent frameworks tend to be a better fit, because the repo context, file contents already sent, and system instructions stay constant across a multi turn session while only the latest user message and tool result change. That’s close to the ideal shape for caching: a large stable prefix, a small delta per call, and many calls per session before anything expires.
The bottom line
Prompt caching cost isn’t a flat discount, it’s a bet that you’ll reuse a stable prefix often enough, soon enough, to make up for paying more on the first call. It works well for session based agent and coding workflows with a large fixed context and frequent turns. It works poorly for pipelines that rebuild the prompt from scratch each time or that see long gaps between calls to the same context. Before you flip it on for a workload, count your actual reuse rate against your provider’s actual cache window and write premium. If you can’t point to a specific number of repeat calls per cached prefix, you’re guessing, and guessing is how a cost optimization quietly turns into a line item that makes the bill worse.
For more breakdowns of what AI tools actually cost to run in production, head back to the AI Tool Gazette homepage.