← all articles

How to read an LLM provider pricing page properly

Why the sticker price is not the price you pay

Every LLM provider publishes a pricing page with a table of dollar amounts per million tokens, and almost nobody reads it the way it needs to be read. You skim the input price, maybe glance at output, and mentally file the model as “cheap” or “expensive.” Then a month later your bill is three or four times what you expected, and you go back to that same table trying to figure out what you missed.

Nothing on the page was wrong. You just didn’t read it as a cost model, you read it as a sticker price. Those are different things. A sticker price is one number. A cost model has inputs you control, multipliers you don’t always notice, and a few line items that only show up once you’re already in production. This piece walks through how to actually read one of these pages before you write a line of code against it, using the mechanics that are common across providers rather than any specific vendor’s current numbers, since prices move and I’m not going to hand you figures that could be stale by the time you read this.

Input tokens and output tokens are not the same product

The first split on every pricing table is input versus output, and the gap between them is usually wide, often two to five times. That’s not an accident of billing design. Generating output requires the model to run a forward pass for every single token it produces, one at a time, autoregressively. Reading input can be processed in parallel across the whole prompt in one pass. The compute profile is different, so the price is different.

This matters for how you architect a feature, not just how you estimate a budget. A summarization tool that ingests a 10,000 word document and returns three sentences is input heavy and cheap per call. A code generation tool that takes a short instruction and returns 2,000 lines is output heavy and expensive per call, even though the prompt looks tiny next to the summarizer’s. If you’re comparing two models by looking only at the input price column, you’re comparing the wrong thing for an output heavy workload. Pull up your own logs, get a rough ratio of input to output tokens for your actual use case, and weight the comparison by that ratio, not by the sticker alone.

Context window size is not a line item, but it acts like one

Context window (the maximum tokens a model can hold in a single request) usually isn’t billed as its own line, but it drives cost in two indirect ways worth catching before you build around it.

First, a bigger window invites you to stuff more into every call. RAG pipelines are the classic case: it’s tempting to retrieve ten chunks instead of three because the window can technically hold ten. Each of those extra chunks is billed at the input rate on every single request, so a retrieval strategy that looked “free” because the model could handle it is still costing you on every call, whether the extra context improved the answer or not.

Second, some providers step the price up at certain context lengths, charging more per token once a request crosses a threshold, because serving long context is genuinely more expensive on their end. If a page has a footnote about pricing tiers by context length, that’s not boilerplate, it’s telling you the cost curve is not flat. Check whether your typical request sits comfortably under that line or regularly spills over it, because spilling over even occasionally can quietly shift your average cost per call.

Cache reads change the math more than people expect

Prompt caching is the line item that trips up the most people, because it looks like a footnote and behaves like a structural change to your bill. The idea: if you send the same prefix (a system prompt, a long set of instructions, a document you’re repeatedly querying) across multiple requests, the provider can reuse the processed representation of that prefix instead of reprocessing it from scratch, and charges a lower rate for the cached portion.

The saving only applies to the part of the prompt that’s identical across calls, byte for byte, up to wherever the match breaks. Two consequences fall out of that:

  • If you’re rebuilding your system prompt dynamically per request (say, injecting a timestamp or a per-user ID near the top), you break the cache on every single call and pay full input price with no benefit. Put dynamic content at the end of the prompt, not the beginning, if you want the static prefix in front of it to stay cacheable.
  • Cache entries expire after a window of inactivity, usually measured in minutes. An agent workflow that calls the same model with the same system prompt every few seconds gets real savings. A background job that fires once an hour gets none, because the cache has already gone cold between calls.

If your workload has a long, stable system prompt and you’re not structuring calls to keep it at the front of the request, you’re leaving a real discount on the table. Go check your prompt construction code, not just the pricing page, to know if this applies to you.

Batch and async endpoints trade latency for a real discount

Most providers now offer a batch or async tier, where you submit a set of requests and get results back within a window of several hours instead of immediately, in exchange for a meaningful discount off the standard rate. This is not a gimmick line item, it reflects a real difference in how the provider schedules compute: batch jobs can be slotted into idle capacity instead of competing for real-time throughput.

The catch is obvious but people miss it anyway: batch only works for workloads where nothing downstream is waiting on the response in real time. Nightly data enrichment, bulk classification, generating a backlog of descriptions or summaries, evaluation runs against a test set. If you’re building a chat interface, batch pricing is irrelevant to you no matter how good the discount looks, because a user isn’t going to wait three hours for a reply. Don’t let a good discount number pull a synchronous product toward an async endpoint it can’t actually use.

Free tiers and credits hide the real per-request economics

Free tiers and signup credits are useful for prototyping and genuinely bad for estimating what a feature will cost once it ships. A free tier usually caps you on requests per minute or tokens per day, not on total spend, so your prototype runs smoothly right up until you turn on real traffic and hit a wall that has nothing to do with the per-token price you budgeted around.

Before you commit to a model based on a prototype that ran entirely on free credits, do the arithmetic on the paid rate at your expected production volume. It sounds obvious written out like that, and it’s still the single most common estimation mistake I see, because the free tier genuinely feels like “the price” while you’re building against it.

Rate limits gate how much you can actually spend, in either direction

Every provider tier comes with a rate limit, expressed as requests per minute or tokens per minute, and it usually rises as your account’s spend history grows. This isn’t a pricing line item exactly, but it belongs in the same read-through, because it determines whether the price on the page is even reachable at the volume you need. A model that’s cheap per token but rate limited to a level below your traffic isn’t a cheap option for you, it’s an option that requires a support ticket and a wait before it’s usable at all. Check the rate limit tier tied to the pricing tier you’re reading, not just the per-token number in isolation.

A five minute checklist before you commit code to a model

Before wiring a model into a feature, walk the pricing page with these questions in hand: what’s my actual input to output token ratio for this workload, does my typical request sit under or over any context length pricing threshold, is my system prompt structured to actually hit the cache, does this workload have a real-time constraint that rules out batch pricing, and does the rate limit at my spend tier support my expected volume. None of these require a benchmark or a vendor’s blessing, they’re just the page read against your own use case instead of read as a single sticker number.

The takeaway

A pricing page is a cost model wearing a table’s clothing. The number that matters isn’t the price per million tokens in isolation, it’s that price multiplied by your actual token ratio, adjusted for whether you’re hitting cache, and checked against whether batch pricing is even on the table for your latency requirements. Read it that way once, and the next bill stops being a surprise.

If you want more breakdowns like this, head back to the AI Tool Gazette homepage.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →