Committed spend deals and the break-even point
Once your monthly LLM bill gets big enough, a sales rep shows up with a spreadsheet and a discount. That’s the moment most teams either save real money or lock themselves into paying for capacity they never use. I’ve been on both sides of that spreadsheet, so here’s how these deals actually work under the hood and how to figure out whether one makes sense for you.
Two different products wearing the same name
“Committed spend discount” gets used loosely, but there are really two distinct mechanisms hiding behind it, and they behave very differently.
The first is a spend commitment. You agree to spend at least some dollar amount over a period, usually monthly or annual, and in exchange you get a percentage off the standard metered token rate. Billing still works the way it always has: you pay per input and output token, per request, same as pay-as-you-go. The only change is the rate card and the fact that you’re on the hook for a minimum, whether you hit it through usage or not.
The second is provisioned capacity, sometimes called reserved throughput units. Azure OpenAI’s PTUs and AWS Bedrock’s provisioned throughput are the two most common examples of this shape. Here you’re not buying tokens at all. You’re renting a fixed slice of model capacity, usually expressed as tokens per minute or requests per minute, for a fixed monthly price. Whether you use 10% of that capacity or 100% of it, the bill is the same. What you get in return isn’t just a discount, it’s a latency and availability guarantee, because you’re no longer sharing a queue with every other tenant on the shared endpoint.
These are not interchangeable decisions. A spend commitment is a bet on your total dollar volume. Provisioned capacity is a bet on your peak throughput. Mixing up which one you’re evaluating is the first mistake I see teams make.
The break-even math for a spend commitment
For a straight spend commitment, the math is simple because the unit economics don’t change, only the rate does.
Say your current pay-as-you-go rate gets you a 15% discount if you commit to a fixed monthly floor for twelve months. Take your trailing three months of actual API spend, not your best month, not your projected month. If your real average is comfortably above that floor, the discount is close to free money: you were going to spend that anyway, and now it costs less. If your average sits right at the floor, run the number that actually matters: is the money you save from the 15% discount bigger than the money you’d lose in months where usage dips below the committed floor and you’re paying for tokens you didn’t use? For most teams with steady, growing usage, that answer is yes. For teams with lumpy usage, a big product launch that spikes traffic then trails off, it’s a much closer call, and I’d want at least two full quarters of stable data before signing anything longer than a quarter.
The trap here isn’t the discount, it’s the commitment term. A 12-month floor based on your current growth rate assumes that growth rate holds. If you’re mid-migration to a cheaper model, or you’re about to ship a caching layer that cuts your token volume by a third, you’re about to commit to spend you’re actively trying to eliminate.
The break-even math for provisioned capacity
Provisioned capacity is a different equation because you’re not comparing rates, you’re comparing a fixed cost against a variable one at a specific volume.
Here’s the shape of it, with made-up numbers to illustrate the mechanic, not to represent any real vendor’s pricing: suppose provisioned capacity costs a flat monthly fee and covers up to some ceiling of tokens per minute. Suppose pay-as-you-go costs a per-token rate. The break-even point is the request volume where flat_fee ÷ per_token_rate equals the tokens you’d actually push through in that period. Below that volume, pay-as-you-go is cheaper. Above it, provisioned capacity is cheaper, and it gets cheaper per token the more you push through, because the fixed cost gets spread over more usage.
The part teams skip is measuring their own tokens-per-minute curve before doing this math. Average usage tells you almost nothing here. What matters is your peak. If your traffic is flat across the day, a provisioned tier sized for your average utilization works fine. If you’ve got a launch-day spike, a batch job that runs at 2am, or an agent pipeline that fans out fifty calls in a burst, you need capacity sized for that peak, not your average, or you’ll blow through the reserved throughput and either queue behind it or fall back to the shared pool anyway, which defeats the point of paying for the reservation in the first place.
I’d pull actual request logs, bucket them by minute, and look at your p95 and p99 tokens-per-minute, not your mean, before sizing any provisioned tier. If you don’t have that data yet, you’re not ready to buy provisioned capacity. Get another month of logs first.
Where these deals quietly go wrong
Three failure modes show up over and over, and none of them are about the vendor being dishonest. They’re about the deal not matching how the workload actually behaves.
Idle capacity. You provision for a peak that turns out to be a one-time spike, not a recurring pattern, and now you’re paying a flat fee every month for headroom you touch twice a year.
Overage stacking. Some spend commitments and provisioned tiers charge on-demand rates, sometimes at a premium, for anything above the committed floor or reserved ceiling. If your usage regularly spills over, you can end up paying the commitment price and the overage price in the same month, which is worse than just staying on pay-as-you-go.
Model lock-in mid-term. A 12-month committed spend deal tied to a specific model family means that if a materially better or cheaper model ships six months in, and it will, you’re weighing “eat the switching cost” against “keep paying for the older model because we already committed the spend.” I’ve seen teams stay on a model a full quarter longer than they should have purely because the commitment made switching feel like it was throwing money away, when the sunk cost was already gone either way.
What I actually check before signing
Three full months of real usage data, not a forecast. A tokens-per-minute peak, not an average, if the deal involves any capacity reservation. And a term length that’s shorter than my confidence in the model I’m currently using, because the discount only pays off if I’m still using that model when the term ends.
The break-even point is never a single number a vendor hands you. It’s your usage curve intersected with their rate card, and you’re the only one who has the first half of that equation.
For more breakdowns like this on the tools and infrastructure behind shipping with LLMs, head back to the AI Tool Gazette homepage.