How to handle API rate limits without losing requests
If you’ve shipped anything on top of an LLM API, you’ve seen the 429. It shows up at the worst time, usually during a traffic spike or a batch job you forgot was still running, and if your error handling is naive, requests just vanish. The user never gets a response, the batch job silently drops rows, and you find out three days later when someone asks why half the dataset is missing.
This isn’t a theoretical problem. Rate limits are a normal, permanent feature of every LLM API, not an edge case you patch around once. Here’s how the limits actually work and what to do so a 429 costs you latency instead of data.
What a rate limit actually is
Most LLM providers enforce limits along a few axes at once: requests per minute (RPM), tokens per minute (TPM), and sometimes concurrent requests or tokens per day. You can hit any of these independently. A batch of short prompts can blow through RPM long before it touches your token budget. A handful of long-context requests can blow through TPM while your RPM count looks fine.
Under the hood this is almost always implemented as a token bucket or a sliding window counter tied to your API key or org. A token bucket refills at a fixed rate and lets you burst up to its capacity, then throttles you once it’s empty. A sliding window just counts requests or tokens in the trailing N seconds and rejects you once you’re over. Which one a given provider uses matters less than the practical result: if you send a burst, you will get some 429s, and the size of the burst you can absorb is smaller than you think, because your usage is also counted against whatever else is running under the same key at the same time.
That last point trips people up constantly. If you have a cron job, a user-facing endpoint, and a background eval script all hitting the same API key, they’re sharing one bucket. A quiet endpoint can start failing because an unrelated batch job is eating the budget.
The retry mistake that makes it worse
The most common bad pattern is retry-immediately-in-a-loop. You catch the 429, sleep a fixed 500ms, and try again. Under light load this looks fine. Under real load it’s the thing that turns a temporary throttle into an outage, because every failed request now retries at roughly the same moment, which means your retries collide with each other and with the next batch of new requests. You get a thundering herd, the provider keeps rejecting you, and your error rate climbs instead of recovering.
The fix has two parts: exponential backoff, and jitter. Exponential backoff means each retry waits longer than the last, usually doubling: 1s, 2s, 4s, 8s, capped at some ceiling like 30 or 60 seconds. Jitter means you don’t wait exactly that long, you wait a random amount within a range around it, so retries from different requests don’t land in the same instant. A simple version is “full jitter”: pick a random wait time between 0 and the current backoff ceiling, rather than the ceiling itself. This spreads retries out instead of syncing them.
Cap your retry count too. Three to five attempts is usually the right range. If you’re still failing after five exponential retries with jitter, the problem isn’t transient congestion, it’s that you’re structurally sending more load than your tier supports, and more retries just burn latency without fixing anything.
Respect Retry-After when you get one
Some 429 responses include a Retry-After header, either as seconds or as a timestamp. When it’s there, use it instead of your own backoff calculation. The provider is telling you exactly when its bucket will have room again, and guessing against that is strictly worse than reading the number they gave you. When it’s absent, fall back to your own exponential-with-jitter logic.
Throttle before you hit the wall, not after
Reactive backoff treats every rate limit as a surprise. A more mature setup treats it as a known constraint and paces requests to stay under it in the first place. This is client-side rate limiting: you implement your own token bucket (a simple counter and timer is enough, you don’t need a library for this) sized just under your actual API tier limit, and every outbound call draws from it. If the bucket is empty, the call waits instead of firing and failing.
This changes your error rate from “occasional 429s under load” to “near zero, because you never send the request that would have been rejected.” It costs you a small amount of added latency during bursts, since calls queue instead of firing immediately, but that’s a better tradeoff than a failed request that has to go through the retry path anyway.
Queue instead of drop
For anything that isn’t a live, synchronous user request, don’t call the API inline at all. Push the work onto a queue (a Postgres table with a status column is enough for most workloads, you don’t need Kafka for this) and have a worker pull from it at a rate that respects your limits. If a call fails after retries, it stays in the queue with a failed or retry_later status instead of disappearing. This is the single biggest lever for “not losing requests”: the request’s existence is now durable, independent of whether the API call that processes it succeeds on the first, fifth, or fiftieth try.
For synchronous, user-facing calls where you can’t queue and wait, the honest move is to degrade gracefully: return a shorter, cheaper response, fall back to a smaller model with more headroom, or tell the user to retry, rather than pretending the call succeeded.
Retries can double-charge you if you’re not careful
Once you’re retrying automatically, you need to make sure a retry doesn’t duplicate an effect. If your endpoint calls the LLM and then writes a result to a database or sends an email based on that result, and the write succeeds but the response never makes it back to your process before a timeout, your retry logic might call the API again for the same logical request. Now you’ve paid for two completions and possibly written the result twice.
The fix is a client-side idempotency key: generate a unique ID per logical request (not per HTTP call) and use it to detect duplicates on your own side, either by checking a processed_requests table before firing the call or by using an idempotency key parameter if the provider’s API supports one. This is orthogonal to rate limiting but comes up in the same code path, so it’s worth building at the same time as your retry logic instead of bolting it on later.
Splitting load across keys is a real lever, with a real catch
If your usage regularly bumps against your tier’s ceiling, running requests across more than one API key or project can raise your effective throughput, since limits are typically scoped per key or per org rather than per account holistically. This is a legitimate scaling technique, not a hack, but read the provider’s terms of service before doing it. Some providers explicitly prohibit using multiple accounts to circumvent rate limits on a single logical workload, and getting that wrong can get keys suspended, which is a much worse outage than the 429s you were trying to avoid. If you’re going to do this, do it because you’ve requested and been granted a higher tier or genuinely separate use cases, not as a workaround for a limit you haven’t asked to be raised.
What to actually monitor
Track your 429 rate as its own metric, separate from general error rate, and track it per endpoint and per API key if you have more than one. A slow climb in 429s over days usually means usage growth is outpacing your tier. A sudden spike usually means a specific job or deploy started sending more traffic than expected. Also track queue depth and oldest-item age if you’re using a queue, since a queue that’s growing faster than it drains is the early warning sign that backoff alone won’t save you and you need either a higher tier or fewer requests.
None of this makes rate limits go away. It just moves the cost from “requests silently disappear” to “requests take a bit longer during bursts,” which is the tradeoff you actually want.
If you’re building on LLM APIs and want more of this kind of production-grade detail instead of vendor demos, check out more breakdowns on the AI Tool Gazette home page.