Batching requests without hurting response time
If you’ve ever gotten an API bill and thought “I’m paying for a lot of idle GPU time here,” batching is the fix everyone points you toward. It’s also the thing that quietly turns a snappy chat feature into one where users stare at a spinner for four extra seconds while your queue fills up. The word “batching” gets used for three different techniques that behave nothing alike, and mixing them up is how teams end up with a cheaper bill and an angry Slack channel from the product team.
Three things people call batching
Async batch APIs are the ones providers sell as a separate product: you submit a file of requests, the provider processes them whenever it has spare capacity, and you get results back within a window, sometimes up to 24 hours. OpenAI’s Batch API and Anthropic’s Message Batches API both work this way. You are explicitly trading latency for a lower per-token price.
Client-side request coalescing is something you build yourself: instead of firing off ten separate API calls for ten small jobs, you stack them into one prompt or one call and split the response after. This cuts fixed overhead per call, but now the whole group waits on whichever item takes longest.
Server-side continuous batching happens inside the inference engine, whether that’s a provider’s backend you never see or a self-hosted stack like vLLM or TGI. The server groups concurrent requests from different users at the token-generation level, not the request level, so it can raise throughput without necessarily raising anyone’s wait time.
Only the third one is close to a free lunch. The other two are real tradeoffs, and the mistake I keep seeing is treating all three as interchangeable ways to “batch and save money.”
Async batch APIs: don’t put them behind a spinner
The async batch endpoints exist because inference providers have capacity that isn’t needed for real-time traffic, and they’d rather sell it cheap than let it sit idle. That’s why the deal works: you agree to wait, they agree to discount. The completion window on these is measured in hours, not seconds, and there’s no guarantee your job finishes early just because the queue looks short right now.
This is the right tool for things like: nightly re-classification of a support ticket backlog, backfilling embeddings for a document store, running an eval suite against a new prompt version, or generating training data overnight. None of these have a human waiting on the other end.
The failure mode is putting a “quick summarize this document” button in your product and routing it through the batch endpoint because the price is better. Users don’t wait 20 minutes for a summary, no matter how cheap it was to generate. If you want the discount for something user-facing, the answer is usually to do the work ahead of time (pre-generate summaries for documents as they’re uploaded, not when someone clicks a button) rather than trying to make an async job feel synchronous.
Client-side batching: you’re waiting for the slowest item in the group
Say you have a pipeline that tags 500 support tickets a day. Firing 500 separate calls means paying 500 times for the system prompt, 500 times for connection overhead, and 500 times for whatever fixed latency the provider adds before the first token comes back. Grouping 20 tickets into one prompt and asking for 20 tagged outputs back cuts all of that by 20x.
The catch: your response time for that batch is now bound by the total generation length for all 20 answers combined, plus whatever time it takes the model to work through a longer input. If 19 tickets are one-liners and one is a rambling five-paragraph complaint, all 20 results arrive only when the model finishes with the rambling one. If this pipeline runs as a background job, fine, nobody’s watching a clock. If it’s inline in a request path, your P95 just inherited the worst-case behavior of your batch size.
The knob that actually matters here isn’t batch size, it’s your latency budget. Decide what response time you can tolerate for the calling code, then size the batch to fit under that ceiling given typical output length, not the other way around. A batch of 20 might be fine for an overnight job and completely wrong for anything with a 2-second SLA.
Continuous batching: the one that isn’t really a tradeoff
This is what’s happening on the inference server itself, and it’s worth understanding even if you never touch it directly, because it explains why concurrent load on an LLM API doesn’t slow everyone down proportionally the way you might expect.
Old-style static batching waited for a full batch of requests to arrive, ran them together start to finish, and only returned once the whole batch was done, which meant a short request got stuck behind a long one exactly like the client-side case above. Continuous batching (the technique behind vLLM, TGI, and most modern serving stacks) instead schedules at the token level: after each generation step, the server checks what’s finished, evicts completed sequences, and slots new requests into the freed capacity. A short request can enter, run, and leave without waiting for a longer one queued around it to finish.
Combined with memory management tricks like paged attention, this is why providers can serve many concurrent users off a shared pool of GPUs without every user’s latency scaling with the number of other users active at that moment. It’s also why your own experience of “the API feels slower at 2pm US time” isn’t your imagination. It’s the server’s batching absorbing real concurrent load, and it absorbs it well up to a saturation point, after which everyone’s tokens per second does drop.
You mostly can’t tune this if you’re calling a hosted provider, but it’s worth knowing it exists, because it’s the reason batching two techniques up the stack (async batch APIs and client-side coalescing) is a deliberate cost/latency tradeoff you’re choosing to make, while this layer is closer to free capacity the provider is already extracting on your behalf.
A playbook for not blowing up your response time
A few habits that keep the cost savings without the latency surprise:
- Keep real-time and batch traffic on separate code paths, and ideally separate API keys, so a batch job spiking your rate limit doesn’t slow down user-facing calls competing for the same quota.
- Size client-side batches against a latency budget, not a throughput target. Ask “what’s the max group size where P95 stays under my SLA” instead of “how big can I make this before the API complains.”
- Watch P95 and time-to-first-token, not average latency. Averages hide the one long item dragging a whole batch, which is exactly the failure mode batching introduces.
- Cap batch wait windows explicitly. If you’re coalescing requests that arrive over time (waiting up to, say, 50ms to see if more requests show up before firing), that wait is pure added latency for whoever arrived first. Make it small and bounded, don’t let it grow to “whatever fills the batch nicely.”
- Use the async batch endpoint only for jobs with no one waiting on the response. If you’re tempted to poll it quickly to fake real-time behavior, that’s a sign you want the regular API, not the batch one, and should eat the higher per-token cost for that specific path.
None of this is exotic. It’s the same latency-budget discipline you’d apply to any queueing system, just applied to a model API instead of a job queue. The providers built batch APIs to be honest about the tradeoff (you get a big discount in exchange for a wide completion window), the mistake is entirely on the calling side when someone tries to use a cost tool to solve a latency problem.
If you’re building with LLM APIs and want more of this kind of practical, no-hype breakdown, check out the rest of the site here.