← all articles

Batch jobs versus realtime calls: when to wait and save

The bill that made me look twice

I run pipelines that call LLMs for classification, tagging, and content generation across a bunch of unrelated jobs. Somewhere around the point where a monthly API bill stopped being a rounding error, I started paying attention to which calls actually needed to come back in under a second and which ones I was hitting synchronously out of habit. That distinction, batch versus realtime, is not a minor implementation detail. It changes your cost structure and your architecture at the same time.

This isn’t a benchmark post. I haven’t run side by side timing tests across providers and I’m not going to pretend I have. What I can walk through is how batch and realtime APIs are actually built, because the mechanics tell you almost everything about when each one makes sense.

The mechanical difference

A realtime, or synchronous, call is what most people mean when they say “call the API.” You send a request, hold a connection open, and the model streams or returns tokens back to you inline. Your code (or your user) is blocked waiting on that response. Latency is the whole point of the interface: the provider’s infrastructure prioritizes getting your tokens back as fast as it can, because someone or something is sitting there waiting.

A batch API works differently at the infrastructure level, not just the pricing level. You don’t send one request and wait. You submit a file, usually JSONL, where each line is a separate request with its own ID. The provider queues the whole file, works through it on infrastructure that isn’t reserved for you, and gives you back a results file once it’s done, or once you poll and it’s ready. Major providers that offer this (OpenAI’s Batch API and Anthropic’s Message Batches API are the two I’ve actually used) both work on a similar model: submit a job, get a window of up to 24 hours for completion, poll or wait for a webhook, then pull the results file. Neither guarantees your job finishes early. It finishes when their scheduler gets to it, which is usually well inside that window but is not something you should build a real-time expectation around.

The reason batch is cheaper isn’t a discount program bolted on top. It’s because the provider can slot your requests into idle capacity instead of reserving hot capacity for you on demand. You’re trading a guarantee (this comes back in two seconds) for a much looser one (this comes back sometime today, at meaningfully lower cost per token). That’s the actual tradeoff. Everything else is a downstream consequence of it.

When realtime is the only option

Anything where a human is looking at a screen waiting for text to appear is realtime by definition. Chat interfaces, coding assistants doing inline completions, voice agents, anything where the model is one step in a live conversation. If a person’s attention is the bottleneck, shaving cost by batching is pointless because you can’t make them wait hours for a reply to a question they asked ten seconds ago.

It’s also realtime when the output of one call determines the very next action in a tight loop and that loop has to finish before something else can happen. An agent that’s mid-task, deciding which tool to call next based on the last tool’s output, needs the model back now, not in a queue. Coding assistants that are doing iterative edit-test-fix cycles fall into this category too. The value of the tool is largely in how fast that loop turns.

When batch is the obvious call

Batch earns its keep whenever the work is decoupled from a person waiting on it. A few patterns I actually use it for:

Bulk classification and tagging. If you’re labeling ten thousand support tickets, product descriptions, or scraped articles with categories or sentiment, nobody is standing over that job. It can run overnight.

Embedding generation for a RAG corpus. Building or rebuilding a vector index from a large document set is a one-shot, offline job. There’s no interactive component until someone actually queries the index afterward.

Evaluation runs. Testing a prompt change against a few hundred or a few thousand held-out examples before you ship it is exactly the kind of workload where you don’t need the result in the next thirty seconds, you need it correct and cheap, because you’re going to run it again after your next prompt tweak.

Data enrichment and backfills. Anything where you’re processing a historical dataset that already exists, as opposed to something arriving live, is a batch candidate by default.

Content generation at scale where a draft sits in a review queue anyway. If a generated article, description, or summary is going to be reviewed by a human before it’s published, the generation step doesn’t need to be instant. It just needs to be done before the reviewer gets to it.

The pattern across all of these: the model’s output is consumed later, by a process or a person who isn’t actively blocked on it arriving right now.

The parts nobody mentions until you build it

Batch sounds simple until you actually wire it into a pipeline, and a few things caught me off guard the first time.

You need an ID scheme you actually trust. Every line in your batch file needs a custom ID you control, because results don’t come back in the order you sent them and you’re the one responsible for matching a result back to the original request. If your ID generation isn’t stable and unique, you’ll spend an afternoon debugging why row 4,812 got the wrong answer.

Partial failure is normal, not exceptional. Some fraction of requests in a batch will fail for reasons that have nothing to do with your prompt: a malformed line, a request that exceeded a token limit you didn’t check for, a transient provider issue on their end. Your pipeline needs to expect a results file with some rows missing or errored, then decide whether those get retried in a follow up batch or dropped. If your code assumes every ID you submitted comes back with a clean result, the first real batch job will break that assumption.

You’re building a state machine, not a function call. A synchronous call is a request and a response, easy to reason about. A batch job is submit, then wait, then poll, then fetch, then parse, then reconcile against what you submitted. That’s five or six states your code has to track instead of one, and if your process restarts mid-job you need to be able to pick up the batch ID and check its status rather than resubmitting the whole thing and paying twice.

The 24 hour window is a ceiling, not a target you should design around. Jobs usually come back faster, but I don’t build anything that assumes a specific turnaround time inside that window. If a downstream process needs the result by a certain hour, submit early enough that the ceiling, not the average, still meets your deadline.

Rate limits and quotas are a separate pool

Batch and realtime usage typically draw against different limits on a provider account, which matters if you’re running both against the same API key. It means a huge batch job usually doesn’t eat into the rate limit your live chat feature depends on, but it also means you can’t reason about your account’s overall throughput using the number you’re used to watching for realtime traffic. Check this per provider before you assume it’s true for whichever one you’re using; it’s an architectural detail, not something guaranteed to be identical everywhere.

A simple rule I use

Before I write a call, I ask one question: is anything, human or process, sitting idle waiting on this specific response before it can do its next thing? If yes, it’s realtime, and I pay for the latency because the latency is the product. If no, and the output is going to sit in a database, a queue, or a review pile until something else gets to it, it goes in a batch file. That one question has moved a meaningful chunk of my own usage out of the expensive, low-latency lane and into the cheap, patient one, without touching anything that actually needed to be fast.

If you’re weighing this tradeoff on your own pipelines, start by listing every LLM call you make and marking which ones have an actual human or live process waiting on the other end. You’ll probably find more candidates for batch than you expected.

For more breakdowns like this on AI tools, LLM infrastructure, and coding assistants, head back to the AI Tool Gazette home page.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →