← all articles

Splitting LLM prompts: when one call should really be two

The question I get asked most in code review

Someone on the team writes a prompt that reasons about a support ticket, decides whether it’s a refund request, and also formats the answer as JSON for the ticketing system. It mostly works. Then one day the model writes three paragraphs of reasoning before it gets to the JSON, the parser chokes on the leading text, and the whole request has to be retried. The fix isn’t a better regex. It’s splitting that one call into two.

This comes up constantly because a single prompt is the path of least resistance. You start with one instruction, keep bolting requirements onto it, and before long you’re asking one model call to route, reason, format, and self-check all at once. Sometimes that’s fine. Sometimes it’s the reason your error rate won’t go below a certain floor no matter how much you tweak the wording. Here’s how to tell which situation you’re in.

What actually happens when you cram tasks into one prompt

A single completion is one continuous generation. The model produces tokens left to right, and every token it writes becomes part of the context for the next one. If you ask it to “think it through, then output JSON,” the free-text reasoning becomes part of what the JSON-formatting instruction has to compete with. The model isn’t running two separate mental processes and merging the results. It’s writing one stream of text where earlier tokens shape later ones.

This matters in three concrete ways:

Instruction-following gets diluted. The more distinct things you ask for in one prompt (summarize, classify, cite sources, format as JSON, stay under 200 words), the more the model has to juggle simultaneously. Long, multi-part system prompts are more prone to the model dropping one instruction to satisfy another, especially near the end of a long response, because attention is spread across everything you asked for instead of concentrated on one job.

Structured output and free reasoning actively fight each other. JSON mode (or function calling, or any strict schema constraint) works best when the model isn’t also trying to produce open-ended prose in the same completion. Ask for both and you’ll see it either skip the reasoning to protect the format, or wrap the JSON in explanation text that breaks your parser. This isn’t a bug you can prompt your way out of reliably. It’s a structural conflict between “write naturally” and “conform to a strict grammar,” happening in the same token stream.

One failure means you retry everything. If your single call does retrieval synthesis, tone adjustment, and formatting, and the formatting step fails, you re-run the entire thing, including the expensive parts that worked fine. You pay for regenerating the reasoning you already had, just to get a fresh shot at the formatting.

Where splitting actually pays off

Reason first, extract second. This is the pattern that fixes the support ticket example above. Call one: give the model the ticket, ask it to think through the situation in plain text, no format constraints. Call two: feed that reasoning back in and ask only for the structured output, nothing else. The second call has one job and no competing instructions, so it’s far more consistent about producing clean JSON. You pay for two completions instead of one, but you stop paying for retries on malformed output, which is often the more expensive failure mode since a retry re-runs the whole thing including the reasoning tokens you already generated once.

Map, then reduce, for anything longer than one context window’s worth of “attention.” If you’re summarizing forty pages, don’t hand the whole thing to one call and hope the model weighs page one and page forty equally. Split into a map step (run the same short prompt against each chunk to pull out key points) and a reduce step (a second call that synthesizes the extracted points into a final answer). Each map call has a small, well-defined job, so it’s less likely to skip material buried in the middle of a long input. The reduce call only has to reason over the condensed output, not the raw document, so it stays focused.

Route cheap, generate expensive. If part of your pipeline is “decide what kind of request this is” and another part is “write a detailed response,” don’t make the expensive model do both. Run a small classification call first, then only invoke the heavier generation call for the branches that actually need it. This is standard triage: a lot of production support-bot and content pipelines route the simple, high-volume cases to a lightweight call and reserve the larger model for cases that need real reasoning. You’re not paying full generation cost for the fifteen percent of requests that turn out to be “what’s your refund policy,” which any short classify-and-template step can handle.

Isolate the part that fails independently. If one piece of your prompt is inherently flaky, tool calls, retrieval that sometimes returns nothing, user input that’s occasionally malformed, put it in its own call. Then you can retry just that piece with backoff logic instead of re-running a prompt that also does three other things correctly on the first try.

Where splitting is just overhead

Splitting isn’t free, and treating it as automatically better is its own mistake.

Two calls means two round trips. Network and inference latency stack. If you’re building something user-facing and latency matters, chaining two sequential calls will feel slower than one, even if each call individually is fast. For a chat interface where someone is watching the response stream in, that added round trip is the kind of thing users notice even when they can’t articulate why.

You often duplicate context. If both calls need the same background (a long system prompt, retrieved documents, conversation history), you’re paying to send those input tokens twice. For a short, cheap task, that duplicated context can end up costing more than whatever you saved by splitting.

More calls means more code to maintain. Two prompts to version, two sets of few-shot examples to keep in sync, two places where a model upgrade might change behavior. If the single-call version is passing your tests and nobody’s complaining about malformed output, splitting it “on principle” is just extra surface area for bugs.

Atomic tasks don’t benefit. Translate this paragraph. Summarize this email in one sentence. Classify this ticket into one of four categories. These are single, well-defined transformations. There’s no competing instruction to untangle and nothing to gain from a second call, because there’s no second job happening.

A rough decision test

Ask three questions about your prompt. Is it asking for free-form reasoning and a strict output format in the same completion? Is the input long enough that the middle of it is likely to get less attention than the start and end? Is there a cheap decision that gates whether an expensive generation even needs to run? If you answer yes to any of these, split it. If the prompt does one clearly scoped thing and produces one kind of output, leave it alone. The extra call is a cost you should be able to point to a specific failure mode to justify, not a default you reach for because it feels more rigorous.

Splitting llm prompts is a tool for untangling competing instructions and isolating failure points, not a general best practice to apply everywhere. Know which job each call is doing, and don’t pay for a second one unless it’s actually solving something the first one can’t.

If you want more breakdowns like this on how these tools actually behave in production, not just what the marketing page says, check out the rest of the site here.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →