← all articles

Multi agent LLM systems: when more agents make things worse

I’ve shipped multi agent LLM pipelines that worked, and I’ve shipped ones that quietly got worse every time I added another agent to “help.” The second kind is more common than the marketing decks suggest. If you’re paying the API bill yourself, you notice the pattern fast: the system gets slower, the cost per task climbs, and the failures get harder to trace, all while the demo still looks impressive.

This isn’t an argument against multi agent systems. It’s a look at the specific mechanisms that make them worse, so you can tell when you’re hitting one of them instead of assuming you just need a better prompt.

What a multi agent LLM system actually is

Strip away the branding and a multi agent setup is usually one of two shapes. Either an orchestrator model calls out to specialized sub-agents and stitches their outputs together, or a set of agents pass a task between them in sequence, each doing one step before handing off to the next. Sometimes there’s a supervisor that reviews outputs and asks for retries. Sometimes agents talk to each other directly through a shared scratchpad or message log.

Every one of those handoffs is a new API call, a new context window, and a new place for the task to drift from what you actually wanted.

The math behind compounding errors

Here’s the part that doesn’t show up in a demo video. If a single agent does its step correctly 90% of the time, and you chain five of those steps together with no error correction between them, the chance the whole chain succeeds is roughly 0.9 to the fifth power, which comes out under 60%. That’s just what happens when independent steps multiply. Add a sixth or seventh agent for “more thoroughness” and the math gets worse, not better, unless each new agent is also catching and fixing errors from the ones before it rather than just adding its own.

Most pipelines I’ve seen don’t build in that correction. They build in more generation. More generation without more verification is how you get a five-agent system that fails more often than the single well-prompted agent it replaced.

Context doesn’t survive the handoff

Each agent in a multi agent system usually gets its own context window, built from whatever the previous agent decided to pass along. That summarization step is lossy by design, you can’t hand a downstream agent the full 30,000 tokens of research a previous agent read, so someone writes a compressed handoff. Whatever nuance didn’t make it into that summary is gone for the rest of the pipeline.

I’ve watched a research agent correctly flag a caveat in a source, only for that caveat to disappear from the three-sentence summary passed to the writing agent, which then produced a confident claim the original source didn’t support. No single agent was wrong. The system was wrong, because the information that would have prevented the mistake never crossed the boundary between agents.

Coordination overhead is a real cost, not a rounding error

Every agent call includes its own system prompt, its own instructions, and often a restatement of the task and constraints so the sub-agent has enough context to act. That’s fixed overhead you pay on every single hop, on top of whatever tokens the actual work needs. A task that would cost one call’s worth of input and output tokens as a single agent can end up costing input tokens for five or six calls once you split it into a pipeline, because each agent re-reads the setup before it does any new work.

That cost shows up directly on the API bill, and it shows up in latency too. Sequential agent calls don’t run in parallel just because you called them “agents.” If agent B needs agent A’s output before it can start, you’re paying the full round-trip time of A before B even begins, and again before C begins. A pipeline that feels architecturally clean on a whiteboard can be five times slower in wall-clock time than a single well-scoped call, for a task the user expected to get back in a few seconds.

Debugging gets harder exactly when you need it easier

With one agent and one prompt, when the output is wrong you read the prompt, read the output, and you usually spot the gap. With five agents passing state between them, the failure could be in the first agent’s research, the second agent’s summarization, the third agent’s interpretation of that summary, or the orchestrator’s decision about what to do with all of it. You end up adding logging between every hop just to find out which step introduced the problem, which is its own engineering cost that never shows up in the “just add another agent” pitch.

This matters more than it sounds like, because debugging time is where a lot of the real cost of these systems hides. A single-agent system that fails is annoying. A five-agent system that fails somewhere in the middle of a chain, silently, and ships a wrong answer to a user, is a debugging session plus a cleanup.

Agents can duplicate or contradict each other

Without a strong arbiter, parallel agents working on related sub-tasks can step on each other. I’ve seen a “researcher” agent and a “fact-checker” agent both independently decide to search for the same source, burning tokens twice, and then disagree about how to interpret it, with the orchestrator having no principled way to resolve the disagreement beyond picking one arbitrarily. That happens by default when you give multiple agents overlapping scope and no clear ownership boundary.

The fix isn’t “add a supervisor agent to referee,” because now you’ve added another agent, another context window, and another place for information to get lost on the way in. The fix is narrowing the scope of each agent so their responsibilities don’t overlap in the first place.

When splitting into agents actually helps

None of this means single-agent-always-wins. Splitting work into separate agents earns its cost when the sub-tasks genuinely need different tools, different context, or different levels of scrutiny, and when the handoff between them is narrow and well-defined rather than “here’s a paragraph, figure out what matters.” A dedicated agent that only writes SQL against a fixed schema, called by an orchestrator that only decides when a query is needed, is a clean split, because the boundary is small and unambiguous. A pipeline where five agents all get “a task and some context” and are expected to figure out their role from a system prompt is where things degrade.

The other place it works is when you put a deterministic check, not another LLM call, between agents. A schema validator, a unit test, a regex, something that either passes or fails without ambiguity. That catches errors before they compound instead of adding another probabilistic step to the chain.

What I actually do before adding an agent

Before splitting a task into more agents, I ask what specific failure the new agent is supposed to fix, and whether a better single prompt, a tool call, or a deterministic check would fix it for less cost and less latency. If I can’t name the specific failure mode, I don’t add the agent. And when I do split something, I try to keep the handoff between agents as narrow as a function signature rather than as loose as a paragraph of prose, because the narrower the handoff, the less there is to lose in translation.

More agents is not a strategy. It’s an architecture decision that has to earn its keep against real overhead in tokens, latency, and debuggability. Sometimes it earns it. Often the honest answer is that one agent with better tools and a tighter prompt would have done the job for a fraction of the cost.

If you want more breakdowns like this on what actually works with AI tools and coding assistants, without the vendor spin, check out the rest of AI Tool Gazette.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →