Why Multi-Agent LLM Systems Fail: The Evidence
Multi-agent LLM systems fail mainly because of how they are specified, coordinated and checked, and less because of the model underneath. A UC Berkeley-led study of the question sorted 1,642 execution traces into 14 failure modes in three groups: system design issues, inter-agent misalignment and task verification. A single agent is the better default for most work. Several agents pay off on broad, parallel, read-heavy tasks where the result is worth a much larger token bill.
Why do multi-agent LLM systems fail?
They fail in three recurring ways: the system is poorly specified, the agents work against each other, or nobody checks the result properly. That is the finding of "Why Do Multi-Agent LLM Systems Fail?" by Mert Cemri and co-authors, most of them at UC Berkeley, first posted on March 17, 2025 and last revised on October 26, 2025.
The authors built the Multi-Agent System Failure Taxonomy (MAST) from "1642 annotated execution traces" collected from seven multi-agent frameworks, among them MetaGPT, ChatDev and AG2. Human annotators developed the taxonomy on 150 traces and reached strong agreement (Cohen's kappa of 0.88). An LLM judge then labeled the full set and matched human experts at "accuracy 94%, Cohen’s Kappa of 0.77".
The headline number is the failure rate itself: "41% to 86.7% failure rate on 7 state-of-the-art (SOTA) open-source MAS".
| Failure category | Share of observed failures | Failure modes and their measured shares |
|---|---|---|
| System design issues | 44.2% | Step repetition 15.7%; unaware of termination conditions 12.4%; disobey task specification 11.8%; loss of conversation history 2.80%; disobey role specification 1.5% |
| Inter-agent misalignment | 32.35% | Reasoning-action mismatch 13.2%; task derailment 7.40%; fail to ask for clarification 6.80%; conversation reset 2.20%; ignored other agent's input 1.90%; information withholding 0.85% |
| Task verification | 23.5% | Incorrect verification 9.10%; no or incomplete verification 8.20%; premature termination 6.20% |
The per-mode shares are quoted from the paper. The category totals are our sums of those figures, and they add to slightly more than 100% because the paper rounds each mode.
The study has limits. The traces come from open-source research frameworks running coding, math and general-agent tasks, and most labels come from the LLM judge. The authors state that they "do not claim it covers every potential failure pattern." A production system with a narrow scope may show a different mix.
What do these failures look like in practice?
They look like ordinary coordination problems: repeated work, lost context, conflicting assumptions and shallow checks.
System design. The three largest modes are repeating steps, not knowing when to stop and ignoring the task specification. Anthropic described the same pattern in its own early agents in an engineering post published June 13, 2025: "spawning 50 subagents for simple queries, scouring the web endlessly for nonexistent sources, and distracting each other with excessive updates."
Inter-agent misalignment. Agents act on assumptions the other agents do not share. Walden Yan of Cognition gave a plain example on June 12, 2025. Asked to build a Flappy Bird clone, one subagent builds a background that looks like Super Mario Bros., another builds a bird that does not match, and a final agent has to combine the two. His rule: "Actions carry implicit decisions, and conflicting decisions carry bad results". The MAST authors add that these errors "occur even when agents within the same framework communicate using natural language", so a shared message protocol does not remove them.
Task verification. A reviewer agent exists but checks the wrong thing. The paper found that "many existing verifiers perform only superficial checks", such as whether code compiles. In one of its examples, a generated chess program passed its review phases and still had runtime bugs, because nothing validated it against the rules of the game.
Is the model or the system design to blame?
Mostly the design, on the evidence so far, although a stronger model still helps. The MAST authors conclude that "many MAS failures arise from the challenges in organizational design and agent coordination rather than the limitations of individual agents."
They tested this by changing one system while keeping the model fixed. Improving agent role specifications "yields a +9.4% success rate increase for ChatDev". Adding a verification step against the high-level task objective "yields a +15.6% improvement in task success". The same section carries a warning: "task completion rates still remain low", and the authors expect that reliable systems will need structural changes as well as better models.
Model choice matters too. Anthropic reported that "upgrading to Claude Sonnet 4 is a larger performance gain than doubling the token budget on Claude Sonnet 3.7."
Single agent vs multi-agent: when is one agent enough?
One agent is enough when the steps depend on each other, the work fits in one context window, or the budget is tight. Vendors that sell multi-agent tooling give the same advice.
OpenAI's practical guide to building agents, published in April 2025, states: "Our general recommendation is to maximize a single agent’s capabilities first." Microsoft's Cloud Adoption Framework guidance, last updated December 10, 2025, observes that multi-agent architectures "are often chosen based on untested assumptions about complexity or performance". It tells teams to move to them "only when testing reveals limitations that cannot be resolved through single-agent optimization."
A controlled study supports the caution. Kim et al., posted December 9, 2025 and revised April 8, 2026, with authors from Google Research, Google DeepMind and MIT, compared one single-agent and four multi-agent architectures across 260 configurations. Performance against the single-agent baseline ranged "from +80.8% on decomposable financial reasoning to -70.0% on sequential planning".
The same paper reports that tasks "where single-agent performance already exceeds 45% accuracy experience negative returns from additional agents". It also found that independent agents with no central check amplified errors 17.2×, against 4.4× under centralized coordination.
| Question | Points to one agent | Points to several agents |
|---|---|---|
| Can the task be split into independent parts? | Steps depend on each other, as in sequential planning and most coding | Subtasks are independent and can run in parallel, as in broad research |
| Does the information fit in one context window? | Yes | No, the material exceeds a single context window |
| Is the work mainly reading or writing? | Writing one coherent artifact | Reading and gathering from many sources |
| Are there hard boundaries? | None | Security, compliance or team boundaries that require separation |
| What is the budget? | Cost or latency is a constraint | The result is valuable enough to pay for many times more tokens |
The reading and writing row comes from Harrison Chase of LangChain, who compared the Anthropic and Cognition posts on June 16, 2025 and concluded that "read actions are inherently more parallelizable than write actions". The boundaries row follows Microsoft's guidance. The context window and budget rows follow Anthropic's post.
How did Anthropic build its multi-agent research system?
Anthropic's Research feature uses "an orchestrator-worker pattern": a lead agent plans the research, starts subagents that search in parallel, and synthesizes what they return. The company described it in "How we built our multi-agent research system", published June 13, 2025.
The post reports the gain and the cost together.
- The gain. A system "with Claude Opus 4 as the lead agent and Claude Sonnet 4 subagents outperformed single-agent Claude Opus 4 by 90.2% on our internal research eval."
- The reason. "Multi-agent systems work mainly because they help spend enough tokens to solve the problem." On the BrowseComp evaluation, "token usage by itself explains 80% of the variance".
- The cost. "agents typically use about 4× more tokens than chat interactions, and multi-agent systems use about 15× more tokens than chats."
- The limit. Domains "that require all agents to share the same context or involve many dependencies between agents are not a good fit for multi-agent systems today." Anthropic names most coding tasks as an example.
The design choices answer the failure categories above. Delegation is written out in detail, effort is capped by rules in the prompt, the lead agent does the synthesis, and a separate step attaches citations to the claims.
How do you build a multi-agent system that fails less?
Treat it as an engineering system with roles, checks, tests and budgets. The published practices line up with the three failure categories.
- Start with one agent and measure it. Keep that result as the baseline any multi-agent design has to beat.
- Write complete briefs. In Anthropic's words, "Each subagent needs an objective, an output format, guidance on the tools and sources to use, and clear task boundaries." Short instructions led its subagents to repeat each other's searches.
- Share context and keep decisions in one place. Cognition's first principle is "Share context, and share full agent traces, not just individual messages". Parallelize the reading and let one agent do the writing.
- Verify against the goal. The MAST authors recommend "using external knowledge, collecting testing output throughout generation, and multi-level checks for both low-level correctness and high-level objectives."
- Build an evaluation set early. Anthropic "started with a set of about 20 queries representing real usage patterns", graded outputs with an LLM judge against a rubric, and kept human testing because "People testing agents find edge cases that evals miss." For agents that change state, it judged the final state instead of each step.
- Set stopping conditions and cost limits. OpenAI lists the common exit conditions as "tool calls, a certain structured output, errors, or reaching a maximum number of turns." Anthropic wrote effort rules into its prompts, for example "Simple fact-finding requires just 1 agent with 3-10 tool calls".
- Make runs resumable. Anthropic combines model adaptability "with deterministic safeguards like retry logic and regular checkpoints", because restarting a long run is expensive.
- Escalate to a person. OpenAI's guide says actions "that are sensitive, irreversible, or have high stakes should trigger human oversight until confidence in the agent’s reliability grows."
What does a multi-agent system cost to run?
It costs several times more per task than a single agent, in tokens and in latency. Anthropic's figures are about 4× the tokens of a chat for one agent and about 15× for a multi-agent system. The post adds that "multi-agent systems require tasks where the value of the task is high enough to pay for the increased performance."
Microsoft's guidance explains where the cost comes from: "each agent processes redundant context and communication overhead multiplies with agent count", and "Latency accumulates at each handoff point".
Three controls keep the bill predictable: a token budget per task, a cap on the number of agents and tool calls, and a rule for which model each agent uses.
Key takeaways
- The MAST study measured failure rates of 41% to 86.7% across seven open-source multi-agent systems.
- By our sum of the paper's figures, 44.2% of failures were system design issues, 32.35% inter-agent misalignment and 23.5% task verification.
- Changes to roles and verification improved one system by up to 15.6% with the same model.
- Multi-agent results depend on the task: from +80.8% to -70.0% against a single agent in one controlled study.
- Anthropic's research system beat a single agent by 90.2% on its internal evaluation and used about 15× the tokens of a chat.
- Start with one agent, an evaluation set and a budget. Add agents when a measured limit requires it.
Frequently asked questions
What is a multi-agent LLM system?
The MAST paper defines it as "a collection of agents designed to interact through orchestration, enabling collective intelligence". Each agent has a prompt, a conversation state and the ability to act through tools. A common form is one lead agent that delegates subtasks to worker agents.
What is the most common failure mode in multi-agent systems?
Step repetition, at 15.7% of observed failures in the MAST dataset. Reasoning-action mismatch follows at 13.2%, then agents that do not recognize the task is finished at 12.4%.
Are multi-agent systems better than a single agent?
On some tasks. Kim et al. measured changes from +80.8% to -70.0% against a single-agent baseline, depending on whether the task could be decomposed. Broad research tasks benefit. Sequential planning did worse in that study, and Anthropic names most coding tasks as a poor fit.
Does a better model fix multi-agent failures?
Partly. A stronger model reduces some errors, but the MAST authors "conjecture that improvements in the base model capabilities will be insufficient to address the full MAST." Role definitions, context sharing and verification are design decisions that a model upgrade does not make for you.
How many test cases do you need to evaluate an agent system?
A small set is enough to start. Anthropic began with about 20 queries drawn from real usage and advises that "it’s best to start with small-scale testing right away with a few examples, rather than delaying until you can build more thorough evals."
Easital Technologies Ltd. designs and builds AI agents, agentic workflows and tool-calling systems, and runs agent-based products of its own. Related pages: AI agent development, AI automation services, LLM cost optimization and our work.
Sources
All sources were opened and checked on October 3, 2026.
- Cemri, Pan, Yang, Agrawal, Chopra, Tiwari, Keutzer, Parameswaran, Klein, Ramchandran, Zaharia, Gonzalez, Stoica, "Why Do Multi-Agent LLM Systems Fail?", arXiv 2503.13657, March 17, 2025 (v3 October 26, 2025). https://arxiv.org/abs/2503.13657
- Anthropic, "How we built our multi-agent research system", June 13, 2025. https://www.anthropic.com/engineering/multi-agent-research-system
- Walden Yan, "Don’t Build Multi-Agents", Cognition, June 12, 2025. https://cognition.ai/blog/dont-build-multi-agents
- OpenAI, "A practical guide to building agents", April 2025. https://cdn.openai.com/business-guides-and-resources/a-practical-guide-to-building-agents.pdf
- Microsoft, "Choosing Between Building a Single-Agent System or Multi-Agent System", Cloud Adoption Framework, last updated December 10, 2025. https://learn.microsoft.com/en-us/azure/cloud-adoption-framework/ai-agents/single-agent-multiple-agents
- Kim et al., "Towards a Science of Scaling Agent Systems", arXiv 2512.08296, December 9, 2025 (v3 April 8, 2026). https://arxiv.org/abs/2512.08296
- Harrison Chase, "How and when to build multi-agent systems", LangChain, June 16, 2025. https://blog.langchain.com/how-and-when-to-build-multi-agent-systems/

