Search⌘ K
AI Features

Containing Failure in Multi-Agent Systems

Explore common failure patterns in multi-agent systems and learn how to design effective containment strategies. Understand how to limit failure propagation using least privilege, structured handoffs, and fail-closed approaches to maintain system safety and reliability.

A single agent fails in one place. When it misreads a document, the mistake sits in one transcript, and one trace shows where it happened. A multi-agent system spreads the same work across a coordinator and several sub-agents, and every handoff between them is a point where a failure can enter, grow, and travel.

Interviewers probe exactly this. A candidate who proposes a coordinator with sub-agents should expect the question, what happens when one of them fails. This lesson covers the common failure modes and the containment that answers that question.

How multi-agent systems fail

Most failures in multi-agent systems fall into six patterns, all of which begin at the boundary between two agents.

  • Lossy handoff: A sub-agent sees only the task the coordinator writes for it. A vague task leaves out the input or the rule to apply, so the sub-agent guesses, repeats another agent’s work, or checks the wrong thing.

  • Error propagation: A sub-agent reports a wrong fact with confidence, such as a value read from the wrong record. The coordinator cannot tell it apart from a correct report, so the error flows into the decision.

  • Runaway work: An agent retries a failing tool call again and again, or the coordinator starts far more sub-agents than the task needs. Every repetition costs tokens across every agent involved.

  • Injection that crosses agents: A document, email, or web page contains an instruction written for the AI. A sub-agent that reads it may repeat the instruction in its report, which carries it to an agent that can act.

  • Privilege creep: Sub-agents receive broad tool access because it was simpler to configure. An agent that only needs to read can then write, and any mistake it makes becomes an action. ...