Skip to slide
Chapter 4 · Why Multi-Agent Systems Fail: The MAST Taxonomy
29 / 41

CHAPTER 04 · Why Multi-Agent Systems Fail: The MAST Taxonomy · 3 / 7

The three categories of failure

MAST sorts its 14 failure modes by the stage of the agent lifecycle where they originate: before execution, during execution, and after execution.

1. Specification issues (about 41.8 percent of failures)

These come from a flawed setup: poor prompt design, missing role constraints, or no clear stopping criteria. Representative modes include:

  • Disobeying the task specification.
  • Repeating steps that were already completed.
  • Losing track of the conversation history.
  • Failing to recognize when the task is actually done.

This is the largest category, which is a striking result on its own: the single biggest source of failure is how the system was specified and set up, before any agent even starts talking to another.

2. Inter-agent misalignment (about 36.9 percent)

These happen during execution, from miscommunication, conflicting assumptions, or context that never gets passed along. Examples include:

  • Ignoring what other agents said.
  • Proceeding without asking for clarification.
  • Resetting conversations unexpectedly.
  • A mismatch between what an agent reasons and what it actually does.

This category is essentially the empirical confirmation of Cognition's Principle 2 from Chapter 3: conflicting hidden decisions and unshared context are not just a theoretical worry, they account for over a third of observed failures.

3. Task verification failures (about 21.3 percent)

These come from weak quality control at the end:

  • Ending the task too early.
  • Skipping validation entirely.
  • Accepting an incorrect solution because the check was shallow.

The article highlights a recurring pattern here: many systems do include a verifier agent, but its checks are superficial. Code is accepted just because it compiles. A program is assumed correct because its comments look consistent. That is not real verification.

An important note: many traces contain several failure modes at once, which is why the paper argues for structured diagnosis rather than ad hoc inspection.

← → arrow keys work too