Skip to slide
Chapter 4 · Why Multi-Agent Systems Fail: The MAST Taxonomy
28 / 41

CHAPTER 04 · Why Multi-Agent Systems Fail: The MAST Taxonomy · 2 / 7

The contributions, in plain terms

The paper makes three contributions:

  1. MAST, a Multi-Agent System Failure Taxonomy: a careful catalog of 14 failure modes grouped into three categories.
  2. An LLM-as-a-judge method that automatically labels failures according to MAST, reaching about 94 percent accuracy and strong agreement with human experts (a Cohen's Kappa of 0.77, which is a statistical measure of agreement beyond chance).
  3. Case studies showing that you can improve real systems by fixing the failure modes MAST reveals, rather than just reaching for a bigger model.

The taxonomy itself was not invented from a whiteboard. It came from manually annotating over 200 execution traces, each averaging more than 15,000 tokens, using a careful method (grounded theory) and refining the categories until annotators agreed with each other. That grounding is what makes the taxonomy credible.

← → arrow keys work too