CHAPTER 04 · Why Multi-Agent Systems Fail: The MAST Taxonomy · 2 / 7
The contributions, in plain terms
The paper makes three contributions:
- MAST, a Multi-Agent System Failure Taxonomy: a careful catalog of 14 failure modes grouped into three categories.
- An LLM-as-a-judge method that automatically labels failures according to MAST, reaching about 94 percent accuracy and strong agreement with human experts (a Cohen's Kappa of 0.77, which is a statistical measure of agreement beyond chance).
- Case studies showing that you can improve real systems by fixing the failure modes MAST reveals, rather than just reaching for a bigger model.
The taxonomy itself was not invented from a whiteboard. It came from manually annotating over 200 execution traces, each averaging more than 15,000 tokens, using a careful method (grounded theory) and refining the categories until annotators agreed with each other. That grounding is what makes the taxonomy credible.