CHAPTER 02 · The Case For Multi-Agent: Anthropic's Research System · 4 / 7
Evaluating a system that has no single right path
Evaluation is hard here because multi-agent systems are non-deterministic: two runs can take completely different but equally valid routes to the same answer. One agent searches three sources, another searches ten. So evaluation has to focus on outcomes, not on whether a specific process was followed. The articles describe a layered approach.
Start small. Early in development, effect sizes were huge (success rates jumping from 30 percent to 80 percent), so a test set of about 20 representative queries was enough to detect changes. This pushes back on the belief that you always need large evaluation sets.
Use LLM-as-a-judge. A separate model graded outputs against a rubric covering factual accuracy, citation accuracy, completeness, source quality, and tool efficiency, producing a 0.0 to 1.0 score and a pass or fail grade. A single well-designed judge prompt proved more consistent than several specialized judges.
Keep humans in the loop. Human testers caught what automation missed: hallucinated answers on unusual queries, subtle biases, and a tendency in early agents to prefer SEO-optimized content farms over authoritative sources like academic PDFs. That finding led to adding source-quality heuristics to the prompts.