CHAPTER 02 · The Case For Multi-Agent: Anthropic's Research System · 7 / 7
Key takeaways
- Anthropic's research system uses an orchestrator-worker design: a lead agent plans and delegates, parallel subagents explore in isolated context windows, and a citation agent verifies sources.
- It beat a single-agent baseline by about 90 percent, mostly because it spends far more tokens across parallel context windows.
- Prompt engineering is the main control lever: teach delegation, scale effort to complexity, design tools carefully, start broad then narrow, and guide the thinking process.
- Evaluate by outcomes, not process: start with small test sets, use a single strong LLM judge, and keep humans in the loop for edge cases.
- Production demands recovery from failures, tracing for non-deterministic debugging, and rainbow deployments; synchronous execution remains a bottleneck.
- Reserve multi-agent for high-value, parallelizable, breadth-heavy tasks, not for interdependent work like coding.
Continue to Chapter 3 for the opposing view.