Skip to slide
Chapter 2 · The Case For Multi-Agent: Anthropic's Research System
11 / 41

CHAPTER 02 · The Case For Multi-Agent: Anthropic's Research System · 3 / 7

Prompt engineering becomes the main control lever

A central lesson from both articles is that in a multi-agent system, prompt engineering gets much harder and much more important, because small wording changes ripple into emergent coordination behavior. Anthropic identified several principles.

Think like your agent. The team ran simulations watching agents work step by step with the real tools and prompts. This revealed failure patterns: agents kept searching after they had enough, repeated the same queries, or picked the wrong tools. Watching the behavior let them predict and fix these problems through wording.

Teach the orchestrator to delegate. The lead's prompt must produce detailed, unambiguous subtask descriptions. Each subagent needs a clear objective, an output format, guidance on which tools to use, and precise task boundaries. Without this, subagents duplicated work or left gaps. The articles give a real example: one subagent investigated the 2021 semiconductor shortage while two others ran nearly identical searches on 2025 supply chains, wasting effort.

Scale effort to query complexity. Agents are bad at judging how much effort a task deserves, so the team wrote explicit effort scaling rules into the prompts: a simple fact check uses one agent and 3 to 10 tool calls; a direct comparison uses 2 to 4 subagents with 10 to 15 calls each; a complex problem might use 10 or more subagents with clearly divided responsibilities. This prevents both over-investing in easy queries and under-investing in hard ones.

Design tools carefully. The interface between an agent and its tools matters as much as a human-computer interface. A poorly described tool can send an agent down completely the wrong path. With MCP servers exposing many external tools of varying quality, this gets worse. The team gave agents explicit heuristics: examine all available tools first, match the tool to the user's intent, and prefer specialized tools over generic ones. They even built a tool-testing agent that repeatedly tried a flawed tool and then rewrote its description to avoid the mistakes, which cut task completion time by about 40 percent for later agents. This is a striking example of agents improving their own environment.

Let agents improve their own prompts. Claude 4 models proved capable of acting as their own prompt engineers: given a failing scenario, they could diagnose what went wrong and suggest better wording.

Start wide, then narrow. Agents tended to jump straight to overly specific queries that returned little. Prompting them to start broad, survey the landscape, and then narrow down mirrors how skilled human researchers work.

Guide the thinking process. The team used extended thinking as a controllable scratchpad for the lead to plan before acting, and interleaved thinking for subagents to evaluate tool results, spot gaps, and refine their next query.

← → arrow keys work too