CHAPTER 15 · Deep Research Agents and How They Are Trained · 10 / 10
Key takeaways
- A deep research agent browses, retrieves, reasons, and writes cited reports. Architecture: a MasterAgent plans and spawns SubAgents, and a ReviewAgent verifies citations before delivery.
- Useful taxonomy: static vs dynamic workflows; planning strategies (planning-only, intent-to-planning, unified); single-agent (easy to train end to end) vs multi-agent (scales but hard to train); and three memory mechanisms (bigger window, compression, external storage).
- Fact-checking is the research form of "an agent cannot mark its own homework": cross-check claims across independent sources and reflect before finalizing.
- Agents are improved by training, not just prompting: SFT (curated examples), RL (reward signals, with GRPO the favored recipe over PPO), and non-parametric continual learning via case-based reasoning (reuse past trajectories without updating weights).
- Today's DR agents show jagged intelligence and operate mostly in digital environments; the surveys conclude that building a reliable harness often matters more than the underlying LLM.
Original sources: "Deep Research Agents: A Systematic Examination and Roadmap" (arXiv:2506.18096) and "Deep Research of Deep Research: from transformer to agent."