Skip to slide
Chapter 15 · Deep Research Agents and How They Are Trained
120 / 142

CHAPTER 15 · Deep Research Agents and How They Are Trained · 10 / 10

Key takeaways

  • A deep research agent browses, retrieves, reasons, and writes cited reports. Architecture: a MasterAgent plans and spawns SubAgents, and a ReviewAgent verifies citations before delivery.
  • Useful taxonomy: static vs dynamic workflows; planning strategies (planning-only, intent-to-planning, unified); single-agent (easy to train end to end) vs multi-agent (scales but hard to train); and three memory mechanisms (bigger window, compression, external storage).
  • Fact-checking is the research form of "an agent cannot mark its own homework": cross-check claims across independent sources and reflect before finalizing.
  • Agents are improved by training, not just prompting: SFT (curated examples), RL (reward signals, with GRPO the favored recipe over PPO), and non-parametric continual learning via case-based reasoning (reuse past trajectories without updating weights).
  • Today's DR agents show jagged intelligence and operate mostly in digital environments; the surveys conclude that building a reliable harness often matters more than the underlying LLM.

Original sources: "Deep Research Agents: A Systematic Examination and Roadmap" (arXiv:2506.18096) and "Deep Research of Deep Research: from transformer to agent."

← → arrow keys work too