Skip to slide
Chapter 15 · Deep Research Agents and How They Are Trained
114 / 142

CHAPTER 15 · Deep Research Agents and How They Are Trained · 4 / 10

How they are trained

This is the genuinely new material. Earlier chapters improved agents by engineering the harness; here you improve the model itself. The DR survey lays out three families.

Supervised fine-tuning (SFT). Train the model on curated examples of good agent behavior: how to formulate search queries, how to structure a report, how to call tools. It improves retrieval quality and reduces hallucination, but it is confined to offline, static data and generalizes only so far.

Reinforcement learning (RL). Let the agent learn from reward signals on real tasks: did the retrieval help, was the answer correct, was the tool call appropriate? RL adapts better to open-ended environments than SFT. The survey highlights GRPO (Group Relative Policy Optimization) as the workhorse for DR systems, contrasting it with the older PPO (Proximal Policy Optimization). The key idea: GRPO drops PPO's separate value network and instead computes advantages relative to a group of responses, which gives richer gradient signal, faster convergence, and fewer conflicting objectives. You do not need the math to take the lesson: RL is how you teach an agent to search and use tools well, and GRPO is the currently favored recipe. Industrial systems (Gemini DR, Grok) use proprietary RL; academic ones favor transparent GRPO-based designs.

Non-parametric continual learning. The newest and most relevant to harness builders. Instead of updating model weights (expensive, slow), the agent improves at runtime by optimizing its external memory, workflows, and tools. The main technique is case-based reasoning (CBR): the agent retrieves, adapts, and reuses past problem-solving trajectories from a "case bank." Unlike RAG (which retrieves static text), CBR retrieves whole reasoning trajectories and adapts them to the new task. This is genuine self-improvement without retraining, and it is well-suited to complex agents precisely because it sidesteps the cost of parameter updates. It is the training-side cousin of the auto-memory and shared-rules ideas from Chapter 9: the agent gets better by accumulating and reusing experience, stored outside the weights.

← → arrow keys work too