Skip to slide
Chapter 15 · Deep Research Agents and How They Are Trained
115 / 142

CHAPTER 15 · Deep Research Agents and How They Are Trained · 5 / 10

Levels of automation and where agents still fall short

The "Deep Research of Deep Research" survey frames progress in levels of automation and notes current DR agents sit around level three; they search and review well but their genuine research capability is still comparatively weak. It also draws a useful map of environments: the IDE (internet/digital environment), the SEE (simulation experimental environment), and the REE (real experimental environment). Today's agents operate mostly in the IDE; reaching the REE (running real experiments) needs better physical perception, more tools, and sometimes embodiment. And it names a failure mode worth remembering: jagged intelligence, the way these systems do some hard things brilliantly while failing at simpler, closely related ones. That unevenness is exactly why the verification habits from Chapter 13 matter so much.

The survey's most quotable line for harness builders: "Harness is an agentic architecture that allows multiple agents to work with shared context across different sessions and context windows. Building reliable harnesses for DR sometimes matters more than the LLMs." After fifteen chapters, that should feel like a homecoming.

← → arrow keys work too