CHAPTER 05 · Recursive Language Models, Thinking Beyond the Memory Limit · 3 / 6
The recursive part
The name comes from the key move: when a chunk is still too big or complex to handle in one go, the model calls itself on that chunk. Recursion means a process that invokes a fresh copy of itself to handle a smaller piece of the same kind of problem. A "root" model orchestrates the work; it spawns sub-calls, each a fresh model instance, to digest individual pieces, then gathers their results back together.
A simple analogy. Imagine a lead researcher handed a thousand-page report and one question. They do not read every page themselves. They skim the table of contents, identify the ten relevant chapters, and hand each chapter to an assistant with the instruction "summarize what this says about the question." Each assistant, if their chapter is still huge, can recruit their own helper for sub-sections. Finally the lead researcher combines the assistants' notes into one answer. The lead never had to hold the whole report in their head at once. That is exactly how an RLM works, with the model playing every role.