CHAPTER 06 · Glossary: Planning and Reasoning · 12 / 12
Working at scale
Inference-time compute (test-time scaling)
Inference-time compute, also called test-time scaling, means spending more computational effort at the moment of answering a question, as opposed to during training. It is the unifying theme of folder 02. Writing a longer chain of thought (Chapter 1), generating and checking many candidate solutions (Chapter 3), and exploring a huge input recursively (Chapter 5) are all ways of spending more effort at answer time to handle harder problems. The core insight is that you can make a model effectively smarter not only by training it more, but by letting it think harder when it actually faces a tough question.
Context window and long context
The context window is the maximum amount of text, measured in tokens, that a model can consider at once, its working memory. "Long context" refers to the challenge of handling very large inputs. The trouble (Chapter 5) is twofold: text beyond the window simply cannot be seen, and even text that fits can overwhelm the model, causing it to miss details as the input grows. Recursive Language Models address this not by enlarging the window but by giving the model tools to navigate inputs larger than its memory.
REPL
REPL stands for Read-Eval-Print Loop, a small interactive programming environment where you can type a piece of code, have it run immediately, and see the result, then type the next piece. In Chapter 5, the giant input is placed into a REPL as a variable, and the model writes code to peek into it, search it, and slice it. The REPL is what turns a too-large input from "text to read" into "data to explore," which is the heart of the Recursive Language Models idea.
Recursion
Recursion is when a process solves a problem by calling a fresh copy of itself on a smaller piece of the same kind of problem, then combining the pieces. In Chapter 5, when a chunk of input is still too big to handle directly, the model calls itself on that chunk, and that call may call itself again, until the pieces are small enough to answer. A "root" model coordinates the whole effort and assembles the results. It is the divide-and-conquer instinct, applied so a model can work through inputs it could never hold all at once.