Skip to slide
Chapter 4 · Tokens, Context Windows, and the Quadratic Problem
20 / 142

CHAPTER 04 · Tokens, Context Windows, and the Quadratic Problem

Tokens, Context Windows, and the Quadratic Problem

Before the model can do anything with your prompt, the prompt is turned into numbers. This step is called tokenization: the text is chopped into tokens (chunks roughly the size of a short word or word-piece), and each token is mapped to an integer that indexes the model's vocabulary. A rough rule of thumb the harnesses use is about 4 characters per token. The model then samples output tokens one at a time, and those get decoded back into text. The token-by-token streaming you see in the terminal is literally this sampling process exposed to you.

Two facts about tokens shape the entire economics of agents.

Fact one: the model must read every input token before it can write a single output token. A bloated AGENTS.md, a huge file you read, a verbose tool result: all of it adds latency and cost to every inference call for the rest of the session, not just the next one.

Fact two: there is a hard ceiling called the context window. This is the maximum number of tokens the model can use for one inference call, and it counts both input and output. An agent that makes hundreds of tool calls in one turn can run straight into that ceiling. Managing it is one of the harness's core jobs.

← → arrow keys work too