Skip to slide
Chapter 4 · Tokens, Context Windows, and the Quadratic Problem
26 / 142

CHAPTER 04 · Tokens, Context Windows, and the Quadratic Problem · 6 / 6

Key takeaways

  • Text becomes tokens (about 4 characters each); the model reads all input tokens before producing any output token, so bloat taxes every later call.
  • The context window is a hard ceiling on input plus output tokens for one inference call.
  • Because every turn re-sends the full history, total tokens over a session grow quadratically with the number of turns. Doubling length roughly quadruples cost.
  • Editing-heavy sessions fill the window fastest. The healthy habit is one focused task per session, with delegation for the rest.
  • You can estimate tokens cheaply (4 chars each) and build budgets on top of that estimate, which is exactly what later chapters do.

Original source: the "Inside the Codex Agent Loop" deep-dive based on Michael Bolin's Unrolling the Codex agent loop.


← → arrow keys work too