CHAPTER 04 · Tokens, Context Windows, and the Quadratic Problem · 6 / 6
Key takeaways
- Text becomes tokens (about 4 characters each); the model reads all input tokens before producing any output token, so bloat taxes every later call.
- The context window is a hard ceiling on input plus output tokens for one inference call.
- Because every turn re-sends the full history, total tokens over a session grow quadratically with the number of turns. Doubling length roughly quadruples cost.
- Editing-heavy sessions fill the window fastest. The healthy habit is one focused task per session, with delegation for the rest.
- You can estimate tokens cheaply (4 chars each) and build budgets on top of that estimate, which is exactly what later chapters do.
Original source: the "Inside the Codex Agent Loop" deep-dive based on Michael Bolin's Unrolling the Codex agent loop.