Skip to slide
Chapter 4 · Tokens, Context Windows, and the Quadratic Problem
23 / 142

CHAPTER 04 · Tokens, Context Windows, and the Quadratic Problem · 3 / 6

How it works: estimating before you spend

You do not need an exact tokenizer to reason about this; a good estimate is enough to build budgets and warnings. Real harnesses use byte-based heuristics (roughly 4 characters per token) for fast counting, and only fall back to a precise tokenizer when it matters. Here is how to think about the cost of a whole session.

If a turn's prompt has P tokens and produces O output tokens, that turn's cost is roughly P * input_rate + O * output_rate. Because P grows each turn, summing over a session gives you that quadratic shape. The estimator below makes this concrete and even shows the difference caching makes (foreshadowing Chapter 5).

← → arrow keys work too