CHAPTER 05 · Prompt Caching and Context Management · 1 / 8
Prompt caching: paying once for the stable part
When the model processes a prompt, it builds up an internal state (the key-value cache, or KV cache) for those tokens. Prompt caching means the provider saves that computed state for a prefix of your prompt the first time it sees it, and reuses it on later calls that share the exact same prefix. You skip recomputing the cached part.
The critical rule, and it is unforgiving: cache hits only work for exact prefix matches. The provider can reuse the work only up to the first token that differs. So the structural advice from Chapter 3 pays off here:
- Put static content (instructions, tool definitions, examples) at the front.
- Put variable content (the latest user message, the newest tool result) at the end.
What caches well in an agent session: the system prompt (identical every turn) and the bulk of the conversation history (unchanged from the previous turn). What never caches: the latest user message and the most recent tool result, because they are always new.
The payoff is large. Cached prefix tokens cost roughly a tenth of fresh input tokens. This is why the quadratic growth in tokens sent does not become quadratic growth in cost: most of what you send each turn is a cached prefix you are billed a tenth for. Codex's lead engineer puts it precisely: with cache hits, sampling the model is linear rather than quadratic. That one optimization is the difference between a usable agent and an unaffordable one.