CHAPTER 05 · Prompt Caching and Context Management · 8 / 8
Key takeaways
- Prompt caching reuses computed state for an exact prompt prefix, at roughly a tenth the cost. It turns the quadratic cost curve back into a linear one.
- Cache hits need exact prefix matches, so put stable content first and variable content last, and never edit earlier messages mid-session; append new ones instead.
- Common cache-busters: changing tools, switching models, changing sandbox or working directory, editing
AGENTS.mdmid-run. Lock instructions before long sessions. - Compaction keeps you under the context window by summarizing old turns (oldest tool outputs first), while preserving the original request and recent exchanges.
- Truncation (head plus tail) tames giant tool outputs, and context diffing avoids re-sending unchanged context. Compaction is lossy, so keep sessions focused.
Original sources: OpenAI's Unrolling the Codex agent loop, Anthropic's How Claude Code works, and the Deep Research Agents survey (arXiv:2506.18096).