Skip to slide
Chapter 5 · Prompt Caching and Context Management
35 / 142

CHAPTER 05 · Prompt Caching and Context Management · 8 / 8

Key takeaways

  • Prompt caching reuses computed state for an exact prompt prefix, at roughly a tenth the cost. It turns the quadratic cost curve back into a linear one.
  • Cache hits need exact prefix matches, so put stable content first and variable content last, and never edit earlier messages mid-session; append new ones instead.
  • Common cache-busters: changing tools, switching models, changing sandbox or working directory, editing AGENTS.md mid-run. Lock instructions before long sessions.
  • Compaction keeps you under the context window by summarizing old turns (oldest tool outputs first), while preserving the original request and recent exchanges.
  • Truncation (head plus tail) tames giant tool outputs, and context diffing avoids re-sending unchanged context. Compaction is lossy, so keep sessions focused.

Original sources: OpenAI's Unrolling the Codex agent loop, Anthropic's How Claude Code works, and the Deep Research Agents survey (arXiv:2506.18096).


← → arrow keys work too