CHAPTER 05 · Prompt Caching and Context Management · 3 / 8
Context management: staying under the ceiling
Caching saves money but does not shrink the prompt. To stay under the context window, harnesses use three moves, listed in the Deep Research survey as the standard menu:
- Extend the window. The blunt approach: use a model with a bigger window (Gemini offers up to a million tokens). Simple, but expensive and not always available.
- Compress intermediate steps. Summarize older parts of the conversation so they take fewer tokens.
- Use external storage. Write intermediate results to files or a database and pull them back only when needed (Chapter 9's file-based memory).
The workhorse is move 2, called compaction. When the token count crosses a threshold, the harness replaces the long history with a shorter representative summary, freeing space while keeping the agent's understanding of what happened. Anthropic's harness clears the oldest tool outputs first (those are the bulkiest and least likely to matter later), then summarizes the conversation if needed; your requests and key code snippets are preserved.
Codex started with a manual /compact command and later moved to automatic compaction when an auto_compact_limit is exceeded. The newest Responses API even has a dedicated compaction endpoint that returns a special compaction item carrying an opaque encrypted blob that preserves the model's latent understanding, more faithful than a plain text summary.
Two related techniques round out the toolkit:
- Truncation. Individual tool outputs can be huge (a 5,000-line test log). Rather than dropping them, harnesses keep the head and tail and elide the middle, with a marker like
... (4,800 lines omitted) .... The beginning and end usually carry the signal; the middle is noise. (We build this in Chapter 6.) - Context diffing. Codex tracks a reference context and, when little has changed between turns, avoids re-injecting the full context. Less to send, better cache hit rates.
One honest caveat the sources raise: compaction can lose detail. A summary is lossy by definition, and on a long, intricate task that loss can hurt later reasoning. So compaction is a necessary evil, not a free lunch; it is one more reason to keep sessions focused.