CHAPTER 03 · How Memory Is Built and Retrieved · 1 / 7
Five families of memory mechanisms
The "Practical Guide" lays out five families of mechanisms. You can think of these as five different engineering strategies for the write and manage phases, each with its own strengths and its own way of failing.
Context-resident compression
These are the "stay in context" strategies: sliding windows, rolling summaries, and hierarchical compression. The idea is to keep memory inside the context window by repeatedly shrinking it. The author warns that rolling summaries feel clean but are not, because each round of summarization throws away detail (the drift problem in Chapter 4). The honest practical note: when a coding assistant compresses a conversation that has grown too large, you are often better off starting a new thread than trusting the compression.
Retrieval-augmented stores
This is RAG applied to the agent's own interaction history rather than to static documents. Past observations are turned into embeddings and retrieved by similarity. It is powerful for long-running agents with deep history, but retrieval quality becomes the bottleneck fast. If the embeddings do not capture intent well, you miss relevant memories and surface stale ones. A telling example: questions like "what happened last Monday" tend to retrieve poor results, because similarity search is bad at time-based queries.
Reflective self-improvement
Here the agent writes verbal post-mortems and stores the conclusions to do better next time. Reflexion and ExpeL are the named examples, and the author groups "dream"-based reflection systems here too. The idea is compelling (agents learn from mistakes), but the failure mode is severe, which Chapter 4 covers: a wrong lesson, once stored, can be reinforced.
Hierarchical virtual context
This is the operating-system-inspired approach, exemplified by MemGPT. The main context window acts like RAM, a recall database acts like disk, and archival storage acts like cold storage, with the agent managing its own paging between tiers. The author's blunt verdict: the overhead of maintaining these separate tiers is burdensome and tends to fail, and despite the idea being around for years, he has not seen it used in production.
Policy-learned management
This is the frontier approach: train the agent, using reinforcement learning, to decide when to store, retrieve, update, summarize, or discard memories. The author sees promise here but notes there are not yet practical, ready-made tools for builders, nor much real production use. The SSGM paper in Chapter 5 discusses this same direction in more formal detail.
The meta-lesson across the five families: there is no single best mechanism. Each trades off differently, and the article's guidance is to choose based on your actual need rather than reaching for the most sophisticated option.