CHAPTER 03 · How Memory Is Built and Retrieved · 4 / 7
How memories get read: retrieval is more than similarity
The "Building Long-Term Memory" article makes the strongest case that naive retrieval is not enough. The obvious approach is pure semantic search: embed the query, find the closest memories, return the top matches. This works for simple cases but misses important dimensions for agents:
- Recency: a memory from yesterday is often more relevant than one from six months ago, even if the older one is semantically closer.
- Context match: memories from the same project or workflow should be weighted higher than unrelated ones.
- Outcome quality: memories of successful interactions are usually more useful than failures (unless you are specifically trying to avoid repeating a mistake).
The recommended answer is multi-factor relevance scoring: combine semantic similarity, temporal relevance (with decay for older memories), context alignment, and success weighting into one composite score, with weights tuned to the use case. On top of that, good systems apply diversity and deduplication: remove near-duplicates, avoid returning ten memories from the same day, and cluster similar memories to return representatives. This gives the agent a broad, useful view instead of repetitive noise.
Finally, retrieved memories must be put back into the loop efficiently. Because the context window is limited, the system retrieves only the top relevant subset, summarizes older context to save tokens, and prioritizes recent high-relevance memories. Once injected, these memories act as grounding: the agent answers based on what it knows about this user and this project, which reduces hallucination and improves consistency.