CHAPTER 03 · How Memory Is Built and Retrieved · 3 / 7
How a production pipeline does write and manage: AgentCore
The AWS "AgentCore" article shows these phases in a real managed service, and it is the best concrete example of the often-neglected manage phase actually being done well.
Extraction. When new events arrive, an asynchronous process uses LLMs to analyze the conversation and pull out meaningful information into a predefined schema, according to whichever memory strategies you configured (semantic, preference, summary). A nice detail: it distinguishes meaningful content from chatter. "I'm vegetarian" should be remembered; "hmm, let me think" should not. Multiple memories can be extracted from one event, and the strategies run in parallel.
Consolidation. This is the heart of the manage phase, and it is memory consolidation done explicitly. Rather than blindly appending each new memory, the system:
- Retrieves the most semantically similar existing memories from the same namespace and strategy.
- Sends the new memory plus those existing ones to an LLM with a consolidation prompt, which decides on an action: ADD (the information is new), UPDATE (it complements or updates an existing memory), or NO-OP (it is redundant). The prompt is designed so that "loves pizza" and "likes pizza" are treated as the same and do not trigger a needless update.
- Applies the action while keeping an immutable audit trail: outdated memories are marked INVALID rather than instantly deleted.
The article gives clear examples of the hard cases this handles. Related facts mentioned at different times ("allergic to shellfish" in January, "can't eat shrimp" in March) get recognized as related and merged without creating duplicates or contradictions. Conflicting information is resolved by prioritizing recency while preserving history: if a budget changes from 500 to 750, the new value becomes active and the old one is marked inactive rather than erased. Out-of-order events are handled through careful timestamp tracking. And if consolidation fails for one memory, it does not break the others; the system retries with exponential backoff, and if it ultimately fails, it still stores the memory to avoid losing information.
There is a useful performance result here too. AgentCore was benchmarked against a RAG baseline that kept the full conversation history. The RAG baseline scored well on factual recall (because it had everything) but poorly on inferring preferences. The memory system gave up a little factual accuracy in exchange for very high compression rates (89 to 95 percent), which means bounded context size, faster inference, and lower cost at scale, and it clearly won on preference inference. The lesson: for tasks that need inference rather than verbatim recall, distilled memory beats raw history, and it scales far better.