CHAPTER 09 · Memory Outside the Weights
Memory Outside the Weights
The model is stateless. Each new session starts with a blank context window and no memory of yesterday (Chapter 2). Yet a good coding agent seems to know your project: that you use pytest, that you indent with four spaces, that the auth module is fragile. Where does that knowledge live, if not in the model? Outside the weights, in files the harness loads. This is the "knowledge" layer (layer 2) and "memory management" (pillar 2), and it is one of the highest-leverage parts of the whole harness.
The "Seven Pillars" breakdown frames the core idea well: harness intelligence lives largely in memory. There are a few distinct kinds, and mixing them up is a common mistake.
The building blocks
| Mechanism | What it is | Loads when |
|---|---|---|
CLAUDE.md / AGENTS.md | Project instructions you write: conventions, test commands, architecture notes | Every session, near the front of the prompt |
| Auto memory | Notes the agent writes itself as it works (project patterns, your preferences) | At session start (a capped amount, e.g. first 200 lines or 25KB) |
Rules (.claude/rules/) | Path-scoped instructions that apply only when matching files are touched | When relevant files are in play |
| File-based scratchpad | Intermediate results the agent writes to disk during a task | On demand, when the agent reads them back |
CLAUDE.md (Anthropic) and AGENTS.md (Codex) are the same idea under different names: a markdown file in your repo that gets injected into the prompt every session. It is where you put the things you would tell a new teammate on day one. Critically, Claude Code reloads CLAUDE.md after compaction, so your conventions survive even when the conversation gets summarized. The lesson the docs repeat: put persistent rules in CLAUDE.md, not in chat, because chat gets compacted away.
There is an empirical reason to care. The "Everything About Codex" guide cites research (an arXiv paper on the impact of AGENTS.md files) that a well-crafted instruction file measurably improves agent efficiency. And from Chapter 5, remember the cost angle: AGENTS.md sits in the cached prefix, so a stable one is also a cheap one. Editing it mid-session throws away your biggest cache discount.
The two-tier pattern
The "Seven Pillars" breakdown describes a sharp pattern for memory that is worth adopting: split it into two tiers.
- Tier 1, shared knowledge (
.claude/rules/, committed to git): facts the whole team should have. Confirmed error patterns and their meaning, known performance traps, API quirks, anti-hallucination constraints. A new teammate gets 80% of the value on day one just by cloning the repo. - Tier 2, personal memory (
~/.claude/..., per user): your own accumulated investigation outcomes and preferences.
The promotion workflow is the elegant bit. The agent discovers something, saves it to personal memory, you validate it, and then it gets promoted to the shared tier with a commit. Memory becomes a real, reviewable team practice instead of a magic black box that silently changes behavior. This connects to pillar 7, portability: because all of this lives in the repo, a colleague who clones it gets the whole configured agent, not just the code.
Why files, not a giant prompt
You might ask why not just stuff everything into one huge system prompt. Two reasons from earlier chapters. First, tokens: everything loaded costs you on every call (Chapter 4), so you load instructions always but defer bulky reference material until needed. Second, external storage is the third context-management move (Chapter 5): writing intermediate work to files and reading it back keeps the live context small. The Deep Research survey notes that frameworks like Manus, OWL, and OpenManus all use external file systems to store intermediate outcomes precisely because the context window cannot hold everything.