A coding agent that only has a session forgets the next time you open the chat. Memory cannot be one store. It has to be a hierarchy that is cheap to inject and cheap to reverse: sessions, atomic facts, compiled knowledge, then weights.
Why one store fails
A coding agent repeats work a person would not: the same grep, the same README, the same constraint the last session already proved. That costs tokens and wall clock, and it fails worse when the workspace is misleading, because a wrong file will beat a fact the agent already earned if that fact is not stored and trusted.
Three common designs collapse because they treat every lesson as the same kind of object.
- The live window mixes current state with old evidence, then truncation or a lossy summary throws the evidence away.
- A flat vector store turns each turn into a fragment, then nearest-neighbor search cannot tell a current constraint from a dead end, or a week-one preference from a week-four decision.
- A fine-tune is slow to run and hard to undo, and it does not transfer cleanly when the base model changes.
Most of what the agent needs tomorrow is a small set of durable facts and procedures. Those belong in text first, and weights come last.
The stack
Memory is a hierarchy of units with different cost, lifetime, and editability. The base is large and cheap to write. The top is small and expensive to change.
- L0, session. The host chat: user turns, tool calls, errors, diffs. This is ground truth and the source for everything above it. Many hosts already persist it, so you do not need a second archive on day one.
- L1, atomic facts. One unit of experience: a fact, a constraint, a gotcha, a decision that proved true. Small enough to retrieve, easy to edit or mark stale. In our work this is a knowledge card, a markdown file with a title and a body.
- L2, compiled knowledge. Many facts, over time, become a procedure or a handbook you query when the task needs a method rather than a single fact. Compile once, keep current, inject the index, and fetch a page when the title is not enough.
- L3, weights. A LoRA or fine-tune can absorb patterns that text memory keeps restating. Open this layer when L1 and L2 stop moving the metric, eval it on the same tasks, and compare no memory, text only, weights only, and both.
Why the split
- The window is scarce. You cannot inject a week of sessions into every prompt. You can inject a short trusted block of titles, an index, or a procedure name, and let the agent fetch a body when the title is not enough.
- Layers have different clocks. A session changes every turn, a fact changes when a task ends, a compiled page changes after many tasks, and weights change after a training job. Mix those clocks in one store and retrieval cannot know what is stable.
- Upper layers hold structure, lower layers hold evidence. The compiled page says how to do the work. The fact says what proved true on a given day. The session shows the exact error. Compression without a drill-down path is amnesia.
- Wrong memory must be cheap to fix. You can mark one fact stale and rebuild a handbook page. You cannot surgically unlearn one sentence from a LoRA, which is why L3 comes last.
Promotion and demotion
Promotion is the write path up the triangle. It is extraction plus write, not a dump of the chat, and not a new training run.
- L0 → L1. Read the session and write the outcome that proved true. Do not write the tour of files, the plan, or the guess. One fact, one card.
- L1 → L2. When the same procedure keeps working across many facts, compile the handbook page. Do this after repetition, not after a single success.
- L2 → L3. Absorb a stable procedure into weights only when the text form is already useful and you need the model itself to act as if it already knew.
Run promotion twice if you can: in flow, when the fact appears, and after the session, for what the in-flow path missed. End-only promotion drops mid-task gotchas. In-flow-only promotion misses the view from the end.
Demotion is the health path down.
- Mark a fact STALE when code or reality has moved, and stop retrieving it.
- Rebuild the compiled page from the facts that still hold, instead of patching a contradiction in place forever.
- Do not put unstable facts into weights. If a LoRA has absorbed a wrong habit, the cheap fix was never to have promoted it.
Keep the stack small enough to trust
A store that only appends will rot. Do this work when the user is not waiting.
- Clean. Merge near-duplicates, drop units that no retrieval uses, cap how many items enter the prompt, and keep an index in context instead of the raw log.
- Repair. Let a later run confirm or flag a fact so retrieval can bias by trust, and lint compiled pages for contradiction, orphans, and gaps. A failure paired with a success for the same task is a sharper patch than a success summary alone.
- Organize. Rebuild instead of appending forever. Keep an index and a chronological log, and file compiled pages so the bank does not grow dead twins.
Build in this order
Make L1 facts small, trusted, and cheap to inject. Compile L2 when facts repeat. Touch L3 when the cheaper layers stop moving the score. Keep a path from a compiled claim back to the session that earned it. If you cannot walk that path, you do not have memory. You have a summary you cannot check.
Appendix
- Packer et al., MemGPT: Towards LLMs as Operating Systems, 2023
- Lin et al., Sleep-time Compute, 2025
- Karpathy, LLM Wiki
- McClelland, McNaughton, and O'Reilly, Complementary Learning Systems
- Tulving, episodic vs semantic memory
- Asawa et al., Continual Learning Bench knowledge cards