A coding agent can rewrite its prompts and still rediscover the same schema the next day. Harness RSI changes how the agent runs. Continual learning changes what it keeps. A strong runtime with weak memory still starts from zero.
Why one loop is not enough
A coding agent repeats work a person would not: the same grep, the same README, the same join key the last session already proved. Recent harness work is a clear map of how that agent should run. It is a weaker map of what the agent should keep.
Lilian Weng's Harness Engineering for Self-Improvement is the best survey of the first path. She treats the harness as the layer around a base model: planning, tools, context, artifacts, and checks. Near-term RSI in that line is not weight rewrite. It is better machinery for getting better answers.
That machinery can still forget the world it just explored. Meta-Harness, ADAS, AFlow, STOP, Self-Harness, Darwin Gödel Machine, ACE, and MCE mostly optimize how one rollout runs, or how the harness itself evolves. Weng names continual learning, self-play, synthetic data, and test-time training as part of the wider RSI story, then leaves that thread for later.
That is the opening. Loop A is mapped. Loop B still needs its own treatment.
The two loops
Agents improve in two ways that people often mix. They can change how they run, and they can keep better state across related work in the same environment.
- Loop A, harness RSI. Change prompts, tools, workflow, permissions, or harness code so future rollouts work better. The question is: how should the agent run?
- Loop B, continual learning. Keep durable knowledge so later work in the same environment scores higher than early work. The question is: what should the agent remember?
Loop A is a search over the runtime. Loop B is a write path over state. Mix them and you will treat a better prompt as if it were a learned map of the repo, or treat a notebook dump as if it were a better harness.
RSI needs both. A better harness with no durable state still starts from zero. Durable state inside a bad harness still cannot act.
Why the split
- They edit different objects. Loop A edits the machine. Loop B edits the notebook, the cards, the weights. A prompt tweak is not a fact about the database.
- They have different clocks. A harness search can run overnight. A fact should be written when a task ends. Weights should move after a training job. Mix those clocks in one memory slot and retrieval cannot tell a current constraint from a dead end.
- The learning signal is not raw reward. Task difficulty varies. Strong models can score high with no learning. Compare stateful performance to a reset or fresh-start baseline. That gap is the learning signal.
- Wrong memory must be cheap to fix. You can mark a fact stale. You cannot surgically unlearn one sentence from a LoRA, and you should not wait for the next harness search to delete a false schema.
What "improve with experience" has to mean
Online improvement needs shared structure across related work, not only a longer transcript. The structure is usually latent: conventions, maps, habits, layouts, dynamics. A system that only reacts to the current prompt cannot use that structure later.
Concept drift is part of the problem. Stale memory must be updated or dropped. Success looks like reuse: fewer rediscovery steps, better late-run score, fewer repeated mistakes.
Why more memory often fails
Dedicated memory modules often fail to beat simple in-context learning. The failure is usually the storage policy, not the absence of a slot.
- Store too much. The window fills with a tour of files, and the constraint that proved true is buried.
- Store the wrong thing. A strategy slogan is not a schema. "Be careful with joins" will not save the next query.
- Store fluent junk. A well-written summary can still be false, and retrieval will prefer fluency over truth.
- Fail to reuse. A fact that is never injected is not memory. It is an archive.
Agents overfit recent observations and underuse older, still-relevant structure. More retrieval does not fix this by itself. Context-engineering research makes the same point for trajectories: dumping everything into context is not a policy.
The companion piece
What a good Loop B surface looks like
Keep durable state outside the growing chat blob, often as files or short structured notes. Write for the next actor, not for the current transcript.
- Prefer entities, parameters, maps, encodings, and counts over strategy slogans.
- Merge conflicts. Drop weak one-offs.
- Make the memory inspectable. If humans and agents cannot read it, they cannot debug it.
- Separate mechanism (how memory is managed) from content (what is stored).
- Do not require weight updates to count as continual learning. Non-parametric state is enough for a large class of agent work.
Reflection is the extract step. Memory is the store.
What a bench forces into the open
Continual Learning Bench is one place that forces Loop B into the open. It is an example, not the whole field. Each task is a sequence of instances in one shared environment with discoverable latent structure: schema conventions, spectrum occupancy, opponent habits, codebase layout, demand patterns.
The bench uses reward plus gain so high base skill does not hide missing online learning. Thin memory systems there, for example end-of-instance reflected notes, show the same pattern as the wider literature: the storage policy matters more than the presence of a memory slot.
Use the bench to test ideas. Do not treat any one task list as the definition of continual learning. For longer horizons and mechanism-neutral curves, see
What stays open
- Horizon. How long must a run be before durable state pays off?
- Drift. When should memory clear, mark stale, or merge?
- Wrong memory. How do we stop systems from writing confident false structure?
- Reward hacking of memory. Notes that help the scorer but not the true latent object.
- Transfer. Does a memory policy learned in one domain help in another?
- Joint path. When should Loop A (edit the harness) and Loop B (edit the notebook) run together?
- Humans. Where should oversight sit when the agent updates its own long-lived state?
Build in this order
Make Loop B a storage policy, not a longer context. Score it against a fresh-start baseline, on the same model, with the rest of the harness fixed. Audit whether stored state matches the true latent structure, or only correlates with score. Report model tier separately from memory method. A stronger model is not the same claim as a better learner.
If later work in the same environment is not cheaper, cleaner, or more correct than early work, you do not have continual learning. You have a harness that still starts from zero.
Appendix
- Weng, Harness Engineering for Self-Improvement, Lil'Log, Jul 2026
- Asawa et al., Continual Learning Bench
- Related harness and context work as summarized in Weng: ACE, MCE, Meta-Harness, Self-Harness, and others