Field note

Reason, act, then reflect

A coding agent can finish a task and still know nothing about the repo tomorrow. ReAct reasons and acts. It does not extract what the environment taught. That extract step belongs in the harness. The reflection prompt is only the knob.

Why ReAct stops too early

Most agent loops still look like ReAct: reason, then act, then repeat until the task ends. That loop can finish a job. It does not, by itself, learn how the environment works.

Yao et al. paired reasoning traces with actions. That improved tool use and transparency. It did not add a learn-the-environment stage after the task. Without that stage, the agent either drops the episode when the chat ends, or dumps the transcript into memory and hopes retrieval will sort it later.

Dumping is not learning. Learning needs a deliberate extract step aimed at environment behavior. Coding agents, analytics agents, and game agents all show the same gap: strong within-task loops, weak across-task learning.

The missing stage

Keep the ReAct core for the task itself. Add one stage after the task completes, or after a clear unit of work.

  • Reason. Decide what to do next inside the task.
  • Act. Call tools, edit files, run checks.
  • Reflect. Read the episode plus prior durable notes, and write an updated model of the environment for the next actor.

During the task, the agent solves. After the task, the agent learns. The next task starts with that model in trusted context, not with a raw history dump. Fail closed if reflection produces noise. Prefer no new notes over wrong notes.

Reflection is not a memory system

Memory systems answer how to store and retrieve: vector databases, graph memories such as Cognee, agent memory layers such as Mem0, file notebooks, card stores, Anthropic memory stores. Reflection answers what to learn from this episode about the environment.

You still need a store. The store is the sink. Reflection is the filter and the teacher. Mix them and three failures show up fast.

  • Retrieve fluent junk. A well-written summary can still be false.
  • Store strategy slogans. "Be careful with joins" is not a schema.
  • Treat memory as learning. A slot that is never filtered is an archive.

The companion piece is the stack: sessions, atomic facts, compiled knowledge, then weights. Reflection is the write path up that stack. is the argument that this write path is a second RSI loop, not a feature of harness search.

Design them apart. Wire them together.

What to reflect depends on the environment

The variable is the learning target, not the presence of a reflect call. Same harness stage. Different learning object.

  • Coding agent. System architecture, module boundaries, external dependencies, build and test quirks, ownership maps.
  • Analytics agent. Database schema, encodings, join keys, traps, drift after migrations.
  • Game agent. Opponent tells, play style, stack tendencies, bluff patterns.
  • Ops agent. Fleet topology, runbook exceptions, alert causes that recur.
  • Support or sales agent. Account facts, constraints, prior commitments.

That is why reflection must be designed per environment class. A generic "summarize the session" prompt will write a tour of files, not a map.

The prompt is a knob, not the stage

People already tune system prompts for role, tools, and style. The reflection prompt is the same kind of knob for the post-task stage. It should name what environment object to extract, what to forbid (transcript dumps, slogans, one-off recipes), how to merge with prior notes, and how to write for the next actor.

Tuning the reflection prompt is useful. Calling the prompt "the harness" is wrong.

  • The harness owns the stage. When it runs, what episode artifacts it sees, what store it writes, how later tasks read the result.
  • The prompt steers extraction. It names the learning target and the bans.
  • The store holds the output. Cards, files, a notebook. Not the transcript.

Dreams as one form of reflection

Anthropic Dreams take an existing memory store and past session transcripts. They produce a new memory store: merge duplicates, replace stale or contradicted entries, surface new insights. The input store is not mutated. You can review and discard the output.

Dreams also take instructions, for example "focus on coding-style preferences; ignore one-off debugging notes." That field is a reflection prompt in practice: it names what to learn and what to ignore.

Dreams sit after work, over many sessions. Per-task reflection sits after one unit of work. Both are reflection. They differ in scope and schedule. Neither is just retrieval.

Why this belongs in the harness

If reflection is optional glue, teams bolt on a memory product and hope. If reflection is a harness stage, the environment class decides the learning target up front.

  • ReAct stays the within-task engine.
  • Reflection becomes the across-task learner.
  • Memory remains the store and retrieve layer.

Without that split, agent-memory work keeps optimizing the wrong surface.

What stays open

  • Schedule. After every task, after a batch, or on a dream schedule?
  • Quality. How do we score reflection without trusting fluency?
  • Specificity. How much should the reflection prompt be domain-specific versus shared?
  • False models. How do we stop reflection from writing confident false environment maps?
  • Search surface. When should harness search edit the reflection stage versus the actor stage?
  • Shared store. How do Dreams-style offline cleanup and per-task reflection share one store without fighting?

Build in this order

Put the stage in the harness first. Name the learning target for the environment class. Keep the prompt as a knob. Write small, inspectable notes for the next actor, and fail closed on noise.

If you only have ReAct plus a memory product, you do not have learning. You have a loop that finishes tasks and an archive you cannot trust.

Appendix