Mac Anderson

Part 11: Orchestration and memory

Agents need working memory, not bigger context windows

You do not reread every book you own before fixing a bug, and your agent should not either.

Mac Anderson4 min read1,015 words3 sources cited
View markdown

At turn 60, your agent's transcript still carries the file it read at turn 3, the dead-end approach it abandoned at turn 12, and the stack trace it resolved at turn 20. You are paying for all of it on every subsequent call, and it is competing for attention with the two hundred lines that matter now. The window is being asked to be four things at once: attention, notebook, library, and diary. Human cognition separates those, with a small working set over vast indexed stores and forgetting as a feature. The append-only transcript is the opposite design.

A memory hierarchy for agents#

  • Working memory (the window). Only what the current step needs: the task frame, the active files, the last failure. Small by policy, not by accident.
  • Task memory. Durable structured state for this job: the plan, the decisions made ("chose option B because of migration risk"), the files touched, the budget consumed. It lives in the workflow engine's state (Part 9), survives a crash, and pages in selectively.
  • Episodic memory. Records of past runs. "We fixed a similar DecimalConversionError in March, and the fix was in the serializer config." Indexed by embeddings and entities, retrieved when relevant, absent otherwise.
  • Semantic memory. The durable knowledge layer: the code graph (Part 4), domain dossiers (Part 5), the schema index (Part 8), test contracts (Part 10).
Evidence · The OS analogy is a working architecture

MemGPT (Packer et al., Berkeley) made this concrete. Treat the context window as RAM and give the system OS-style virtual memory, paging information between the window and external storage with explicit memory-management operations. The approach let bounded-context models sustain coherent long-horizon behavior, including document analysis and multi-session conversation beyond the raw window, by managing what is resident rather than growing it.17 The same conclusion arrives from the failure side. Chroma's results imply that even within a large window, keeping stale content resident harms performance,7 and Anthropic's context-engineering guidance formalizes the remedies, naming compaction, structured note-taking outside the window, and just-in-time retrieval as the standard toolkit for long-horizon agents.19

Packer et al., arXiv:2310.08560 · Chroma, 2025 · Anthropic engineering, 2025

Context aging, eviction, retrieval scheduling#

Once memory is tiered, management becomes ordinary systems engineering with ordinary policies. Aging: a file read 30 turns ago decays to its signature, and a resolved sub-task collapses to a one-line decision record. Eviction: superseded attempts and dead-end explorations leave working memory, recoverable from task memory if needed, no longer taxing attention. Retrieval scheduling: each workflow node declares its context contract, so the patch node gets code and contracts, the migration node gets schema and lineage, and the review node gets the diff and the invariants. Nothing rides along just in case, because just in case is the distractor mass the research says hurts most.

A context contract is a few lines of configuration per node. It is worth writing down because it turns an argument about what the model needs into a diff a reviewer can read.

nodes:
  patch:
    include: [task_frame, target_symbols, covering_tests, last_failure]
    max_tokens: 24000
  migrate:
    include: [task_frame, schema_slice, column_lineage, prior_migrations]
    max_tokens: 16000
  review:
    include: [task_frame, diff, invariants, covering_tests]
    max_tokens: 12000

on_overflow: drop lowest rank first, then fail the node

Aging needs a policy table rather than a paragraph, because each rule has to hold for a run you did not watch.

Item in working memoryPolicyWhere it goes
File read, untouched for 20 turnsDecay to signature and docstringFull text stays in semantic memory
Resolved sub-taskCollapse to a one-line decision recordFull transcript stays in task memory
Abandoned approachEvict, keep the reason it failedTask memory, one line
Superseded stack traceEvict on the next green runRun log, outside the window
Active target symbolPin until the node completesStays resident

Every one of those policies is deterministic. What pages in, what decays, and what leaves are decidable by code against declared contracts. The window stops being a place things accumulate and becomes the small, hot set of what this step needs. A structured memory of entities, relations, and decisions is also most of the way to a knowledge graph, which is where this series goes next.

Do this week
  1. Measure what is resident. Instrument one long run to log, per turn, the token count by category: task frame, active files, stale files, resolved failures. Done looks like a chart showing how much of turn 60 is turn 3.
  2. Write a context contract for one node. Pick your patch node, declare its include list and a token ceiling, and reject anything outside the list. Done looks like a node whose input you can predict from its config.
  3. Add an aging rule. Decay any file untouched for N turns to its signature, keeping the full text one lookup away. Done looks like the same run finishing with a smaller peak window and the same outcome.
  4. Move decisions into task memory. When a sub-task resolves, write a one-line record to the workflow engine's state and evict the transcript from the window. Done looks like a crashed run resuming from its decisions rather than replaying the conversation.
  5. Add episodic recall behind a relevance gate. Index past run summaries by entity and error type, and retrieve one only when the current error matches. Done looks like recall firing on a repeat incident and staying quiet otherwise.
Measure it
  • resident_token_p95: the 95th percentile window size across turns in a run. Log the prompt size per call. Down is good while outcomes hold.
  • stale_token_share: tokens in the window belonging to items untouched for 20 or more turns, divided by total window tokens. Tag each item with its last-touch turn. Down is good.
  • context_contract_violations: node inputs containing an item the node's include list does not name. Count them per run. Down is good, and zero means the contracts are enforced rather than advisory.
  • resume_success_rate: crashed or halted runs that resume from task memory and finish, divided by all crashed runs. Up is good, and it is the clearest proof state lives outside the window.
Takeaway

Scale memory, not context. The window is RAM: small, fast, and expensive. Everything else belongs in indexed storage with deterministic paging policies.

Cite this

Anderson, M. (2026). Agents need working memory, not bigger context windows. In Engineering Deterministic AI Coding Agents (2nd ed., Part 11). Oxagen Inc. https://macanderson.com/manual/working-memory-not-bigger-windows

BibTeX
@incollection{anderson2026workingmemorynot,
  author    = {Anderson, Mac},
  title     = {Agents need working memory, not bigger context windows},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  chapter   = {11},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/working-memory-not-bigger-windows}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.