Part 5: Deterministic retrieval
Build semantic memory once
A run should start from what the last one learned instead of repeating it.
On Monday a run spends about 40,000 tokens learning that your payments code lives in billing/, that LedgerEntry is the money-movement primitive, and that notifications go through an outbox pattern. On Tuesday a new run spends about 40,000 tokens learning the same three facts. In the traces this book draws on, that rediscovery is most of the startup cost, for knowledge that changes monthly.
Repositories have natural semantic structure#
Any codebase that has survived contact with production organizes into domains (Payments, Authentication, Billing, Notifications, Infrastructure) whether or not the directory tree admits it. You can recover those domains automatically. Embed every file and module summary, cluster the embeddings, then cross-check the clusters against the dependency graph from part 4. Community detection over the import graph (Louvain, Leiden) finds the same boundaries from pure structure.
◌ Clusters and graph disagree
- A module imports across what the text says is a boundary
- One directory holds two unrelated vocabularies
- Treat it as a debt list, ranked by how many edges cross
● Clusters and graph agree
- Semantics and structure name the same set of files
- That set is a real architectural boundary
- Give it a dossier and an owner, and retrieve it as one unit
This runs Conway's law backwards: the code's semantic clusters recover the team and domain boundaries that produced it. Once you infer a domain, give it a compact, versioned dossier. Entry points, core types, invariants ("money is integer cents everywhere"), owned tables, and the suites that cover it. Generate it once, refresh it incrementally on merge, and retrieve it at the start of a run by lookup rather than by exploration.
Microsoft Research's GraphRAG shows the economics of amortized semantic indexing. An LLM pass extracts an entity graph from the corpus once, the Leiden algorithm detects semantic communities, and community summaries are pre-generated at index time. At query time, answering corpus-level questions from those summaries won 70 to 80% of head-to-head comparisons against naive RAG on comprehensiveness and diversity, while using roughly 2 to 3% of the tokens per query that hierarchical source-text summarization would require, because the expensive understanding was paid for once, up front.10,11 The same trade applies to code. Indexing is a fixed cost. Rediscovery inside every run is a tax.
Edge et al., arXiv:2404.16130 · Microsoft Research, 2024
What a dossier looks like#
A dossier is the briefing a senior engineer gives a new hire, generated from code and stamped with its commit.
domain: billing
sha: 9f31c2
entry_points: [billing/api.py, billing/tasks.py]
core_types: [LedgerEntry, Invoice, RefundRequest]
invariants:
- money columns are integer cents (lint: no-float-money)
- outbound mail goes through the outbox, not the mailer
owned_tables: [ledger_entries, invoices, refunds]
covering_tests: [tests/billing/, tests/contract/test_ledger.py]
may_call: [payments, notifications]
# 280 tokens. Rebuilt when any path under billing/ merges.
What goes in the memory layer#
- Embeddings: file, symbol, and doc-chunk vectors for fuzzy entry ("where do we throttle webhooks?") when the prompt carries no hard anchors.
- Domain dossiers: the 300-token expert briefing per domain, as above.
- Architectural boundary map: which domains may call which, with violations flagged by a check before the model proposes one.
- Convention registry: error-handling idioms, naming schemes, "we use the repository pattern here", extracted once and injected only when the diff is in scope.
This memory is infrastructure with a freshness contract, not a cache. It rebuilds incrementally from the merge queue, it carries the commit SHA it was derived from, and the same CI that validates your code invalidates it. That contract is what separates semantic memory from a stale wiki.
- Build the import graph and cut it into communities. Parse the repository with tree-sitter or your language's own import resolver, load the edges into networkx or python-igraph, and run Leiden. The artifact is
domains.json, one community id per module. - Cross-check the communities against embeddings. Summarize each file, embed the summaries, cluster them, and print the modules where the two methods disagree. That list is your boundary debt, and it is worth reading before you write a single dossier.
- Generate one dossier for your busiest domain. Fill the fields above from the index, cap it at 300 tokens, and stamp it with the commit SHA. Done looks like a file a new engineer could read in a minute and act on.
- Rebuild dossiers in CI on merge. Recompute only the domains whose paths changed. Done looks like the dossier SHA matching
HEADwithin one merge. - Make the start of a run a lookup. Resolve the touched paths to a domain, inject that dossier, and remove the instruction that told the agent to go exploring.
run_startup_tokens: tokens consumed before the first file edit. Sum the input tokens of every request that precedes the first write. Lower is better, and the drop after step 5 is the whole return on this chapter.dossier_staleness_commits: merges between the dossier's stamped SHA andHEADfor that domain. Lower is better, with 0 to 1 as the working target.domain_agreement_rate: share of modules where the embedding cluster and the graph community name the same domain. Higher is better, and the remainder is a ranked debt list.dossier_hit_rate: share of runs that retrieved a dossier for the domain they actually edited. Higher is better, and a low number usually means your path-to-domain resolver is wrong, not that the dossiers are.
Understanding your repository is a build artifact, not a conversation. Compute it once, version it, refresh it incrementally, and let every run start from knowledge instead of archaeology.
Cite this
Anderson, M. (2026). Build semantic memory once. In Engineering Deterministic AI Coding Agents (2nd ed., Part 5). Oxagen Inc. https://macanderson.com/manual/build-semantic-memory-once
BibTeX
@incollection{anderson2026buildsemanticmemory,
author = {Anderson, Mac},
title = {Build semantic memory once},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
chapter = {5},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/build-semantic-memory-once}
}Updates by email
Get the next edition of the field manual and new research when it is published.