Part 12: Orchestration and memory
Knowledge graphs beat prompt engineering
Knowledge should be connected, not concatenated.
Your system prompt has grown to four thousand words because every time the agent missed a relationship, someone added a sentence describing it. "RefundProcessor writes ledger_entries, which ReconciliationJob reads nightly, which the finance dashboard queries." In prose, that chain is three sentences the model has to re-derive by attention on every call. In a graph, it is three edges the system traverses before the model wakes up.
One ontology, many sources#
The previous eleven parts each built a specialized index. The knowledge graph is where they federate into a single ontology over your engineering reality: code graphs (symbols, calls, imports, Part 4), schema graphs (tables, lineage, Part 8), documentation graphs (architecture decision records, or ADRs, and dossiers linked to the entities they govern, Parts 5 to 6), test and contract graphs (behaviors linked to covered symbols, Part 10), and runtime graphs (services, traces, and incidents from your tracing and incident systems). Entity relationships cross the boundaries: this function ↔ writes this table ↔ governed by this ADR ↔ covered by these tests ↔ implicated in that incident. No single tool holds those edges today, which is why an incident review opens with an hour of archaeology.
Why graph retrieval reduces reasoning complexity#
When an answer requires multi-hop connection, flat retrieval makes the model do the hops. It retrieves chunks, holds them in attention, and infers the links: probabilistic work, paid in tokens, degraded by every distractor (Part 7). Graph retrieval does the hops in the system by deterministic traversal, then hands the model a pre-connected subgraph. You have converted reasoning, which is expensive and variable, into lookup, which is cheap and exact. That is the thesis of this book applied to knowledge.
Microsoft's GraphRAG is the flagship result: an LLM-extracted entity graph, Leiden community detection, and pre-computed community summaries. On global sensemaking questions over corpora of roughly 1M tokens, it beat vector RAG with win rates of about 70 to 80% on comprehensiveness and diversity, and it matched full hierarchical source-text summarization while using roughly 2 to 3% of the tokens per query, because connection-finding moved from query time to index time.10,11 Vector search finds things that sound like the question. Graph traversal finds things related to the answer.
Edge et al., arXiv:2404.16130 · Microsoft Research, 2024
Honest engineering requires the caveat. Naive graph retrieval can inflate prompts instead of shrinking them. Comparative studies measured some graph-RAG variants stuffing 40k to 100k tokens per query, orders of magnitude above vector baselines, and found that piling on more retrieved subgraph stopped improving answers while adding noise.21 Indexing also costs real money up front and has to be maintained. The lesson is not that graphs win everywhere. It is that a graph is an index for precise traversal, not a license to dump neighborhoods. Budgets and slicing (Part 7) apply to subgraphs the way they apply to logs. For code you are in the luckiest domain: unlike a prose corpus, this graph does not need LLM extraction, because compilers, parsers, and migrations give you most of the edges deterministically and nearly free.
A small ontology beats a large prompt#
The practical migration path is short. Define a small ontology. Populate it from the deterministic sources you already built in Parts 4, 5, 8, and 10. Add embedding-based entry points for fuzzy queries, but make edges rather than similarity the backbone of expansion. Retrieval becomes anchor (parsed, Part 3) → traverse (typed edges, bounded hops) → rank (centrality plus task relevance) → slice to budget. Every step is deterministic except the final generation.
entities: Function, Table, Service, ADR, Test, Incident, Domain
edges: calls, writes, covered_by, governed_by, caused
anchor("process_refund") # parsed from the ticket
.out("writes") hops=1 → ledger_entries
.in("reads") hops=1 → ReconciliationJob
.out("governed_by") hops=1 → ADR-031 (money is integer cents)
.out("covered_by") hops=1 → 3 tests
.rank(centrality + task_relevance)
.slice(budget=8000) # subgraphs get budgets too
The difference that buys is worth stating as a comparison, because it is the reason the token bill moves.
| Question: what breaks if I change the refund amount type? | Flat retrieval | Graph traversal |
|---|---|---|
Finding the writer of ledger_entries | Hope a chunk mentions both | One writes edge |
| Finding the nightly reader | A second similarity query, unlinked | One reads edge |
| Finding the rule that governs the column | The ADR ranks low on lexical overlap | One governed_by edge |
| Who does the joining | The model, in attention, per call | The index, once, at write time |
| Cost of a wrong answer | A silent miss you find in production | A missing edge you can go add |
- Write the ontology on one page. Name seven entity types and five edge types, and reject anything you cannot populate deterministically this quarter. Done looks like a schema file, not a diagram.
- Load the two cheapest sources. Emit
callsandimportsedges from tree-sitter or your compiler's index, andwritesedges from your object-relational mapper's (ORM) introspection or from migration history. Done looks like a graph you can query for one real symbol. - Link the governing documents. Add a front-matter field to each ADR listing the entities it governs, and load
governed_byedges from it. Done looks like a traversal that returns the integer-cents rule without anyone naming it. - Attach your tracing and incident systems. Map service names from OpenTelemetry resource attributes to graph nodes, and add a
causededge from each incident's postmortem to the symbols it names. Done looks like an incident reachable in one hop from the function that caused it. - Put a budget on the traversal. Cap hops, cap returned nodes, and slice to a token ceiling before the subgraph reaches the model. Done looks like a subgraph whose size you can state in advance.
edge_coverage: symbols with at least one outbound typed edge, divided by all symbols in the repository. Up is good, and a low number tells you the graph is a demo rather than an index.hops_to_answer: the traversal depth at which the node the patch actually needed appeared. Log the anchor, the path, and the edited symbols. Down is good, and a rising number means an edge type is missing.subgraph_tokens_per_query: tokens of retrieved subgraph handed to the model per call. Down is good, and the counterweight above is the reason to watch it rather than assume it.graph_query_share: graph queries divided by graph queries plus exploratory file reads. Up is good, because it means the agent looks things up instead of foraging.
Concatenation forces the model to rediscover relationships you already know. Store knowledge as a graph and retrieval becomes traversal, which moves connection-finding out of the token bill and into the index.
Cite this
Anderson, M. (2026). Knowledge graphs beat prompt engineering. In Engineering Deterministic AI Coding Agents (2nd ed., Part 12). Oxagen Inc. https://macanderson.com/manual/knowledge-graphs-over-concatenation
BibTeX
@incollection{anderson2026knowledgegraphsover,
author = {Anderson, Mac},
title = {Knowledge graphs beat prompt engineering},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
chapter = {12},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/knowledge-graphs-over-concatenation}
}Updates by email
Get the next edition of the field manual and new research when it is published.