# Knowledge graphs beat prompt engineering

> Knowledge should be connected, not concatenated.

Part 12: Orchestration and memory. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/knowledge-graphs-over-concatenation

Your system prompt has grown to four thousand words because every time the agent missed a relationship, someone added a sentence describing it. "`RefundProcessor` writes `ledger_entries`, which `ReconciliationJob` reads nightly, which the finance dashboard queries." In prose, that chain is three sentences the model has to re-derive by attention on every call. In a graph, it is three edges the system traverses before the model wakes up.

### One ontology, many sources

The previous eleven parts each built a specialized index. The knowledge graph is where they federate into a single ontology over your engineering reality: **code graphs** (symbols, calls, imports, Part 4), **schema graphs** (tables, lineage, Part 8), **documentation graphs** (architecture decision records, or ADRs, and dossiers linked to the entities they govern, Parts 5 to 6), **test and contract graphs** (behaviors linked to covered symbols, Part 10), and **runtime graphs** (services, traces, and incidents from your tracing and incident systems). Entity relationships cross the boundaries: this function ↔ writes this table ↔ governed by this ADR ↔ covered by these tests ↔ implicated in that incident. No single tool holds those edges today, which is why an incident review opens with an hour of archaeology.

### Why graph retrieval reduces reasoning complexity

When an answer requires multi-hop connection, flat retrieval makes the model do the hops. It retrieves chunks, holds them in attention, and infers the links: probabilistic work, paid in tokens, degraded by every distractor (Part 7). Graph retrieval does the hops in the system by deterministic traversal, then hands the model a pre-connected subgraph. You have converted reasoning, which is expensive and variable, into lookup, which is cheap and exact. That is the thesis of this book applied to knowledge.

> **Evidence · Structure wins where connection matters**
>
> Microsoft's **GraphRAG** is the flagship result: an LLM-extracted entity graph, Leiden community detection, and pre-computed community summaries. On global sensemaking questions over corpora of roughly 1M tokens, it beat vector RAG with **win rates of about 70 to 80% on comprehensiveness and diversity**, and it matched full hierarchical source-text summarization while using roughly **2 to 3% of the tokens per query**, because connection-finding moved from query time to index time.[10](https://macanderson.com/manual/sources#r10 "Edge et al. (Microsoft Research). \"From Local to Global: A Graph RAG Approach to Query-Focused Summarization.\" 2024. arXiv:2404.16130"),[11](https://macanderson.com/manual/sources#r11 "Microsoft Research. \"GraphRAG: New tool for complex data discovery.\" July 2024. microsoft.com/research blog") Vector search finds things that sound like the question. Graph traversal finds things related to the answer.
>
> Edge et al., arXiv:2404.16130 · Microsoft Research, 2024

> **Counterweight, graphs are not free**
>
> Honest engineering requires the caveat. Naive graph retrieval can inflate prompts instead of shrinking them. Comparative studies measured some graph-RAG variants stuffing 40k to 100k tokens per query, orders of magnitude above vector baselines, and found that piling on more retrieved subgraph stopped improving answers while adding noise.[21](https://macanderson.com/manual/sources#r21 "Zhu et al.. \"When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation.\" 2025. arXiv:2506.05690") Indexing also costs real money up front and has to be maintained. The lesson is not that graphs win everywhere. It is that a graph is an index for precise traversal, not a license to dump neighborhoods. Budgets and slicing (Part 7) apply to subgraphs the way they apply to logs. For code you are in the luckiest domain: unlike a prose corpus, this graph does not need LLM extraction, because compilers, parsers, and migrations give you most of the edges deterministically and nearly free.

### A small ontology beats a large prompt

The practical migration path is short. Define a small ontology. Populate it from the deterministic sources you already built in Parts 4, 5, 8, and 10. Add embedding-based entry points for fuzzy queries, but make edges rather than similarity the backbone of expansion. Retrieval becomes anchor (parsed, Part 3) → traverse (typed edges, bounded hops) → rank (centrality plus task relevance) → slice to budget. Every step is deterministic except the final generation.

```
entities: Function, Table, Service, ADR, Test, Incident, Domain
edges:    calls, writes, covered_by, governed_by, caused

anchor("process_refund")                  # parsed from the ticket
  .out("writes")          hops=1          → ledger_entries
  .in("reads")            hops=1          → ReconciliationJob
  .out("governed_by")     hops=1          → ADR-031 (money is integer cents)
  .out("covered_by")      hops=1          → 3 tests
  .rank(centrality + task_relevance)
  .slice(budget=8000)                     # subgraphs get budgets too
```

The difference that buys is worth stating as a comparison, because it is the reason the token bill moves.

| Question: what breaks if I change the refund amount type? | Flat retrieval | Graph traversal |
| --- | --- | --- |
| Finding the writer of `ledger_entries` | Hope a chunk mentions both | One `writes` edge |
| Finding the nightly reader | A second similarity query, unlinked | One `reads` edge |
| Finding the rule that governs the column | The ADR ranks low on lexical overlap | One `governed_by` edge |
| Who does the joining | The model, in attention, per call | The index, once, at write time |
| Cost of a wrong answer | A silent miss you find in production | A missing edge you can go add |

> **Do this week**
>
> 1.  **Write the ontology on one page.** Name seven entity types and five edge types, and reject anything you cannot populate deterministically this quarter. Done looks like a schema file, not a diagram.
> 2.  **Load the two cheapest sources.** Emit `calls` and `imports` edges from tree-sitter or your compiler's index, and `writes` edges from your object-relational mapper's (ORM) introspection or from migration history. Done looks like a graph you can query for one real symbol.
> 3.  **Link the governing documents.** Add a front-matter field to each ADR listing the entities it governs, and load `governed_by` edges from it. Done looks like a traversal that returns the integer-cents rule without anyone naming it.
> 4.  **Attach your tracing and incident systems.** Map service names from OpenTelemetry resource attributes to graph nodes, and add a `caused` edge from each incident's postmortem to the symbols it names. Done looks like an incident reachable in one hop from the function that caused it.
> 5.  **Put a budget on the traversal.** Cap hops, cap returned nodes, and slice to a token ceiling before the subgraph reaches the model. Done looks like a subgraph whose size you can state in advance.

> **Measure it**
>
> -   `edge_coverage`: symbols with at least one outbound typed edge, divided by all symbols in the repository. Up is good, and a low number tells you the graph is a demo rather than an index.
> -   `hops_to_answer`: the traversal depth at which the node the patch actually needed appeared. Log the anchor, the path, and the edited symbols. Down is good, and a rising number means an edge type is missing.
> -   `subgraph_tokens_per_query`: tokens of retrieved subgraph handed to the model per call. Down is good, and the counterweight above is the reason to watch it rather than assume it.
> -   `graph_query_share`: graph queries divided by graph queries plus exploratory file reads. Up is good, because it means the agent looks things up instead of foraging.

> **Takeaway**
>
> Concatenation forces the model to rediscover relationships you already know. Store knowledge as a graph and retrieval becomes traversal, which moves connection-finding out of the token bill and into the index.
