Code graphs
Graph-grounded retrieval vs vector search
Vector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference.
A support agent is asked whether a customer's contract includes the uptime credit. The agent searches its documents and gets back a chunk, a short piece of text cut from a larger document. The chunk is about uptime credits. It names the right product, and it reads as if it was written for this question. So the agent answers yes.
The chunk came from the standard contract template. This customer negotiated the credit out of their contract fourteen months ago. Each part of the search pipeline worked as designed, and the answer is still wrong.
The chunk that matched, and the fact that was true
| Fact | Holds | Status (Today) |
|---|---|---|
| Standard template includes the uptime credit | Throughout | Holds |
| This customer's contract includes it | Until 14 months ago | No longer holds |
| Amendment removes it for this customer | Since 14 months ago | Holds |
A chunk can look right and still hold the wrong fact#
Retrieval-augmented generation (RAG) gives a model a searchable index of documents to draw from. Before RAG, a model could rely only on what its weights had stored in training.1 RAG works. But as it is usually set up, it aims at the wrong target.
The standard survey of hallucination grades generated text against the source the model was given. It splits failures into two kinds. In intrinsic failures, the output contradicts the source. In extrinsic failures, the output cannot be checked against the source.2 Neither kind covers a bad source. If the search step supplies the wrong source, the model can repeat it faithfully and still state something false. It will then score as faithful.
Text that is about your question is different from text that answers it. Vector search ranks text by cosine distance, a measure of how close two pieces of text are in meaning. That measure cannot tell the two kinds of text apart.
What dense retrieval does well#
Dense retrieval turns each piece of text into an embedding, a list of numbers that stands for its meaning. It then finds the pieces whose embeddings sit closest to the question's. The numbers show it works.
Dense Passage Retrieval trained two encoders on a modest number of question and passage pairs. It beat BM25, a standard keyword ranking method, by 9 to 19 points absolute. The measure was top-20 accuracy, whether a right passage appears in the top 20 results, across open-domain question answering benchmarks.3 On the task of finding text about a topic, that margin is large and can be reproduced. Dense retrieval is also cheap. You cut the documents into chunks and embed them. In an afternoon, you can search documents that nobody ever organized.
Keep using it. Dense retrieval is the right tool for most of what a company holds: prose, tickets, transcripts, notes, and the rest of the text that would never fit a schema. The question is which questions it should answer.
Where similarity search falls short#
Research documents two problems. Both come from how the method works, so a better embedding model cannot fix them.
The first problem is questions that need several steps, called hops. HotpotQA built 113,000 questions that need facts from several documents. Systems had to name those supporting facts along with the answer.4 When a question needs facts from two places joined together, no single passage holds the answer. Search that returns the top few passages turns each hop into a separate guess. The chance of error grows with each guess.
The second problem shows up when you add more text to the model's input to make up for it. Liu and colleagues measured how models use long inputs. Models answered best when the needed information sat at the start or end of the input. Accuracy dropped a lot when the answer sat in the middle.5 This held even for models built for long inputs. So retrieving fifty passages instead of five does not fix poor precision. It puts the answer where the model reads worst.
Here is how the two methods compare.
| Dense vector retrieval | Graph-grounded retrieval | |
|---|---|---|
| What it returns | Passages ranked by how close their embeddings are | Nodes and edges that match a typed pattern |
| How it matches | Closeness of meaning, as learned by a model | Stated relationships and rules |
| Questions with several hops | Each hop is a new guess, and errors add up | One query follows all the hops |
| Time | Present only if the text mentions it | Stored on each edge, so "as of" a date is a query input |
| Source of the answer | The passage is the evidence | Each edge stores its source, the extraction run, and a confidence score |
| Setup cost | Hours to chunk and embed | Weeks for the schema, extraction, and matching duplicate entities |
| Free-form text | Works on anything | Covers only what was extracted into the schema |
| Typical failure | A chunk that looks right but holds the wrong fact | A missing edge, and no answer at all |
The last row shows the trade-off. When vector search fails, it gives a confident answer that is wrong. When graph-grounded retrieval fails, it returns nothing. You can see an empty result, so you can recover from it.
How graph-grounded retrieval works#
A knowledge graph stores facts as nodes (things, such as a customer) and edges (relationships, such as "has term"). Graph-grounded retrieval first turns the question into the things and relationships it asks about. Next it retrieves the facts. Then the model writes the answer.
Find the facts first, then write the answer
- ResolveThe question becomes things and relationships
- TraverseA typed pattern, limited to one customer organization and one date
- Return factsEach row carries its source document
- WriteThe model answers from those rows
The simplest version already improves results. KAPING finds the entities named in the question. It writes their stored facts out as sentences and puts them before the prompt. It needs no fine-tuning. It reports gains of up to 48 percent on average over similar zero-shot baselines, across models of several sizes.6 The method is plain. The model gets the specific facts that bear on the question, stored as triples (subject, relationship, object), instead of paragraphs that look similar.
Think-on-Graph goes further. The model acts as an agent that explores the graph one relationship at a time, keeping several of the best paths as it goes (a method called beam search).7 The path it follows serves as the citation. You can read why the system reached its answer. A similarity score does not give you that. The paper reports that this lets smaller models beat GPT-4 on some of these benchmarks.
If cost is your limit, look at HippoRAG. It builds a graph index over the documents and retrieves in one step. On multi-hop question answering, it matches methods that retrieve in many steps, with gains of up to 20 percent. It costs 10 to 30 times less and runs 6 to 13 times faster than those methods.8 Here, structure is both more accurate and cheaper. The graph does the join, so you stop paying for repeated searches.
GraphRAG handles the questions that chunking handles worst: questions about a whole set of documents. It pulls out a graph of entities and finds groups of related entities. It summarizes each group, then builds an answer from the partial answers across groups. On document sets of around a million tokens, this made answers more complete and more varied.9
In all four systems, the unit of retrieval is a fact with a type, not a window of text. So you can put limits on it. The query below is written in Cypher, the query language for the Neo4j graph database. It finds the contract terms in force for one customer on one date, within one workspace:
// Contract terms in force for one customer on a given date,
// scoped to the caller's workspace.
MATCH (c:Customer { publicId: $customerId, workspaceId: $workspaceId })
-[r:HAS_TERM]->(t:ContractTerm)
WHERE r.validFrom <= $asOf
AND (r.validTo IS NULL OR r.validTo > $asOf)
RETURN t.displayName AS term,
t.value AS value,
r.sourceDocId AS source,
r.validFrom AS since
ORDER BY r.validFrom DESC
That query does three things a vector search does not. The workspace is part of the pattern, so it cannot return another customer's data. The dates are part of the filter, so it cannot return a term that was replaced. And each row it returns carries the document it came from. This query would not have returned the wrong uptime credit from the start of this post.
Where graphs do worse#
Building the schema is real work. Hogan and eighteen co-authors wrote a survey of knowledge graphs. Most of it covers schema, identity, context, quality, and refinement. Much less covers querying.10 That balance matches practice. Deciding what your types of things are takes people who disagree, and the work does not end.
Matching duplicate entities is the hardest part, and it is never fully solved. This task is called entity resolution. It decides whether two records describe the same real thing. It is a research field of its own. One survey in ACM Computing Surveys covers each step, from grouping likely candidates through matching and clustering.11 An error in one direction merges two customers into one node. An error in the other direction points half your edges at a duplicate that no query reaches. Neither error shows up on its own.
The graph holds only what extraction put in it. A vector index covers whatever you embedded. A graph covers only what you pulled out into the schema, which is always less. Anything subtle, hedged, or new in a document is lost when it becomes a triple.
Turning a question into a graph query can also go wrong. The Cypher above is correct because a person wrote it. If a model writes that query from a plain-language question, the hard part moves to the model. A query that looks right but follows the wrong relationship returns rows that look certain. The fix is to offer a fixed set of queries with fill-in inputs. The model picks one and fills it in, and it does not write graph queries freely. This limits what the model can ask on purpose. That is a real cost. It is the right trade for anything that touches contracts or money.
How Oxagen uses graph-grounded retrieval#
Oxagen is the agent control plane for the agents you run. It does not run them. The knowledge an agent is given comes from a Neo4j graph and an ontology. An ontology is a model of the business's types of things and how they relate. So the unit an agent retrieves is a typed fact with its source and its dates, not a chunk that looked right. The nodes and edges behind an answer stay in the run's record. So you can trace a wrong answer back to the fact that produced it. This does not make the model more accurate. It lets a person check whether an answer that looks right is correct.
Footnotes#
-
Lewis et al. (2020). Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks. NeurIPS 2020. https://arxiv.org/abs/2005.11401 ↩
-
Ji et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). https://arxiv.org/abs/2202.03629 ↩
-
Karpukhin et al. (2020). Dense Passage Retrieval for Open-Domain Question Answering. EMNLP 2020. https://arxiv.org/abs/2004.04906 ↩
-
Yang et al. (2018). HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. EMNLP 2018. https://arxiv.org/abs/1809.09600 ↩
-
Liu et al. (2024). Lost in the Middle: How Language Models Use Long Contexts. Transactions of the Association for Computational Linguistics. https://arxiv.org/abs/2307.03172 ↩
-
Baek, Aji, & Saffari (2023). Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answering. https://arxiv.org/abs/2306.04136 ↩
-
Sun et al. (2024). Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph. ICLR 2024. https://arxiv.org/abs/2307.07697 ↩
-
Gutiérrez et al. (2024). HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models. NeurIPS 2024. https://arxiv.org/abs/2405.14831 ↩
-
Edge et al. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. https://arxiv.org/abs/2404.16130 ↩
-
Hogan et al. (2021). Knowledge Graphs. ACM Computing Surveys 54(4), 71:1-71:37. https://arxiv.org/abs/2003.02320 ↩
-
Christophides, Efthymiou, Palpanas, Papadakis, & Stefanidis (2020). An Overview of End-to-End Entity Resolution for Big Data. ACM Computing Surveys 53(6), Article 127. https://dl.acm.org/doi/10.1145/3418896 ↩
Cite this
Oxagen Research. (2026, September 9). Graph-grounded retrieval vs vector search. oxagen.sh. https://oxagen.sh/blog/graph-grounded-retrieval-vs-vector-search
BibTeX
@online{anderson2026graphgroundedretrieval,
author = {{Oxagen Research}},
title = {Graph-grounded retrieval vs vector search},
year = {2026},
date = {2026-09-09},
url = {https://oxagen.sh/blog/graph-grounded-retrieval-vs-vector-search}
}Related research
- Governing an agent that rewrites itselfAn agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts.
- What an Ontology Buys an AgentAn agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers.
Updates by email
New research reaches subscribers first.