Code graphs, AI agents
What an Ontology Buys an Agent
An agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers.
Ask an agent which pricing tier a customer is on, and it will answer. It will sound certain and quote a paragraph. Some of the time, the tier it gives is wrong. It may be the tier the customer had in March. It may come from an unsigned template. It may belong to a different company with a similar name.
The answer reads well, but it came from the wrong context. Nothing in the pipeline was asked to check the context, so nothing noticed.
Why a faithful answer can still be wrong#
Researchers have a name for this failure. The standard survey of hallucination in text generation splits it into two kinds.1 An intrinsic hallucination contradicts the source the model was given. An extrinsic hallucination cannot be checked against that source at all. Both kinds are judged against the source, and that is the problem. Suppose retrieval picked the wrong source. A model can follow that source perfectly and still give a false answer. Every faithfulness test you run will pass, because those tests only compare the answer to the source.
Retrieval was meant to ground answers in facts, and it does part of that job. It puts a document in the context window. It does not show that the document is about the entity you asked about. It does not show that the fact was true on the date you care about. It does not show who is accountable for stating the fact.
An ontology lets a machine check whether the context is the right one. Without it, the system can only judge how similar the text looks.
What an ontology is#
An ontology has three parts: classes, relations, and constraints.
Classes are the kinds of things in your domain, such as Customer, Contract, Subscription, and LegalEntity. Relations are the allowed links between classes. Each relation says which class it starts from (its domain) and which class it points to (its range). For example, SIGNED_BY goes from a Contract to a LegalEntity and nowhere else. Constraints are rules that must hold. A Contract has exactly one counterparty. A Subscription tier is one of five fixed values. A LegalEntity cannot be its own parent.
Classes, relations, and the constraints on them
- ContractSIGNED_BYexactly one counterpartyLegalEntity
- CustomerHAS_SUBSCRIPTIONtier from a closed set of fiveSubscription
- LegalEntityPARENT_OFnever itselfLegalEntity
Underneath sits a branch of logic called description logics. These are limited forms of first-order logic, designed so that a reasoner always finishes its work. The standard reference is the Description Logic Handbook.2 It covers the main tradeoff in the field. Each feature you add to the language makes it harder to work out what follows from your facts. You cannot have a fully open modelling language and cheap reasoning at once. You must choose.
The web standard for this is OWL 2. Its primer shows how to model with classes, properties, and individuals.3 OWL 2 comes with three profiles, called EL, QL, and RL. Each profile drops some features so that reasoning stays fast. So the description logic tradeoff shows up as a product choice.
Two properties matter more than the syntax.
First, you can check constraints. Say the graph records that a Person signed a Contract, when only a LegalEntity may sign one. A machine can detect that violation. So you can catch one kind of wrong answer before anyone reads it.
Second, OWL uses the open world assumption, and this often surprises teams. Under this assumption, a missing fact is not a false fact. Say your graph has no HAS_SUBSCRIPTION edge for a customer. An OWL reasoner concludes nothing from that. It does not conclude that the customer has no subscription. Most application code assumes the opposite. If you want a missing fact to count as false, say so in the model. Use a cardinality constraint, which is a rule on how many links an entity must have, or a validation shape. Do not assume the reasoner thinks the way you do.
A knowledge graph holds the data an agent queries#
An ontology alone is a schema. What an agent queries is a knowledge graph. It has nodes for entities and edges for relations, filled in with real instances.
The 2021 ACM Computing Surveys paper by Hogan and eighteen co-authors is the nearest thing the field has to a shared definition.4 Its order of topics is useful on its own. It covers graph data models first. Then it covers schema, identity, and context. Then it covers checking and improving quality, and last, publishing. That order shows where the difficulty lies. Most of the hard work in a knowledge graph is not the query language. It is identity, which means deciding that two records are the same entity. It is also context, which means knowing when a fact holds and under what assumptions.
A graph has one practical advantage over a document index. Its edges are records in their own right. An edge does not exist just because two nodes are close together. You can look it up, give it a type, and attach properties to it. The next two sections depend on that.
A graph moves the hard work to a different place. It does not remove it. A document index is cheap to fill and costly to trust. A knowledge graph is the opposite. You pay up front to decide that "Acme Corp", "Acme Corporation", and a row keyed on a tax identifier are one node. You pay again each time a source changes its data. In return, the links between records already exist, so the system does not have to guess them at query time.
Graphs answer questions that take several steps#
Take this question: "Which of our contracts with subsidiaries of Acme expire before the renewal date on the parent agreement?" No single passage answers it. The answer links four types of entity, and any one chunk of text holds only part of it.
Researchers have measured this kind of question since HotpotQA.5 HotpotQA built 113,000 questions that need facts from more than one document. It asked systems to name the supporting facts as well as give the answer. These are called multi-hop questions, because the answer takes several steps from one fact to the next. Once multi-hop questions could be measured on their own, the gap between retrieving text and reasoning over structure could be shown with data.
Survey work describes how the field responded. In their roadmap, Pan and colleagues describe three ways to combine graphs and language models.6 Knowledge graphs can improve language models. Language models can build and fill in graphs. Or both can work together in one system. Their diagnosis is direct, and it still holds. Language models generalise well but are weak at grounding answers in facts. Graphs are strong at grounding but weak at handling anything new.
Two results stand out. Think-on-Graph treats the model as an agent that searches the graph.7 It uses beam search, which keeps the few best paths at each step. It follows relation paths one step at a time, instead of retrieving once. The path it follows serves as the citation, so you can see why the answer came out as it did. The paper reports that this can let smaller models beat GPT-4 on some of these benchmarks. That result suggests the difficulty lies in how knowledge is structured for retrieval, not in model size.
GraphRAG works on the same problem from the side of the document collection.8 It pulls an entity graph out of the source documents. It finds communities of related entities and writes a summary for each one. To answer a question about the whole collection, it combines partial answers from those communities. The test was query-focused summarisation over collections of about a million tokens. There, GraphRAG gave more complete and more varied answers than standard retrieval. It helps most with questions about the whole collection. The usual method splits documents into chunks and searches them by similarity. That method handles these questions worst, because no one chunk covers the whole collection.
Facts need a date range and a source#
A triple stores one fact as a subject, a relation, and an object, such as (Acme, subscribed_to, Enterprise). That triple lacks two things an auditor will ask for first: since when, and who says so.
Time#
Real facts are true only for a period. A job ends. A price changes. A new contract term replaces the one before it. If you store only the current value and overwrite the old one, you lose the data that answers "what did we bill them in March."
Research on completing knowledge graphs has worked on this problem. Lacroix and colleagues extended ComplEx, a tensor factorisation method, to facts with timestamps.9 They decompose an order-4 tensor, a four-dimensional array of numbers. So time is part of the model itself, not extra data added on. For question answering, Saxena and colleagues built CRONQUESTIONS.10 It is a question answering dataset over a temporal knowledge graph, about 340 times larger than earlier work. They reported about a 120 percent gain in accuracy over the baselines of the time. Both results point the same way. A system that treats time as part of each fact does measurably better on questions that depend on time.
In practice, each edge carries validFrom and validTo. "As of" becomes a query parameter, not an assumption.
One customer's tier, with a date range on each fact
| Fact | Holds | Status (As of mid-March) |
|---|---|---|
| Acme subscribed_to Growth | January to March | Holds |
| Acme subscribed_to Enterprise | April onward | Not yet |
Provenance#
The second missing field is provenance, a record of where a fact came from. The W3C standardised terms for this in PROV-O. PROV-O describes provenance in OWL with three core types: Entity, Activity, and Agent.11 Each statement says that something was derived from something else, by some process, and credits someone.
On an edge, provenance becomes a source document, an extraction run, a timestamp, and a confidence score. Then a citation is no longer a paragraph the model chose to quote afterward. The citation is stored with the fact when the fact is written. It stays there whether or not the model mentions it.
So you can check the answer instead of only trusting it. The check does not depend on what the model chooses to say.
Where this meets Oxagen#
Oxagen is the control plane for the agents you run. It does not run them. Two clauses of an agent's mandate apply here. The knowledge an agent receives comes from a Neo4j graph database and an ontology. So each answer carries citations and dates because of how the data is stored, not because a prompt asked for them. The record keeps the lineage of an answer: the nodes, the edges, and the sources behind it. So you can look it up later. Neither clause makes a model smarter. Both make a wrong answer visible. That matters when someone asks why the agent said what it said.
Footnotes#
-
Ji et al. (2023). Survey of Hallucination in Natural Language Generation. ACM Computing Surveys 55(12). https://arxiv.org/abs/2202.03629 ↩
-
Baader, Calvanese, McGuinness, Nardi, & Patel-Schneider, eds. (2003). The Description Logic Handbook: Theory, Implementation, and Applications. Cambridge University Press. https://www.inf.unibz.it/~calvanese/papers-html/DLHB-2003-ed.html ↩
-
W3C (2012). OWL 2 Web Ontology Language Primer (Second Edition). W3C Recommendation, 11 December 2012. https://www.w3.org/TR/owl2-primer/ ↩
-
Hogan et al. (2021). Knowledge Graphs. ACM Computing Surveys 54(4), 71:1-71:37. https://arxiv.org/abs/2003.02320 ↩
-
Yang et al. (2018). HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering. EMNLP 2018. https://arxiv.org/abs/1809.09600 ↩
-
Pan et al. (2024). Unifying Large Language Models and Knowledge Graphs: A Roadmap. IEEE Transactions on Knowledge and Data Engineering. https://arxiv.org/abs/2306.08302 ↩
-
Sun et al. (2024). Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph. ICLR 2024. https://arxiv.org/abs/2307.07697 ↩
-
Edge et al. (2024). From Local to Global: A Graph RAG Approach to Query-Focused Summarization. https://arxiv.org/abs/2404.16130 ↩
-
Lacroix, Obozinski, & Usunier (2020). Tensor Decompositions for Temporal Knowledge Base Completion. ICLR 2020. https://arxiv.org/abs/2004.04926 ↩
-
Saxena, Chakrabarti, & Talukdar (2021). Question Answering Over Temporal Knowledge Graphs. ACL 2021. https://arxiv.org/abs/2106.01515 ↩
-
W3C (2013). PROV-O: The PROV Ontology. W3C Recommendation, 30 April 2013. https://www.w3.org/TR/prov-o/ ↩
Cite this
Oxagen Research. (2026, September 9). What an Ontology Buys an Agent. oxagen.sh. https://oxagen.sh/blog/what-an-ontology-buys-an-agent
BibTeX
@online{anderson2026whatanontology,
author = {{Oxagen Research}},
title = {What an Ontology Buys an Agent},
year = {2026},
date = {2026-09-09},
url = {https://oxagen.sh/blog/what-an-ontology-buys-an-agent}
}Related research
- Steering a run you are not watchingWhen a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.
- The agent time horizon is doublingMETR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.
- The problem is not slop, it is your processThe worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work.
Updates by email
New research reaches subscribers first.