AI agents
The science of AI agents: from ReAct to tool use
Planning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do.
Most systems shipped as agents are a while loop around a chat completion. The loop calls the model and parses anything that looks like a tool call. It runs the call, pastes the result back, and repeats until the token budget runs out. This works in a demo. Then it fails in production, and nobody on the team can say why. No one had a theory of what the loop was supposed to do. So the team changes the prompt, because the prompt is the only part anyone can see.
A theory does exist. Between 2022 and 2023, a short line of papers took the loop apart and named its pieces: planning, tool use, memory, and reflection. Two surveys, published a month apart, arrived at about the same split. One describes profile, memory, planning, and action modules.1 The other describes a brain, perception, and action architecture.2 Each piece traces back to a paper that measured something specific. If you know which paper measured what, you can debug an agent instead of rewriting the prompt and hoping.
ReAct puts reasoning and acting in one trace#
Before ReAct, researchers studied the two halves separately.3 Chain-of-thought prompting has the model reason step by step, but the model cannot check a claim against the outside world. Action-only policies act in the world, but they give no running explanation of why. Yao and colleagues combined the two. The model writes a thought, then takes an action, then reads the result, then thinks again. All of this goes into one trace, a single log of the run.
ReAct: one trace, three kinds of line
- ThoughtCarries the plan to the next step
- ActionA tool call that acts on the environment
- ObservationWhat came back
Repeat think again with the result
The gains were large. On ALFWorld, a text-based household-task environment, ReAct beat imitation and reinforcement learning baselines by 34 absolute percentage points of success rate. On WebShop it gained 10 points. On the question-answering tasks HotpotQA and Fever, the main change was in how the model failed. Each step was checked against a Wikipedia lookup. That cut the model's habit of inventing a fact and then reasoning confidently from it. ReAct did all of this with one or two in-context examples.
Two lessons carry over. First, the thought has a job. It carries the plan from one step to the next. That is why deleting it hurts performance, not just readability. Second, a person can read the trace afterward and find the step where the agent went wrong. That property, more than the benchmark score, is why the format spread.
Toolformer learned when to call a tool#
A prompt can tell a model that a calculator exists. It cannot tell the model when a call is worth the round trip. Toolformer worked on that second problem.4
The method is self-supervised, which means the model builds its own training data. First, sample candidate API calls into a corpus of ordinary text. Then run each call. Keep only the calls whose result makes the next tokens easier to predict, measured as lower perplexity. Then train on the calls you kept. Schick and colleagues connected a calculator, a question-answering system, a search engine, a translation system, and a calendar. The trained model's zero-shot performance was competitive with much larger models. It did not lose its core language modelling ability. Each API needed only a handful of demonstrations.
The main finding is about when to call a tool. The filter checks whether calling the tool at this point makes the next tokens more predictable. No prompt can check that. Most tool-use bugs in production are about timing. The agent has the right tool, but it calls it one step too late, or calls it three times, or describes using it without using it.
Tree of Thoughts treats planning as search#
When a model samples one chain of thought, it follows the first path it picks and never looks at the other options. Tree of Thoughts lays the options out.5 Each partial solution becomes a node. The model rates each node as sure, maybe, or impossible. A search procedure with backtracking then decides which node to expand next.
On the Game of 24, a puzzle that needs arithmetic planning with dead ends, GPT-4 with standard chain-of-thought prompting solved 4 percent of instances. The same model with Tree of Thoughts solved 74 percent.
Game of 24, GPT-4
| Chain-of-thought | Tree of Thoughts | |
|---|---|---|
| Instances solved | 4% | 74% |
That gap is the most quoted number in agent planning, and people usually quote it without its cost. Tree of Thoughts uses many more model calls. Each node the model rates is a call. Each branch it drops is compute already spent. A 70-point gain that takes an order of magnitude more inference is a real result and a real bill. Every planner built since makes that trade. Almost none of the papers report the cost side.
Memory and reflection carry lessons across attempts#
Reflexion asked what an agent can learn between attempts without changing any model weights.6 Its answer is to write a post-mortem in words. The agent fails and reflects in words on why. It stores that reflection in an episodic buffer, a memory of past attempts. It reads the reflection before the next attempt. Shinn and colleagues call this verbal reinforcement learning. On the HumanEval coding benchmark it reached 91 percent pass@1 (solved on the first try), against the 80 percent reported for GPT-4.
Reflexion has a precondition. It works because a unit test tells it clearly that the last attempt was wrong. The reflection is written in language, but the signal behind it comes from outside the model and is a plain pass or fail.
Generative Agents studied memory at a different scale.7 Twenty-five agents lived in a sandbox town. Each kept a memory stream, a running log in plain language of everything it observed. Retrieval scored each memory on recency, importance, and relevance. Reflection ran on top of that and turned raw observations into higher-level conclusions from time to time. Planning then turned those conclusions into a plan for the day. Park and colleagues ran ablations, tests that remove one part at a time. Removing memory, reflection, or planning each lowered how believable the agents' behaviour was. In the best-known result, the agents organised a Valentine's Day party and spread the invitation by word of mouth. That behaviour came from the retrieval function, not from a party subroutine.
Voyager added what the others left out, which is skills that last beyond one episode.8 When the agent works out how to do something in Minecraft, it writes the behaviour as code. It checks the code against the game and stores it in a skill library to use later. Wang and colleagues report 3.3 times more unique items obtained than prior methods and 2.3 times longer distances travelled. They also report key tech-tree milestones reached up to 15.3 times faster. The skills carried over to new worlds, where other methods stalled.
| Piece | Paper | What it measured |
|---|---|---|
| Reasoning plus acting | ReAct | 34 points of absolute success over baselines on ALFWorld, 10 points on WebShop |
| Tool use | Toolformer | Zero-shot performance competitive with much larger models, core language ability retained |
| Planning as search | Tree of Thoughts | Game of 24: 4 percent with chain-of-thought, 74 percent with ToT |
| Reflection | Reflexion | HumanEval pass@1 of 91 percent, against 80 percent reported for GPT-4 |
| Memory | Generative Agents | Ablating memory, reflection, or planning each lowered believability ratings |
| Skill retention | Voyager | 3.3x unique items, 2.3x distance, up to 15.3x faster tech-tree milestones |
The rows work together as one design. Each row names a failure the loop has when that piece is missing. An agent with no memory repeats work. An agent with no reflection repeats mistakes. An agent with no planner commits to its first idea. An agent with no skill library relearns the same procedure every run and pays for it every time.
Four problems are still open#
These four parts split the problem into pieces, but they do not solve it. Four problems in the research are still open, and each one shows up in production.
Reflection needs a grader. Huang and colleagues tested intrinsic self-correction, where the model revises its own answer with no outside feedback. They found that models struggle to correct themselves. Sometimes performance got worse after the revision.9 Both this result and Reflexion's hold. Reflexion's gains depend on a unit test. Without the test, the reflection has nothing to check against. If your agent reflects only against its own opinion of its output, the loop can convince itself of anything.
Errors add up over long tasks. Dziri and colleagues studied transformers on compositional tasks, where each step builds on earlier ones, such as multi-digit multiplication and dynamic programming. They found that models reduce multi-step reasoning to linearised subgraph matching. In plain terms, the models match patterns they have seen instead of learning a general procedure.10 The authors give a theoretical argument and empirical evidence that accuracy falls quickly as the number of dependent steps grows. An agent taking twenty dependent steps has the same problem. A per-step accuracy of 95 percent sounds high. But the twentieth step depends on all nineteen before it, so an error in any of them carries forward.
What to keep in memory is still a rule of thumb. Recency, importance, and relevance are three weights set by hand. They happened to work in one sandbox. No one has a principled account of what an agent should forget, and long-running agents go wrong in what they forget. The survey literature lists memory management as an open problem.1
Most papers do not report cost. Some of the results above gain accuracy by spending more inference. Those results do not report what the extra inference cost. That makes the numbers hard to compare and easy to misread. The evaluation literature argues about this separately, and that argument needs its own post.
Where this fits in Oxagen#
Oxagen is the control plane for the agents you run. It does not run them. The research above explains the shape of the mandate. One object ties together identity, knowledge scope, permitted actions, commercial terms, outcome, and the audit record. That record gives reflection the outside signal it needs. It also shows when the agent called each tool, which debugging needs. The knowledge an agent is given sits in a Neo4j graph with an ontology, a defined set of entity types and relationships. So an answer cites structured data that tracks time, not whatever the retriever happened to return. Planners gain accuracy by spending more inference. So Oxagen prices each governed action and attributes it to the run, the turn, and the step. The cost of the extra search then shows up as a line item.
Footnotes#
-
Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., & Wen, J.-R. (2023). A Survey on Large Language Model based Autonomous Agents. arXiv. https://arxiv.org/abs/2308.11432 ↩ ↩2
-
Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Cheng, W., Zhang, Q., Qin, W., Zheng, Y., Qiu, X., Huang, X., & Gui, T. (2023). The Rise and Potential of Large Language Model Based Agents: A Survey. arXiv. https://arxiv.org/abs/2309.07864 ↩
-
Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2022). ReAct: Synergizing Reasoning and Acting in Language Models. ICLR 2023. https://arxiv.org/abs/2210.03629 ↩
-
Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). Toolformer: Language Models Can Teach Themselves to Use Tools. NeurIPS 2023. https://arxiv.org/abs/2302.04761 ↩
-
Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS 2023. https://arxiv.org/abs/2305.10601 ↩
-
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. https://arxiv.org/abs/2303.11366 ↩
-
Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. https://arxiv.org/abs/2304.03442 ↩
-
Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv. https://arxiv.org/abs/2305.16291 ↩
-
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798 ↩
-
Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Le Bras, R., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., & Choi, Y. (2023). Faith and Fate: Limits of Transformers on Compositionality. NeurIPS 2023. https://arxiv.org/abs/2305.18654 ↩
Cite this
Oxagen Research. (2026, September 9). The science of AI agents: from ReAct to tool use. oxagen.sh. https://oxagen.sh/blog/the-science-of-ai-agents-from-react-to-tool-use
BibTeX
@online{anderson2026thescienceof,
author = {{Oxagen Research}},
title = {The science of AI agents: from ReAct to tool use},
year = {2026},
date = {2026-09-09},
url = {https://oxagen.sh/blog/the-science-of-ai-agents-from-react-to-tool-use}
}Related research
- Steering a run you are not watchingWhen a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.
- The agent time horizon is doublingMETR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.
- The problem is not slop, it is your processThe worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work.
Updates by email
New research reaches subscribers first.