Mac Anderson

AI agents, Coding agents

Why agents fail: measuring reliability and cost

AgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides.

Oxagen Research8 min readFirst published on oxagen.sh
View markdown

The agent works in the demo, so the team ships it. Two weeks later, support sees a pattern. About one run in four does something wrong, and the mistake is different each time. Finance asks what a successful run costs, and nobody knows. There is a token bill, but it is not broken down by outcome. In the ledger, a run that worked looks the same as a run that failed three times and was retried.

Both problems come from how agents are measured, and research has studied them since 2023. The benchmarks that made agents look good measure one try at a task in an environment that never changes. Almost none of them report cost. If you read them closely, the failure rate in production is no surprise.

What the benchmarks measure#

People quote five benchmarks as if they measure the same thing. Each one measures something different.

AgentBench placed language models in 8 different interactive environments, from operating-system tasks to databases to card games. It scored task success in each one.1 The main result was a wide gap between top commercial models and open models under 70B parameters. The diagnosis matters more. Liu and colleagues traced agent failure mainly to weak long-term reasoning, weak decisions, and weak instruction following. They did not trace it to a missing skill.

WebArena took the opposite approach and built one environment in full.2 It runs real, self-hosted, working sites: an online store, a social forum, a software development platform, and a content management system. It also includes supporting tools and documentation. WebArena checks success by looking at the resulting state of the site, not by matching the text of a summary. The best agent Zhou and colleagues tested finished 14.41 percent of tasks from start to finish. People on the same tasks finished 78.24 percent.

GAIA has 466 questions that a person finds simple in concept.3 Answering them takes reasoning, handling more than one kind of input, web browsing, and tool use. The answers are short and can be checked exactly. Human respondents scored 92 percent. GPT-4 with plugins scored 15 percent. Mialon and colleagues state their design goal directly. They argue that research should focus on how reliably agents handle tasks people find easy, not on scores for tasks people find hard.

The best agent tested against people, on the same tasks

Best agent testedHumans
WebArena14.41%78.24%
GAIA (GPT-4 with plugins)15%92%
Sources: Zhou et al. (2023) for WebArena and Mialon et al. (2023) for GAIA, each at publication. A person finds the tasks in both routine.

SWE-bench covers coding, and its grader is the strictest of the five.4 It has 2,294 real GitHub issues and the pull requests that fixed them, drawn from 12 popular Python repositories. A submission succeeds when the repository's own test suite passes after the model's patch is applied. When the paper came out, the best system Jimenez and colleagues measured was Claude 2. It resolved 1.96 percent of tasks.

The fifth benchmark, tau-bench, measures what the other four skip.5

BenchmarkTask typeWhat success meansCost reported
AgentBench8 interactive environments (OS, DB, web, games)Task success in each environment, then combinedNo
WebArenaMulti-step tasks on real self-hosted websitesA check of the site's resulting stateNo
GAIA466 real-world assistant questionsExact match against a short correct answerNo
SWE-bench2,294 real GitHub issues in 12 Python reposThe repository's own test suite passes after the patchNo
tau-benchConversations between a user and a tool-using agent under a domain policyFinal database state matches the goal state, over repeated trialsNo

Look at the right-hand column. The five benchmarks define success in five ways. None of them counts the cost of a try as part of the result.

Two more points go with that table. First, a real grader stays useful after the scores move on. Published SWE-bench resolve rates have climbed far above 1.96 percent. Those rates can still be compared across years, because the test suites that grade them did not change. So a benchmark that checks success for real stays useful after its leaderboard is out of date. Second, each of these environments is frozen on purpose, which means it does not change during testing. That makes the comparison fair, but it also makes the results optimistic. In production, sites change their markup, APIs drop fields, and someone edits a policy document on a Tuesday. An agent that scored 14 percent against a fixed copy of a web stack was not tested on what breaks it most often in real use.

One try does not measure reliability#

Yao and colleagues built tau-bench around a case the others do not simulate.5 An agent talks to a user, uses domain tools, and follows a written domain policy, all at the same time. tau-bench compares the final database state with the goal state. So the agent must change the data correctly. Describing the change is not enough.

The metric is their main contribution. The usual practice reports pass@k, the chance that at least one of k tries succeeds. That suits a research leaderboard. It does not suit a production system, where you get one try and a customer is waiting. So the authors defined pass^k, the chance that all k separate tries at the same task succeed.

The results were clear. State-of-the-art function-calling agents, including GPT-4o, succeeded on under 50 percent of tasks. In the retail domain, pass^8 fell below 25 percent. An agent that is right half the time on one try is right on all eight tries only about a quarter of the time. That number matches what support sees, and nobody publishes it.

Errors add up over many steps, and self-checks do not undo them#

Inconsistency is not bad luck. It follows from simple arithmetic when a task has many steps.

An agent doing a real task takes a chain of steps, and each step depends on the ones before it. It reads the ticket, queries the database, decides, calls the API, and checks the result. Suppose each step is right 95 percent of the time, independently. Then twenty such steps all come out right only 36 percent of the time. At 99 percent per step, twenty steps still come out right only 82 percent of the time. This is multiplication, not a flaw in the model. So per-step accuracy is a misleading target on its own.

Chance a chain of dependent steps finishes correctly

Dependent steps95% right per step99% right per step
0100%100%
195%99%
290.3%98%
385.7%97%
481.5%96.1%
577.4%95.1%
673.5%94.1%
769.8%93.2%
866.3%92.3%
963%91.4%
1059.9%90.4%
1156.9%89.5%
1254%88.6%
1351.3%87.8%
1448.8%86.9%
1546.3%86%
1644%85.1%
1741.8%84.3%
1839.7%83.5%
1937.7%82.6%
2035.8%81.8%
Per-step accuracy raised to the power of the number of steps. This is not model data. It is the multiplication the section describes, and it assumes accuracy stays the same at every step. Dziri et al. find that it does not.

The way models behave makes this worse. Dziri and colleagues studied transformers on compositional tasks, which are tasks built from smaller steps. Their examples were multi-digit multiplication, logic puzzles, and dynamic programming.6 They concluded that the models do not learn a general procedure for these tasks. Instead, the models reduce many-step reasoning to matching pieces of patterns they have seen, which the authors call linearised subgraph matching. The authors argue, in theory and in experiments, that models which generate one token at a time lose accuracy quickly as tasks grow more complex. So per-step accuracy does not even stay the same across a long task. It gets worse as the task gets longer.

The obvious fix is to let the agent check its own work, but that fix is weaker than it looks. Huang and colleagues found that models struggle to correct themselves without outside feedback.7 Performance sometimes got worse after a round of self-correction with no outside help. Self-correction works when something outside the model can judge the last try. A test suite can do that, and so can a check of the database state. The model's own confidence cannot.

Cost is half of the result#

Kapoor and colleagues made an argument the field had avoided.8 They argued that an accuracy number should not be reported without its cost. Their analysis of agent benchmarks makes four claims that belong together.

First, when a benchmark scores only accuracy, spending more carries no penalty. So researchers are pushed toward agents that are more complex and costly than they need to be. Second, accuracy and cost should be optimised together. When the authors did that, they cut cost by a large amount and kept accuracy the same. Third, many agent benchmarks lack proper holdout sets, which are tasks kept back for final testing. Without them, agents can overfit and use shortcuts that will not exist in production. Fourth, the field mixes up two audiences. A model developer compares systems, and a downstream developer chooses one to deploy. The two need different evaluations, and a single leaderboard column serves neither well.

The previous post's numbers show the same problem. Tree of Thoughts raised success on Game of 24 from 4 percent to 74 percent.9 It searched over many candidate thoughts, which means many model calls per task. Reflexion reached 91 percent pass@1 on HumanEval.10 After each failure it wrote a short review of what went wrong and tried again, which means several tries per task. Both results are real. Both pay for accuracy with extra model calls. Neither headline says how much accuracy the extra calls bought. A team deciding whether to ship needs exactly that figure.

Overfitting is the less visible part of their argument, and it explains the gap between the demo and the deployment. An agent tuned on a benchmark with no holdout set can learn the benchmark itself. It learns the shape of the tasks, the quirks of the environment, and the shortcuts the grader accepts. None of that carries over. A team may pick an agent design from a leaderboard and then be surprised in production. Often that is not a regression. It is the first accurate measurement of the agent.

This has a practical form for anyone running an agent today. Accuracy per try is the wrong basis. The number that matters is cost per verified outcome. Add up the spend across every try, retry, and abandoned branch. Then divide it by the number of outcomes that something outside the agent confirmed were correct. That number is usually several times the simple one, and it moves when reliability moves. It is the only figure that makes the cost of a retry loop clear on a finance dashboard.

Where this meets Oxagen#

Oxagen is the control plane for the agents you run. It does not run them. Two of its parts follow from the research above. The first is the meter. Oxagen prices every governed action and assigns it to the person, the agent, the run, the turn, and the step. So you can divide total spend by checked outcomes instead of by tries. That is the pass^k problem written as a bill. The second is the record. Oxagen keeps every run against the agent's mandate, one frame at a time, where a frame is one recorded event. Each run shows what the agent asked for, which rule answered, and what checked the outcome. So the outside check that self-correction needs is a stored row, not the model's own opinion. None of this makes an agent reliable. It lets you put a number on the agent's reliability and its cost.

Footnotes#

  1. Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., & Tang, J. (2023). AgentBench: Evaluating LLMs as Agents. ICLR 2024. https://arxiv.org/abs/2308.03688 ↩

  2. Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2023). WebArena: A Realistic Web Environment for Building Autonomous Agents. ICLR 2024. https://arxiv.org/abs/2307.13854 ↩

  3. Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., & Scialom, T. (2023). GAIA: a benchmark for General AI Assistants. ICLR 2024. https://arxiv.org/abs/2311.12983 ↩

  4. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues?. ICLR 2024. https://arxiv.org/abs/2310.06770 ↩

  5. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv. https://arxiv.org/abs/2406.12045 ↩ ↩2

  6. Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Le Bras, R., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., & Choi, Y. (2023). Faith and Fate: Limits of Transformers on Compositionality. NeurIPS 2023. https://arxiv.org/abs/2305.18654 ↩

  7. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798 ↩

  8. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). AI Agents That Matter. arXiv. https://arxiv.org/abs/2407.01502 ↩

  9. Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). Tree of Thoughts: Deliberate Problem Solving with Large Language Models. NeurIPS 2023. https://arxiv.org/abs/2305.10601 ↩

  10. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. https://arxiv.org/abs/2303.11366 ↩

Cite this

Oxagen Research. (2026, September 9). Why agents fail: measuring reliability and cost. oxagen.sh. https://oxagen.sh/blog/why-agents-fail-measuring-reliability-and-cost

BibTeX
@online{anderson2026whyagentsfail,
  author  = {{Oxagen Research}},
  title   = {Why agents fail: measuring reliability and cost},
  year    = {2026},
  date    = {2026-09-09},
  url     = {https://oxagen.sh/blog/why-agents-fail-measuring-reliability-and-cost}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.