Coding agents
What SWE-bench Measures, and What It Misses
Coding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you.
You are choosing a coding agent. The only number you can compare across the options is a SWE-bench resolve rate, so you choose by it. Then the agent starts work in your repository, on your backlog, under your review process. It behaves nothing like the number suggested.
This gap is real, and it can be measured. You do not need to throw out benchmarks. You need to know what the score is made of. A resolve rate measures four things at once: a model, a scaffold, a budget, and a grading harness. The scaffold is the code around the model that decides what the model can see and do. The grading harness is the code that runs the tests and scores the patch. Yet the rate is reported as if it measured the model alone.
Jimenez and colleagues introduced SWE-bench in 2023.1 It has 2,294 tasks built from real GitHub issues and the pull requests that fixed them, across 12 popular Python repositories. Each task gives a system the issue text and the repository at the commit before the fix. The system must write a patch, and tests then grade it. When the paper came out, the best system tested resolved 1.96% of the issues.1 Frontier systems now solve most of the curated subset. That is a fast climb, and it is why every model launch quotes the benchmark.
What "resolved" means#
A task counts as resolved when two things happen. The patch makes a specific set of failing tests pass. It also leaves the tests that passed before still passing. That is the full definition.
It does not say whether a maintainer would accept the patch. It does not cover slowdowns the tests miss, conventions the patch ignores, or whether a reviewer could follow the change. Passing tests stand in for correctness, and the two can drift apart.
The original benchmark had a worse problem. Some tasks could not be graded fairly at all. OpenAI worked with the SWE-bench authors and 93 professional Python developers to review the test split.2 In August 2024 they released SWE-bench Verified. It has 500 tasks that reviewers confirmed have clear problem statements, correct test patches, and enough information to solve. The removed tasks show what went wrong. Some issue descriptions left out too much, so no reader could tell what behaviour was wanted. Some tests were written so closely around the original fix that a different correct fix would fail them. Verified is now the standard reported figure. It exists because a large share of the original test split was not fit for the ranking the leaderboard built on it.
Where the grading goes wrong#
Two problems matter, and each works in a different way.
The first problem is tests that are too weak. Yu and colleagues built UTBoost, a tool that adds tests, and ran it on SWE-bench submissions. They found 345 patches marked as passing that were in fact wrong. The pull request's own tests never checked the failing case.3 Once the missing tests were added, the ranking changed for 40.9% of SWE-bench Lite leaderboard entries and 24.4% of SWE-bench Verified entries.3 So the true share of correct patches can be lower than the resolve rate, but not higher.
The second problem is contamination, which means the test tasks were in the model's training data. These are public repositories with public issues and public fixes. All of them are older than the training cutoff of every system on the board. Liang and colleagues tested this directly. They gave models only the issue text, with no access to the repository, and asked which file held the bug. Models named the correct path up to 76% of the time on SWE-bench. On tasks from repositories outside SWE-bench, they named it up to 53% of the time.4 The authors also asked models to reproduce the fixed function exactly. They scored this with consecutive 5-gram accuracy, the share of five-token runs that match the real fix. Models reached up to 35% on SWE-bench Verified and up to 18% elsewhere.4 So part of the score comes from memory, not reasoning. Nobody can say exactly how large that part is.
What models know from the issue text alone
| Repositories outside SWE-bench | SWE-bench | |
|---|---|---|
| Names the file that holds the bug | 53% | 76% |
| Reproduces the fixed function, 5-gram accuracy | 18% | 35% |
Most of the score comes from the scaffold#
A model alone does not produce the number you are comparing. A model inside a scaffold produces it, and the scaffold decides what the model can see and do.
SWE-agent made this case directly. Yang and colleagues built a custom agent-computer interface for moving through a repository, editing files, and running tests.5 They reported a 12.5% pass@1 rate on SWE-bench, which means 12.5% of tasks were solved on the first try. Approaches that did not interact with the repository had done far worse. The authors argue that language model agents are a new kind of end user. People need an IDE to work well, and agents likewise need interfaces built for them. The paper's contribution was the interface, not the model weights.
OpenHands turned the same idea into an open platform.6 On it, agents write code, use a command line, and browse the web inside a sandbox. It comes with code to run the standard benchmarks. It matters here because it changed how results are produced. Many reported numbers now depend on this one platform.
Agentless took the opposite approach. Xia and colleagues dropped tool use and the agent's own decisions entirely. They used three fixed phases instead: localise, repair, and validate. Agentless resolved 32.00% of SWE-bench Lite at about $0.70 per issue. That beat every open-source agent of the time on both score and cost.7 This result should change how you read a leaderboard. Much of what looked like agent skill was finding the right code. A fixed pipeline could do that for less.
Agentless: three fixed phases, no autonomous tool use
- LocaliseFind the files and functions to change
- RepairSample candidate patches
- ValidateRun tests and keep a patch that passes
Retrieval quality matters on its own too. RepoCoder showed that context from the whole repository beats context from the current file by more than 10%.8 This held for line, API, and function-body completion. RepoCoder uses a loop in which each generated draft guides the next retrieval. So where the relevant code lives, and whether the scaffold can find it, explains a large share of any repository-scale score.
In practice, two numbers printed side by side may not be comparable at all. One may come from a fixed pipeline with a strict budget. The other may come from an agent allowed to run for an hour with unlimited retries. A report needs to give the scaffold, the model, the retry policy, and the spend together. Without them, a resolve rate is closer to a headline than a measurement.
The variants, and what each one still misses#
The benchmark family has grown. Most new versions fix one of the weaknesses above.
| Benchmark | Instances | Domain | What it fixed | What it still misses |
|---|---|---|---|---|
| SWE-bench (2023) | 2,294 | 12 Python repositories | Set up the task: a real issue in a real repository, graded by tests1 | Vague issues, unfair tests, contamination |
| SWE-bench Verified (2024) | 500 | Same Python repositories | Human review of statements, tests, and whether each task can be solved2 | Contamination, thin test coverage |
| SWE-bench Multimodal (2024) | 617 | 17 JavaScript libraries | Visual, user-facing bugs outside Python9 | Whether the fix looks right to a user |
| SWE-Bench Pro (2025) | 1,865 | 41 repositories, some private | Long multi-file tasks, and licences chosen to resist contamination10 | Your codebase, your conventions, your reviewers |
The newer variants show that scores do not carry over well. SWE-bench Multimodal has 617 tasks across 17 visual JavaScript libraries. On it, the best system resolved 12% of tasks and the next best resolved 6%, though the same systems already scored far higher on Python.9 SWE-Bench Pro has 1,865 tasks across 41 repositories. It uses copyleft and proprietary code on purpose, so that models are less likely to have trained on the tasks. Its reference patches change 107.4 lines across 4.1 files on average. Widely used models solve under 25% of its tasks on the first try (pass@1).10
Taken together, the two results show one thing. A high Verified score carries over poorly to a different language, a different way of working, and a longer task. Your repository differs in at least one of those three ways.
One more mismatch remains, and no variant has fixed it. It comes from the shape of the task itself, not from a flaw in any dataset. A SWE-bench task is a closed question. The issue is already triaged, already reproducible, and already scoped to one repository. A correct answer already exists in a merged pull request. A real backlog looks different. Its tickets can be vague, duplicated, or simply wrong. Some span two services. Some describe a symptom whose cause sits in a dependency. Many are not bugs at all. They are migrations, deprecations, and cleanups, where no failing test exists to make pass. The benchmark chose tasks it could grade. As a result, it left out most of the work your team has.
What to measure instead#
The benchmark is still useful, as one input among several.
A leaderboard number cannot tell you four things, because of how it is built. Each is cheap to measure on your own code:
- Grading fidelity. SWE-bench grades a patch with tests written by the person who fixed the bug. Your agent will be graded by tests written before the bug existed. If your test suite would not catch the regression, a passing result is not evidence.
- Cost and variance per outcome. Agentless reported $0.70 per issue for a reason.7 Cost per resolved task can be compared across systems. A resolve rate alone hides a scaffold that made fifty tool calls to get there. Run the same task several times and record the spread, not just the best try.
- Reviewability. If a reviewer cannot follow a patch, the patch costs more than it saves. A resolve rate says nothing about diff size, how much of the system the change touches, or whether the change came with a reason.
- Task admissibility. First measure what share of your backlog is shaped like a task the agent can attempt. Then ask how often it succeeds. If a quarter of your tickets have no reproducible failure, the resolve rate applies to three quarters of the work at most.
In short, SWE-bench checks whether a system can turn an issue description into a patch that passes a test written in advance. The code is Python, in a repository whose history the model has probably seen. That is useful to know. It does not tell you whether the agent should be allowed near your main branch.
Where this meets Oxagen#
Oxagen does not run coding agents, and it does not publish a benchmark score. It is the control plane for the agents you run, and it records what they did. Two parts of it matter here. First, the record keeps outcomes that something outside the agent checked, not the agent's own report of success. That follows the same logic as refusing to take a resolve rate at face value. Second, the meter prices every governed action. That gives the cost per outcome that a leaderboard leaves out. A fair judgment of a governed agent can be built from the record it leaves.
Footnotes#
-
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? ICLR 2024. https://arxiv.org/abs/2310.06770 ↩ ↩2 ↩3
-
OpenAI (2024). Introducing SWE-bench Verified. OpenAI, 13 August 2024. https://openai.com/index/introducing-swe-bench-verified/ ↩ ↩2
-
Yu, B., Zhu, Y., He, P., & Kang, D. (2025). UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench. https://arxiv.org/abs/2506.09289 ↩ ↩2
-
Liang, S., Garg, S., & Zilouchian Moghaddam, R. (2025). The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason. https://arxiv.org/abs/2506.12286 ↩ ↩2
-
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. https://arxiv.org/abs/2405.15793 ↩
-
Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., et al. (2024). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. ICLR 2025. https://arxiv.org/abs/2407.16741 ↩
-
Xia, C. S., Deng, Y., Dunn, S., & Zhang, L. (2024). Agentless: Demystifying LLM-based Software Engineering Agents. https://arxiv.org/abs/2407.01489 ↩ ↩2
-
Zhang, F., Chen, B., Zhang, Y., Keung, J., Liu, J., Zan, D., Mao, Y., Lou, J.-G., & Chen, W. (2023). RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation. EMNLP 2023. https://arxiv.org/abs/2303.12570 ↩
-
Yang, J., Jimenez, C. E., Zhang, A. L., Lieret, K., Wu, X., Muennighoff, N., Synnaeve, G., Narasimhan, K. R., Yang, D., Wang, S. I., & Press, O. (2024). SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains? ICLR 2025. https://arxiv.org/abs/2410.03859 ↩ ↩2
-
Deng, X., Da, J., Pan, E., He, Y. Y., Ide, C., Garg, K., et al. (2025). SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks? https://arxiv.org/abs/2509.16941 ↩ ↩2
Cite this
Oxagen Research. (2026, September 9). What SWE-bench Measures, and What It Misses. oxagen.sh. https://oxagen.sh/blog/what-swe-bench-measures-and-what-it-misses
BibTeX
@online{anderson2026whatswebench,
author = {{Oxagen Research}},
title = {What SWE-bench Measures, and What It Misses},
year = {2026},
date = {2026-09-09},
url = {https://oxagen.sh/blog/what-swe-bench-measures-and-what-it-misses}
}Related research
- Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's tracesEvery coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.
- Agents are waiting on a process built for peopleAn agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour.
- Deterministic coding agents: every turn on the recordA coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit.
Updates by email
New research reaches subscribers first.