AI agents, Coding agents
Deterministic coding agents: every turn on the record
A coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit.
On this page
- Coding agents improve when they get the results of running their code
- Random sampling makes the output change from one run to the next
- Output that changes between runs makes review harder in three ways
- What a turn record should contain
- Which parts of a run you can hold fixed
- How Oxagen records a coding agent's run
An agent changed fourteen files overnight. The diff is on the branch and the tests pass. In review, someone asks why the agent changed one of those files. No one can answer. The agent's reasoning and tool calls were not saved. Only the output is left.
If you run the agent again on the same ticket, it writes a different patch. So there is no earlier run to go back to, and the question has no answer at all.
This problem can be solved. The fix is not to make the model deterministic, meaning it gives the same output every time. The fix is to record the run so a person can replay it and check each turn. The research on what makes coding agents work points to the same records that review needs. So one set of records can serve both.
Coding agents improve when they get the results of running their code#
The results that moved code repair forward share one ingredient. A program runs the code, and the result goes back to the model.
Chen and colleagues showed this with Self-Debugging. In it, a model reads the results of running its code and explains the code back to itself. They compare this to rubber duck debugging, where you explain code out loud to find the bug. It gained 2 to 3% on the Spider text-to-SQL benchmark, with a 9% gain on the hardest difficulty level. It gained up to 12% on TransCoder and MBPP when unit tests were available.1 It also needed fewer tries. It matched baselines that generated over ten times as many candidates.1
Xia and Zhang applied the same idea to automated program repair with ChatRepair. The older method generates a batch of patches and then tests them. ChatRepair tests each patch right away and gives the result to the model before the next try. It fixed 162 of 337 bugs at roughly $0.42 each. It learned from failed attempts as well as successful ones.2 Reflexion used the pattern outside code. It stored written notes on failed attempts in memory. On HumanEval, 91% of problems passed on the first answer (pass@1), against 80% for the prior result.3
Read together, these three papers show the mechanism. The model does not get smarter between attempts. It gets new information: an exit code, a stack trace, or a diff that did not apply. That information exists as concrete records at one point in time. Most agent setups throw it away as soon as the loop moves on. Those are the records a reviewer would want.
The loop each repair result shares
- GenerateThe model writes a candidate patch
- ExecuteRun the tests against the code
- ResultExit code, stack trace, the part of the patch that failed
- ReviseThe next attempt reads the result
Repeat repeat until the tests pass
Random sampling makes the output change from one run to the next#
Benchmark papers have reported this for five years, in the metric they use.
Codex solved 28.8% of HumanEval problems with one sample (pass@1). It solved 70.2% when it could take 100 samples and count a problem solved if any passed (pass@100).4 Most of the measured skill sits in the gap between those two numbers. Much of what a model can do shows up only when you sample it many times. In a benchmark, you keep the sample that passed. In your repository, you keep the sample that ran. Nothing tells you whether it was the good one.
Temperature is the setting that controls how random the model's choices are. Lowering it does not fix this. Ouyang and colleagues sent the same requests many times across 829 code generation problems in three benchmarks. The share of tasks with no identical test output across the repeated requests was 75.76% on CodeContests, 51.00% on APPS, and 47.56% on HumanEval.5 They also found that setting temperature to zero reduces this variation but does not remove it.5 The same prompt, model, and repository can still produce a different patch.
So the variation comes with the system, and changing the model cannot engineer it away. You cannot predict which sample you will get. So record the one you got.
Output that changes between runs makes review harder in three ways#
First, you cannot narrow down a decision you cannot reproduce. A common way to debug is to run again with one thing changed, and repeat until you find the cause. This is called bisecting. It only works if the earlier run stays the same. If the agent writes a different patch each time, each step measures the random sampling instead of your change.
Second, the tests that grade the agent can also change between runs. Luo and colleagues did the first large study of flaky tests, which pass or fail without any code change. They examined 201 commits that fixed flaky tests across 51 open-source projects. The most common causes were async waits, concurrency, and test order dependence.6 When both the agent and its tests vary, one passing result is a matter of chance in two ways.
Third, a passing result shows less than it seems. Yu and colleagues studied SWE-Bench. In 345 patches, the tests from the original pull request never covered the failure. Those patches were marked as passing but were wrong.7 So a claim that the tests passed needs its evidence attached before anyone trusts it.
What a turn record should contain#
Review needs four nested records: the run, the turn, the tool call, and the gate result. A gate result is the pass or fail of a required check, such as the unit tests. Each record is added to the end of a log and never edited. This design is called event sourcing. Each event stays fixed once it is written, and you rebuild the current state by replaying the log. So the log is the audit trail. You do not have to keep a second record in step with it.
Here is one turn record, cut down to the fields review asks about:
{
"run": "r_01J8QK4M2ZT",
"turn": 7,
"model": "claude-opus-5",
"prompt_sha256": "9f2c1ab0…",
"tool_calls": [
{ "seq": 1, "name": "read_file", "path": "src/billing/grants.ts", "result_sha256": "1a7b…" },
{ "seq": 2, "name": "apply_patch", "diff_sha256": "c40e…", "files": 2, "added": 31, "removed": 4 },
{ "seq": 3, "name": "run_tests", "cmd": "pnpm --filter billing test:unit grants.test.ts", "exit": 0 }
],
"gate": { "name": "unit", "verdict": "pass", "log_sha256": "77d1…" }
}
The record stores hashes instead of full content. A hash is a short fingerprint computed from the content, and it changes when the content changes. Hashes keep the record small, and anyone can still check it. But this only works if every service hashes the same record in the same way. RFC 8785, the JSON Canonicalization Scheme, sets that rule. It fixes the serialization, the property order, and the encoding, so a JSON document has one standard form for signing and hashing.8 Without such a rule, two services can agree on the content and still compute different hashes. Then the audit trail stops checking out, and nothing warns you.
Here is what each record holds, and the review question it answers:
| Record | What it holds | What review can now ask |
|---|---|---|
| Run | The agent's identity, repository, branch, model id, tool versions | Who ran this, on what code, with which tools |
| Turn | Prompt hash, sampling settings, timestamp | What was asked, and with which settings |
| Tool call | Name, arguments, output hash, exit code | What the agent did to the code |
| Diff | Content hash, file count, lines added and removed | What changed, and how much code it touched |
| Gate result | Command, result, log hash | Whether "it passed" has evidence behind it |
None of these records needs the model to give the same output each time. They need the tool that runs the agent to record what happened in full.
Which parts of a run you can hold fixed#
Some parts of a run can be reproduced and some cannot. It helps to separate them.
You can hold fixed the toolchain version, the container image, the repository commit, the prompt text, the model identifier, the sampling settings, and the recorded input and output of every tool call. If you record those, a colleague can replay the same actions on the same code. This works even if a new generation would come out different.
Software supply chains faced this problem first. Lamb and Zacchiroli describe reproducible builds: you build the same source and get bit-for-bit identical output. They say plainly that this is hard in real projects. Timestamps, build paths, and ordering all leak into the output.9 The benefit is practical. A third party can check that a binary matches its source instead of trusting whoever built it. Coding agents cannot reach bit-for-bit output today. But the same goal applies: someone other than the author must be able to check the claim.
Design also matters. The agent-computer interface is the set of commands an agent uses to work with the computer. The SWE-agent authors argue that this interface decides what an agent can do at all. Their interface had purpose-built commands for navigation, editing, and testing. It reached 12.5% pass@1 on SWE-bench, where non-interactive approaches did far worse.10 A fixed set of named tools is also easy to log. An open-ended stream of shell commands is harder to record cleanly.
Agentless goes further. Xia and colleagues replaced the agent loop with three fixed steps: find the fault, repair it, and validate the fix. It resolved 32.00% of SWE-bench Lite at about $0.70 per issue. That was ahead of the open-source agents of the time.11 A pipeline with fixed steps is much easier to replay than an open-ended loop. This result does not argue against agents. It shows that a system can be both reproducible and capable. So pick an open-ended agent loop for a reason you can explain, and not by default.
How Oxagen records a coding agent's run#
Oxagen is the agent control plane for the agents you run. It does not run them. It cannot make a model's samples come out the same. It keeps the record complete. Oxagen records each run frame by frame. A frame is one recorded event in a run. Each frame carries a hash of the frame before it, so a changed frame breaks the chain. When the run ends, Oxagen signs the whole run, which is called the seal. Anyone with the export can then check a turn instead of taking it on trust. The coding agent does the work and sends the frames. Oxagen ties each step to the agent's mandate (its access, budget, tools, and rules) and to the rule that answered its request. The claim is narrow. A second run can still give a different patch. But you can say exactly what happened on the run you shipped.
Footnotes#
-
Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). Teaching Large Language Models to Self-Debug. ICLR 2024. https://arxiv.org/abs/2304.05128 ↩ ↩2
-
Xia, C. S., & Zhang, L. (2023). Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT. ISSTA 2024. https://arxiv.org/abs/2304.00385 ↩
-
Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. https://arxiv.org/abs/2303.11366 ↩
-
Chen, M., Tworek, J., Jun, H., et al. (2021). Evaluating Large Language Models Trained on Code. https://arxiv.org/abs/2107.03374 ↩
-
Ouyang, S., Zhang, J. M., Harman, M., & Wang, M. (2023). An Empirical Study of the Non-determinism of ChatGPT in Code Generation. https://arxiv.org/abs/2308.02828 ↩ ↩2
-
Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). An Empirical Analysis of Flaky Tests. FSE 2014, 643-653. https://doi.org/10.1145/2635868.2635920 ↩
-
Yu, B., Zhu, Y., He, P., & Kang, D. (2025). UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench. https://arxiv.org/abs/2506.09289 ↩
-
Rundgren, A., Jordan, B., & Erdtman, S. (2020). JSON Canonicalization Scheme (JCS). RFC 8785, IETF. https://www.rfc-editor.org/rfc/rfc8785.html ↩
-
Lamb, C., & Zacchiroli, S. (2021). Reproducible Builds: Increasing the Integrity of Software Supply Chains. IEEE Software. https://arxiv.org/abs/2104.06020 ↩
-
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. https://arxiv.org/abs/2405.15793 ↩
-
Xia, C. S., Deng, Y., Dunn, S., & Zhang, L. (2024). Agentless: Demystifying LLM-based Software Engineering Agents. https://arxiv.org/abs/2407.01489 ↩
Cite this
Oxagen Research. (2026, September 9). Deterministic coding agents: every turn on the record. oxagen.sh. https://oxagen.sh/blog/deterministic-coding-agents-every-turn-on-the-record
BibTeX
@online{anderson2026deterministiccodingagents,
author = {{Oxagen Research}},
title = {Deterministic coding agents: every turn on the record},
year = {2026},
date = {2026-09-09},
url = {https://oxagen.sh/blog/deterministic-coding-agents-every-turn-on-the-record}
}Related research
- Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's tracesEvery coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.
- Agents are waiting on a process built for peopleAn agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour.
- Steering a run you are not watchingWhen a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.
Updates by email
New research reaches subscribers first.