# Site blueprint > How macanderson.com is built to be read by people and by AI agents, and how to build a site like it. Markdown twins, llms.txt, an MCP server, a guarded assistant, and an installable app. Canonical page: https://macanderson.com/blueprint This site is written for two kinds of readers: people, and the AI agents that read the web for them. Every choice below serves one of them, and most serve both. Copy any of it. ## Every page has a markdown twin Add `.md` to any address and you get the same page as plain markdown. A request that sends `Accept: text/markdown` gets the markdown version at the normal address. Agents read markdown faster and more accurately than a page full of layout, and each twin names its canonical page so search engines know which one to rank. ## llms.txt and llms-full.txt [/llms.txt](/llms.txt) is a short map of the site in the [llms.txt format](https://llmstxt.org): who Mac is, then a link and a one-line description for every page. [/llms-full.txt](/llms-full.txt) holds the whole site in one file, so an agent can load all of it in one request. ## An MCP server The site runs a [Model Context Protocol](https://modelcontextprotocol.io) server at `https://macanderson.com/mcp`. It is read-only and needs no key. Its tools search the site, read any page as markdown, list the field manual and research, and return Mac's profile and GitHub activity. To add it to Claude Code: ```sh claude mcp add --transport http macanderson https://macanderson.com/mcp ``` In browsers that support WebMCP, the page also registers its tools with the browser, so an agent in the tab can search and open pages without scraping them. ## Structured data on every page Each page carries schema.org data that says what it is. The home page describes Mac as a `Person` with his profiles. The field manual is a `Book` whose parts are chapters. Each part and essay is a `TechArticle` that lists its cited sources. Search engines and agents read these facts directly instead of guessing them from the layout. ## A guarded assistant The assistant answers questions about Mac and his work. It follows the method the field manual teaches: decide in code what code can decide, and spend the model last. 1. Code finds the passages first. A keyword index (BM25) ranks passages from the site's own pages before any model call. The same question always gets the same passages. 2. The model sees only those passages, under a fixed token budget. It is told to answer from them and to say so when they do not cover the question. 3. Every answer lists its sources, so you can check it. 4. Limits sit in code, not in the prompt. Each question has a length cap. Each network address has a request limit. The whole site has a daily cap, and Vercel BotID screens out automated traffic. The assistant has no tools that write anything. 5. When the model is unavailable or a limit is reached, the assistant still answers with the best matching passages and links. ## An installable app The site is a progressive web app. On a phone, add it to your home screen and it opens full screen with its own icon. A service worker keeps the pages you have read available offline. ## Fast by default Pages are built ahead of time and served as static files. The only code that runs on a server is the assistant, the subscribe form, and the MCP endpoint. Fonts load from the same domain. Animations use the browser's own scroll and view timelines, and they stop when your device asks for reduced motion. ## Build your own You need three things: your writing in plain files, a build step that turns them into pages and markdown twins, and a short list of facts about you that every page can reuse. Start with llms.txt and the markdown twins. They take an afternoon and help every agent that reads your site. Add the MCP server and the assistant when you have enough writing for them to search. --- # Mac Anderson > Mac Anderson is the founder and CEO of Oxagen, the author of Engineering Deterministic AI Coding Agents, and the creator of Stella, an open-source coding agent. He researches and builds coding agents that ship software on their own. Canonical page: https://macanderson.com/ Mac Anderson researches agentic coding systems and builds coding agents. Some of them are self-evolving: they write their own tools and get better at the job. His goal is software that agents deliver on their own, and that is useful, scales, and stays easy to maintain. He lives in Los Angeles. He is the founder and CEO of [Oxagen](https://oxagen.sh), workforce management for autonomous agents. Each agent gets its own identity, a mandate that sets what it may do and spend, the tools it needs, and a record of what it did. He also created [Stella](https://github.com/macanderson/stella), an open-source coding agent written in Rust, and wrote *Engineering Deterministic AI Coding Agents*, a field manual in 21 parts. He wrote software by hand for 16 years. Before Oxagen he co-founded Fonteva, a Salesforce-native software company for associations that Togetherwork acquired in 2021, and founded inTown Technologies, which built software for city governments. In 2026 he handed all of his coding to agents and kept the job of directing them, reading their work, and deciding what merged. Between 26 February and 20 September 2026 his GitHub account recorded 21,861 contributions, against 38 in all of 2025. ## What he builds - **[Oxagen](https://oxagen.sh)** gives each AI agent an identity, a mandate, equipment, and a record. For actions routed through Oxagen, the agent asks, a rule the team wrote answers, and Oxagen records the answer with its cost. It wraps the agents teams already run, such as Claude Code and Codex. - **[Stella](https://github.com/macanderson/stella)** is a fast, model-agnostic terminal coding agent in Rust. You bring your own model key. On Terminal-Bench 2.1, with both agents on the same model, Stella solved 58 of 89 tasks and Claude Code solved 44. - **[Arena](https://arena.oxagen.sh)** runs a coding agent head to head against Claude Code, Gemini, and other agents, with regression checks. - **[Context Graph Protocol](https://contextgraphprotocol.org)** is an open specification for representing, exchanging, and accounting for an agent's context as a typed graph of frames, with JSON schemas, a conformance suite, and SDKs for TypeScript, Python, and Go. ## What he writes - **[Engineering Deterministic AI Coding Agents](/manual)** is a free field manual in 21 parts with 27 cited sources. Parts 1 to 13 show how to build the deterministic system around a coding agent. Parts 14 to 20 show how to operate the agents you run. Every part ends with steps for this week and metrics to track. - **[Research](/research)** holds essays on how agents work, fail, and should be operated, each with its sources. ## Where to start - If an agent costs too much or behaves differently each time, read [part 1](/manual/agents-are-not-expensive-bad-architecture-is) and [part 2](/manual/stop-making-agents-read-your-logs). - If you answer for several agents, read [part 14](/manual/give-every-agent-its-own-identity), [part 15](/manual/write-the-mandate), and [part 16](/manual/the-agent-asks-a-rule-decides). - To find out what to fix first, take the [self-assessment](/manual/self-assessment). - To ask a question about Mac or his work, use the [assistant](/ask). ## Contact Email [mac@oxagen.sh](mailto:mac@oxagen.sh). Mac is [macanderson on GitHub](https://github.com/macanderson), [macanderson on LinkedIn](https://www.linkedin.com/in/macanderson), [@macandersoncto on X](https://x.com/macandersoncto), and [macanderson on Hugging Face](https://huggingface.co/macanderson). His ORCID iD is [0009-0005-1646-9676](https://orcid.org/0009-0005-1646-9676). New research, field manual updates, and his upcoming video series go to [subscribers](/subscribe) first. --- # Privacy > What this site collects, why, and how to remove it. It sets no tracking cookies and runs no analytics. Canonical page: https://macanderson.com/privacy This site sets no tracking cookies and runs no analytics scripts. ## Subscribing When you subscribe, the site stores your email address, the name you give (if any), the topics you pick, and the time you signed up. Mac uses them only to send you new research, field manual updates, and news about his video series. Nobody else gets them, and nothing is sold. To be removed, use the [unsubscribe form](/subscribe#unsubscribe) or email [mac@oxagen.sh](mailto:mac@oxagen.sh). Removal deletes your record. ## The assistant The assistant sends your question to a language model through Vercel's AI Gateway to write an answer. The site does not store your conversations. It counts requests per network address for a short time to stop abuse, and then forgets them. Do not type personal or secret information into the assistant. ## Saved in your browser Your theme choice, the checklist items you tick in the field manual, and your self-assessment answers stay in your own browser's storage. They never leave your device. ## Hosting The site runs on Vercel. Vercel keeps standard server logs, such as request times and network addresses, to run the service. --- # The story > In 2026 Mac Anderson stopped writing code by hand and let AI agents write all of it. This is what the seven months from 26 February to 20 September looked like, from his own record. Canonical page: https://macanderson.com/story Mac Anderson wrote software by hand for 16 years. In 2026 he made a decision: agents would write all of the code he shipped. His job would be to direct them, read their work, and decide what merged. Some of it went well. A lot of it went wrong. The hardest part turned out to be managing the agents, and he built Oxagen to do that job. This page follows those seven months from the record he left: his GitHub history, the git history of 13 repositories, and 8,133 prompts he sent to Claude Code. The interactive original, with the charts, is at [oxagen.sh/story](https://oxagen.sh/story). ## Nine chapters ### Before the log *26 February to 28 June.* GitHub shows the work starting on 26 February. March brought 689 contributions, April 1,846, and May 2,167. On 28 May the first commits land in the Oxagen repository: the foundations and the agent runtime. His prompts from these months were on an earlier laptop. So this chapter uses commits only. ### Beat Claude Code *29 June to 9 July.* He sets up a new laptop and works fast on the TypeScript Oxagen command-line agent. He adds caching, effort levels, a code graph, and an SWE-bench test setup. On 3 July he sets the goal: "Stripe for agents." Four days later he drops metered billing and moves to governance, traceability, budgets, and permissions. He also writes "if you see it you own it" into his agent instructions. > i do not want to falsely promote that oxagen is better ### Stella is born *10 July to 19 July.* One night of "crazy ideas" leads to the mutation verifier. It checks that each test fails when its fix is removed. The next day the Rust rewrite gets a name, Stella, and a public repository. The witness test becomes the definition of done: a test that fails on the old code and passes on the new code. By 17 July you can install Stella 0.3.0 from Homebrew, and fleets run with no coordinator at all. > determinism is fast, cheap, reproducable, and better than intelligence 10/10 times. ### Protocols and proof *20 July to 31 July.* The context protocol is renamed the Context Graph Protocol, and its SDK ships on npm. ArenaBench goes public. Steering records store the agent's standing instructions in git, as TOML files. On 31 July he runs the first head-to-head test against Claude Code on Terminal-Bench 2.1, with its rules published before the run. Both agents use the same model. Stella solves 58 of 89 tasks, and he publishes the run. > verification can not lie ### The hard weeks *1 August to 11 August.* Benchmark runs had hidden caps. A verifier model spent four minutes and never wrote its test. On 8 August he found that the code graph had been empty, with no error. Stella's "losses" came from a $1-per-task budget cap. So proving the agent worked took more effort than building it. He made rules from these problems, and they still hold. Prove a defect with a failing test first. Check results with tests, not with a model. Never cap a benchmark. On 11 August he also published a result that made Stella look worse: on Sonnet 5, Stella cost about four times what Claude Code did. > WE HAVE MOVED TO FAST AND I AM WORRIED I WILL HAVE TO THROW AWAY THIS PROJECT ### Less is more *12 August to 19 August.* Semantic code search starts working, and Stella's 72 tools shrink to five. Verification moves out of Stella into Vera, a plugin that can block a turn. The core becomes one loop. Everything else becomes a plugin, written in any language. On 19 August he deletes the pipeline crate. Self-driving Stella starts working through its own list of issues. > less is more and way less is best the 5 tools kicked ass ### Many agents at once *20 August to 30 August.* He runs ten to fifteen agent conversations at once, each on a long goal that needs little input. So most of his prompts are short: "go," "resume," and "merge." He types fewer prompts, and the output climbs. He moves everything to AWS. He writes down how he steers his agents as standing decisions. He has Stella build a supply-chain demo app on its own. On 30 August a file-by-file review of Oxagen opens 1,059 issues. GitHub counts 1,144 contributions that day, the most of any day. > it is supposed to be hands off autonomous with controls to rollback ### Oxagen becomes the control plane *31 August to 10 September.* He removes the generative AI features from Oxagen and makes Stella the engine inside it. Agents get their own identities and permissions, the same way people do. The Terminal-Bench page goes up: Stella solves 58 of 89 tasks, and Claude Code solves 44. He adds wrappers, hooks installed beside an agent, so Oxagen can govern Claude Code and Codex. Oxagen starts pricing by the governed action: one call routed through Oxagen that it checked and recorded. > I think stella has the best agent. I think oxagen has the best control plane. ### Starting over from a new spec *11 September to 20 September.* He restarts Oxagen from a new spec. The spec says: wrap any agent, decide what each tool call may do, replay every run, and match spend to the cent. About twelve agent conversations, running side by side, build the demo mockups in one day. Within the week, a macOS installer wraps Claude Code and Codex. Oxagen becomes a gateway for model calls, and every model call on his own laptop goes through it. > they are worried about agents and the damage is done at tool call level. ## Learning to manage a fleet With one agent, you give it a task, watch it work, and read the diff. With many agents at once, the job changes. It was the hardest thing he learned in seven months, and none of it was about prompting. He started with one agent on one branch, and he watched it work. That setup cannot grow, because a person can only watch one thing at a time. So he added more agents. In the setup he ended with, six or more agents work in parallel, each in its own git worktree. More agents caused problems he had not expected. Two agents edited the same branch. A branch merged into main with no conflicts, but it still did not work with code that had merged beside it. He had to repair a broken main three separate times before he understood the cause. A merge with no conflicts says nothing about whether the code works. Agents also finished, reported success, and left their work uncommitted. What running many agents needed: - A separate worktree for each agent, so two agents cannot edit the same files. - A written definition of done. Without one, an agent decides for itself what finished means. - Checks that run without him. There are 21 of them, so he does not have to review every agent's work by hand. - A trial merge before the real one, because a merge with no conflicts can still break the code. What he still could not answer: - What each agent did, in order, and what it read. - What that cost, and which piece of work it paid for. - Who allowed each action, and under which rule. - Which agent is stuck right now, and on what. He could build the first list himself. The second list is a management problem, and it got worse with every agent he added. He built Oxagen to answer it. ## The code his agents wrote Counted from tracked files on 21 September 2026: - **2,258 test files.** Unit, integration, and end to end. - **272 architecture tests.** They fail the build when code breaks a structure rule, such as a layer, import, or boundary rule. - **21 checks in CI**, including one that fails on an em dash. - **136 decision records**, plus 5 schema change records. Each one says what was decided and why. A human reviewer notices when a file imports across a boundary it should not. Many agents working in parallel do not. Each may break the same rule in a different place on the same afternoon. A test that fails the build checks every change, however many agents make them. So with many agents, the architecture tests matter most. ## The counts He did not type this code. Agents wrote it. He directed them, read their work, and merged it. - **21,861 GitHub contributions**, 26 February to 20 September 2026. The same account shows 38 in all of 2025. - **6,341 commits** on the default branch of 13 repositories, 28 May to 20 September. He made commits on 114 of those 116 days. - **4,427 pull requests merged**, 28 May to 20 September. - **8,133 prompts to Claude Code**, 29 June to 20 September, across 1,769 conversations on one laptop. - **3,906,359 source lines added and 1,529,699 removed**, counting source files on default branches only. Read these counts with three facts beside them. The repositories were four months old, with no release schedule, no on-call rotation, and no other engineers. When he opened and merged a pull request himself, no second person reviewed it. The prompt count is the least complete number, because it covers 81 days on one laptop. ## How the counts were made GitHub contributions come from GitHub's contribution calendar for the account macanderson, private repositories included. Commits and lines come from the default branch of 13 local clones, with each commit counted once by its hash, across the 10 name and email pairs he has used. Prompts come from Claude Code's prompt history on one laptop, loaded into SQLite and read in order. The 58 of 89 result is Stella against Claude Code on Terminal-Bench 2.1, both on the same GLM model, on 31 July 2026. --- # Work > What Mac Anderson builds and has built. Oxagen, Stella, Arena, the Context Graph Protocol, his career from Fonteva to Oxagen, and his GitHub activity. Canonical page: https://macanderson.com/work Mac Anderson is the founder and CEO of Oxagen, and its principal engineer and architect. He is also the creator of Stella, Arena, and the Context Graph Protocol. Before Oxagen he co-founded Fonteva, which Togetherwork acquired in 2021, and founded inTown Technologies. ## Oxagen [Oxagen](https://oxagen.sh) is workforce management for autonomous agents. You give each agent its own identity. You set its authority, budget, tools, and skills in one mandate. For actions routed through Oxagen, the agent asks, a rule your team wrote answers, and Oxagen records the answer with its cost. A mandate has four parts: - **Access**: the identity the agent acts as, and what it may request. - **Budget and rules**: what it may spend, and the rules it works under. - **Equipment**: the tools, skills, and business context it may use. - **Record**: what each governed run read, changed, and cost, and which rule answered each request. Oxagen does not run agents. It wraps the ones teams already use, such as Claude Code and Codex, through hooks installed beside the agent and a gateway for model calls. ## Stella [Stella](https://github.com/macanderson/stella) is a fast, model-agnostic terminal coding agent written in Rust. You bring your own model key. Its definition of done is a witness test: a test that fails on the old code and passes on the new code. Its core is one loop with five tools, and everything else is a plugin. On Terminal-Bench 2.1 on 31 July 2026, with both agents on the same GLM model, Stella solved 58 of 89 tasks and Claude Code solved 44. Stella is open source under the AGPL. Install it with `brew install oxageninc/stella/stella`, and read the docs at [stella.oxagen.sh](https://stella.oxagen.sh). ## Arena [Arena](https://arena.oxagen.sh) ([source](https://github.com/macanderson/arena)) runs your coding agent head to head against Claude Code, Gemini, and other agents, with regression checks that keep it improving instead of drifting. ## Context Graph Protocol The [Context Graph Protocol](https://contextgraphprotocol.org) (CGP) is an open specification for representing, exchanging, and accounting for an agent's context as a typed graph of frames instead of loose text. It ships with JSON schemas, a conformance suite, and SDKs for TypeScript, Python, and Go. ## Other open source - **[stella-lang](https://github.com/macanderson/stella-lang)**: a programming language written by AI agents for AI agents. - **[token-tax-meter](https://github.com/macanderson/token-tax-meter)**: measures what each model actually bills for the same sentence in six languages. - **[rainforest](https://github.com/macanderson/rainforest)**: a proving ground. A supply-chain app specified in markdown and issues, built on its own by Stella. - **[arena-bench](https://github.com/macanderson/arena-bench)**: a loop-integrity benchmark for coding agents, with kill-and-resume chaos tests and execution-trace checks. ## Career - **2026 to now. Oxagen, founder and CEO.** Mac founded Oxagen in 2026 and is its principal engineer and architect. He builds the agent control plane, Stella, and the Context Graph Protocol. - **2024 to 2026. The Unnatural Group, director of AI and machine learning engineering.** In Los Angeles he built the firm's AI and machine learning team and its trading systems, including LSTM pricing models and retrieval over a Neo4j knowledge graph. - **2021 to 2024. inTown Technologies, founder and CTO.** In Washington, DC, he built software for city governments, including a multi-agent LLM system that answers residents' requests. - **2010 to 2021. Fonteva, co-founder and CTO.** Mac co-founded Fonteva in 2010 with Jerry Huskins and Paul Lundy. Fonteva built membership, events, and commerce software for associations, native to Salesforce. As CTO he built the engineering team and the platform. In 2017 Fonteva's revenue grew 71% to $12.8 million, and it had more than 100 employees. The Washington Business Journal ranked it 24th among the region's fastest-growing companies in 2018, and the U.S. Small Business Administration named it Small Business Exporter of the Year for the Washington area. Togetherwork acquired Fonteva in February 2021. - **Before 2010. Wipro, senior technical architect.** He led architecture for enterprise Salesforce Service Cloud implementations. ## Talks - **Dreamforce 2014**: "Design Patterns Every ISV Needs to Know," with Andrey Volosevich of Salesforce and Ross Belmont of Appiphony. --- # Engineering Deterministic AI Coding Agents, second edition > A field manual in 21 parts, by Mac Anderson of Oxagen. Parts 1 to 13 show how to build the deterministic system around a coding agent, with published evidence in each part. Parts 14 to 20 show how to operate the agents you now run: identity, authority, budget, equipment, and a record another person can read. Every part ends with steps for this week and metrics to track. By Mac Anderson. 21 parts and 27 cited sources. - [How to use this book](https://macanderson.com/manual/how-to-use-this-book.md): Read the two parts that match your problem this week. Come back for the rest. - [Part 1: AI agents are not expensive. Bad architecture is.](https://macanderson.com/manual/agents-are-not-expensive-bad-architecture-is.md): Most of an agent's token bill comes from the system around the model, not from the model. - [Part 2: Stop making agents read your entire logs](https://macanderson.com/manual/stop-making-agents-read-your-logs.md): Logs are structured data, and treating them as prompt text costs tokens you do not have to spend. - [Part 3: Parse the user prompt before the model ever sees it](https://macanderson.com/manual/parse-the-prompt-before-the-model-sees-it.md): Your prompt already carries structured information, so extract it before you pay for inference. - [Part 4: Why grep is the wrong retrieval engine](https://macanderson.com/manual/why-grep-is-the-wrong-retrieval-engine.md): Searching is not understanding, and code already carries the structure you need. - [Part 5: Build semantic memory once](https://macanderson.com/manual/build-semantic-memory-once.md): A run should start from what the last one learned instead of repeating it. - [Part 6: A rule written in prose cannot fail CI](https://macanderson.com/manual/rules-belong-in-data-not-in-prose.md): Constraints belong in a representation a validator can check. - [Part 7: Context compression is worth more than a bigger model](https://macanderson.com/manual/compression-beats-a-bigger-model.md): Removing irrelevant information beats buying a larger model. - [Part 8: Agents guess at schema they were not shown](https://macanderson.com/manual/show-the-agent-the-schema.md): A model adds a duplicate column when the context did not carry the existing one. - [Part 9: Why coordinator agents don't scale](https://macanderson.com/manual/why-coordinator-agents-do-not-scale.md): Most multi-agent systems pay a coordination tax and call it architecture. - [Part 10: Tests should be first-class retrieval objects](https://macanderson.com/manual/tests-as-retrieval-objects.md): Tests describe behavior in executable form, and a test runner's verdict is computed rather than inferred. - [Part 11: Agents need working memory, not bigger context windows](https://macanderson.com/manual/working-memory-not-bigger-windows.md): You do not reread every book you own before fixing a bug, and your agent should not either. - [Part 12: Knowledge graphs beat prompt engineering](https://macanderson.com/manual/knowledge-graphs-over-concatenation.md): Knowledge should be connected, not concatenated. - [Part 13: Measuring agent intelligence](https://macanderson.com/manual/measuring-an-agent-in-production.md): Benchmarking on coding challenges misses what matters in production. - [Part 14: Give every agent its own identity](https://macanderson.com/manual/give-every-agent-its-own-identity.md): You cannot set authority for, bill, or review something you cannot name. - [Part 15: Write the mandate](https://macanderson.com/manual/write-the-mandate.md): Four decisions govern an agent. Most teams have made all four. Few have them in one place. - [Part 16: The agent asks, a rule decides](https://macanderson.com/manual/the-agent-asks-a-rule-decides.md): Authority is a decision at the moment of use, written down by the team that owns the system. - [Part 17: Spend you can attribute](https://macanderson.com/manual/spend-you-can-attribute.md): A total is not an answer. The answer is which agent spent what, and on whose behalf. - [Part 18: Equip agents on purpose](https://macanderson.com/manual/equip-agents-on-purpose.md): A tool in the catalog is not a tool in the agent's hands. Assign equipment the way you assign access. - [Part 19: Keep a record another person can read](https://macanderson.com/manual/keep-a-record-another-person-can-read.md): A log answers the engineer who wrote it. A record answers the person who was not there. - [Part 20: Bounded tasks and ongoing work](https://macanderson.com/manual/bounded-tasks-and-ongoing-work.md): Some work has an endpoint. Other work continues. Manage each in its own way. - [Part 21: Better systems, not better models](https://macanderson.com/manual/better-systems-not-better-models.md): You rent the model. You own the system around it, and the way you operate it. - [Closing thought](https://macanderson.com/manual/closing-thought.md): Intelligence is expensive. Determinism is cheap. Spend the first only where the second cannot do the job. - [Field kit](https://macanderson.com/manual/field-kit.md): Templates and worksheets to copy. Print this section and fill it in with the people who own each line. - [Glossary](https://macanderson.com/manual/glossary.md): The words this book uses in a fixed sense. - [About the author](https://macanderson.com/manual/about-the-author.md): Oxagen is workforce management for autonomous agents: give each agent an identity, set its authority and budget, equip it with tools and skills, and review… - [Sources](https://macanderson.com/manual/sources.md): Every empirical claim in this book traces to one of the sources below. Read the primary literature. It is better than any summary of it, including this one.… --- # Research by Mac Anderson > Essays on how AI agents work, fail, and should be operated. Each one cites its sources. - [Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models.md) (2026-10-09): Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week. - [Agents are waiting on a process built for people](https://macanderson.com/research/agents-are-waiting-on-a-process-built-for-people.md) (2026-09-30): An agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour. - [Steering a run you are not watching](https://macanderson.com/research/steering-a-run-you-are-not-watching.md) (2026-09-16): When a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record. - [The agent time horizon is doubling](https://macanderson.com/research/the-agent-time-horizon-is-doubling.md) (2026-09-16): METR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision. - [The problem is not slop, it is your process](https://macanderson.com/research/the-problem-is-not-slop-it-is-your-process.md) (2026-09-16): The worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work. - [What an agent should be told before it starts](https://macanderson.com/research/what-an-agent-should-be-told-before-it-starts.md) (2026-09-16): A prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things. - [A model cannot grade its own homework](https://macanderson.com/research/the-limits-of-self-correction-and-model-collapse.md) (2026-09-09): Without outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead. - [Deterministic coding agents: every turn on the record](https://macanderson.com/research/deterministic-coding-agents-every-turn-on-the-record.md) (2026-09-09): A coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit. - [From STaR to DeepSeek-R1: what self-improvement means](https://macanderson.com/research/self-improving-models-from-star-to-self-rewarding.md) (2026-09-09): A vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one. - [Governing an agent that rewrites itself](https://macanderson.com/research/governing-an-agent-that-rewrites-itself.md) (2026-09-09): An agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts. - [Graph-grounded retrieval vs vector search](https://macanderson.com/research/graph-grounded-retrieval-vs-vector-search.md) (2026-09-09): Vector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference. - [Self-evolving agents: what the evidence shows](https://macanderson.com/research/self-evolving-agents-what-the-evidence-shows.md) (2026-09-09): What changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve. - [The science of AI agents: from ReAct to tool use](https://macanderson.com/research/the-science-of-ai-agents-from-react-to-tool-use.md) (2026-09-09): Planning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do. - [What an Ontology Buys an Agent](https://macanderson.com/research/what-an-ontology-buys-an-agent.md) (2026-09-09): An agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers. - [What SWE-bench Measures, and What It Misses](https://macanderson.com/research/what-swe-bench-measures-and-what-it-misses.md) (2026-09-09): Coding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you. - [Why agents fail: measuring reliability and cost](https://macanderson.com/research/why-agents-fail-measuring-reliability-and-cost.md) (2026-09-09): AgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides. --- # How to use this book > Read the two parts that match your problem this week. Come back for the rest. Field manual. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/how-to-use-this-book You run coding agents, or you are about to, and two questions keep coming up. The first is from your own team: why does this cost so much and behave differently each time? The second is from everyone else: what are these agents allowed to do, what did they spend, and who answers for them? Parts 1 to 13 answer the first question. Parts 14 to 20 answer the second. Part 21 explains why both have the same answer. ### How each part is laid out - **The situation.** Each part opens on a problem you have probably seen. - **The mechanism.** What to build, in order, with a short sample. - **Evidence.** A published result, cited. The sources are listed at the back, and the primary papers are better than any summary of them. - **Counterweight.** Where the argument stops applying. Not every part has one. - **Do this week.** Three to five steps, cheapest first. Each says what done looks like. - **Measure it.** The readings that tell you whether the step worked. Parts 14 to 20 add a short box, "Where Oxagen fits". Oxagen publishes this book and builds a product for that half of the job. The practice in each part stands without the product, and the box says what the product does and where its boundary is. Parts 1 to 13 mention no product. ### Three reading paths | If you are | Start with | Then | | --- | --- | --- | | An engineer building or tuning an agent | Parts 1, 2, and 3 | Part 13 to set up measurement, then the rest of 4 to 12 in the build order below | | A platform lead or operator answering for several agents | Parts 14, 15, and 16 | Parts 17 and 20, then part 13 for the engineering scorecard | | A security or finance lead reviewing agent use | Part 16 (security) or part 17 (finance) | Parts 15 and 19, which cover the object you review and the record you review it from | ### A build order The parts are numbered by topic, not by the order to build them. If you are starting from an agent that works and costs too much, this order pays back fastest, because each step makes the next one measurable. | Step | Build | Part | Why now | | --- | --- | --- | --- | | 1 | Token and cost logging per step, with run, agent, and person | 13, 17 | Nothing later can be verified without it | | 2 | An inventory of agents, each with an operator | 14 | One afternoon, and every later step needs the names | | 3 | Trace and log slicing | 2 | The largest reduction in context for the least code | | 4 | Prompt parsing and a symbol index | 3, 4 | Replaces the exploratory calls at the start of a run | | 5 | Schema slice and test lookup | 8, 10 | Removes two common causes of wrong patches | | 6 | A mandate for each agent, and rules for the top ten write actions | 15, 16 | By now other teams are asking | | 7 | A workflow skeleton and tiered memory | 9, 11 | Worth doing once runs are long enough to loop | | 8 | Per-agent tool assignment | 18 | Needs the offered-against-called data from step 1 | | 9 | Domain dossiers and the knowledge graph | 5, 12 | The largest build. It federates the indexes from steps 4 and 5 | | 10 | The run record, chained, and completion checks for bounded tasks | 19, 20 | Turns the rows you already write into something a reviewer can use | ### A self-assessment Answer yes or no. Count a yes only if you could show the evidence today. Each no points at a part. | # | Question | If no, read | | --- | --- | --- | | 1 | Can you split one run's tokens into context construction and reasoning? | 1 | | 2 | Does code, not the model, reduce a stack trace before it reaches the window? | 2 | | 3 | Are file paths and symbols in a prompt resolved against an index before the first model call? | 3, 4 | | 4 | Does a new run start from stored knowledge of the repository? | 5 | | 5 | Can every standing instruction fail a check in CI? | 6 | | 6 | Does every context assembler take a token budget? | 7 | | 7 | Does the agent see the live schema, with lineage, before it writes a migration? | 8 | | 8 | Is control flow between steps written as code? | 9 | | 9 | Can you look up the tests that cover a function? | 10 | | 10 | Does old content leave the window by policy? | 11 | | 11 | Can you answer a three-hop question about your system with one query? | 12 | | 12 | Do you report cost per completed task, including retries? | 13 | | 13 | Can you list every agent that can write to a system, with a named operator for each? | 14 | | 14 | Is each agent's authority, budget, and equipment written in one reviewed place? | 15 | | 15 | For a given write action, can you name the rule that allowed it and its author? | 16 | | 16 | Can you attribute last month's spend to the agent, the run, and the person? | 17 | | 17 | Do you know how many tools each agent is offered, and how many it calls? | 18 | | 18 | Could someone outside your team reconstruct one agent action in 30 minutes? | 19 | | 19 | Are completion checks for bounded tasks fixed before the run starts? | 20 | | 20 | Does each ongoing agent have a review date on a calendar? | 20 | Score yourself now and write the number down. Part 21 asks you to do it again. --- # AI agents are not expensive. Bad architecture is. > Most of an agent's token bill comes from the system around the model, not from the model. Part 1: The economics. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/agents-are-not-expensive-bad-architecture-is Your agent works, and the invoice is bigger than the team that uses it. The first suspects are the obvious ones: the per-token price, the reasoning overhead, the size of the context window. Then you instrument the agent, log every request, and look at where the tokens went. The model did what it was asked to do. The system asked it to do far too much. ### Measure where tokens actually go Run an agent framework against a real repository and log every request. In the runs this book draws on, the dominant cost is not *reasoning about the problem*. It is *rebuilding context the system already had*: re-reading files it read two turns ago, carrying 400 lines of stack trace when 6 frames mattered, pushing the whole conversation history into every tool call, and walking the filesystem with `ls` and `cat` because nothing indexed the repository up front. The published numbers point the same way. Anthropic's engineering team, describing their production multi-agent research system, reported that agents consume roughly **4× the tokens of a chat interaction**, and multi-agent systems roughly **15×**, and that on their internal evaluations **token usage alone explained about 80% of performance variance**.[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") The largest lever they found was not model choice or prompt wording. It was how many tokens the architecture decided to spend. > **Evidence · Simplicity wins the leaderboard** > > The clearest data point in this book is **Agentless** (Xia et al., UIUC). Instead of an autonomous agent choosing its own actions with tools, Agentless runs a fixed three-phase pipeline: hierarchical localization, patch generation, validation. On SWE-bench Lite it became the top open-source approach of its time at **27.3% solved for $0.34 per issue** (later 32.0% at $0.70), while contemporary agent-based systems spent roughly 10× more per issue for comparable or worse results.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") The model did not choose the next action. The pipeline chose, and the model filled in the parts that needed judgment. > > Xia et al., "Agentless: Demystifying LLM-based Software Engineering Agents," FSE 2025 ### Context construction against reasoning cost Split your agent's token bill into two buckets. **Reasoning tokens** are the model thinking about the actual problem: the diff, the root cause, the design decision. **Context-construction tokens** are everything spent getting the model to the point where reasoning can begin: file dumps, search results, log output, retries after the model lost the thread. In an unindexed architecture the second bucket dominates, often by an order of magnitude. Most of what sits in that second bucket is computable without a model. Parsing, indexing, slicing, and graph traversal are deterministic operations. They run on CPU, return the same answer on the same input, and you can test them like any other code. You cannot split the bill you do not record, so record it at the call site. The caller knows why it is making the call, and the caller is the only part of the system that does. ``` # One row per model call, written when the call returns. log_call( task_id=task.id, purpose="context", # or "reasoning", set by the caller source="file_read:billing/processors.py", input_tokens=usage.input, cached_input_tokens=usage.cache_read, output_tokens=usage.output, ) # context_token_share = input where purpose == "context" / all input # Group by source to rank what is filling the window. ``` > **Evidence · Cost must be a first-class metric** > > Princeton's "AI Agents That Matter" (Kapoor, Stroebl, Siegel, Nadgir, Narayanan) showed that agent research had been ignoring cost, and that when you evaluate on a joint cost and accuracy Pareto frontier, simple baselines such as retrying a model call matched or beat elaborate agent architectures on HumanEval at a fraction of the price. Their conclusion: accuracy alone cannot identify progress, and state of the art (SOTA) agents are frequently more complex and more costly than the task requires.[2](https://macanderson.com/manual/sources#r2 "Kapoor, Stroebl, Siegel, Nadgir & Narayanan (Princeton). \"AI Agents That Matter.\" TMLR 2025. arXiv:2407.01502") > > Kapoor et al., "AI Agents That Matter," TMLR 2025 ### Thinking harder against needing less thinking The default answer to agent failure is to think harder: a bigger model, a longer reasoning chain, more retries, more sub-agents. Each of those multiplies token spend and adds another probabilistic step, which is another place for variance to enter. The engineering answer runs the other way. Shrink the problem before the model sees it. Every ambiguity you resolve deterministically is reasoning the model no longer has to do, tokens you no longer pay for, and variance you no longer ship. Determinism is not a philosophical preference. It is a cost model. Attributing that spend across the teams who run the agents is part 17. > **Do this week** > > 1. **Log usage on every model call.** Write one row per call with `task_id`, `purpose`, `source`, input tokens, cached input tokens, and output tokens. The artifact is a table you can group by, not a log line you can read. > 2. **Tag each call reasoning or context.** Set the tag at the call site, where the caller knows why the call is happening. Done looks like every call carrying a tag and none defaulting to unknown. > 3. **Rank your context sources.** Group last week's rows by `source` and sum input tokens. The top three sources are your work queue for parts 2 to 4. > 4. **Turn on prompt caching for the stable prefix.** Put the system prompt, tool definitions, and repository map ahead of anything that changes per turn, then watch the cache-read share rise. > 5. **Put a token budget in code.** Give each task a ceiling on input tokens and fail the run when it trips, so a runaway loop ends in a test failure instead of an invoice. > **Measure it** > > - `context_token_share`: context-construction input tokens divided by all input tokens, per task. Down is good. Above 0.8 means the model is paying to rediscover what your code already knows. > - `tokens_per_completed_task`: input plus output tokens divided by tasks that passed their check. Down is good, and it is the only token metric that survives a change in task mix. > - `cache_read_share`: cached input tokens divided by all input tokens. Up is good. A falling share usually means something volatile moved into the prefix. > - `cost_per_completed_task`: dollars divided by tasks that passed. Down is good, and it is the number to put next to accuracy when you compare two architectures. > **Takeaway** > > Reduce uncertainty before you invoke intelligence. Every token of context your system constructs deterministically is a token the model does not spend rediscovering it probabilistically. --- # Stop making agents read your entire logs > Logs are structured data, and treating them as prompt text costs tokens you do not have to spend. Part 2: Deterministic retrieval. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/stop-making-agents-read-your-logs Open a run where your agent debugged a failing test and you will find the same four steps every time: `cat` the log file, `grep` for "error", `tail -n 200`, paste the result into context. The most expensive component in the stack is now doing the job of a parser, and doing it probabilistically. ### Signal to noise is a system property A production stack trace from a Python web request can run hundreds of frames, and most of them are framework internals: WSGI plumbing, middleware, object-relational mapper (ORM) dispatch. The frames that matter for the bug are usually the handful that touch *your* code. An irrelevant frame is not neutral padding. It costs tokens on the way in, and it competes for attention once it is there. Chroma's "Context Rot" study evaluated 18 frontier models and found that performance degrades as input length grows, even on simple tasks, and that it degrades faster when the context holds distractors that are related to the target without being the target.[7](https://macanderson.com/manual/sources#r7 "Hong, Troynikov & Huber (Chroma). \"Context Rot: How Increasing Input Tokens Impacts LLM Performance.\" July 2025. research.trychroma.com/context-rot") A raw log is a distractor-dense document by construction: hundreds of lines that look like the one line you need. Retrieval research reports the same shape. Cuconasu et al. ("The Power of Noise," SIGIR 2024) found that adding semantically related but non-answer-bearing documents to a model's context degrades accuracy, and that related noise hurts more than unrelated noise, because the model cannot cheaply dismiss it.[9](https://macanderson.com/manual/sources#r9 "Cuconasu et al.. \"The Power of Noise: Redefining Retrieval for RAG Systems.\" SIGIR 2024. arXiv:2401.14887") Framework frames in a stack trace are that kind of high-similarity distractor. ### Execution slicing: preprocess, do not prompt The deterministic alternative treats program output as structured data and slices it before inference: - **Stack trace reduction**: parse the trace, keep frames inside your package roots, collapse framework frames to one-line markers, keep the exception type, the message, and the innermost user frame's local variables. - **Dynamic trace extraction**: run the failing test under a tracer (`sys.settrace`, `coverage.py`, eBPF, OpenTelemetry) and extract the execution path that actually ran, with the values that flowed through it. - **Log windowing**: anchor on the failure timestamp or request id and take a bounded, correlated window instead of the file. ##### Reach for stack trace reduction when The failure raises. You have a traceback, and what you need is the frames in your own packages, the exception type, and the locals at the innermost user frame. ##### Reach for dynamic trace extraction when The test fails without raising, or a legal-looking wrong value arrives somewhere downstream. You need the path that ran and the values that moved along it. ##### Reach for log windowing when The failure happened in another process or another service. You have a request id or a timestamp, and you need the correlated window around it, bounded in lines. ``` # Instead of handing the model a shell and hoping: # grep -rn "error" logs/ | tail -500 ← the model pays for every candidate line result = run_code_with_trace( # deterministic, runs in milliseconds cmd="pytest tests/test_billing.py::test_refund -x", package_roots=["app/"], # slice to code you own capture=["exception", "locals", "executed_lines", "sql"], max_tokens=1200, # hard budget, enforced by code ) context = result.to_prompt_block() # 1.2k tokens of signal ``` The difference is architectural, not cosmetic. `grep | cat | tail` makes the *model* the filter. You pay tokens for every candidate line, and the filtering step itself is probabilistic. `run_code_with_trace()` makes *code* the filter. Filtering costs no tokens, returns the same slice on the same input, and you can test it like any other function. Bash is a fine escape hatch. It is a poor primary retrieval engine, because each round trip is another model call that carries the whole conversation history with it. > **Evidence · The interface is the bottleneck** > > SWE-agent (Yang et al., Princeton) showed that agent performance depends heavily on the **agent-computer interface (ACI)**. Replacing raw shell interaction with purpose-built commands that return compact, structured views of files and search results improved resolution rates on SWE-bench.[14](https://macanderson.com/manual/sources#r14 "Yang, Jimenez et al. (Princeton). \"SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.\" NeurIPS 2024. arXiv:2405.15793") The model did not get smarter. The interface got stricter about what reached the context window. Anthropic's guidance on tool design for agents makes the same point: a tool should return a concise, high-signal representation and enforce token-efficient defaults, because an agent inherits every inefficiency of its tools.[20](https://macanderson.com/manual/sources#r20 "Anthropic Engineering. \"Writing effective tools for agents.\" 2025. anthropic.com/engineering") > > Yang et al., "SWE-agent," NeurIPS 2024 · Anthropic engineering, 2025 > **Do this week** > > 1. **Write a stack trace reducer.** Parse the traceback with your language's own grammar, keep frames whose file path sits under a configured package root, and collapse the rest to one line each. The artifact is a function that takes a traceback string and returns a block under 1,500 tokens. > 2. **Put a token budget on every tool result.** Each tool returns at most N tokens and says how much it dropped and why. Done looks like a truncation counter you can graph, not a silent cut. > 3. **Trace the failing test instead of reading its output.** Run it under `coverage.py` or `sys.settrace`, keep executed lines inside your package roots, and hand the model the path rather than the log. > 4. **Correlate logs by request id.** Emit a request id on every log line, then serve a bounded window around a failure instead of a file. If the id is missing today, adding it is the whole task. > **Measure it** > > - `context_tokens_per_tool_call`: tokens returned by a tool, at p50 and p99. Down is good. The p99 is where an unbounded `cat` hides. > - `frames_kept_ratio`: frames sent divided by frames in the raw trace. Down is good, as long as the innermost user frame survives every time. > - `tool_calls_per_task`: tool calls between the prompt and the first edit. Down is good, because each one carries the full history with it. > - `truncation_rate`: share of tool results that hit the budget. Watch both directions. Zero means the budget is slack, and a high rate means the slicer is not slicing. > **Takeaway** > > Every line your system filters in code is a line the model does not pay to read and cannot be distracted by. Deterministic preprocessing moves work off the model and onto a CPU you already own. --- # Parse the user prompt before the model ever sees it > Your prompt already carries structured information, so extract it before you pay for inference. Part 3: Deterministic retrieval. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/parse-the-prompt-before-the-model-sees-it Someone files a bug against your agent and it reads like this: "The `POST /api/v2/refunds` endpoint throws `DecimalConversionError` in `billing/processors.py` after #4821 merged." That string holds an API route, an exception class, a file path, and an issue reference. Four retrieval keys, sitting in plain sight. Most agent stacks hand the raw string to the model and let it decide what to search for. ### Deterministic extraction is a solved problem Long before you need a language model, thirty lines of parsing gets you: - **File paths**: anything matching path grammar, validated against the actual file tree, with near misses fuzzy-matched. - **Symbols and signatures**: CamelCase and snake\_case identifiers, `module.func()` call syntax, function signatures, validated against your symbol index (part 4). - **Stack traces**: a pasted traceback has a rigid grammar in every language. Parse it fully rather than summarizing it. - **Issue and PR references**: `#4821`, `JIRA-123`, resolved through an API into titles, diffs, and linked commits. - **URLs and API endpoints**: a route pattern maps straight onto a router definition in the codebase. - **Version and config literals**: package names, versions, environment variable names, feature flags. Every extracted entity is a *verified anchor*. It either resolves against your index or it does not. That check is binary, and it is the check a generative approach cannot give you, so an invented file path or a misremembered function name survives into the next call instead of being dropped at the door. ``` # Thirty lines, no model, runs before the first call. PATH = re.compile(r"\b[\w./]+\.(py|ts|go|rs|java)\b") SYMBOL = re.compile(r"\b([A-Z][A-Za-z0-9]+|[a-z_][a-z0-9_]{2,})\(?\)?") ISSUE = re.compile(r"#(\d+)|\b([A-Z]{2,}\-\d+)\b") def anchors(prompt, index): found, dropped = [], 0 for raw in PATH.findall(prompt) + SYMBOL.findall(prompt): hit = index.resolve(raw) # None when the tree has no such path or symbol if hit: found.append(hit) else: dropped += 1 metrics.record("anchor_resolution_rate", len(found), dropped) return found ``` ### From entities to a retrieval plan The output of parsing is not decoration. It is an execution plan that runs *before inference begins*: ``` plan = build_retrieval_plan(prompt) # { # anchors: [File("billing/processors.py"), Symbol("DecimalConversionError"), # Route("POST /api/v2/refunds"), Issue(4821)] # expand: graph_neighbors(anchors, hops=1) ← Part 4 # tests: tests_covering(anchors) ← Part 10 # schema: tables_touched(anchors) ← Part 8 # budget: 8_000 tokens, ranked by relevance # } ``` The model's first call now holds the right file, its direct dependencies, the covering tests, and the issue diff, assembled by code, ranked by graph centrality, and cut to budget. Compare that to the common pattern, where the model spends its first five calls, each carrying a growing history, rediscovering what a regex had at millisecond zero. > **Evidence · Structured localization beats free exploration** > > This is the design that let Agentless outperform autonomous agents. Its first phase is **hierarchical localization**, a fixed funnel from repository to files to classes and functions to edit locations, rather than a model wandering with tools. The paper's ablations show that staged narrowing keeps ground-truth edit locations in the candidate set while shrinking the code the model must read, and it is a large part of why the pipeline reached top open-source results at roughly a tenth of the cost of agent baselines.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") Constrained input, focused attention, fewer irrelevant tokens: the localization step is deterministic scaffolding doing the model's foraging for it. > > Xia et al., "Agentless," FSE 2025 #### ◌ Prompt as opaque string - The model infers what to look for, probabilistically, once per run - Search terms may be invented, and paths go unverified - 3 to 8 exploratory tool calls before the real work starts - Cost scales with how long the model explores #### ● Prompt as structured input - A parser extracts entities the same way on every run - Every anchor is validated against a real index - The retrieval plan executes before the first model call - Cost scales with the complexity of the task > **Do this week** > > 1. **Write the entity extractor.** Start with four patterns: path grammar, identifiers, issue references, and route patterns. The artifact is a function from prompt string to a list of typed entities, with unit tests over ten real prompts from your own queue. > 2. **Validate every entity against an index.** Resolve paths against the file tree and symbols against your symbol index. Drop what does not resolve, fuzzy-match near misses, and count both. > 3. **Emit the retrieval plan as an artifact.** Write the plan to the task record before the first model call, with anchors, expansions, and the token budget. When a run goes wrong, the plan is the first thing you read. > 4. **Seed the first call from the plan.** Assemble the first message from resolved anchors instead of letting the model open with a search, then compare tool-call counts against the previous week. > **Measure it** > > - `anchor_resolution_rate`: entities that resolved against an index divided by entities extracted. Up is good. A drop usually means the index went stale, not that the prompts got worse. > - `first_call_hit_rate`: share of tasks where the file the final diff edited was already in the first call's context. Up is good, and this is the metric that tells you whether localization works. > - `preflight_tool_calls`: tool calls between the prompt and the first edit. Down is good. > - `unresolved_anchor_count`: entities the model was handed that no index could confirm. Down is good, because each one is a chance to act on a path that does not exist. > **Takeaway** > > Treat the prompt as the first document your system parses, not the first thing your model reads. Extraction is cheap, checkable against an index, and identical on every run. Inference is none of those things. --- # Why grep is the wrong retrieval engine > Searching is not understanding, and code already carries the structure you need. Part 4: Deterministic retrieval. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/why-grep-is-the-wrong-retrieval-engine You ask your agent to change a function signature, and it starts grepping. Grep answers one question: which lines contain this string? The questions the change actually turns on are structural. Who calls this function? What implements this interface? What breaks if I change the signature? Where does this value come from? Text search approximates those answers over many noisy round trips. A code graph returns them in one query. ### Code is already a graph, so stop flattening it Compilers have modeled code as structure for decades: ASTs, symbol tables, call graphs, dependency graphs, cross-reference indexes. Your IDE resolves "go to definition" in milliseconds with that machinery, and no model is involved. The default agent loop discards all of it and hands the model a shell, which forces it to rebuild structural knowledge out of string matches: ``` # The grep loop (typical agent run, abbreviated): grep -rn "process_refund" . # 47 matches, 31 in tests and comments cat billing/processors.py # 800 lines into context grep -rn "RefundProcessor" . # which of these are real call sites? cat api/handlers/refunds.py # another 600 lines # ... more calls, and the call graph is still not in context # The graph query: g.callers("process_refund", depth=2) # ranked call sites, 0 tokens g.slice(symbol="process_refund", budget=6000) # minimal closed subgraph ``` One hundred greps approximate what one graph traversal computes. The grep version is not only slower. Each round trip is a model call that re-sends the conversation history, and most of the lines it returns are irrelevant to the change, which part 2 showed is the kind of noise that costs accuracy rather than just tokens. It also gives the model one more chance to take a wrong turn. The graph version is deterministic: the same query against the same index returns the same answer. ### The index stack - **AST layer** (tree-sitter): language-aware parsing, with definitions, references, and signatures for most mainstream languages. - **Symbol index**: every definition and every reference site, linked in both directions, which is what the Language Server Protocol (LSP) and the SCIP Code Intelligence Protocol (SCIP) already standardize. - **Dependency graph**: imports, package boundaries, build targets. - **Call graph**: who invokes whom, with edge weights from reference counts or runtime traces. - **Centrality ranking**: PageRank over the reference graph tells you which symbols matter most for understanding any region of the code. > **Evidence · Graph-ranked context in production tools** > > **Aider's repo map** is the clearest proof this works at tool scale. It parses the whole repository with tree-sitter into definitions and references, builds a graph where files are nodes and symbol references are edges, then runs a **PageRank-style algorithm** to pick the most-referenced identifiers that fit a fixed token budget (around 1k tokens for an entire repository by default). The model receives a compressed structural map, with key signatures and the most-referenced symbols, instead of raw file dumps. Aider's own benchmarking found that this improves code-editing performance on large repositories.[12](https://macanderson.com/manual/sources#r12 "Aider (Paul Gauthier). \"Building a better repository map with tree-sitter\" & repo-map docs. 2023 onward. aider.chat/2023/10/22/repomap.html · aider.chat/docs/repomap.html") > > **AutoCodeRover** (Zhang et al., NUS) made the same bet in research. It replaced plain text retrieval with AST-based search APIs (search-class, search-method, search-code-in-file), letting the model query program structure instead of strings, and reported improved issue-resolution efficacy on SWE-bench against text-based navigation.[13](https://macanderson.com/manual/sources#r13 "Zhang, Ruan, Fan & Roychoudhury (NUS). \"AutoCodeRover: Autonomous Program Improvement.\" ISSTA 2024. arXiv:2404.05427") > > Aider, "Building a better repository map with tree-sitter" · Zhang et al., ISSTA 2024 ### Filesystem traversal against graph traversal The reason graphs win is precision per token. A filesystem is organized for humans and build tools, by directory rather than by meaning. The code relevant to one change is scattered across handlers, models, migrations, and tests, and filesystem traversal makes the agent page through that scatter linearly. Graph traversal follows *edges of actual relationship*, and a two-hop neighborhood around a symbol is close to the minimal sufficient context for editing it. You are not searching a haystack. You are following a wire. ##### Query the graph when The question names a symbol and asks about relationships: callers, implementers, overrides, the blast radius of a signature change, or which tests reach this line. ##### Search text when The target is a string rather than a symbol: a log message, an error copy string, a feature flag name, a config key, a TODO, or a value in a file no parser understands. > **Do this week** > > 1. **Index the repository with tree-sitter.** Parse every file into definitions and references, and write them to a table keyed by symbol. The artifact is an index you can rebuild from a clean checkout in one command. > 2. **Build the reference graph and rank it.** Files or symbols are nodes, references are edges. Run PageRank over it and emit a repository map that fits a fixed token budget, then put that map in the cached prompt prefix. > 3. **Expose three queries as tools.** `callers(symbol, depth)`, `implementers(interface)`, and `slice(symbol, budget)`. Each returns a ranked, budgeted block, not a file. > 4. **Keep text search, and scope it.** Leave a search tool for string literals and config values, and say so in its description, so the model reaches for the graph when the question is structural. > 5. **Reindex on change.** Update the index from the files a commit touched rather than rebuilding it, and record how old the index was at query time. > **Measure it** > > - `index_staleness_seconds`: age of the index at query time, at p99. Down is good. A stale index returns confident answers about code that moved. > - `context_precision`: lines in the assembled context that the final diff touched, divided by lines sent. Up is good, and it is the number that tells you whether the two-hop neighborhood is the right radius. > - `graph_query_share`: graph queries divided by all retrieval calls. Up is good while `context_precision` holds, and a fall usually means a query the graph cannot answer yet. > - `tokens_to_first_edit`: input tokens spent between the prompt and the first diff. Down is good. > **Takeaway** > > Grep ×100 approximates what one graph query computes. Build the graph once, deterministically, and let the model spend its tokens on the change instead of on reconstructing the compiler's knowledge by string matching. --- # Build semantic memory once > A run should start from what the last one learned instead of repeating it. Part 5: Deterministic retrieval. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/build-semantic-memory-once On Monday a run spends about 40,000 tokens learning that your payments code lives in `billing/`, that `LedgerEntry` is the money-movement primitive, and that notifications go through an outbox pattern. On Tuesday a new run spends about 40,000 tokens learning the same three facts. In the traces this book draws on, that rediscovery is most of the startup cost, for knowledge that changes monthly. ### Repositories have natural semantic structure Any codebase that has survived contact with production organizes into domains (Payments, Authentication, Billing, Notifications, Infrastructure) whether or not the directory tree admits it. You can recover those domains automatically. Embed every file and module summary, cluster the embeddings, then cross-check the clusters against the dependency graph from part 4. Community detection over the import graph (Louvain, Leiden) finds the same boundaries from pure structure. #### ◌ Clusters and graph disagree - A module imports across what the text says is a boundary - One directory holds two unrelated vocabularies - Treat it as a debt list, ranked by how many edges cross #### ● Clusters and graph agree - Semantics and structure name the same set of files - That set is a real architectural boundary - Give it a dossier and an owner, and retrieve it as one unit This runs Conway's law backwards: the code's semantic clusters recover the team and domain boundaries that produced it. Once you infer a domain, give it a compact, versioned dossier. Entry points, core types, invariants ("money is integer cents everywhere"), owned tables, and the suites that cover it. Generate it once, refresh it incrementally on merge, and retrieve it at the start of a run by lookup rather than by exploration. > **Evidence · Index once, query cheaply forever** > > Microsoft Research's **GraphRAG** shows the economics of amortized semantic indexing. An LLM pass extracts an entity graph from the corpus once, the Leiden algorithm detects semantic communities, and community summaries are pre-generated at index time. At query time, answering corpus-level questions from those summaries won 70 to 80% of head-to-head comparisons against naive RAG on comprehensiveness and diversity, while using roughly **2 to 3% of the tokens per query** that hierarchical source-text summarization would require, because the expensive understanding was paid for once, up front.[10](https://macanderson.com/manual/sources#r10 "Edge et al. (Microsoft Research). \"From Local to Global: A Graph RAG Approach to Query-Focused Summarization.\" 2024. arXiv:2404.16130"),[11](https://macanderson.com/manual/sources#r11 "Microsoft Research. \"GraphRAG: New tool for complex data discovery.\" July 2024. microsoft.com/research blog") The same trade applies to code. Indexing is a fixed cost. Rediscovery inside every run is a tax. > > Edge et al., arXiv:2404.16130 · Microsoft Research, 2024 ### What a dossier looks like A dossier is the briefing a senior engineer gives a new hire, generated from code and stamped with its commit. ``` domain: billing sha: 9f31c2 entry_points: [billing/api.py, billing/tasks.py] core_types: [LedgerEntry, Invoice, RefundRequest] invariants: - money columns are integer cents (lint: no-float-money) - outbound mail goes through the outbox, not the mailer owned_tables: [ledger_entries, invoices, refunds] covering_tests: [tests/billing/, tests/contract/test_ledger.py] may_call: [payments, notifications] # 280 tokens. Rebuilt when any path under billing/ merges. ``` ### What goes in the memory layer - **Embeddings**: file, symbol, and doc-chunk vectors for fuzzy entry ("where do we throttle webhooks?") when the prompt carries no hard anchors. - **Domain dossiers**: the 300-token expert briefing per domain, as above. - **Architectural boundary map**: which domains may call which, with violations flagged by a check before the model proposes one. - **Convention registry**: error-handling idioms, naming schemes, "we use the repository pattern here", extracted once and injected only when the diff is in scope. This memory is infrastructure with a freshness contract, not a cache. It rebuilds incrementally from the merge queue, it carries the commit SHA it was derived from, and the same CI that validates your code invalidates it. That contract is what separates semantic memory from a stale wiki. > **Do this week** > > 1. **Build the import graph and cut it into communities.** Parse the repository with tree-sitter or your language's own import resolver, load the edges into networkx or python-igraph, and run Leiden. The artifact is `domains.json`, one community id per module. > 2. **Cross-check the communities against embeddings.** Summarize each file, embed the summaries, cluster them, and print the modules where the two methods disagree. That list is your boundary debt, and it is worth reading before you write a single dossier. > 3. **Generate one dossier for your busiest domain.** Fill the fields above from the index, cap it at 300 tokens, and stamp it with the commit SHA. Done looks like a file a new engineer could read in a minute and act on. > 4. **Rebuild dossiers in CI on merge.** Recompute only the domains whose paths changed. Done looks like the dossier SHA matching `HEAD` within one merge. > 5. **Make the start of a run a lookup.** Resolve the touched paths to a domain, inject that dossier, and remove the instruction that told the agent to go exploring. > **Measure it** > > - `run_startup_tokens`: tokens consumed before the first file edit. Sum the input tokens of every request that precedes the first write. Lower is better, and the drop after step 5 is the whole return on this chapter. > - `dossier_staleness_commits`: merges between the dossier's stamped SHA and `HEAD` for that domain. Lower is better, with 0 to 1 as the working target. > - `domain_agreement_rate`: share of modules where the embedding cluster and the graph community name the same domain. Higher is better, and the remainder is a ranked debt list. > - `dossier_hit_rate`: share of runs that retrieved a dossier for the domain they actually edited. Higher is better, and a low number usually means your path-to-domain resolver is wrong, not that the dossiers are. > **Takeaway** > > Understanding your repository is a build artifact, not a conversation. Compute it once, version it, refresh it incrementally, and let every run start from knowledge instead of archaeology. --- # A rule written in prose cannot fail CI > Constraints belong in a representation a validator can check. Part 6: Context engineering. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/rules-belong-in-data-not-in-prose It starts with one file that tells the agent to stop repeating a mistake. Then a second for architecture, a third for conventions, a fourth for the things the first three left out. Six months later every run loads 30,000 tokens of prose, nobody can say which rules are still true, and the agent has started ignoring half of them. What you have is a second codebase written in a language with no compiler, no tests, and no dead-code detection. ### Prompt bloat is measurable harm, not just cost The instinct says instructions are free, and that the worst case is a model skimming the irrelevant ones. The research points the other way. Performance degrades as input grows even when the added material is task-relevant in spirit. Liu et al. showed models lose information positioned in the middle of long contexts,[6](https://macanderson.com/manual/sources#r6 "Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni & Liang. \"Lost in the Middle: How Language Models Use Long Contexts.\" TACL 2024. arXiv:2307.03172") and Chroma's 18-model study showed degradation with input length across every frontier model tested, at different rates.[7](https://macanderson.com/manual/sources#r7 "Hong, Troynikov & Huber (Chroma). \"Context Rot: How Increasing Input Tokens Impacts LLM Performance.\" July 2025. research.trychroma.com/context-rot") Anthropic's context-engineering guidance is direct about the mechanism: attention is a finite budget, every token in context draws it down, and the goal is the *smallest set of high-signal tokens* that achieves the outcome.[19](https://macanderson.com/manual/sources#r19 "Anthropic Engineering. \"Effective context engineering for AI agents.\" 2025. anthropic.com/engineering") Thirty thousand tokens of standing instructions are not a safety net. They are attention spent before the task begins, and billed per request. ### The caching economics make it worse Prompt caching looks like it rescues the large standing prompt. Cache reads cost about 10% of base input price on Anthropic's API, with a 25% premium on writes, cutting cost by up to 90% and latency by up to 85% for long stable prefixes.[16](https://macanderson.com/manual/sources#r16 "Anthropic. \"Prompt caching with Claude\" (up to 90% cost and 85% latency reduction, and a cache read costs about 0.1 times the input price). 2024. anthropic.com/news/prompt-caching") The catch is that caching is prefix-exact. Edit one character of an early block and everything after it recomputes at full price, plus the write premium. A pile of frequently edited prose concatenated into the prompt is the worst shape for this, because each edit invalidates the cache across your whole fleet. Caching rewards small, stable, ordered context. That is an argument for structure, not for volume. #### ◌ Rules as prose - Loaded whole, in every run, relevant or not - No validation, so rules rot silently as the code moves - Conflicting rules resolved probabilistically by the model - Every edit invalidates the prompt cache prefix - Grows monotonically, and nobody dares delete a line #### ● Rules as data - Retrieved selectively, only the rules matching the touched domain - Machine-checkable, because rules reference real symbols and run in CI - Conflicts detected at build time - Stable structured core caches cleanly, variable tail stays small - Dead rules detected the way dead code is ### What replaces the prose - **JSON Schema or typed config** for anything that is a constraint: allowed dependencies, naming patterns, review requirements. A constraint in schema form is enforced by a validator rather than suggested to a model. - **Tool metadata**: capabilities, costs, and preconditions declared on the tools themselves, the direction MCP standardizes, so capability discovery is a lookup instead of a paragraph. - **Rule objects with scopes**: `{rule, applies_to: "billing/**", verified_by: "lint:no-float-money", since: "a1b2c3"}`. The retrieval layer injects the three rules that bear on this diff, not the three hundred that exist. - **Generated docs**: the part 5 dossiers, derived from code, carrying their source SHA. ### One rule, both ways ``` # before: a line in conventions.md, loaded in full in every run # "Money is cents. Do not use floats for money." # after: a rule object, injected when the diff touches billing/ {"id": "no-float-money", "applies_to": "billing/**", "verified_by": "lint:no-float-money", "since": "a1b2c3", "text": "Money columns and money variables are integer cents."} # 40 tokens in the prompt, and CI fails on a float literal under billing/. ``` The second form costs less, arrives only when it applies, and has a witness. If the lint rule stops firing because the code moved, you learn that the rule is dead instead of carrying it for another year. ### Order the prompt so the cache survives your edits ``` prompt = [ schema_snapshot, # stable for days domain_dossier, # stable until that domain merges matched_rules, # 3 of 300, changes per task task_slice, # changes per task ] # Edits land in the tail. The cached prefix survives them. ``` Keep prose for what prose is good at: intent, history, and taste. Anything that functions as a rule belongs in a representation you can validate, scope, version, and retrieve selectively. If a rule cannot fail CI, it is a preference. > **Do this week** > > 1. **Count what you are loading.** Run every standing prose file through a tokenizer (tiktoken, or your provider's token-count endpoint) and print a table sorted by tokens. The number at the bottom is what each run pays before it reads any code. > 2. **Sort each line into constraint, context, or history.** Constraints move to a rule file. Context moves to a generated dossier. History stays in prose and leaves the prompt. > 3. **Give your three most repeated constraints a check.** An ESLint or Ruff rule, a JSON Schema validator, or a dependency-cruiser rule. Done looks like the check failing on a violation you introduce on purpose, then passing when you revert it. > 4. **Inject rules by glob.** Match `applies_to` against the paths in the diff and pass only the matches. Done looks like a prompt carrying 3 rules on a billing change rather than the full registry. > 5. **Reorder the prompt.** Stable structured blocks first, task slice last, so an edit lands in the tail and the cached prefix holds. > **Measure it** > > - `prompt_tokens_static`: tokens loaded before the task-specific slice. Read it from the request log. Lower is better, and it should fall in steps as files move out. > - `rules_injected_per_task`: rules in the prompt divided by rules in the registry. Lower is better, and a ratio near 1 means your scoping is not working. > - `rule_enforcement_rate`: share of rules whose `verified_by` check runs in CI. Higher is better. Rules without a check are the ones that rot. > - `cache_read_ratio`: cached input tokens divided by total input tokens, taken from the provider's usage fields. Higher is better, and a sudden drop means somebody edited an early block. > **Takeaway** > > Natural-language configuration is configuration without a type system. Move constraints into structured, validated, selectively retrieved data, and let prose go back to being prose. --- # Context compression is worth more than a bigger model > Removing irrelevant information beats buying a larger model. Part 7: Context engineering. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/compression-beats-a-bigger-model An agent underperforms on your repository, and the available levers look like this: a bigger model, a longer window, a larger reasoning budget. All three cost money and none of them change what you are sending. The question that predicts success is not how much the model can hold. It is what fraction of what it holds bears on the task, and that is a question you can answer with instrumentation you already own. ### The Shannon view of a context window Shannon's framework[18](https://macanderson.com/manual/sources#r18 "Shannon, C.E. \"A Mathematical Theory of Communication.\" Bell System Technical Journal, 1948. The original framework for signal and noise") gives the right vocabulary. The task has an intrinsic information requirement, the minimal set of facts needed to produce the correct patch. Everything else in the window is noise relative to that task, and attention is the finite channel the signal has to pass through. Two consequences follow. First, mutual information matters more than volume: a 4k-token context where every line bears on the task carries more usable information than 100k tokens where the answer is diluted 25 to 1. Second, compression is safe when it is lossless with respect to the task, which is what deterministic slicing (parts 2 to 4) gives you and what generic truncation does not. A stack-trace parser that keeps your frames and drops framework frames is a task-aware compressor with a retention property you can state. `head -c` has no such property. ### Four results point the same way U-curve Liu et al.: accuracy collapses for facts mid-context, long-context models included[6](https://macanderson.com/manual/sources#r6 "Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni & Liang. \"Lost in the Middle: How Language Models Use Long Contexts.\" TACL 2024. arXiv:2307.03172") 18 / 18 Chroma: every frontier model tested degrades as input grows, even on trivial tasks[7](https://macanderson.com/manual/sources#r7 "Hong, Troynikov & Huber (Chroma). \"Context Rot: How Increasing Input Tokens Impacts LLM Performance.\" July 2025. research.trychroma.com/context-rot") −dozens of pts NoLiMa: when matches are semantic rather than literal, most models fall far below short-context accuracy by 32k[8](https://macanderson.com/manual/sources#r8 "Modarressi et al. (LMU Munich & Adobe). \"NoLiMa: Long-Context Evaluation Beyond Literal Matching.\" 2025. arXiv:2502.05167") Related > random Power of Noise: semantically related distractors hurt accuracy most[9](https://macanderson.com/manual/sources#r9 "Cuconasu et al.. \"The Power of Noise: Redefining Retrieval for RAG Systems.\" SIGIR 2024. arXiv:2401.14887") Read the four together for coding agents specifically. A repository dump is long, which triggers length degradation. It is full of near-duplicate distractors, which is the most harmful noise type the fourth result identifies. Using it requires semantic rather than lexical matching, the NoLiMa condition where the drop is steepest. A raw-context design sits at the intersection of all three. > **Counterweight · long-context is improving** > > Be honest about the trend. Newer models post high scores on simple needle-in-a-haystack retrieval at extreme lengths, and vendors are attacking positional bias directly. The gains are strongest on literal retrieval, the easiest case, while semantic reasoning over long, distractor-dense context remains the weak point, which is the regime coding agents live in. And a model that holds its accuracy at length still bills you for every irrelevant token. Compression wins on cost even where it stops winning on accuracy. ### Compression as an engineering discipline - **Retrieval precision over recall by default**: start from verified anchors (part 3) and expand along graph edges (part 4) instead of dumping candidates. - **Representation changes**: signatures instead of bodies (a repo map is roughly a 100 to 1 compressor[12](https://macanderson.com/manual/sources#r12 "Aider (Paul Gauthier). \"Building a better repository map with tree-sitter\" & repo-map docs. 2023 onward. aider.chat/2023/10/22/repomap.html · aider.chat/docs/repomap.html")), schemas instead of migrations, sliced traces instead of logs. - **Budgets enforced in code**: every context assembler takes a token budget and ranks content to fit it, and going over budget fails the build. - **Context utilization as a metric**: what fraction of the provided tokens does the final patch depend on? In the traces this book draws on it often sits below 10%, and it is the clearest single indicator of wasted spend (part 13, and part 17 for attributing that spend to the team that caused it). ### A budget the assembler cannot talk its way out of ``` def assemble(task, budget=8000): blocks = rank(candidates(task)) # anchors first, then graph neighbours kept, used = [], 0 for b in blocks: if used + b.tokens > budget: break kept.append(b) used += b.tokens missing = required(task) - {b.id for b in kept} if missing: # a required anchor did not fit raise BudgetError(task, missing, used, budget) return kept # Over budget is a failure with a name, not a truncation nobody sees. ``` The failure matters more than the cap. When a required anchor does not fit, you want a stack trace naming the task and the anchor, because that is a retrieval bug. Silent truncation turns the same bug into a wrong patch and a confusing review. > **Do this week** > > 1. **Record what you sent and what got used.** For each run, log the ids and token counts of every block you provided, then log the files the accepted patch touched. One JSON line per run is enough to compute utilization. > 2. **Replace one file dump with signatures.** Parse with tree-sitter and emit declarations rather than bodies for anything outside the edit site. The artifact is a repo map you can diff against the old prompt, token for token. > 3. **Give the assembler a budget.** Add the cap and the named failure above. Done looks like a test that asks for an over-budget assembly and asserts `BudgetError`. > 4. **Drop near-duplicates before they ship.** Run minhash or simhash over retrieved chunks and collapse anything above your similarity threshold. Related distractors are the ones the fourth result says cost you most. > 5. **Re-run last week's failures at the new budget.** If accuracy holds with 40% fewer tokens, you have found the model upgrade you were about to buy. > **Measure it** > > - `context_utilization`: tokens of blocks the accepted patch depended on, divided by tokens provided. Higher is better. Below 10% means most of your bill is noise. > - `tokens_per_accepted_patch`: total input tokens across the run, divided by patches a human accepted. Lower is better, and it is the number to quote when somebody proposes a larger window. > - `distractor_rate`: share of retrieved chunks the final patch does not reference. Lower is better, and a rising value usually means recall crept back into your retriever. > - `budget_overflow_rate`: assemblies that raised `BudgetError`, divided by assemblies. Lower is better, and each one is a retrieval bug with an address. > **Takeaway** > > A bigger model reads your noise at a higher price. A better system removes the noise. Only one of the two compounds: every point of retrieval precision you engineer carries over to every model you run after this one. --- # Agents guess at schema they were not shown > A model adds a duplicate column when the context did not carry the existing one. Part 8: Context engineering. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/show-the-agent-the-schema A bug report that teams running coding agents against a real product tend to meet: The agent added a `user_email` column to a table that already had `email`. Or wrote a migration duplicating an index. Or built a query against a column renamed two quarters ago. The postmortem instinct is to call it a hallucination. The mechanism is duller: the model was not shown the schema, so it did what an engineer with no view of the database would do, which is guess. ### Why schema is the worst-served context type Schema knowledge resists text retrieval. The live truth is not in any one file. It is the fold of an ordered migration history: 400 migration files where the current shape of `orders` is the sum of 23 of them, interleaved with everything else. Grep for the table name and you get all 23 plus every query that touches it, which are closely related distractors, the kind of noise part 7 showed does the most damage. Object-relational mapper (ORM) models drift from the database. Column semantics (`amount` is integer cents, not a float) live in tribal memory. Telling the agent to read the migrations folder does not fix this, because what it needs is a computed state rather than a document. ### Schema becomes another graph node The deterministic fix treats schema the way part 4 treats code, as structure to be indexed rather than text to be searched. - **Schema index**: introspect the live database (or replay migrations) into a canonical, versioned snapshot of tables, columns, types, constraints, indexes, and foreign keys. Most ORMs ship the machinery. - **Migration graph**: each column carries its lineage, created in `0042`, renamed in `0187`, backfilled in `0201`. "Why is this nullable?" becomes a graph query with commit-linked answers. - **Code and schema edges**: static analysis of ORM models and query sites links `LedgerEntry` to `ledger_entries` to the 14 call sites that write it. "What breaks if I drop this column?" becomes a traversal rather than a guess. - **Domain mapping**: tables attach to the part 5 domains, so a billing task retrieves billing tables and only billing tables. ``` ctx = schema.slice_for(anchors=["RefundProcessor"]) # orders(id, user_id→users.id, amount_cents INT NOT NULL, status ENUM,...) # refunds(id, order_id→orders.id, amount_cents INT, reason TEXT,...) # lineage: refunds.amount_cents renamed from amount in 0187 (PR #3122) # invariant: money columns are integer cents (lint: no-float-money) # ~600 tokens. Verified against production at SHA 9f31c2. ``` A slice of a few hundred verified tokens replaces the migration files the agent would otherwise read. Because the slice comes from the live database, the duplicate-column failure also gets a second line of defence. The deterministic layer diffs the proposed migration against the index and rejects collisions before a human reviews them. That is the pattern of this book at small scale. Do not ask the model to remember reality. Show it reality, then check its output against reality. ### The check that catches the duplicate column ``` $ schema-check migrations/0203_add_user_email.py # reject: users.user_email collides with users.email (created 0042) # lineage: users.email created 0042, indexed 0091, no rename since # hint: 3 call sites already read users.email # 1 collision, 0 warnings. exit 1 ``` Run it in the same CI job that lints migrations. The failure names the column, the migration that created it, and the call sites that depend on it, which is enough for the agent to correct itself on the next turn. ### Introspect the database, or replay the migrations - **Introspect a restored copy** when production is the source of truth and hand-run DDL exists. You get the real shape, including the index somebody added by hand. The cost is a restore in CI. - **Replay migrations into a scratch database** when the migration history is authoritative and complete. It is cheaper and reproducible from the repository alone, but drift between the files and production stays invisible. Pick one and stamp the snapshot with the SHA it came from. An unstamped snapshot becomes the stale wiki of part 5, with worse consequences. > **Evidence · Grounding beats recall** > > The general principle holds even where the schema-specific study does not exist yet. Retrieval grounding reduces fabrication compared with parametric recall, the precision of what is retrieved dominates its volume (the distractor results in parts 2 and 7[7](https://macanderson.com/manual/sources#r7 "Hong, Troynikov & Huber (Chroma). \"Context Rot: How Increasing Input Tokens Impacts LLM Performance.\" July 2025. research.trychroma.com/context-rot"),[9](https://macanderson.com/manual/sources#r9 "Cuconasu et al.. \"The Power of Noise: Redefining Retrieval for RAG Systems.\" SIGIR 2024. arXiv:2401.14887")), and structured representations of a system outperform prose descriptions of it for machine consumption (parts 4 to 6[12](https://macanderson.com/manual/sources#r12 "Aider (Paul Gauthier). \"Building a better repository map with tree-sitter\" & repo-map docs. 2023 onward. aider.chat/2023/10/22/repomap.html · aider.chat/docs/repomap.html"),[13](https://macanderson.com/manual/sources#r13 "Zhang, Ruan, Fan & Roychoudhury (NUS). \"AutoCodeRover: Autonomous Program Improvement.\" ISSTA 2024. arXiv:2404.05427")). Schema is the highest-stakes instance. It is the context type where a fabricated column becomes a migration that someone has to reverse. > > Synthesis. See refs 7, 9, 12, 13 > **Do this week** > > 1. **Produce one snapshot.** Run your ORM's introspection (SQLAlchemy reflection, Django `inspectdb`, Prisma or Drizzle introspect) against a restored copy or a replayed scratch database. The artifact is `schema.json` with tables, columns, types, constraints, indexes, foreign keys, and the SHA it was built from. > 2. **Slice it by anchor.** Given a symbol or a path, return the tables that symbol touches plus their direct foreign keys. Done looks like the 600-token slice above for your busiest table. > 3. **Add the collision check.** Diff any proposed migration against the snapshot and fail on a column whose name, or an obvious variant of it, already exists on that table. Done looks like the check rejecting a migration you write on purpose. > 4. **Link code to tables.** A static pass over ORM model definitions and raw query sites gives you table to symbol to call site. That index answers the drop-column question without anyone reading 400 files. > 5. **Refresh the snapshot where migrations run.** Same job, same commit, so the stamp stays honest. > **Measure it** > > - `schema_slice_tokens`: tokens of schema context per task. Lower is better, with a few hundred as the working target. Compare it against the token count of the migration files the agent used to read. > - `schema_snapshot_age`: migrations applied since the snapshot's stamped SHA. Lower is better, and 0 is reachable once step 5 lands. > - `collision_rejections`: proposed migrations the check rejects per week. Watch the direction. A count that falls while agent volume holds steady means the slice is doing its job upstream. > - `unmapped_query_sites`: query sites the static pass could not link to a table. Lower is better, and each one is a place where the drop-column answer is still a guess. > **Takeaway** > > Your database's current shape is computable, versionable, and sliceable. If the agent guesses about schema, that is a missing index in your system rather than a failure of the model. --- # Why coordinator agents don't scale > Most multi-agent systems pay a coordination tax and call it architecture. Part 9: Orchestration and memory. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/why-coordinator-agents-do-not-scale You split the work. A manager agent decomposes the ticket, three specialists execute, and the manager stitches the results back together. Then the token bill climbs, the wall clock stretches, and when the patch comes back wrong you read four transcripts to find which hand-off dropped the constraint. The org chart is not an architecture. ### The tax, itemized - **Token duplication.** Each sub-agent needs enough context to act, so you re-send and re-bill the shared background once per agent, then re-serialize every result back through the coordinator's window. - **Latency amplification.** Coordinator → worker → coordinator turns serialize model calls. A hand-off costs seconds, not microseconds. - **Lossy hand-offs.** Each hand-off compresses intent into natural language. Errors compound across hops. - **Conflicts found at merge time.** Parallel agents cannot see each other's decisions. You pay for both branches, then pay a model call to adjudicate the collision. - **Fan-out with no budget.** A misbehaving branch multiplies cost recursively, and an instruction written in prose does not stop it. > **Evidence · Three independent sources, one conclusion** > > **Anthropic** runs multi-agent systems in production, so it is not a hostile witness. Its report: multi-agent systems consume ~**15× the tokens** of chat, are economical only where task value justifies it, and are explicitly "not a good fit" for domains with tight dependencies between agents. Their example of a poor fit is coding.[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") > > **Cognition** (Devin) published "Don't Build Multi-Agents." In their production experience, parallel sub-agents with fragmented context make conflicting assumptions, and reliability comes from a single agent with continuous, fully-shared context. Context engineering, not agent proliferation.[4](https://macanderson.com/manual/sources#r4 "Cognition (Walden Yan). \"Don't Build Multi-Agents.\" June 2025. cognition.ai/blog/dont-build-multi-agents") > > **Berkeley** (Cemri et al., NeurIPS 2025) annotated 1,600+ traces from 7 popular multi-agent frameworks. Performance gains over single agents were often minimal. The team catalogued **14 recurring failure modes** in 3 classes: specification failures, inter-agent misalignment, and verification failures. Most stem from *system design*, not model capability, and better models will not fix them.[5](https://macanderson.com/manual/sources#r5 "Cemri et al. (UC Berkeley). \"Why Do Multi-Agent LLM Systems Fail?\" (MAST). NeurIPS 2025 Datasets & Benchmarks. arXiv:2503.13657") > > Anthropic engineering, 2025 · Cognition, 2025 · Cemri et al., MAST (the Berkeley failure taxonomy), NeurIPS 2025 ### Deterministic orchestration, the boring alternative Look at what the coordinator agent does all day: routing, sequencing, retrying, aggregating. None of those four needs judgment. All four have mature, deterministic machinery. #### ◌ Coordinator agent - Control flow decided per run by inference, so it differs every time - State lives in a context window, and a crash loses it - Retries, timeouts, and budgets improvised in prose - Debugging means reading several transcripts and reconstructing the order #### ● State machine / workflow engine - Control flow is code: explicit states, typed transitions, reviewable in a PR - State is durable and replayable (Temporal-class engines, LangGraph-style graphs) - Retries, timeouts, budgets, and circuit breakers are first-class primitives - Debugging means inspecting a state machine's history The pattern that works is *deterministic skeleton, probabilistic muscles*. A workflow engine owns the graph, localize → retrieve → patch → test → review, and calls the model only inside the nodes where judgment is required. Event-driven orchestration handles fan-out where sub-tasks are independent, the one regime where even Anthropic's data says parallelism pays.[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") Where sub-agents do earn their keep, the framing to keep is Cognition's: the reliable ones are mostly read-only context gatherers, closer to expensive tool calls than to collaborating colleagues.[4](https://macanderson.com/manual/sources#r4 "Cognition (Walden Yan). \"Don't Build Multi-Agents.\" June 2025. cognition.ai/blog/dont-build-multi-agents") Written down, the skeleton fits on a page. Every line below is code you can read in a diff, and the budget is a number the engine enforces rather than a sentence the model may ignore. ``` workflow fix_ticket: budget tokens=120_000 wall_clock=8m attempts=3 localize : deterministic # parser and code graph, no model call retrieve : deterministic # graph traversal, sliced to budget patch : model # judgment lives here test : deterministic # test runner, structured failure review : model # judgment lives here on test.fail → retry patch with the sliced failure, up to attempts on budget.out → halt, persist state, page a person ``` That leaves one question worth asking per node: does this step need a model at all? | Use a workflow node when | Use a sub-agent call when | | --- | --- | | The step has a decidable rule (route by file type, retry on exit code 1) | The step needs judgment over ambiguous text | | The step's output is typed and checkable | The output is prose a person or a test will grade | | Later steps depend on this one's exact result | The sub-task is independent and read-only | | You want the same answer twice | You can afford variance and will verify the result | > **Do this week** > > 1. **Draw the graph your coordinator improvises.** Take ten recent multi-agent runs, write down the node sequence each one actually followed, and count how many distinct sequences you get. More than two or three for the same task class tells you the routing is inference, not design. > 2. **Move routing into code.** Replace the manager's decomposition prompt with a function that maps task type to a fixed node list. Done looks like a run whose node sequence you can predict before you start it. > 3. **Port one task class to a workflow engine.** Pick a durable engine (Temporal-class) or a graph library (LangGraph-style) and encode localize, retrieve, patch, test, and review as typed states. Done looks like a replayable run history you can open after a crash. > 4. **Put the budget in the engine.** Set a token ceiling, a wall-clock ceiling, and an attempt ceiling per run, and make the engine halt on breach. Done looks like a halted run with persisted state, not a surprise invoice. > 5. **Demote the survivors.** Any sub-agent still standing should be read-only. If it writes files or decides control flow, it is a node in disguise. > **Measure it** > > - `coordination_token_share`: tokens spent on hand-off messages and re-sent background, divided by total tokens for the run. Tag each model call with its node, then sum the routing and aggregation nodes. Down is good. > - `handoff_count`: model calls whose input is mostly another model's output. Count them per completed task. Down is good, and a run with zero is a run with no telephone game in it. > - `path_entropy`: the number of distinct node sequences observed across runs of the same task class. Log the sequence on every run. Down is good, and 1 means the control flow is code. > - `budget_halt_rate`: runs the engine stopped on a ceiling, divided by all runs. You want this small and nonzero. Zero usually means no ceiling is set. > **Takeaway** > > Coordination is a solved, deterministic problem: state machines, workflow engines, queues. Spend inference on judgment inside the nodes, and keep it off the edges between them. --- # Tests should be first-class retrieval objects > Tests describe behavior in executable form, and a test runner's verdict is computed rather than inferred. Part 10: Orchestration and memory. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/tests-as-retrieval-objects Your agent writes a patch, then runs the suite to find out whether it broke something. The tests that would have told it what the function was supposed to do were sitting in the repository the whole time, and it read none of them before it started. Most stacks treat tests as a thing you run after generating a patch, not a thing you retrieve before. That is backwards. The test suite is the best-maintained behavioral specification your repository has. ### Behavioral indexing Make tests queryable along the axes that matter: - **Coverage mapping.** Run the suite under coverage once per merge and invert the result: for every function, the exact tests that execute it. `tests_covering("process_refund")` becomes a deterministic lookup. - **Contract extraction.** Parse assertions into behavior statements. "A refund of a settled order returns `RefundResult(status=PENDING)`. A refund exceeding the order total raises `InvalidRefund`." Signatures told the model the shape (Part 4). Assertions tell it the meaning. - **Fixture graph.** Fixtures encode how valid domain objects are constructed. Link them into the code graph and the model gets canonical object construction for free. Now the retrieval plan from Part 3 gains a behavioral layer. Touching `process_refund` pulls its three covering tests, a precise executable account of current behavior, usually far smaller than the docs that would attempt to describe it. The inversion is a few lines in CI, and the artifact it produces is a map from symbol to tests. ``` # once per merge, in CI pytest --cov=app --cov-context=test --cov-report=json # invert contexts: {symbol: [tests that executed it]} index = defaultdict(list) for fn, arcs in coverage["files"].items(): for line, contexts in arcs["contexts"].items(): for test in contexts: index[symbol_at(fn, line)].append(test) tests_covering("process_refund") → ["test_refund_settled_order", "test_refund_exceeds_total", "test_refund_is_idempotent"] ``` > **Evidence · Tests as selection machinery** > > Agentless does more than run tests at the end. Tests are load-bearing in its pipeline. It generates *reproduction tests* from the issue, uses existing *regression tests* to filter candidate patches, and re-ranks the survivors, a deterministic selection stage that the paper's ablations credit for a substantial share of final accuracy.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") SWE-bench makes the broader point: the industry's shared yardstick for agent capability defines "solved" as nothing more or less than "the fail-to-pass tests now pass."[15](https://macanderson.com/manual/sources#r15 "Jimenez et al.. \"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?\" ICLR 2024. arXiv:2310.06770") If tests are how you judge agents, they should be how agents see. > > Xia et al., FSE 2025 · Jimenez et al., ICLR 2024 ### What an assertion says that a docstring does not Contract extraction is worth a worked example, because the output is smaller than the input and carries more. #### ◌ The docstring - "Processes a refund for an order." - Written once, at creation - Silent about the settled case, the over-total case, and idempotency - No signal when it drifts from the code #### ● The extracted contract - settled order → `RefundResult(status=PENDING)` - amount > order.total → raises `InvalidRefund` - same idempotency key twice → one ledger entry - CI reruns it, so a drift turns the build red ### Execution-driven, contract-first loops Tests also give the workflow skeleton (Part 9) its cheapest validators. A contract-first loop runs like this: retrieve the covering tests, have the model write or extend the failing test first so the intent becomes executable, have it patch, run the deterministic runner, then feed the structured failure (sliced per Part 2) into the next attempt. A compiler and a test runner grade each iteration. Their verdicts are computed from the code that ran, not inferred from a diff by another model emitting an opinion. The loop converges on green, and green is a fact about a process exit code. > **Do this week** > > 1. **Turn on per-test coverage contexts.** Run the suite once with coverage.py recording a context per test (or your language's equivalent) and write the JSON to a build artifact. Done looks like one file that records which test touched which line. > 2. **Invert it into a symbol index.** Map each covered line to its enclosing symbol with tree-sitter and store `{symbol: [tests]}` next to your code graph. Done looks like `tests_covering("process_refund")` answering in milliseconds. > 3. **Put covering tests in the retrieval bundle.** When a node targets a symbol, attach its covering tests ahead of any prose documentation, under the same token budget. Done looks like a bundle you can diff against the old one and see docs give way to tests. > 4. **Extract assertion contracts for your ten hottest symbols.** Parse the assert statements into one-line behavior statements and store them as graph properties. Done looks like ten symbols whose contracts a reviewer agrees with. > 5. **Make the loop write the test first.** Change the patch node to emit a failing test before it emits a patch, and reject attempts that skip it. Done looks like every merged agent patch carrying a test that failed before it. > **Measure it** > > - `symbol_test_coverage_index_freshness`: hours since the coverage inversion last ran, compared against commits merged since. Down is good. A stale index points an agent at tests that no longer exist. > - `covering_tests_in_bundle`: share of retrieval bundles for an edit task that include at least one covering test for the target symbol. Compute it from the bundle logs. Up is good. > - `first_attempt_green_rate`: patches that pass the suite on attempt 1, divided by all patches. Up is good, and it moves when retrieval improves rather than when the model changes. > - `test_first_compliance`: agent patches that shipped with a test that failed before the patch, divided by all agent patches. Up is good. > **Takeaway** > > Index your tests the way you index your code. They are the documentation in your repository that executes, and the ones running in CI are current as of the last green run. That makes them the highest-signal retrieval objects you own. --- # Agents need working memory, not bigger context windows > You do not reread every book you own before fixing a bug, and your agent should not either. Part 11: Orchestration and memory. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/working-memory-not-bigger-windows At turn 60, your agent's transcript still carries the file it read at turn 3, the dead-end approach it abandoned at turn 12, and the stack trace it resolved at turn 20. You are paying for all of it on every subsequent call, and it is competing for attention with the two hundred lines that matter now. The window is being asked to be four things at once: attention, notebook, library, and diary. Human cognition separates those, with a small working set over vast indexed stores and forgetting as a feature. The append-only transcript is the opposite design. ### A memory hierarchy for agents - **Working memory** (the window). Only what the current step needs: the task frame, the active files, the last failure. Small by policy, not by accident. - **Task memory.** Durable structured state for this job: the plan, the decisions made ("chose option B because of migration risk"), the files touched, the budget consumed. It lives in the workflow engine's state (Part 9), survives a crash, and pages in selectively. - **Episodic memory.** Records of past runs. "We fixed a similar DecimalConversionError in March, and the fix was in the serializer config." Indexed by embeddings and entities, retrieved when relevant, absent otherwise. - **Semantic memory.** The durable knowledge layer: the code graph (Part 4), domain dossiers (Part 5), the schema index (Part 8), test contracts (Part 10). > **Evidence · The OS analogy is a working architecture** > > **MemGPT** (Packer et al., Berkeley) made this concrete. Treat the context window as RAM and give the system OS-style virtual memory, paging information between the window and external storage with explicit memory-management operations. The approach let bounded-context models sustain coherent long-horizon behavior, including document analysis and multi-session conversation beyond the raw window, by managing what is resident rather than growing it.[17](https://macanderson.com/manual/sources#r17 "Packer et al. (UC Berkeley). \"MemGPT: Towards LLMs as Operating Systems.\" 2023. arXiv:2310.08560") The same conclusion arrives from the failure side. Chroma's results imply that even within a large window, keeping stale content resident harms performance,[7](https://macanderson.com/manual/sources#r7 "Hong, Troynikov & Huber (Chroma). \"Context Rot: How Increasing Input Tokens Impacts LLM Performance.\" July 2025. research.trychroma.com/context-rot") and Anthropic's context-engineering guidance formalizes the remedies, naming compaction, structured note-taking outside the window, and just-in-time retrieval as the standard toolkit for long-horizon agents.[19](https://macanderson.com/manual/sources#r19 "Anthropic Engineering. \"Effective context engineering for AI agents.\" 2025. anthropic.com/engineering") > > Packer et al., arXiv:2310.08560 · Chroma, 2025 · Anthropic engineering, 2025 ### Context aging, eviction, retrieval scheduling Once memory is tiered, management becomes ordinary systems engineering with ordinary policies. **Aging:** a file read 30 turns ago decays to its signature, and a resolved sub-task collapses to a one-line decision record. **Eviction:** superseded attempts and dead-end explorations leave working memory, recoverable from task memory if needed, no longer taxing attention. **Retrieval scheduling:** each workflow node declares its context contract, so the patch node gets code and contracts, the migration node gets schema and lineage, and the review node gets the diff and the invariants. Nothing rides along just in case, because just in case is the distractor mass the research says hurts most. A context contract is a few lines of configuration per node. It is worth writing down because it turns an argument about what the model needs into a diff a reviewer can read. ``` nodes: patch: include: [task_frame, target_symbols, covering_tests, last_failure] max_tokens: 24000 migrate: include: [task_frame, schema_slice, column_lineage, prior_migrations] max_tokens: 16000 review: include: [task_frame, diff, invariants, covering_tests] max_tokens: 12000 on_overflow: drop lowest rank first, then fail the node ``` Aging needs a policy table rather than a paragraph, because each rule has to hold for a run you did not watch. | Item in working memory | Policy | Where it goes | | --- | --- | --- | | File read, untouched for 20 turns | Decay to signature and docstring | Full text stays in semantic memory | | Resolved sub-task | Collapse to a one-line decision record | Full transcript stays in task memory | | Abandoned approach | Evict, keep the reason it failed | Task memory, one line | | Superseded stack trace | Evict on the next green run | Run log, outside the window | | Active target symbol | Pin until the node completes | Stays resident | Every one of those policies is deterministic. What pages in, what decays, and what leaves are decidable by code against declared contracts. The window stops being a place things accumulate and becomes the small, hot set of what this step needs. A structured memory of entities, relations, and decisions is also most of the way to a knowledge graph, which is where this series goes next. > **Do this week** > > 1. **Measure what is resident.** Instrument one long run to log, per turn, the token count by category: task frame, active files, stale files, resolved failures. Done looks like a chart showing how much of turn 60 is turn 3. > 2. **Write a context contract for one node.** Pick your patch node, declare its include list and a token ceiling, and reject anything outside the list. Done looks like a node whose input you can predict from its config. > 3. **Add an aging rule.** Decay any file untouched for N turns to its signature, keeping the full text one lookup away. Done looks like the same run finishing with a smaller peak window and the same outcome. > 4. **Move decisions into task memory.** When a sub-task resolves, write a one-line record to the workflow engine's state and evict the transcript from the window. Done looks like a crashed run resuming from its decisions rather than replaying the conversation. > 5. **Add episodic recall behind a relevance gate.** Index past run summaries by entity and error type, and retrieve one only when the current error matches. Done looks like recall firing on a repeat incident and staying quiet otherwise. > **Measure it** > > - `resident_token_p95`: the 95th percentile window size across turns in a run. Log the prompt size per call. Down is good while outcomes hold. > - `stale_token_share`: tokens in the window belonging to items untouched for 20 or more turns, divided by total window tokens. Tag each item with its last-touch turn. Down is good. > - `context_contract_violations`: node inputs containing an item the node's include list does not name. Count them per run. Down is good, and zero means the contracts are enforced rather than advisory. > - `resume_success_rate`: crashed or halted runs that resume from task memory and finish, divided by all crashed runs. Up is good, and it is the clearest proof state lives outside the window. > **Takeaway** > > Scale memory, not context. The window is RAM: small, fast, and expensive. Everything else belongs in indexed storage with deterministic paging policies. --- # Knowledge graphs beat prompt engineering > Knowledge should be connected, not concatenated. Part 12: Orchestration and memory. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/knowledge-graphs-over-concatenation Your system prompt has grown to four thousand words because every time the agent missed a relationship, someone added a sentence describing it. "`RefundProcessor` writes `ledger_entries`, which `ReconciliationJob` reads nightly, which the finance dashboard queries." In prose, that chain is three sentences the model has to re-derive by attention on every call. In a graph, it is three edges the system traverses before the model wakes up. ### One ontology, many sources The previous eleven parts each built a specialized index. The knowledge graph is where they federate into a single ontology over your engineering reality: **code graphs** (symbols, calls, imports, Part 4), **schema graphs** (tables, lineage, Part 8), **documentation graphs** (architecture decision records, or ADRs, and dossiers linked to the entities they govern, Parts 5 to 6), **test and contract graphs** (behaviors linked to covered symbols, Part 10), and **runtime graphs** (services, traces, and incidents from your tracing and incident systems). Entity relationships cross the boundaries: this function ↔ writes this table ↔ governed by this ADR ↔ covered by these tests ↔ implicated in that incident. No single tool holds those edges today, which is why an incident review opens with an hour of archaeology. ### Why graph retrieval reduces reasoning complexity When an answer requires multi-hop connection, flat retrieval makes the model do the hops. It retrieves chunks, holds them in attention, and infers the links: probabilistic work, paid in tokens, degraded by every distractor (Part 7). Graph retrieval does the hops in the system by deterministic traversal, then hands the model a pre-connected subgraph. You have converted reasoning, which is expensive and variable, into lookup, which is cheap and exact. That is the thesis of this book applied to knowledge. > **Evidence · Structure wins where connection matters** > > Microsoft's **GraphRAG** is the flagship result: an LLM-extracted entity graph, Leiden community detection, and pre-computed community summaries. On global sensemaking questions over corpora of roughly 1M tokens, it beat vector RAG with **win rates of about 70 to 80% on comprehensiveness and diversity**, and it matched full hierarchical source-text summarization while using roughly **2 to 3% of the tokens per query**, because connection-finding moved from query time to index time.[10](https://macanderson.com/manual/sources#r10 "Edge et al. (Microsoft Research). \"From Local to Global: A Graph RAG Approach to Query-Focused Summarization.\" 2024. arXiv:2404.16130"),[11](https://macanderson.com/manual/sources#r11 "Microsoft Research. \"GraphRAG: New tool for complex data discovery.\" July 2024. microsoft.com/research blog") Vector search finds things that sound like the question. Graph traversal finds things related to the answer. > > Edge et al., arXiv:2404.16130 · Microsoft Research, 2024 > **Counterweight, graphs are not free** > > Honest engineering requires the caveat. Naive graph retrieval can inflate prompts instead of shrinking them. Comparative studies measured some graph-RAG variants stuffing 40k to 100k tokens per query, orders of magnitude above vector baselines, and found that piling on more retrieved subgraph stopped improving answers while adding noise.[21](https://macanderson.com/manual/sources#r21 "Zhu et al.. \"When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation.\" 2025. arXiv:2506.05690") Indexing also costs real money up front and has to be maintained. The lesson is not that graphs win everywhere. It is that a graph is an index for precise traversal, not a license to dump neighborhoods. Budgets and slicing (Part 7) apply to subgraphs the way they apply to logs. For code you are in the luckiest domain: unlike a prose corpus, this graph does not need LLM extraction, because compilers, parsers, and migrations give you most of the edges deterministically and nearly free. ### A small ontology beats a large prompt The practical migration path is short. Define a small ontology. Populate it from the deterministic sources you already built in Parts 4, 5, 8, and 10. Add embedding-based entry points for fuzzy queries, but make edges rather than similarity the backbone of expansion. Retrieval becomes anchor (parsed, Part 3) → traverse (typed edges, bounded hops) → rank (centrality plus task relevance) → slice to budget. Every step is deterministic except the final generation. ``` entities: Function, Table, Service, ADR, Test, Incident, Domain edges: calls, writes, covered_by, governed_by, caused anchor("process_refund") # parsed from the ticket .out("writes") hops=1 → ledger_entries .in("reads") hops=1 → ReconciliationJob .out("governed_by") hops=1 → ADR-031 (money is integer cents) .out("covered_by") hops=1 → 3 tests .rank(centrality + task_relevance) .slice(budget=8000) # subgraphs get budgets too ``` The difference that buys is worth stating as a comparison, because it is the reason the token bill moves. | Question: what breaks if I change the refund amount type? | Flat retrieval | Graph traversal | | --- | --- | --- | | Finding the writer of `ledger_entries` | Hope a chunk mentions both | One `writes` edge | | Finding the nightly reader | A second similarity query, unlinked | One `reads` edge | | Finding the rule that governs the column | The ADR ranks low on lexical overlap | One `governed_by` edge | | Who does the joining | The model, in attention, per call | The index, once, at write time | | Cost of a wrong answer | A silent miss you find in production | A missing edge you can go add | > **Do this week** > > 1. **Write the ontology on one page.** Name seven entity types and five edge types, and reject anything you cannot populate deterministically this quarter. Done looks like a schema file, not a diagram. > 2. **Load the two cheapest sources.** Emit `calls` and `imports` edges from tree-sitter or your compiler's index, and `writes` edges from your object-relational mapper's (ORM) introspection or from migration history. Done looks like a graph you can query for one real symbol. > 3. **Link the governing documents.** Add a front-matter field to each ADR listing the entities it governs, and load `governed_by` edges from it. Done looks like a traversal that returns the integer-cents rule without anyone naming it. > 4. **Attach your tracing and incident systems.** Map service names from OpenTelemetry resource attributes to graph nodes, and add a `caused` edge from each incident's postmortem to the symbols it names. Done looks like an incident reachable in one hop from the function that caused it. > 5. **Put a budget on the traversal.** Cap hops, cap returned nodes, and slice to a token ceiling before the subgraph reaches the model. Done looks like a subgraph whose size you can state in advance. > **Measure it** > > - `edge_coverage`: symbols with at least one outbound typed edge, divided by all symbols in the repository. Up is good, and a low number tells you the graph is a demo rather than an index. > - `hops_to_answer`: the traversal depth at which the node the patch actually needed appeared. Log the anchor, the path, and the edited symbols. Down is good, and a rising number means an edge type is missing. > - `subgraph_tokens_per_query`: tokens of retrieved subgraph handed to the model per call. Down is good, and the counterweight above is the reason to watch it rather than assume it. > - `graph_query_share`: graph queries divided by graph queries plus exploratory file reads. Up is good, because it means the agent looks things up instead of foraging. > **Takeaway** > > Concatenation forces the model to rediscover relationships you already know. Store knowledge as a graph and retrieval becomes traversal, which moves connection-finding out of the token bill and into the index. --- # Measuring agent intelligence > Benchmarking on coding challenges misses what matters in production. Part 13: Measurement. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/measuring-an-agent-in-production Someone asks how good your agent is, and the number you have is a benchmark score. It tells you one thing about one distribution of tasks under lab conditions. It does not tell you what a task costs, how long it takes, how much collateral change it makes, or whether it will do the same thing twice. Benchmarks measure capability ceilings. Production needs efficiency, precision, and repeatability, and those take different instruments. > **Evidence · Accuracy-only evaluation misleads** > > The Princeton team showed this formally. Without cost as a co-equal axis, leaderboards reward scientifically meaningless spend, including retry loops and complexity that buy accuracy at any price, and the community drew wrong conclusions about why agents improved. Their prescription, jointly optimizing on a **cost-accuracy Pareto frontier**, revealed simple designs matching complex ones at a fraction of the cost.[2](https://macanderson.com/manual/sources#r2 "Kapoor, Stroebl, Siegel, Nadgir & Narayanan (Princeton). \"AI Agents That Matter.\" TMLR 2025. arXiv:2407.01502") Agentless made the same point empirically by reporting *% resolved, average cost, and average tokens together*, and separately audited SWE-bench Lite itself, finding problem-quality issues (such as issues whose text leaks the solution) that inflate naive readings of benchmark scores.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") Anthropic's variance analysis completes the picture. If token usage explains ~80% of performance variance,[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") then a score reported without token counts is mostly measuring budget, not intelligence. > > Kapoor et al., TMLR 2025 · Xia et al., FSE 2025 · Anthropic, 2025 ### The production scorecard ##### files\_touched / lines\_changed Surgical or scattered. Two agents both solve the ticket. One edits 2 files, the other rewrites 14. Blast radius is review cost, merge risk, and regression surface. ##### tokens\_consumed per task The direct cost of thinking. Tracked per phase (retrieval, generation, repair), it tells you which part of the architecture is wasteful. ##### tool\_calls and graph\_queries The exploration ratio. A high tool-call count with a low graph-query count means the agent is foraging probabilistically for what an index should hand it. ##### time\_to\_green Wall clock from task start to passing CI. The metric people feel, and the one a latency-amplifying architecture (Part 9) quietly destroys. ##### test\_pass\_rate and first\_attempt\_rate Not just eventually green but green in how many attempts. Each repair loop multiplies cost and proxies for context quality. ##### cost\_per\_completed\_task The number finance cares about, inclusive of retries, abandoned runs, and human rework. Report it like cost of goods sold (COGS), because it is. ##### retrieval\_precision Of the context supplied, how much did the final patch depend on? Intersect provided spans with edited and read spans. The best single health metric for your deterministic layer. ##### context\_utilization The inverse view: tokens supplied against tokens that mattered. Utilization under 10% means you are paying attention-degradation costs (Part 7) for nothing. ##### variance / repeatability Run the same task 10 times. Determinism-heavy systems cluster tightly on cost and outcome. Agent-heavy systems scatter. Variance is risk, and risk is cost. These do not replace benchmarks. They are the instrument panel a benchmark cannot be. A capability score tells you which model to buy. The scorecard tells you whether your system is improving: whether this quarter's retrieval work moved precision from 8% to 40%, whether graph queries are displacing exploratory tool calls, whether cost per task is falling while pass rates hold. Optimizing these numbers is engineering. Optimizing a single leaderboard score tells you about the model you bought, not the system you built. ### The scorecard on one page Copy this into your runbook. Each row is a metric, the computation behind it, and the failure it usually points at. | Metric | How to compute it | What a bad reading usually means | | --- | --- | --- | | `files_touched` | Count distinct files in the final diff, per completed task | Retrieval is too broad, so the agent edits what it read rather than what the ticket named | | `tokens_per_task` | Sum prompt plus completion tokens per run, tagged by phase | One phase dominates. Repair usually means bad context, retrieval usually means no budget | | `graph_query_ratio` | Graph queries divided by graph queries plus exploratory file reads | The index is missing edges, so the agent forages instead of looking up | | `time_to_green` | Wall clock from task start to the first passing CI run | Hand-offs between model calls dominate, which is the Part 9 tax | | `first_attempt_rate` | Tasks green on attempt 1 divided by all completed tasks | The context bundle lacks the contract, so the model guesses the behavior | | `retrieval_precision` | Supplied spans that the final patch read or edited, divided by all supplied spans | The retriever is shipping neighborhoods rather than slices | | `cost_per_completed_task` | Total spend including retries and abandoned runs, divided by tasks merged | Abandoned runs are uncounted, and the true figure is higher than the one you quote | | `outcome_variance` | Run one task 10 times, take the spread of cost and of pass or fail | Control flow is decided by inference, so the same input takes different paths | Cost per task is the number that travels furthest outside engineering, and it raises a question this part does not answer: which agent, which run, and which person does a given dollar belong to. Part 17 takes that up. > **Do this week** > > 1. **Emit one span per run.** Wrap each agent run in an OpenTelemetry span carrying task id, model, prompt tokens, completion tokens, tool calls, and outcome. Done looks like a trace you can group by task class. > 2. **Log the bundle and the diff together.** Record the spans you supplied and the spans the patch read or edited, so `retrieval_precision` becomes an intersection rather than an estimate. Done looks like one precision number for last week. > 3. **Count abandoned runs in the cost.** Add halted, timed-out, and human-rewritten runs to the denominator of `cost_per_completed_task`. Done looks like a figure higher than the one you were quoting, and true. > 4. **Run a repeatability probe.** Pick three representative tasks, run each 10 times on the same commit, and record the spread of cost and outcome. Done looks like a variance number you can put next to your pass rate. > 5. **Publish the eight-row table weekly.** Post it where the team reads it, with last week's values beside this week's. Done looks like an argument about one row instead of an argument about the agent. > **Measure it** > > - `retrieval_precision`: supplied context spans the final patch depended on, divided by all supplied spans. Up is good, and it moves when your index improves rather than when the model changes. > - `cost_per_completed_task`: all spend, including retries and abandoned runs, divided by tasks merged. Down is good. Compute it weekly, because a monthly figure hides a bad week. > - `first_attempt_rate`: tasks green on the first attempt, divided by all completed tasks. Up is good, and each repair loop you remove pays twice, in tokens and in wall clock. > - `outcome_variance`: the spread of cost and pass or fail across 10 runs of one task on one commit. Down is good, and a wide spread tells you inference is deciding something code should decide. > **Takeaway** > > A benchmark score measures the model's ceiling. Tokens, tool calls, blast radius, time to green, and retrieval precision measure your architecture. One of those is under your control, so instrument it. --- # Give every agent its own identity > You cannot set authority for, bill, or review something you cannot name. Part 14: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/give-every-agent-its-own-identity A pull request lands overnight. The commit author is a shared bot account that six agents and two scripts use. The change is fine, but the review question is not about the diff. It is: which agent did this, who started it, and what was it allowed to do? If the answer takes an afternoon of reading logs, the agents are doing work that nobody in particular is accountable for. ### What parts 1 to 13 left open The first thirteen parts moved decisions out of the model and into code: what to retrieve, what to keep in the window, which step runs next. One decision is still open, and it is the one a second team will ask about. What may this agent do in systems other people own? The same principle answers it. Decide in the system, before the model is involved, and keep the decision where a person can read it. That starts with a name. ### The familiar model and what it cannot answer The familiar model is a shared credential. One token carries everything any agent might need, and every agent that holds it looks the same to the system on the other end. It is quick to set up. It also means three questions have no answer in the data: - **Which agent acted?** The target system saw the token, not the agent. - **On whose behalf?** The person who started the task is not in the call. - **Under what authority?** The token's scope is the union of every task anyone ever planned, so it says nothing about this task. Zero trust architecture gives the general form of the fix. NIST's definition assumes no implicit trust from network location or asset ownership, and it evaluates each request against policy for the specific subject and resource.[22](https://macanderson.com/manual/sources#r22 "Rose, Borchert, Mitchell and Connelly (NIST). \"Zero Trust Architecture.\" NIST Special Publication 800-207, 2020. doi.org/10.6028/NIST.SP.800-207") An agent is a subject. It needs to be one the policy can tell apart from the other agents. ### What an agent identity holds An identity is a record, not a credential. It answers who this is and who answers for it: | Field | What it says | Example | | --- | --- | --- | | `agent_id` | A stable principal, one per agent, never reused | `agt_refund_triage` | | `operator` | The person accountable for this agent and its runs | Dana Okafor, payments platform | | `harness` | What runs it, and at which version | Claude Code 2.x, wrapped | | `purpose` | One sentence a reviewer can check a run against | Triage refund failures and open a fix PR | | `environment` | Where it may work | staging, `billing/**` | | `status` | Enrolled, held, or retired, with a date | enrolled 2026-08-04 | Three properties matter more than the fields. The identity is **separate from any connection credential**, so rotating a GitHub token does not change who the agent is. It **carries the initiator through**: each run records the person who started it, so "on whose behalf" is data, not inference. And it can be **revoked alone**: holding one agent does not stop the other five. #### ◌ One shared credential - The target system sees one actor for every agent - The initiator is not recorded - Scope is the union of all planned tasks - Revoking it stops everything that uses it #### ● One identity per agent - Each call names the agent that made it - Each run names the person who started it - Authority attaches to the identity (part 15) - One agent can be held while the others work > **Evidence · Excess authority is a design property** > > OWASP's Top 10 for LLM applications lists excessive agency as its own category. It breaks the cause into three parts: excessive functionality, excessive permissions, and excessive autonomy. Each is a property of how the system around the model was built, not of the model.[23](https://macanderson.com/manual/sources#r23 "OWASP. \"Top 10 for LLM Applications 2025,\" LLM06: Excessive Agency. genai.owasp.org") The Berkeley failure taxonomy from part 9 points the same way: most recorded multi-agent failures traced to system design, including specification and verification, and the authors do not expect better models alone to fix them.[5](https://macanderson.com/manual/sources#r5 "Cemri et al. (UC Berkeley). \"Why Do Multi-Agent LLM Systems Fail?\" (MAST). NeurIPS 2025 Datasets & Benchmarks. arXiv:2503.13657") > > OWASP, Top 10 for LLM applications · Cemri et al., NeurIPS 2025 ### Start with an inventory Most teams cannot list their agents, because agents arrive one team at a time. The inventory is a spreadsheet before it is a system. One row per agent, the six fields above, and two more columns: which credentials it can reach today, and who would notice if it stopped. The rows with no operator are the first finding. The rows that share a credential are the second. > **Where Oxagen fits** > > In Oxagen each enrolled agent is a registered principal with its own identity and an accountable operator. That identity is distinct from any connection credential. Agents built by different groups, on different harnesses, appear on one Runs page with their open requests, their spend, and their last run. Oxagen does not run the agent. A supported wrapper sits beside the harness, and the agent keeps running where it runs today. > **Do this week** > > 1. **List every agent that can write to a system.** Ask each team. Include scheduled scripts that call a model. Done is one row per agent with an operator's name in it. > 2. **Mark each shared credential.** For every credential, list the agents that can reach it. Any credential with more than one agent is a row in your backlog. > 3. **Give each agent a stable id and a one-sentence purpose.** Put the id in the commit trailer, the user agent string, or the request header, whichever the target system records. > 4. **Record the initiator.** Pass the id of the person who started the task into the run, and write it next to the agent id on every outbound call you log. > 5. **Test a single revocation.** Hold one agent in staging. Confirm that it stops and that nothing else does. > **Measure it** > > - `agents_with_operator`: agents with a named, current operator, over all agents. Anything under 1.0 is an agent nobody answers for. > - `agents_per_credential`: the largest number of agents that share one credential. The target is 1. > - `attributable_actions`: logged write actions that carry both an agent id and an initiator, over all logged write actions. Track it weekly. It should rise as wrappers roll out. > **Takeaway** > > Authority, budget, and review all attach to a name. Give each agent its own identity and a person who answers for it, and keep both apart from the credentials it works through. --- # Write the mandate > Four decisions govern an agent. Most teams have made all four. Few have them in one place. Part 15: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/write-the-mandate Ask what one of your agents is allowed to do and you will get four answers from four places. Security points at an IAM role. Finance points at a cloud budget alert. Engineering points at a prompt file and a tool list. The record of what happened is in a log platform under a fourth team's account. Each answer is right. None of them is about the agent, so nobody can read the whole picture, and nobody can change one part without asking who owns the others. ### One object per agent Set an agent's authority, budget, tools, and skills in its mandate. A mandate is one object per agent with four clauses. Three of them are written by the people who already own that decision. The platform keeps the fourth. | Clause | Who sets it | What it says | | --- | --- | --- | | Access | Security | The identity the agent acts as, which systems it may request, which data it may read, and which actions it may request | | Budget and rules | Finance | What it may spend, under which commercial terms, and the rules it must obey: approval thresholds, allowed vendors, decision rules | | Equipment | Engineering | The tools and skills it may use, the business context it may read, and the steering it runs under | | Record | The platform | What each run read, what it changed, what it cost, and which rule answered each request | The three writers can be three teams or one person with three hats. The structure does not assume an org chart. What it assumes is that each clause has one owner who can change it without a meeting, and that a reader can see all four at once. ### Why it is one object and not four Part 6 argued that a rule which cannot fail CI is a hope. The same holds here. Four settings in four systems cannot be checked against each other. A tool in the equipment clause that writes to a system the access clause does not list is a contradiction, and only a single object makes it visible. A budget with no rule for what happens at the limit is incomplete, and only a schema can say so. A mandate is typed configuration. It lives in version control, it changes by pull request, and a validator runs on every change. ``` # mandates/agt_refund_triage.yaml agent: agt_refund_triage operator: dana.okafor access: # owner: security acts_as: agt_refund_triage may_request: - system: github scope: "repo:acme/billing" actions: [read, open_pull_request] - system: postgres_staging scope: "schema:billing" actions: [read] budget_and_rules: # owner: finance monthly_limit_usd: 400 at_limit: hold_and_notify_operator rules: [no-push-main, release-branch-needs-owner] equipment: # owner: engineering tools: [run_code_with_trace, schema_slice, tests_covering] skills: [billing-conventions] context_scope: ["domain:billing"] record: # kept by the platform, not writable here retention_days: 400 ``` The validator checks what a reviewer would otherwise check by eye: every tool's target system appears under `may_request`, every rule named exists, the operator is a current employee, and the budget has an `at_limit` action. Those are four deterministic checks, and none of them needs a model. ### Lead with the clause the reader owns A mandate is read by different people for different reasons. The security lead reads access first and wants to know which rule answers a request. The finance lead reads the budget and wants the number attributed. The engineering lead reads equipment, then the record. The operator reads all four, because the operator is the one who gets asked. Write each clause so its owner can review it without reading the other three, and so the operator can read them together. > **Counterweight · Scope is part of the claim** > > A mandate governs the actions that pass through the layer that checks it. An agent with a second path to a system, such as a token in an environment variable, is not governed on that path, whatever the mandate says. State the boundary when you describe the control: "for actions routed through the gateway." If a mode only records calls and does not block them, say recorded, not enforced. An auditor tests the claim as it is written, so write it as it applies. > **Where Oxagen fits** > > The mandate is the unit Oxagen manages. Identity, knowledge scope, permitted action, commercial terms, and the audit record are one typed contract. For actions routed through Oxagen, that contract is checked when the agent calls, not reconstructed after an incident. An agent outside that boundary, or a call that does not pass through Oxagen, is not governed by it. Observe mode is recorded, not enforced. > **Do this week** > > 1. **Pick one agent and fill in the template.** The field kit has a blank mandate. Write what is true today, not what should be true. Done is four clauses with an owner's name on each. > 2. **Find the contradictions.** List every tool the agent can call and the system it touches. Any system missing from the access clause is either an access gap or a tool to remove. > 3. **Add an `at_limit` action.** Decide what happens when the budget is spent: hold the agent, notify the operator, or continue and flag. Any of the three beats finding out from the invoice. > 4. **Put it in version control.** One file per agent, a `CODEOWNERS` entry per clause, and a schema check in CI. > 5. **Write down the boundary.** One sentence: which paths to which systems this mandate governs, and which it does not. > **Measure it** > > - `agents_with_mandate`: agents with all four clauses filled and owned, over all agents from the part 14 inventory. > - `mandate_contradictions`: validator failures on the current set of mandates. The target is 0, and a rising number after a tooling change is the early signal. > - `ungoverned_paths`: known routes from an agent to a system that bypass the checking layer. Count them and list them, even if you cannot close them yet. > **Takeaway** > > Access, budget and rules, equipment, and the record are four clauses of one object. Give each clause an owner, keep them in one validated file per agent, and state which actions the mandate covers. --- # The agent asks, a rule decides > Authority is a decision at the moment of use, written down by the team that owns the system. Part 16: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/the-agent-asks-a-rule-decides An agent needs to push a fix to a release branch at 2 a.m. Someone decided, at some point, whether that is fine. Where is that decision? If it is "the token has write scope," then the decision was made once, for every branch, for every task, by whoever created the token. When a team adds an agent, where does its authority get set, and who sets it? ### A request, not a standing grant The alternative to a broad credential is a request at the moment of use. When a task needs a system, a scope, or an action, the agent asks for it. The request carries four facts, and each one is data your system already has: - **Who started the task.** The initiator from part 14. - **Which agent is asking.** Its identity. - **What it wants.** The tool and the action: `github.push`. - **What the action would reach.** The resource: `acme/billing`, branch `release/2026.09`. This is the prompt parsing of part 3 applied to authority. The request is structured input. A rule is a deterministic function of it. No model decides whether the agent may act, in the same way no model decides which files to retrieve. ### Three answers The team that owns the system writes the rule, and a rule has three possible answers. | Answer | What happens | Example | | --- | --- | --- | | Allowed | The action proceeds inside the same call. The record keeps the rule id. | Read a public repository | | Denied | The action does not reach the code that would have done it. The agent gets a reason it can act on. | Push to `main` | | Routed to a person | The run waits. A named person answers. The run resumes with the answer in the record. | Push to a release branch | A rule that allows everything is still a rule. Someone wrote it, someone can read it, and someone can change it. That is the difference between an allowed request and a standing grant: the first has an author and a row. ``` # rules/release-branch.yaml owner: release engineering id: release-branch when: action: github.push resource.branch: "release/*" effect: route route_to: role:release-owner timeout: 4h on_timeout: deny # rules/no-push-main.yaml id: no-push-main when: { action: github.push, resource.branch: main } effect: deny reason: "Push to a branch and open a pull request." ``` ### Write the denial for the agent A denied request is input to the next step, so write the reason as an instruction. "Request denied by rule `no-push-main`. Push to a branch and open a pull request." An agent that receives that sentence opens a pull request. An agent that receives `403` retries, then tries another path. Part 2 made the same point about tools: what the interface returns decides what the model does next.[20](https://macanderson.com/manual/sources#r20 "Anthropic Engineering. \"Writing effective tools for agents.\" 2025. anthropic.com/engineering") ### The default is a decision too Every rule set has a request that matches nothing. Decide what happens then, and say it on the same page as the rules. "When no rule matches, deny" and "when no rule matches, identity alone answers" are both defensible. An unstated default is the one to avoid, because the first person to learn it will be learning it during a review. ### Where the credential lives For a connection the platform mediates, the agent does not need the connection credential at all. The platform holds it, checks the request, and makes the call on the agent's behalf. The agent still has its own identity and whatever reaches its context, so this is not a promise about leaks. It is a statement about custody: the secret for that connection stays in one place, and each use of it is a row. > **Evidence · Per-request decisions are the established pattern** > > NIST's zero trust architecture places a policy decision point and a policy enforcement point between every subject and every resource. Access is granted per request, under dynamic policy, and limited to what the task needs.[22](https://macanderson.com/manual/sources#r22 "Rose, Borchert, Mitchell and Connelly (NIST). \"Zero Trust Architecture.\" NIST Special Publication 800-207, 2020. doi.org/10.6028/NIST.SP.800-207") OWASP's guidance for excessive agency recommends the same controls for LLM applications: limit the permissions an extension holds, execute actions in the context of the specific user, and require human approval for high-impact actions.[23](https://macanderson.com/manual/sources#r23 "OWASP. \"Top 10 for LLM Applications 2025,\" LLM06: Excessive Agency. genai.owasp.org") The three answers above are those recommendations in one mechanism. > > NIST SP 800-207 · OWASP, Top 10 for LLM applications > **Counterweight · Routing has a cost** > > Every routed request makes a person the latency. Route too much and approvals become a reflex, which is worse than a written rule. Review the routed requests monthly. A request the same person approves every time, without changes, is a candidate for an allow rule with a narrower condition. A request that is often rejected is a candidate for a deny rule with a better reason. > **Where Oxagen fits** > > In Oxagen a decision rule names a capability, a condition, and one effect: allow, deny, or require approval. For governed calls, the kernel checks the rule after identity and entitlement and before the handler runs, so a denied action does not reach the code that would have done it. A routed request pauses the run. The person answers from the Access page, and the run resumes with the answer in the record. For mediated connections, Oxagen uses the connection credential on the agent's behalf and the agent does not receive it. When a workspace has no rules, identity alone answers requests, and the Access page says so. > **Do this week** > > 1. **List the ten write actions your agents take most.** Pull them from the logs of part 14. For each, write the answer you want: allowed, denied, or routed, and to whom. > 2. **Write the three rules that matter most as files.** Use the shape above. Give each an id, an owner, and a reason. > 3. **State the default.** One sentence at the top of the rules directory that says what happens when nothing matches. > 4. **Rewrite one denial message.** Pick the denial your agents hit most and make its text the next step. > 5. **Route one action to a named person.** Set a timeout and an on-timeout answer. Run it once in staging and time how long the run waits. > **Measure it** > > - `requests_by_answer`: allowed, denied, and routed, per agent, per week. A sudden rise in denials after a deploy is a tooling change nobody told security about. > - `time_to_answer_p50` and `p95`: how long routed requests wait. This is the latency you are adding to `time_to_green` from part 13. > - `unmatched_requests`: requests answered by the default. Each one is a rule nobody has written yet. > - `retries_after_denial`: calls an agent makes after a denial before it changes course. A good reason brings this toward 0. > **Takeaway** > > Do not hand your agents the keys. Let the agent ask at the moment of use, let a rule the owning team wrote answer allowed, denied, or routed to a person, and keep the answer. --- # Spend you can attribute > A total is not an answer. The answer is which agent spent what, and on whose behalf. Part 17: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/spend-you-can-attribute The model invoice arrives and the number is larger than last month. You can see the total by API key and by day. The question from finance is different: which agent spent it, for which person, on which task, under which budget? Can you attribute last month's agent spend to the agent, the run, and the person who started the run? If the honest answer is an estimate built from a spreadsheet, the data to answer it was never recorded. ### Why the invoice cannot answer A provider bills the account that made the call. It has no field for your agent, your run, or your colleague, because you never sent them. Attribution cannot be reconstructed afterward from a total. It is recorded at the call, by the layer the call passes through, or it does not exist. Part 13 made cost per completed task a first-class metric, and the Princeton work behind it showed why accuracy reported without cost misleads.[2](https://macanderson.com/manual/sources#r2 "Kapoor, Stroebl, Siegel, Nadgir & Narayanan (Princeton). \"AI Agents That Matter.\" TMLR 2025. arXiv:2407.01502") That chapter measured the architecture. This one measures the organization: the same rows, with four more keys on each. ### The meter is a row Write one row per priced step. A step is one model call or one tool call. Each row carries the keys that every later question groups by. | Column | Why it is there | | --- | --- | | `person` | Who started the task. The "on whose behalf" column. | | `agent` | The identity from part 14. | | `run`, `turn`, `step` | Where in the work the cost fell. A run is one agent on one task. A turn is one prompt through to the point the agent stops. A step is one call. | | `model`, `input_tokens`, `cached_tokens`, `output_tokens` | The quantities the price applies to. Cached reads price differently, so they need their own column.[16](https://macanderson.com/manual/sources#r16 "Anthropic. \"Prompt caching with Claude\" (up to 90% cost and 85% latency reduction, and a cache read costs about 0.1 times the input price). 2024. anthropic.com/news/prompt-caching") | | `price_version` | Which price list was applied. Prices change, and a rate negotiated in March should not reprice February. | | `cost_usd`, `cost_basis` | The amount, and whether it was `measured`, `reported`, or `estimated`. | | `rule_id` | The rule that answered the request behind this step, from part 16. | The `cost_basis` column carries more weight than it looks. Measured means your layer saw the tokens and applied the price. Reported means the harness or the provider told you a figure. Estimated means you inferred it. A finance lead can work with all three as long as each one is labelled. A blended number with no label is the one that fails review. ### Budgets belong beside the agent A budget in a cloud console alerts someone after the money is spent, and it alerts on the account, not the agent. A budget in the mandate is checked before the next step and names the agent it stops. Three settings are enough to start: - **A monthly limit per agent,** in currency, set by the person who owns the number. - **A per-run ceiling,** so one run that loops cannot spend the month. Part 9 called this a circuit breaker, and this is where it gets its number. - **An `at_limit` action:** hold the agent and notify the operator, route the next step to a person, or continue and flag. The finance lead and the operator then read the same rows. One reads them grouped by cost center, the other grouped by run. > **Evidence · Allocation comes before optimization** > > The FinOps Foundation's framework treats allocation as a core capability: assign cost to the teams and workloads that incur it, using consistent metadata, so that owners are accountable for their own usage.[24](https://macanderson.com/manual/sources#r24 "FinOps Foundation. \"FinOps Framework,\" the Allocation capability. finops.org/framework") The order matters. A saving is hard to verify when nobody can say whose spend fell. Anthropic's production report gives agent teams a reason to care about the same order: on their internal evaluations, token usage alone explained about 80% of performance variance.[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") Spend is a description of what the architecture did. Attributed spend says which part did it. > > FinOps Foundation, FinOps Framework · Anthropic engineering, 2025 ### A worked example One agent, one week, read three ways from the same rows: | Grouped by | What it shows | Who acts | | --- | --- | --- | | `agent` | `agt_refund_triage` spent $212 of a $400 limit | The operator: on pace, no change | | `run` | Two runs cost $61 together. The median run costs $1.90. | The engineer: both runs looped on a failing migration, so set a per-run ceiling and fix the schema slice from part 8 | | `person` | One person started 70% of the runs | Finance: charge the payments cost center, not platform | The figures are illustrative. The point is that three people took three different actions from one table, and none of them needed the invoice. > **Counterweight · What this does not claim** > > Attribution is not savings. Recording who spent what does not lower the bill. It shows where the bill comes from, which is the precondition for any change you can verify afterward. If you later claim a reduction, state the workload, the baseline, the model and price versions, the sample size, the quality measure, and the costs including retries. Without those, the claim you can support is visibility. > **Where Oxagen fits** > > In Oxagen every governed action is priced and attributed to the person, the agent, the run, the turn, and the step. The Spend page and the Run page read the same rows, and measured, reported, and estimated costs are labelled as such. The budget an agent runs under, what it has spent this month, and the rule that stops it sit beside the agent. This covers calls that pass through Oxagen. A harness whose model calls do not pass through the gateway reports its spend, or is not metered, and the page says which. > **Do this week** > > 1. **Add four keys to your model-call log.** `person`, `agent`, `run`, and `step`. If you proxy model calls, add them there. If you do not, add them in the wrapper that starts the run. > 2. **Separate cached input tokens.** Log them in their own column and price them at the cached rate. > 3. **Version your price list.** A small table keyed by model and effective date. Stamp each row with the version used. > 4. **Label the cost basis.** Go through each source of cost data and mark it measured, reported, or estimated. > 5. **Set one per-run ceiling.** Pick the agent with the widest spread between median and maximum run cost. Set the ceiling at ten times the median and route to the operator when a run reaches it. > **Measure it** > > - `attributed_spend_ratio`: spend carrying person, agent, and run, over total model spend on the invoice. The gap is what you cannot explain. > - `cost_per_run_p50` and `p99`, per agent. A wide spread points at loops, and part 9 and part 11 are where to look. > - `unpriced_steps`: steps with tokens and no price, usually a new model missing from the price list. > - `budget_holds`: times an `at_limit` action fired, per agent, per month. > **Takeaway** > > Record the person, the agent, the run, and the step on every priced call, label how each cost was obtained, and keep the budget beside the agent it limits. Then the question of who spent what has a query for an answer. --- # Equip agents on purpose > A tool in the catalog is not a tool in the agent's hands. Assign equipment the way you assign access. Part 18: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/equip-agents-on-purpose Your platform team has built the tools this book describes: a trace slicer, a code graph, a schema index, a test lookup. They are registered once and every agent gets all of them, along with forty others that accumulated. The refund agent can see the deploy tool. The docs agent can see the database tool. Nobody decided that. It is what happens when the catalog and the assignment are the same list. ### The tool list is context, and it is the largest part Every tool an agent can see is sent with every request: its name, its description, and its input schema. Part 7 showed what irrelevant context costs, and tool definitions are context like any other. One measurement from Oxagen's own in-app agent, taken by hand because nothing reported it: 271 tools came to 45,007 tokens, which was 92.4% of the cacheable prefix. Providers also cap the count. OpenAI's limit is 128 tools per request, and a request over the cap fails at the provider with an error about a payload nobody on your side can inspect. Anthropic's guidance on writing tools makes the design point: more tools do not always lead to better outcomes, and it recommends building a few tools aimed at the workflows that matter most.[20](https://macanderson.com/manual/sources#r20 "Anthropic Engineering. \"Writing effective tools for agents.\" 2025. anthropic.com/engineering") The assignment is an engineering decision. Make it in the mandate, where it can be reviewed. ### Three kinds of equipment | Kind | What it is | Assigned by | | --- | --- | --- | | Tools | Callable actions with typed inputs: `schema_slice`, `tests_covering`, `open_pull_request` | A list in the equipment clause | | Skills | Packaged instructions for one kind of task, loaded when that task comes up, not on every request | A list in the equipment clause | | Business context | The knowledge the work needs: domain dossiers from part 5, the graph from part 12, conventions, decisions | A scope, such as `domain:billing` | Skills follow the argument of part 6. A rule that applies to billing migrations should reach the agent when it touches a billing migration, and stay out of the window otherwise. A skill is that rule with a trigger. Context scope follows the argument of part 12. A graph makes retrieval exact, and it also makes permission exact. If knowledge is nodes with domains attached, then "this agent may read `domain:billing`" is a filter on a traversal. The same filter serves relevance and authority, and you maintain it once. ### Assignment is not permission Equipment and access are different clauses for a reason. Assigning `open_pull_request` to an agent means the tool is in its hands. Whether a particular pull request may be opened is still a request that a rule answers (part 16). Keep both checks. The first keeps the window small and the agent on task. The second decides each action at the moment of use. ``` equipment: # owner: engineering tools: - run_code_with_trace # part 2 - schema_slice # part 8 - tests_covering # part 10 - open_pull_request # each call still meets the rules skills: - billing-conventions # loads when a diff touches billing/** - migration-review context_scope: - "domain:billing" - "adr:*" tool_budget: max_tools: 24 # a build failure over this, not a warning max_tokens: 6000 ``` > **Evidence · Interfaces decide outcomes** > > SWE-agent's central result was that the agent-computer interface changes what the same model can do: purpose-built commands with compact, structured output raised resolution rates on SWE-bench over raw shell access.[14](https://macanderson.com/manual/sources#r14 "Yang, Jimenez et al. (Princeton). \"SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.\" NeurIPS 2024. arXiv:2405.15793") Anthropic's tool-writing guidance adds the operating advice: choose tools deliberately, namespace them, return high-signal output, and keep token use low by default, because an agent inherits the inefficiency of its tools.[20](https://macanderson.com/manual/sources#r20 "Anthropic Engineering. \"Writing effective tools for agents.\" 2025. anthropic.com/engineering") The Model Context Protocol moves in the same direction by having each server declare its tools, their schemas, and their descriptions, so that discovery is a lookup.[25](https://macanderson.com/manual/sources#r25 "Model Context Protocol. Specification: servers, tools, and schemas. modelcontextprotocol.io/specification") A declared catalog is what makes a per-agent assignment possible. > > Yang et al., NeurIPS 2024 · Anthropic engineering, 2025 · Model Context Protocol specification ### Check that the equipment reached the work An assignment is a claim until a run shows it. For each run, record which tools were offered, which were called, which skills loaded, and which context scopes were read. Two findings come out of that data quickly. Tools that are offered on every run and called on none are candidates to remove. Tools that an agent asks for and does not have show up as failed attempts or as workarounds through a shell. > **Counterweight · Context is scoped, not universal** > > Giving an agent business context does not mean it recalls it correctly, and it does not mean every agent should have it. Storing a fact is not the same as retrieving it at the right moment, which is why parts 3 to 5 exist. Scope context to what the agent's work requires and what its mandate permits, and measure retrieval precision for the context layer the same way you measure it for code. > **Where Oxagen fits** > > In Oxagen, tools and skills reach an agent through its mandate. A tool in the catalog does not grant every agent permission to use it, and each request still meets the decision rules. Business context is served within the scope the mandate permits. Oxagen wraps four harnesses: Claude Code, Codex, Cursor, and Stella. What a wrapper can gate and meter differs by harness, and the product says which. > **Do this week** > > 1. **Count the tools and the tokens.** For one agent, serialize the tool list exactly as it is sent and count both. Done is two numbers you did not have yesterday. > 2. **Log offered against called.** Thirty days of runs is enough. Sort tools by calls per run. > 3. **Remove the tools with no calls.** Take them out of that agent's assignment, not out of the catalog. Watch `first_attempt_rate` from part 13 for a week. > 4. **Move one always-loaded rule into a skill.** Pick the longest block of instructions that applies to one kind of task. Give it a trigger and load it on demand. > 5. **Add a tool budget to CI.** Fail the build when an agent's assignment exceeds a count or a token limit, or a provider's cap. > **Measure it** > > - `tool_tokens_per_request`: tokens spent on tool definitions, per agent. Compare it with the agent's median task context. > - `tool_use_ratio`: distinct tools called over tools offered, per run. A low ratio means the assignment is wider than the work. > - `skill_load_rate`: runs in which a skill loaded, over runs where its trigger matched. Under 1.0 means the trigger is wrong. > - `out_of_scope_reads`: context requests outside the agent's scope. Each one is either a scope to widen or a task the agent should not have. > **Takeaway** > > Give each agent the tools and skills its work requires, and the business context its mandate permits. Keep the catalog and the assignment as separate lists, and check the assignment against what the runs used. --- # Keep a record another person can read > A log answers the engineer who wrote it. A record answers the person who was not there. Part 19: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/keep-a-record-another-person-can-read Something an agent did is being reviewed. It might be an incident, an audit sample, or a customer question. The person reviewing was not on the run and does not know your logging conventions. They ask four things: what did the agent read, what did it change, who or what allowed it, and what did it cost? When an agent requests an action in a system another team owns, which rule answers it, and where is that answer recorded? If those answers live in four tools and a chat transcript, the review becomes an investigation. ### Logs and records are different artifacts Logs are for debugging. They are verbose, unstructured at the edges, sampled under load, and deleted after a few weeks. All of that is correct for debugging. A record has a different reader and a different job. It is complete for the actions it covers, structured so a query can answer a question, ordered so the sequence is not in doubt, and kept as long as the accountability lasts. #### ◌ A log - Written for the author of the code - Free text with some fields - Sampled and rotated - Answers "why did this break?" #### ● A record - Written for a reader who was not there - One typed row per event, with stable keys - Complete for governed actions, and retained - Answers "what happened, and under whose authority?" ### What a run's record holds Parts 14 to 18 each produced rows. The record is those rows in one sequence, keyed by the run. | Event | Fields | From | | --- | --- | --- | | Run opened | agent, operator, initiator, task, mandate version | Parts 14 and 15 | | Context read | what was retrieved, from which scope, at which source commit | Parts 3 to 5 and 18 | | Request | tool, action, resource, the answer, the rule id, and the approver when routed | Part 16 | | Step | model or tool call, tokens, cost, cost basis | Part 17 | | Change | files, rows, or resources written, with a diff or a reference | The tool that made it | | Run closed | outcome, totals, and the checks that ran if the task was bounded | Part 20 | The important property is that the request, the rule that answered it, and the cost of the step sit in the same row family under the same keys. A reviewer follows one action back to its authority without joining systems by timestamp. ### Order you can check A record that can be edited quietly is an account, not evidence. Two inexpensive mechanisms raise the bar. First, chain the rows: each row stores a hash of the row before it, so a removed or altered row breaks the chain at a known point. Second, hash a canonical form. Two serializers can emit the same JSON object with keys in a different order, and the hashes will differ. RFC 8785 defines a canonical JSON serialization for this purpose, so that the same data yields the same bytes and the same digest everywhere.[26](https://macanderson.com/manual/sources#r26 "Rundgren, Jordan and Erdtman. \"JSON Canonicalization Scheme (JCS).\" RFC 8785, 2020. rfc-editor.org/rfc/rfc8785") ``` row = { "run": "run_7f3a", "seq": 41, "kind": "request", "agent": "agt_refund_triage", "initiator": "dana.okafor", "action": "github.push", "resource": "acme/billing@release/2026.09", "answer": "routed", "rule": "release-branch", "approver": "priya.n", "prev": "sha256:9c1e..." # digest of row 40 } digest = sha256(canonicalize(row)) # RFC 8785, then hash # At close, sign the last digest. A reviewer with the export can # recompute the chain without access to the system that wrote it. ``` This shows integrity, and only integrity. A valid chain means the rows are the rows that were written, in that order. It does not mean the agent's work was correct, and it does not cover an action that never passed through the recording layer. > **Evidence · A shared written record changes what a review can do** > > Google's site reliability practice treats the written postmortem as the unit of learning from an incident: a record of the impact, the actions taken, the root causes, and the follow-up, kept blameless so that people contribute what they know.[27](https://macanderson.com/manual/sources#r27 "Beyer, Jones, Petoff and Murphy (eds.). \"Site Reliability Engineering,\" chapter 15: Postmortem culture, learning from failure. O'Reilly, 2016. sre.google/sre-book") The practice depends on a timeline that people accept as accurate. With agents, the timeline is the hard part, because the actor cannot be interviewed and the transcript is long. Part 13 instrumented the run to improve the architecture. The same instrumentation, kept and ordered, is what a review reads. > > Beyer et al., Site Reliability Engineering, chapter 15 ### Write the record for its three readers - **The operator** wants the run as a sequence: what it is waiting on now, and what it did before that. - **The security reviewer** wants requests grouped by answer and by rule, with every routed request showing who signed. - **The finance reader** wants the same rows grouped by person and cost center. One set of rows serves all three when the keys are consistent. Three separate exports, built at three different times, will disagree with each other, and the review will be about the disagreement. > **Counterweight · A record holds what passed through it** > > State what the record covers. It holds governed activity: the calls that went through the layer that writes it. It does not hold what an agent did on a path around that layer, and it records an observe-mode call without having enforced anything on it. It also holds data that may be sensitive. Decide what is stored in full, what is stored as a digest, who can read which rows, and how long each kind is kept, before the first review asks. > **Where Oxagen fits** > > For governed activity, Oxagen keeps what the run read, what it changed, what it cost, and which rule answered each request. Every governed request is a row: who started the task, which agent asked, which tool it wanted, what data it would reach, which rule answered, and who signed. It is the same row the meter prices. Oxagen calls each recorded event in a run a frame. Every frame is hash-chained to the one before it, and the seal signs the close of the run, so a reviewer with the export can check its integrity outside the product. > **Do this week** > > 1. **Run a tabletop review.** Pick one agent action from last week. Give a colleague who was not involved 30 minutes to answer the four questions in the opening. Write down where they got stuck. > 2. **Pick the run id and thread it through.** One id, created when the run opens, present on every log line, request, and cost row. > 3. **Separate the record from the logs.** Give it its own store, schema, and retention, even if the first version is one append-only table. > 4. **Chain the rows.** Add `seq` and `prev`, and canonicalize before hashing. Write the 20-line verifier at the same time, and run it nightly. > 5. **Write the coverage sentence.** "This record covers X. It does not cover Y." Put it at the top of the export. > **Measure it** > > - `time_to_answer_review`: minutes for someone outside the team to answer the four questions for a sampled action. Measure it quarterly. This is the number the whole part exists to lower. > - `record_coverage`: write actions present in the record, over write actions seen by the target systems' own audit logs. The gap is the ungoverned path count from part 15, observed. > - `chain_breaks`: verifier failures per night. The expected value is 0, and any other value is worth a person's morning. > **Takeaway** > > Keep one ordered, typed record per run that holds the request, the rule that answered it, the change, and the cost under the same keys. Make its order checkable, and say what it covers. --- # Bounded tasks and ongoing work > Some work has an endpoint. Other work continues. Manage each in its own way. Part 20: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/bounded-tasks-and-ongoing-work You operate two kinds of agent and they tend to get one kind of management. The first fixes a bug, runs a migration, or writes a report. It starts, works, and stops. The second triages an inbox, watches a queue, or keeps dependencies current. It has no last step. Ask "is it done?" of the first and the question has an answer. Ask it of the second and the honest reply is that the question does not apply. ### Tell the two apart first | | Bounded task | Ongoing responsibility | | --- | --- | --- | | Shape | Has a start and an end state you can describe | Runs until someone stops it | | Examples | Fix issue 4821, migrate a table, draft the quarterly report | Triage refunds, keep dependencies patched, answer tier-1 tickets | | The question | Did the checks hold? | Is it inside its authority and budget, and what is it waiting on? | | Managed through | Completion criteria set before the work starts | Authority, budget, and review points | | Reviewed | At the end of the run | On a schedule, and when a threshold trips | The common mistake runs in one direction: inventing a finish line for ongoing work so that it fits a task tracker. An inbox agent with a "done" state will report done. That tells you nothing about the inbox. ### Bounded tasks: define completion before the work starts Part 10 made tests the ground truth of behavior and had the agent write the failing test first. Generalize that. Before a bounded run begins, write down the checks that decide completion, and fix them so the run cannot revise them. - **A command that must pass.** `pytest tests/test_billing.py` exits 0. - **A file or a diff condition.** The migration exists. No file outside `billing/**` changed. - **A human decision.** The release owner approves the pull request. Fixing the checks first matters for a plain reason. An agent that judges its own work can stop early or skip verification. The Berkeley taxonomy names this family of failures: premature termination, and no or incomplete verification, both recorded across the frameworks they studied.[5](https://macanderson.com/manual/sources#r5 "Cemri et al. (UC Berkeley). \"Why Do Multi-Agent LLM Systems Fail?\" (MAST). NeurIPS 2025 Datasets & Benchmarks. arXiv:2503.13657") A check the run cannot edit removes the agent's opinion from the verdict. The verdict becomes a function of recorded results, and another person can recompute it from the record (part 19). ``` # done/issue-4821.yaml written before the run, hashed when it opens task: "Fix DecimalConversionError on POST /api/v2/refunds" checks: - run: "pytest tests/test_billing.py::test_refund -x" - run: "ruff check billing/" - diff: { only_paths: ["billing/**", "tests/**"] } - human: { role: release-owner, question: "Safe to ship in 2026.09?" } on_fail: { retries: 3, then: route_to_operator } ``` > **Counterweight · What a passing verdict means** > > A passing verdict means the specified checks held. It does not establish that every requirement was captured or that every action was correct. Your team decides whether those checks are enough for the task, and a thin set of checks passes thin work. When checks fail, the useful record is not just "failed". It is the reasons, and whether the agent continued, waited for a person, or reached its retry limit. ### Ongoing work: authority, budget, and review points An ongoing agent is managed the way you manage a standing responsibility held by a person. You do not ask whether they are done. You look at what they are allowed to do, what they have spent, what they are waiting on, and what they did since you last looked. - **Authority.** The mandate from part 15, reviewed on a date, not only when something goes wrong. - **Budget.** The monthly limit and the per-run ceiling from part 17, with the `at_limit` action stated. - **Open requests.** The routed requests from part 16 that a person has not yet answered. This is the list an operator reads first each morning. - **Review points.** A schedule (weekly for a new agent, monthly once it is stable) and a few thresholds that trigger a review early. | Review trigger | Reading | Usual cause | | --- | --- | --- | | Denials rise week over week | `requests_by_answer` | A tool or a task changed and the rules did not | | Run cost p99 leaves its band | `cost_per_run_p99` | A loop, often on a failing check or a stale index | | Routed requests wait longer | `time_to_answer_p95` | The named approver changed roles, or the rule routes too much | | Tool use ratio falls | `tool_use_ratio` | The work drifted away from the equipment | | The operator changes | Inventory | Nobody has re-read the mandate since | ### The operator's week Put together, the operating job is small and regular. Daily, answer what is waiting. Weekly, read the spend and the denials for each agent, and review the runs that tripped a threshold. Monthly, re-read each mandate with its three owners, retire the agents nobody can justify, and turn repeated approvals into narrower rules. The field kit has this as a one-page checklist. None of it requires reading transcripts, which is the test of whether parts 14 to 19 are in place. > **Where Oxagen fits** > > Oxagen is where that week happens. Every agent enrolled in Oxagen is on the Runs page with its mandate, its open requests, its spend, and its last run. An operator answers routed requests from the Access page, and an ongoing agent shows its authority, its open requests, and its spend this month with no finish line. Oxagen treats completion checks as an optional control for bounded tasks. They are not the definition of the product, and ongoing work does not need them. > **Do this week** > > 1. **Sort your agents into the two columns.** Use the inventory from part 14. Anything you cannot place is doing both jobs and should be split. > 2. **Write completion checks for one bounded task type.** Start with the one that reaches review most often. Three checks are enough: a command, a diff condition, and a human decision if the change ships. > 3. **Fix the checks before the run.** Hash the file when the run opens and record the digest as the first row. Reject a run whose checks changed midway. > 4. **Set two review triggers for one ongoing agent.** Pick from the table. Send the alert to the operator by name. > 5. **Put the monthly mandate review on a calendar.** Thirty minutes, three owners, one agent at a time. > **Measure it** > > - `checks_held_first_attempt`: bounded runs whose checks held without a retry, over all bounded runs. This is `first_attempt_rate` from part 13, against criteria the run could not edit. > - `runs_closed_without_checks`: bounded runs that ended with no completion criteria at all. Each is a verdict by assertion. > - `open_requests_age`: the oldest unanswered routed request, per operator. > - `days_since_mandate_review`: per agent. Over 90 is a review that is due. > **Takeaway** > > Define completion when the work has an endpoint, and fix the checks before the run starts. When the work continues, do not invent a finish line. Manage its authority, its requests, and its spend on a schedule. --- # Better systems, not better models > You rent the model. You own the system around it, and the way you operate it. Part 21: The endgame. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/better-systems-not-better-models A new model ships and your team spends a week deciding whether to switch. The benchmark says it is better. Whether it is better for you depends on things the benchmark did not measure: how much context you send, how you verify a patch, what each run costs, and what the agent is allowed to touch. If those are engineered, switching is a configuration change and a week of comparing scorecards. If they are not, switching means re-tuning prompts and hoping. ### Where the advantage moved Capability gaps between frontier models have tended to close within months. Prices per token have fallen, and open-weight models keep climbing the same benchmarks. Whatever you gain from using the strongest model today, a competitor can rent with one configuration change. Model choice is becoming procurement. What compounds is the work in this book: - **Retrieval.** Parsed prompts, code graphs, schema indexes, and test contracts (parts 3, 4, 8, and 10). Each point of precision carries over to every model you run later. - **Orchestration.** A deterministic skeleton with the model inside the nodes (part 9), and budgets and circuit breakers as primitives. - **Memory.** Tiered, indexed, and paged by policy (parts 5 and 11), so knowledge accumulates across runs. - **Tooling.** Interfaces that return sliced signal (part 2), assigned per agent (part 18). - **Measurement.** The part 13 scorecard on every run, with cost attributed to the agent, the run, and the person (part 17). - **Authority.** An identity per agent, a mandate with four owned clauses, and a rule that answers each request (parts 14 to 16). - **The record.** One ordered account per run that another person can read (part 19). ### Two halves of one argument Parts 1 to 13 took decisions away from the model because code makes them cheaper and the same way each time: what to retrieve, what to keep, what runs next. Parts 14 to 20 took a second set of decisions away from the model for a different reason. What an agent may do, what it may spend, and whether it is finished are not the agent's decisions to make. They belong to the teams accountable for it. Both halves use one method. Make the decision in the system, before the model is involved, and keep it where a person can read it. The asymmetry is what makes this a durable strategy. Model improvements reach everyone at once, your competitors included. System improvements are yours. They encode your repository, your schema lineage, your test contracts, your rules, and your history of runs. They also carry over: when a cheaper or stronger model ships, a system built this way swaps it in and keeps everything else. > **Evidence · The numbers from this book, side by side** > > A fixed pipeline solved SWE-bench Lite issues at $0.34 each while agent-based systems of the same period spent roughly ten times as much for comparable results.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") On Anthropic's internal evaluations, token usage alone explained about 80% of performance variance.[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") Evaluated on cost and accuracy together, simple baselines matched elaborate agent architectures at a fraction of the price.[2](https://macanderson.com/manual/sources#r2 "Kapoor, Stroebl, Siegel, Nadgir & Narayanan (Princeton). \"AI Agents That Matter.\" TMLR 2025. arXiv:2407.01502") Each result points at the system, not at the model inside it. > > Xia et al., FSE 2025 · Anthropic engineering, 2025 · Kapoor et al., TMLR 2025 ### A precedent Databases stopped competing on raw storage and started competing on query planners. Networks moved from bandwidth to protocols. Compute moved from clock speed to scheduling. In each case the commodity layer kept improving, and the value moved to the systems that used it most efficiently and could account for what they did. Inference looks to be following the same curve. > **Counterweight · Models still matter** > > None of this says the model is irrelevant. A stronger model raises the ceiling of what any system can do, and some tasks are out of reach below a capability threshold. The claim is narrower. Between two teams with the same model, the published results above favor the one with the better system, and that team can also say what its agents did. > **Do this week** > > 1. **Score yourself.** Take the self-assessment in "How to use this book" again. Compare it with your first pass. > 2. **Pick one part from each half.** One from parts 2 to 12 that lowers tokens per task, and one from parts 14 to 20 that answers a question someone outside your team has asked. > 3. **Make the model a configuration value.** If switching models means editing prompts, list the prompts. Each one is a decision that belongs in the system. > 4. **Book the 30, 60, and 90 day reviews.** The field kit has the plan. Put the three dates on a calendar now. > **Measure it** > > - `model_swap_cost`: engineer-days to evaluate and adopt a new model. It falls as decisions move out of prompts. > - `cost_per_completed_task` and `retrieval_precision` from part 13, tracked quarter over quarter. These two show whether the system is improving independent of the model. > - `time_to_answer_review` from part 19. It shows whether you can account for what the system did. > **Takeaway** > > Models are rented. Systems are owned. Build the layer that makes any model cheaper and more precise, and operate it so that you can say who did what, under which rule, at what cost. --- # Closing thought Field manual. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/closing-thought Intelligence is expensive. Determinism is cheap. Spend the first only where the second cannot do the job. Twenty-one parts, one discipline. Parse before you prompt. Index before you search. Traverse before you reason. Validate before you accept. Decide authority before the agent asks. Record what happened in a form someone else can read. Measure all of it. The model is the most capable component in the stack and the most expensive one to call. Treat each call the way you treat any scarce resource: schedule it deliberately, feed it only what it needs, and account for what it did. If this book changes one habit, make it this one. Before each model call your system makes, ask what part of it code could have done. Before each action your agent takes, ask who decided it may. In both cases the answer should already be written down. --- # Field kit > Templates and worksheets to copy. Print this section and fill it in with the people who own each line. Field manual. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/field-kit Every template here is referenced from a part of the book. They are starting points. Change the fields to match your systems, and keep the owners. Sheet 1 · Agent inventory (part 14) One row per agent that can write to a system. A row without an operator is the first thing to fix. | Agent id | Purpose, one sentence | Operator | Harness | Bounded or ongoing | Credentials it can reach | Shared with | | --- | --- | --- | --- | --- | --- | --- | | | | | | | | | | | | | | | | | | | | | | | | | Sheet 2 · Mandate (part 15) One file per agent. Each clause has one owner who can change it by pull request. ``` agent: operator: boundary: "Governs actions routed through ____. Does not govern ____." access: # owner: acts_as: may_request: - system: scope: actions: [] budget_and_rules: # owner: monthly_limit_usd: per_run_ceiling_usd: at_limit: # hold_and_notify | route_to_operator | continue_and_flag rules: [] equipment: # owner: tools: [] skills: [] context_scope: [] tool_budget: { max_tools: , max_tokens: } record: covers: retention_days: last_reviewed: # date, and the three owners who were present ``` Sheet 3 · Decision rule (part 16) ``` id: owner: when: action: resource: effect: # allow | deny | route route_to: # a role or a named person, when effect is route timeout: on_timeout: # deny | allow reason: "" # written as the agent's next step # At the top of the rules directory: default_when_no_rule_matches: # deny | identity_alone ``` | Write action | Calls per week | Answer you want | Routed to | Rule id | | --- | --- | --- | --- | --- | | | | | | | | | | | | | | | | | | | Sheet 4 · Scoped rule object (part 6) ``` { "rule": "Money is integer cents. No float arithmetic on amounts.", "applies_to": "billing/**", "verified_by": "lint:no-float-money", "since": "a1b2c3", "owner": "payments-platform" } ``` A rule with no `verified_by` is a preference. Keep it in prose, or write the check. Sheet 5 · Domain dossier (part 5) ``` domain: billing source_sha: # the commit this was generated from entry_points: [] # routes, jobs, and consumers core_types: [] # with one line each invariants: [] # each links to the check that holds it owned_tables: [] covering_tests: [] may_call: [] # other domains must_not_call: [] recent_decisions: [] # ADR ids budget_tokens: 300 ``` Sheet 6 · Cost row (part 17) ``` person, agent, run, turn, step, model, input_tokens, cached_tokens, output_tokens, price_version, cost_usd, cost_basis, # measured | reported | estimated rule_id, started_at ``` Sheet 7 · Production scorecard (part 13) | Metric | This quarter | Last quarter | Owner | | --- | --- | --- | --- | | `cost_per_completed_task` | | | | | `tokens_consumed` per task, by phase | | | | | `retrieval_precision` | | | | | `first_attempt_rate` | | | | | `time_to_green` | | | | | `files_touched` per task | | | | | Variance across 10 repeats of one task | | | | | `attributed_spend_ratio` | | | | | `agents_with_mandate` | | | | | `time_to_answer_review` | | | | Sheet 8 · Completion checks for a bounded task (part 20) ``` task: "" checks: - run: "" # a command that must exit 0 - diff: { only_paths: [] } - file: { exists: "" } - human: { role: , question: "" } on_fail: { retries: , then: route_to_operator } # Hash this file when the run opens. Record the digest as the first row. ``` Sheet 9 · The operator's week (part 20) | When | Do | Read | | --- | --- | --- | | Daily | Answer routed requests. Look at anything held at a budget limit. | `open_requests_age`, `budget_holds` | | Weekly | Read spend and denials per agent. Review runs that tripped a threshold. | `requests_by_answer`, `cost_per_run_p99` | | Monthly | Re-read each mandate with its owners. Turn repeated approvals into narrower rules. Retire agents nobody can justify. | `days_since_mandate_review`, `tool_use_ratio` | | Quarterly | Run a tabletop review with someone outside the team. Fill in the scorecard. | `time_to_answer_review`, sheet 7 | Sheet 10 · A 30, 60, and 90 day plan | By day | Engineering half | Operating half | You can now answer | | --- | --- | --- | --- | | 30 | Per-step token and cost logging. Trace slicing in one agent. | Inventory complete. Every agent has an operator and an id. | Where do the tokens go, and who answers for each agent? | | 60 | Prompt parsing, a symbol index, a schema slice, and a test lookup. | One mandate per agent. Rules for the top ten write actions. A stated default. | What is each agent allowed to do, and which rule said so? | | 90 | A workflow skeleton with budgets. Per-agent tool assignment. The scorecard filled in twice. | Attributed spend. A chained run record. Completion checks for one bounded task type. Review dates booked. | What did it cost, per agent and per person, and can someone else verify what happened? | One task through the whole system To see how the parts connect, follow one prompt. "The `POST /api/v2/refunds` endpoint throws `DecimalConversionError` in `billing/processors.py` after #4821 merged." 1. **The run opens.** The record stores the agent, the operator, the initiator, and the mandate version (parts 14, 15, and 19). The completion checks are hashed (part 20). 2. **Code parses the prompt** into a route, an exception class, a file path, and an issue, and resolves each against an index (part 3). 3. **The graph expands the anchors** one hop. The schema index adds the two tables. The test lookup adds three covering tests. The billing dossier adds its invariants. Everything is cut to an 8,000 token budget (parts 4, 5, 7, 8, and 10). 4. **The workflow engine runs the failing test** under a tracer, and the slicer returns about 1,200 tokens of frames and locals (parts 2 and 9). 5. **The model writes the patch.** This is the first step where judgment is needed, and the first large model call. 6. **The runner executes the checks.** On a failure, the sliced result feeds one more attempt, and working memory drops the superseded one (parts 10 and 11). 7. **The agent requests `github.push`** to a release branch. The rule routes it to the release owner, who approves. The push goes through a mediated connection (part 16). 8. **Every step wrote a cost row** with the person, the agent, and the run (part 17). The run closes, the checks held, and the chain is signed (parts 19 and 20). 9. **The scorecard updates:** tokens by phase, retrieval precision, attempts, time to green, and cost (part 13). The model made one kind of decision in that sequence: what the patch should be. Code, rules, and people made the rest, and each of those decisions is in the record. --- # Glossary > The words this book uses in a fixed sense. Field manual. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/glossary | Term | Meaning here | Part | | --- | --- | --- | | Agent | A program that calls a model in a loop and takes actions through tools. In parts 14 to 20, a registered principal with its own identity. | 14 | | Agent-computer interface | The commands and output formats an agent works through. SWE-agent's term. | 2 | | Anchor | An entity parsed from a prompt and resolved against an index: a path, a symbol, a route, an issue. | 3 | | Bounded task | Work with an end state you can describe before it starts. | 20 | | Check | One completion criterion for a bounded task: a command, a file or diff condition, or a human decision. | 20 | | Clause | One of the four parts of a mandate, owned by one team. | 15 | | Connection | A system reached on the agent's behalf. For a mediated connection, the platform holds the credential and the agent does not receive it. | 16 | | Context utilization | Tokens the final output depended on, over tokens supplied. | 7, 13 | | Cost basis | How a cost figure was obtained: measured, reported, or estimated. | 17 | | Domain dossier | A short, generated briefing on one area of a codebase, stamped with its source commit. | 5 | | Equipment | The tools, skills, and permitted business context assigned to an agent through its mandate. | 18 | | Governed action | A call that passed through the layer that checks the mandate and writes the record. | 15, 19 | | Identity | The agent's own principal, with an accountable operator. Separate from any connection credential. | 14 | | Initiator | The person who started the task a run is working on. | 14, 17 | | Mandate | The one object an agent works under: access, budget and rules, equipment, and the record. | 15 | | Observe mode | Calls are recorded and not blocked. Recorded, not enforced. | 15 | | Ongoing responsibility | Work with no last step, managed through authority, budget, and review points. | 20 | | Operator | The person accountable for an agent and its runs. | 14 | | Record | The ordered, typed account of a run, kept for a reader who was not there. | 19 | | Request | An agent asking for a system, scope, or action at the moment of use. | 16 | | Retrieval precision | Of the context supplied, the share the final patch depended on. | 13 | | Rule | What answers a request: allowed, denied, or routed to a person. | 16 | | Run | One agent, under one operator, on one task. | 13 | | Skill | Packaged instructions for one kind of task, loaded when that task comes up. | 18 | | Step | One model call or one tool call. | 17 | | Turn | One prompt through to the point the agent stops. | 17 | | Working memory | What is in the context window for the current step, kept small by policy. | 11 | --- # About the author Field manual. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/about-the-author ![Mac Anderson, founder and CEO of Oxagen Inc.](/portrait/mac-720.jpg) ### Mac Anderson Founder and CEO, Oxagen Inc. · Los Angeles, California **Oxagen** is workforce management for autonomous agents: give each agent an identity, set its authority and budget, equip it with tools and skills, and review what it did and what its operators spent, through a shared agent control plane. Oxagen does not run agents. It decides what the agents you run may do for the actions routed through it, hands them what they need, and keeps the record. Mac has built coding agents since the first release of OpenAI's Codex. The first thirteen parts of this book come from that work and from the published research it cites. The operating parts come from a later lesson: once agents work, the questions change from how well they code to who answers for them. [oxagen.sh](https://oxagen.sh) --- # Sources Field manual. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/sources Every empirical claim in this book traces to one of the sources below. Read the primary literature. It is better than any summary of it, including this one. Sources 1 to 21 support parts 1 to 13. Sources 22 to 27 were added for the operating parts in the second edition. 1. **Xia, Deng, Dunn & Zhang**. "Agentless: Demystifying LLM-based Software Engineering Agents." FSE 2025. [arXiv:2407.01489](https://arxiv.org/abs/2407.01489) · [github.com/OpenAutoCoder/Agentless](https://github.com/OpenAutoCoder/Agentless) 2. **Kapoor, Stroebl, Siegel, Nadgir & Narayanan** (Princeton). "AI Agents That Matter." TMLR 2025. [arXiv:2407.01502](https://arxiv.org/abs/2407.01502) 3. **Anthropic Engineering**. "How we built our multi-agent research system." June 2025. [anthropic.com/engineering/multi-agent-research-system](https://www.anthropic.com/engineering/multi-agent-research-system) 4. **Cognition (Walden Yan)**. "Don't Build Multi-Agents." June 2025. [cognition.ai/blog/dont-build-multi-agents](https://cognition.ai/blog/dont-build-multi-agents) 5. **Cemri et al.** (UC Berkeley). "Why Do Multi-Agent LLM Systems Fail?" (MAST). NeurIPS 2025 Datasets & Benchmarks. [arXiv:2503.13657](https://arxiv.org/abs/2503.13657) 6. **Liu, Lin, Hewitt, Paranjape, Bevilacqua, Petroni & Liang**. "Lost in the Middle: How Language Models Use Long Contexts." TACL 2024. [arXiv:2307.03172](https://arxiv.org/abs/2307.03172) 7. **Hong, Troynikov & Huber** (Chroma). "Context Rot: How Increasing Input Tokens Impacts LLM Performance." July 2025. [research.trychroma.com/context-rot](https://research.trychroma.com/context-rot) 8. **Modarressi et al.** (LMU Munich & Adobe). "NoLiMa: Long-Context Evaluation Beyond Literal Matching." 2025. [arXiv:2502.05167](https://arxiv.org/abs/2502.05167) 9. **Cuconasu et al.**. "The Power of Noise: Redefining Retrieval for RAG Systems." SIGIR 2024. [arXiv:2401.14887](https://arxiv.org/abs/2401.14887) 10. **Edge et al.** (Microsoft Research). "From Local to Global: A Graph RAG Approach to Query-Focused Summarization." 2024. [arXiv:2404.16130](https://arxiv.org/abs/2404.16130) 11. **Microsoft Research**. "GraphRAG: New tool for complex data discovery." July 2024. [microsoft.com/research blog](https://www.microsoft.com/en-us/research/blog/graphrag-new-tool-for-complex-data-discovery-now-on-github/) 12. **Aider (Paul Gauthier)**. "Building a better repository map with tree-sitter" & repo-map docs. 2023 onward. [aider.chat/2023/10/22/repomap.html](https://aider.chat/2023/10/22/repomap.html) · [aider.chat/docs/repomap.html](https://aider.chat/docs/repomap.html) 13. **Zhang, Ruan, Fan & Roychoudhury** (NUS). "AutoCodeRover: Autonomous Program Improvement." ISSTA 2024. [arXiv:2404.05427](https://arxiv.org/abs/2404.05427) 14. **Yang, Jimenez et al.** (Princeton). "SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering." NeurIPS 2024. [arXiv:2405.15793](https://arxiv.org/abs/2405.15793) 15. **Jimenez et al.**. "SWE-bench: Can Language Models Resolve Real-World GitHub Issues?" ICLR 2024. [arXiv:2310.06770](https://arxiv.org/abs/2310.06770) 16. **Anthropic**. "Prompt caching with Claude" (up to 90% cost and 85% latency reduction, and a cache read costs about 0.1 times the input price). 2024. [anthropic.com/news/prompt-caching](https://www.anthropic.com/news/prompt-caching) 17. **Packer et al.** (UC Berkeley). "MemGPT: Towards LLMs as Operating Systems." 2023. [arXiv:2310.08560](https://arxiv.org/abs/2310.08560) 18. **Shannon, C.E.** "A Mathematical Theory of Communication." Bell System Technical Journal, 1948. The original framework for signal and noise 19. **Anthropic Engineering**. "Effective context engineering for AI agents." 2025. [anthropic.com/engineering](https://www.anthropic.com/engineering/effective-context-engineering-for-ai-agents) 20. **Anthropic Engineering**. "Writing effective tools for agents." 2025. [anthropic.com/engineering](https://www.anthropic.com/engineering/writing-tools-for-agents) 21. **Zhu et al.**. "When to use Graphs in RAG: A Comprehensive Analysis for Graph Retrieval-Augmented Generation." 2025. [arXiv:2506.05690](https://arxiv.org/abs/2506.05690) 22. **Rose, Borchert, Mitchell and Connelly** (NIST). "Zero Trust Architecture." NIST Special Publication 800-207, 2020. [doi.org/10.6028/NIST.SP.800-207](https://doi.org/10.6028/NIST.SP.800-207) 23. **OWASP**. "Top 10 for LLM Applications 2025," LLM06: Excessive Agency. [genai.owasp.org](https://genai.owasp.org/llmrisk/llm062025-excessive-agency/) 24. **FinOps Foundation**. "FinOps Framework," the Allocation capability. [finops.org/framework](https://www.finops.org/framework/capabilities/allocation/) 25. **Model Context Protocol**. Specification: servers, tools, and schemas. [modelcontextprotocol.io/specification](https://modelcontextprotocol.io/specification) 26. **Rundgren, Jordan and Erdtman**. "JSON Canonicalization Scheme (JCS)." RFC 8785, 2020. [rfc-editor.org/rfc/rfc8785](https://www.rfc-editor.org/rfc/rfc8785) 27. **Beyer, Jones, Petoff and Murphy (eds.)**. "Site Reliability Engineering," chapter 15: Postmortem culture, learning from failure. O'Reilly, 2016. [sre.google/sre-book](https://sre.google/sre-book/postmortem-culture/) **How to use this book.** Each part stands alone, and the whole reads in order. Share the file, or excerpt a part with its evidence block and takeaway intact. The numbers cited reflect the referenced publications. Check them against the primary source before you quote them, because sources change. Figures in worked examples are illustrative and are labelled as such. © 2026 Mac Anderson. Published with Oxagen Inc., Los Angeles. Second edition. Each edition is a single HTML file with no tracking. The field manual prints on US Letter, and the reader exports any part as a PDF. --- # Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces > Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week. Published 2026-10-09 by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models Your organization runs coding agents every day. Each session produces a record: the task, every file the agent read, every command it ran, every edit it made, and whether the work held up. Most teams throw that record away the moment the session ends. This book argues that the record is the most valuable asset the session produces, and that a team which keeps it can, within a year, fine-tune an open-weight model that does its own work as well as the rented frontier model does today. The book is long because the argument has many parts and each part has evidence behind it. It is written for the engineer who will build the pipeline and for the person who has to approve it. Chapters 1 and 2 make the case. Chapters 3 to 6 are about the oracle, the program that decides whether a trace counts. Chapters 7 to 9 are about collecting traces from the agent harness you already use, with a working Claude Code plugin you can install today. Chapters 10 to 12 are about turning traces into training data and how much you need. Chapters 13 to 16 are about the delivery pipeline that ships new weights on a schedule. Chapter 17 lists the earlier systems that tried the same idea, so you can see what held up. Three words carry most of the weight, so here they are up front. - A **trace** is the full record of one agent session: prompts, tool calls, tool results, edits, and the final state of the repository. - An **oracle** is a program that reads the result of a session and returns one verdict, PASS or FAIL, without the agent's help. In this book the oracle is a hidden test that the agent never sees, run in a container the agent cannot reach. - A **flip** is the event the whole pipeline is built around. The oracle said FAIL before the agent started. The oracle says PASS after the agent finished. The trace between those two verdicts is a worked example of the task being done right, graded by something the model did not control. > **The reference plugin** > > The plugin this book describes is public at [github.com/macanderson/oracle-flip](https://github.com/macanderson/oracle-flip). It installs into Claude Code in two commands, writes every tool call to a trace file, and runs an oracle before it lets the agent say the work is done. Chapter 8 walks through every file. ## The argument in one page Rented tokens are a cost that never ends and a capability you never own. The price per token has fallen fast, and it will keep falling, but the thing you are paying for is a model that knows nothing about your systems on the day you start and nothing more on the day you stop. Every improvement happens in the vendor's weights, under the vendor's terms, on the vendor's schedule. Open-weight models are now close enough to the frontier that the gap no longer decides the outcome. In late 2024, the best open model trailed the best closed model by about a year.[1](#user-content-fn-epoch-open) In the year to early 2025, the gap between the best open and closed models on a public human-preference leaderboard narrowed from about 8 percent to under 2 percent.[2](#user-content-fn-ai-index-2025) Models released under permissive licenses by DeepSeek, Alibaba, Meta, Mistral, and in August 2025 by OpenAI itself, run on hardware you can rent by the hour or buy outright.[3](#user-content-fn-gpt-oss) A general open model is a floor. What raises the floor on your tasks is training on your tasks. Traces are that training data, but only when something outside the model grades them. The research on models grading their own work is consistent and unkind. Left to judge themselves, models prefer their own output, cannot reliably find their own errors, and often get worse when they try to self-correct without an outside signal.[4](#user-content-fn-huang-self-correct)[5](#user-content-fn-panickssery) The methods that do produce lasting gains all share one feature: a verifier the model does not control, such as an answer key, a compiler, or a test.[6](#user-content-fn-star)[7](#user-content-fn-deepseek-r1) For software, that verifier already exists. It is the test suite, the type checker, the build, and the hidden test you write before the agent starts. The oracle has to be deterministic, air-gapped, and silent. Deterministic, because a flaky verdict is noise in the training set, and noise at this stage is expensive to remove later. Air-gapped, because frontier models have been caught editing tests, stubbing functions, and exiting early to make a check pass, and a reward signal the agent can reach is a signal it will eventually reach.[8](#user-content-fn-baker-monitoring)[9](#user-content-fn-denison-subterfuge) Silent, meaning it says only PASS or FAIL, because a verdict that explains itself leaks the hidden test into the agent's context, and a test the agent can see is a test it can overfit. The theory for this is older than language models: a holdout that answers with low information stays valid under many adaptive queries, and one that answers with high information does not.[10](#user-content-fn-ladder)[11](#user-content-fn-reusable-holdout) You do not have to build the agent. Every serious coding harness now exposes hooks: programs that run before and after each tool call, and when the agent tries to stop. A hook that appends each event to a file is a trace collector. A hook that runs the oracle on Stop and refuses to let the agent finish on FAIL is the flip detector. The plugin in Chapter 8 is about 400 lines. You need fewer flips than you think to start, and more than you think to finish. Published results put the first large gains at a few hundred verified trajectories on a 32-billion-parameter open model, and the best open results at several thousand.[12](#user-content-fn-swe-gym)[13](#user-content-fn-swe-smith)[14](#user-content-fn-skywork-swe) A team of fifty engineers running agents daily reaches the first number in a week and the second in a quarter. The chapters on volume give the arithmetic. The pipeline is ordinary continuous delivery with weights as the artifact. Ingest, redact, validate, version, train, evaluate against held-out flips the model never saw, gate, package, canary in shadow mode against the live oracle, promote, roll back. None of these stages is new. What is new is that the artifact learns. ## Chapter 1: Rented tokens A team that uses a frontier model through an API is renting. The rent has three parts, and only the first shows up on the invoice. The first part is the token bill. It is real and it is falling. Epoch AI measured the price of reaching a fixed level of performance on a set of benchmarks and found it falling by somewhere between 9 and 900 times a year, depending on the task.[1](#user-content-fn-epoch-prices) That is good news for anyone who pays the bill, and it is the number vendors point to when someone asks why a team should not train its own model. The number is also beside the point. The question is not whether tokens will get cheaper. They will. The question is what the team owns at the end of the year. The vendors in question are the frontier labs: Anthropic, OpenAI, Google, and the others whose models you reach only through an API and whose weights you never see. A black box is the right name for the product, whatever you think of the company. The second part of the rent is the capability you never keep. When a vendor ships a better model, your agents get better. When the vendor changes the model, deprecates it, raises the price, or changes the terms, your agents change with it. Nothing the model learned about your systems during a year of sessions stays with you, because the model learned nothing. It cannot. Inference does not update weights. Every session starts from the same checkpoint the vendor shipped, and every piece of context your team has written to steer it, every CLAUDE.md and every rule file, is a workaround for a model that does not know your code. The third part is the data you discard. A session produces a record. The record shows which files the agent opened to understand a feature, which commands it ran to check its work, what it tried that failed, and what finally passed. If the agent was working in your repository on your task against your tests, that record is a worked example of your work being done. Nobody else has it. Nobody else can produce it. And in most teams, it is gone when the terminal closes. > **Where the value of a session goes today** > > 1. **Task**: An engineer or a work queue dispatches a task to a coding agent > 2. **Session**: The agent reads, edits, runs commands, and reports done > 3. **Tokens billed**: The vendor bills every input and output token > 4. **Record discarded**: The transcript is deleted or buried in a local cache > > *The two emphasized steps are the rent. The money leaves and the record leaves. The model learned nothing, and neither did your organization's data.* ## What the open-weight trend changes The case for keeping traces rests on one bet: that open-weight models will stay close enough to the frontier that a model trained on your traces beats a rented model on your tasks. Three measurements say the bet is reasonable. Epoch AI compared the best open-weight model to the best closed model across benchmarks and found the open model trailing by about a year, measured by the date at which a closed model first reached the same score.[2](#user-content-fn-epoch-open) A year behind the frontier on general benchmarks is not a year behind on your codebase, because the frontier model is also starting from zero on your codebase. Stanford's AI Index reported that the gap between the best open and best closed models on the Chatbot Arena leaderboard narrowed from 8.0 percent to 1.7 percent in one year.[3](#user-content-fn-ai-index-2025) Chatbot Arena measures human preference on chat, not software engineering, but the direction holds across benchmarks the Index tracks. On the benchmark that matters most for this book, SWE-bench Verified, open-weight models went from under 10 percent to above 40 percent in eighteen months. Meta's SWE-RL took a 70-billion-parameter Llama to 41.0 percent.[4](#user-content-fn-swe-rl) Mistral's Devstral, a 24-billion-parameter model that runs on a single workstation card, reported 46.8 percent.[5](#user-content-fn-devstral) The SWE-smith team took a 32-billion-parameter Qwen to 40.2 percent with about five thousand trajectories.[6](#user-content-fn-swe-smith) The Skywork team took the same base model from 6.4 percent to 38.0 percent with about eight thousand, and found the gain still growing with each doubling of the data.[7](#user-content-fn-skywork-swe) These are not the frontier numbers. The frontier was above 70 percent at the time. They are the numbers that a team can run, inspect, and fine-tune. > **Open-weight models on SWE-bench Verified** > > | Model and method | Resolved | > | --- | --- | > | Qwen2.5-Coder-32B, no agent training (SWE-Gym baseline) | 7% | > | Same model after fine-tuning on 491 SWE-Gym trajectories | 20.6% | > | Same model with a verifier and 16 samples per task | 32% | > | Skywork-SWE-32B, 8,209 trajectories | 38% | > | SWE-agent-LM-32B, 5,016 SWE-smith trajectories | 40.2% | > | Llama3-SWE-RL-70B, reinforcement learning | 41% | > | Devstral-Small-2505, 24B | 46.8% | > | DeepSWE-Preview, Qwen3-32B, RL with test-time scaling | 59% | > > *Sources: Pan et al. (2024), Yang et al. (2025), Wei et al. (2025), Mistral (2025), Agentica and Together AI (2025). Each row is a different base model and training recipe, so the bars show the range open models reached, not a controlled comparison.* Two things about these numbers matter for the argument. First, the jump from 7.0 to 20.6 percent came from fine-tuning on 491 trajectories.[8](#user-content-fn-swe-gym) That is not a large dataset. It is about a week of sessions for a mid-sized team. The further jump to 32.0 percent came from training a verifier on the same trajectories and letting it pick the best of 16 attempts, which is a preview of Chapter 9. Second, every model in the table was trained on public repositories solving public issues. None of them had seen the training team's own code. A model trained on your traces starts from these numbers and climbs on your distribution. ## Licenses that permit it The models in the table ship under licenses that allow fine-tuning and commercial use. DeepSeek-R1 and its distilled variants are MIT licensed.[9](#user-content-fn-deepseek-r1) The Qwen2.5 and Qwen3 families are Apache 2.0, with a small number of size variants under a Qwen license.[10](#user-content-fn-qwen3) The Llama 3 family uses Meta's community license, which permits commercial use below a very large monthly-user threshold.[11](#user-content-fn-llama3) OpenAI's gpt-oss models are Apache 2.0.[12](#user-content-fn-gpt-oss) The point of listing them is not legal advice. It is that the permission to do what this book describes is ordinary and granted. ## The asset that compounds Consider two teams of the same size doing the same work for one year. Team A uses a rented frontier model. It spends on tokens, writes rule files, and gets better at prompting. At the end of the year it has a set of rule files, a bill, and a vendor relationship. If the vendor's next model is worse at the team's tasks, the team has no recourse. If a competitor uses the same vendor, the competitor has the same model. Team B uses the same rented model, and also installs a trace collector and an oracle. It spends the same on tokens. Each session that flips an oracle from FAIL to PASS is saved as a verified trajectory. At the end of the year, Team B has thousands of worked examples of its own work being done correctly, each graded by a test the model did not write. It has fine-tuned an open model on them three or four times and has a model that resolves its own tasks at a rate the rented model cannot match on the same distribution, running on hardware it controls, at a marginal cost per token that is a fraction of the rent. If the vendor changes terms, Team B's model does not change. If a competitor wants the same model, the competitor needs Team B's traces. The cost of being Team B instead of Team A, for the first several months, is close to zero. The collector is a plugin. The oracle is a test. Storage is cheap. The training runs come later and are cheaper than one month of a mid-sized team's token bill. The only thing Team A has to do to become Team B is to stop deleting the record. ## Limits of the claim It does not claim that a fine-tuned 32-billion-parameter model will match the best frontier model on every task next quarter. It will not. It does not claim that the open-weight gap will close to zero. It may not. It does not claim that collecting traces is free of risk; Chapter 16 is about secrets, memorization, and the ways a pipeline like this goes wrong. It claims that the traces are the asset, that the oracle is what makes them an asset, and that a team which starts collecting today will be in a different position in a year from a team that does not. The rest of the book is about how to do it carefully. ## Chapter 2: The trace A trace is the record of one agent session. The word is borrowed from systems tracing, where a trace is the record of one request as it passes through many services. The analogy is close. An agent session passes through many tool calls, and the trace records all of them in order with their inputs and outputs. ## Anatomy of a trace A coding-agent session has a regular structure. The agent receives a task. It takes a turn: it thinks, then it either calls a tool or writes a message. A tool call runs and returns a result. The agent takes another turn. This continues until the agent decides it is done, or a person stops it, or a limit is reached. The trace records each of these events: 1. The **task** as the agent received it, including the system prompt, the rule files loaded into context, and the user's instruction. 2. Each **tool call**: the tool name, the arguments the agent passed, the result the tool returned, and how long it took. 3. Each **message** the agent wrote to the person, including the final message where it says the work is done. 4. The **repository state** at the start and end: the base commit, the final diff, and the list of files touched. 5. The **oracle verdicts**: the verdict before the agent started, and the verdict at each point the agent tried to stop. 6. **Metadata**: the model, the harness and its version, the session id, timestamps, token counts, and cost. Of these, the tool calls carry the learning signal. They show the model how an expert in this codebase finds the relevant file, which command verifies a change, what a failing test looks like here, and how a fix is shaped. Training on them teaches the model to take those actions on this codebase. The final diff alone would not teach that. A diff is an answer; a trace is a worked solution. > **The records in a trace and how they relate** > > - Session has one Task: System prompt, rule files, user instruction; Base commit of the repository > - Session has many Turn: One model response; Ends in a tool call or a message > - Turn has zero or one Tool call: Name, input, output, duration; Output redacted before storage > - Session has many Oracle verdict: PASS or FAIL and nothing else; One at the start, one per stop attempt > - Session has one Outcome: flipped, no flip, or aborted; Final diff and files touched > > *The oracle verdicts are kept beside the turns, not inside them. The agent never sees the verdict's reason, so the reason is not in any turn.* ## The episode and the flip In reinforcement learning, an episode is one complete attempt at a task from start to a terminal state. A trace is an episode. What makes an episode useful for training is a reward: a number that says how well it went. For most of what organizations do, there is no clean reward. For software, there is one, and it has a name in the benchmark literature. SWE-bench, the benchmark most coding-agent papers report against, grades a candidate patch with two sets of tests.[1](#user-content-fn-swe-bench) The FAIL\_TO\_PASS tests fail on the repository before the fix and pass after the reference fix. The PASS\_TO\_PASS tests pass both before and after; they are there to catch a patch that fixes the issue by breaking something else. A patch resolves the task when every FAIL\_TO\_PASS test passes and every PASS\_TO\_PASS test still passes. That is the flip. Before the agent starts, the oracle runs the hidden tests and records FAIL. After the agent says it is done, the oracle applies the agent's diff to a clean copy, runs the same hidden tests, and records PASS or FAIL. A session whose verdict goes from FAIL to PASS, with no regression in the tests that already passed, has flipped. The trace of that session is a verified trajectory. A session that ends on FAIL is also kept, because a failed attempt beside a successful one on the same task is a preference pair, and preference pairs are training data too. Chapter 10 covers that. The flip is the unit of value for three reasons. It is binary, so it cannot be argued with. There is no rubric, no score from a judge model, no partial credit for a patch that "looks right." Chapter 6 is about why that matters. It is cheap, so it can be run on every session. A test suite that runs in minutes grades a session for cents. A human review of the same session would cost more than the session. It is yours. The hidden test encodes what your organization means by correct for this task. A public benchmark encodes what a benchmark author meant. ## Other records A trace is not a transcript. Harnesses keep transcripts for the person to scroll back through, and those files are useful, but they mix the agent's messages with display formatting, and they may not include the final message when the session ends.[2](#user-content-fn-claude-hooks) A trace is written by hooks as events happen, in a schema you control, with redaction applied before anything touches disk. A trace is not a log. Logs are for debugging a system. Traces are for training a model. A log can be lossy and unstructured. A trace that drops the tool output for one call has a hole in the worked example at exactly the point the model needs to learn from. A trace is not the diff. The diff is the answer. A model trained only on diffs learns to produce patches that look like your patches. It does not learn to find the file, run the test, read the failure, and try again. Pan and colleagues trained on full agent trajectories, each a sequence of tool calls and observations, and moved their model from 7.0 to 20.6 percent with 491 of them.[3](#user-content-fn-swe-gym) The actions are the data. ## The schema Appendix A gives a full schema. The shape is a JSON Lines file per session, one event per line, with a small set of event types. ```text session_start session_id, task, base_commit, model, harness, started_at tool_call turn, tool_name, tool_input, tool_output_redacted, duration_ms message turn, role, text oracle_verdict attempt, verdict (PASS|FAIL), tests_hash, ran_at session_end outcome (flipped|no_flip|aborted), final_diff_hash, ended_at ``` Two design choices in the schema deserve a sentence each. The verdict record carries a hash of the hidden test list, not the list. That lets you prove later which oracle graded a trace without putting the test names anywhere the agent's process can read. And the tool output is stored after redaction, never before. A trace store is a secret store if you let it be one, and Chapter 16 explains why that is the failure that ends programs like this. ## Why the harness should write it You could write a coding agent and have it write its own traces. Several groups have, and their agents are good. But the harness your team already uses runs thousands of sessions a week today, and it exposes hooks. A hook is a program the harness runs at a fixed point in the session: before a tool call, after a tool call, when the agent tries to stop, when the session starts and ends. Each hook receives the event as JSON on standard input. A hook that appends that JSON to a file is a trace collector, and it took the author of this book an afternoon to write. Chapter 7 covers hooks in detail. The reason to mention them here is that the most common objection to collecting traces, "we would have to build our own agent," is false. ## Chapter 3: The oracle problem Software testing has a name for the thing that decides whether a program's output is correct: the test oracle. Barr, Harman, McMinn, Shahbaz, and Yoo surveyed the field in 2015 and called the difficulty of building one "the oracle problem."[1](#user-content-fn-barr-oracle) Generating inputs to a program is easy. Knowing what the program should have done with them is hard. Every automated check, from a unit test to a type checker, is a partial answer to that problem. This book uses the word in the same sense, narrowed. An oracle here is a program that takes the state of a repository after an agent has worked on it and returns PASS or FAIL. It runs without the agent. It does not take the agent's word for anything. The quality of everything downstream, every training set and every evaluation, is bounded by the quality of the oracle, so this chapter is about what makes a good one. ## Deterministic A deterministic oracle returns the same verdict every time it runs on the same input. That sounds like a low bar. In practice, most test suites do not clear it. Luo, Hariri, Eloussi, and Marinov studied flaky tests, tests that pass and fail on the same code, across 51 open-source projects and classified the causes in 201 fixing commits. The leading causes were waiting on asynchronous work, concurrency, and dependence on test order.[2](#user-content-fn-luo-flaky) At Google, Micco reported that about 1.5 percent of all test runs gave a flaky result, that almost 16 percent of tests showed some flakiness, and that about 84 percent of transitions from pass to fail involved a flaky test.[3](#user-content-fn-micco-flaky) A flaky test in a benchmark is an annoyance. A flaky test in an oracle that labels training data is corruption: a session is recorded as a flip when the agent did nothing, or as a failure when the agent succeeded. The fixes are the fixes the testing literature has recommended for a decade, applied strictly because the stakes are higher. - Run the oracle in a fresh container from a pinned image, so the environment is identical every time. Reproducible builds give the same guarantee for the artifact under test: the same source produces the same binary, bit for bit, so a verdict is about the code and not the build machine.[4](#user-content-fn-reproducible-builds) - Pin every dependency by hash, with a local mirror, and give the container no network. A test that downloads anything is not deterministic. - Fix the clock, the random seed, the locale, and the time zone inside the container. A test that depends on the wall clock is a test that depends on when the oracle ran. - Run the hidden tests twice on the base commit before the agent starts. If the two verdicts disagree, the task is not fit for an oracle. Quarantine it. - Record the test list as a hash with the verdict, so a later change to the tests cannot be confused with a change in the code. The cost of determinism is that some tests cannot be oracles. Integration tests against a live service, tests that depend on timing, and tests that depend on data that changes are excluded. That is a loss, and Chapter 4 is partly about what you can use instead. ## Independent An oracle must be independent of the agent in two senses. It must not run in the agent's environment. The agent's working copy is under the agent's control. The agent can edit any file there, including the tests, the test configuration, the build file, and the environment variables the test runner reads. An oracle that runs `pytest` in the agent's directory is asking the agent whether it passed. Chapter 5 describes the air gap that fixes this. It must not be a model grading itself. A judge model that reads the diff and decides whether the task is done is not an oracle. It is a second opinion from a system with the same blind spots as the first. Chapter 6 reviews the evidence. ## Hidden The hidden test is the version of the oracle this book recommends. The team writes a test that fails on the current code and passes when the task is done correctly. The agent never sees it. The agent sees the task description and the existing test suite, and can write and run its own tests freely. When the agent says it is done, the oracle applies the agent's changes to a clean copy, runs the hidden test and the existing suite, and records the verdict. This is the SWE-bench construction.[5](#user-content-fn-swe-bench) The FAIL\_TO\_PASS tests are hidden from the agent during the attempt. The agent reads the issue, not the test. The construction is also how the Defects4J database of real Java bugs has been used for a decade: each bug comes with at least one test that exposes it and passes after the developer's fix.[6](#user-content-fn-defects4j) Hidden tests matter because visible tests are a weaker oracle than they look. Qi, Long, Achour, and Rinard examined patches that three automated repair systems had reported as fixing bugs, meaning the patches made the visible tests pass. Most of the patches were wrong.[7](#user-content-fn-qi-kali) They passed the tests by deleting the functionality the tests happened not to cover. Smith, Barr, Le Goues, and Brun showed the same thing with a controlled experiment and gave it a name: patch overfitting.[8](#user-content-fn-smith-cure) A patch that passes the tests the agent can see has been fitted to those tests, and nothing more has been shown. A patch that passes a test the agent could not see has been shown something. ## The flip, formally With the pieces named, the flip can be stated as a rule. Let `H` be the hidden tests for the task and `S` be the existing suite, both frozen as a list with a hash. Let `base` be the commit the agent started from and `diff` be the agent's final change. 1. On `base`, run `H` and `S` in a fresh container. Require that `H` fails and `S` passes. If `H` passes, there is no task; if `S` fails, the base is broken and the task is not fit for an oracle. Record the verdict as the baseline. 2. On `base` with `diff` applied to source paths only, run `H` and `S` in a fresh container. Record PASS if every test in `H` passes and every test in `S` passes. Record FAIL otherwise. 3. A session has flipped when the baseline is FAIL and the final verdict is PASS. Step 2 says "source paths only." The diff the agent produced may include changes to test files, test configuration, continuous-integration files, and build scripts. The oracle does not apply those. It applies the agent's changes to the code under test, then runs its own copies of the tests against them. Chapter 5 explains the allowlist that does this. > **One oracle run** > > 1. **Fresh container**: Pinned image, no network, fixed clock and seed > 2. **Clean checkout**: The base commit, from the oracle's own mirror > 3. **Apply source diff**: Only paths on the allowlist > 4. **Inject hidden tests**: From the oracle's store, never from the workspace > 5. **Run H and S**: Hidden tests and the existing suite > 6. **One bit out**: PASS or FAIL, plus a hash of the test list > > *Everything the agent could have touched is replaced before the tests run. The agent's only contribution to the oracle run is the source diff.* ## Tests are a weak oracle too Hidden tests are the best oracle most teams can build quickly. They are not a perfect one. Inozemtseva and Holmes showed that code coverage, the share of lines a test suite runs, is not strongly correlated with how many faults the suite detects once suite size is accounted for.[9](#user-content-fn-inozemtseva-coverage) A hidden test that runs a line is not a hidden test that checks it. The same study of patch overfitting that argues for hidden tests also argues for better ones. Mutation testing is the standard way to measure how good a test is. A mutation tool makes small changes to the code, such as flipping a comparison or deleting a statement, and checks whether the tests fail. A test that does not fail on a mutant is not checking the mutated behavior. The idea is from 1978 and has been used at Google at scale since at least 2018, where mutants are shown to developers during code review.[10](#user-content-fn-demillo-mutation)[11](#user-content-fn-petrovic-mutation) A hidden test that kills the mutants near the code the task touches is a stronger oracle than one that merely runs. Chapter 9 uses mutation the other way around, to manufacture tasks. The practical rule is: a hidden test must fail on the base commit for the right reason. Before you accept a task into the oracle pool, read the failure. If the test fails because of an import error or a fixture that is missing, the agent can flip it by fixing the fixture. If it fails because the feature is missing, the agent has to build the feature. ## Chapter 4: Organizational oracles The hidden test is one oracle. An organization has many. This chapter is a catalogue, ordered from the oracles that give the cleanest signal to the ones that give the noisiest, with a note on how each can be used. The ordering matters because the cleaner the oracle, the more directly its verdict can be used as a training label. The noisier the oracle, the more it belongs in a preference pair or a filter rather than as a label. ## Hard oracles A hard oracle is deterministic, runs in minutes, and returns a verdict a program can read. These can label a trace directly. **Hidden unit and integration tests.** The reference oracle, covered in Chapter 3. Written by a person or by a separate model before the task is dispatched. The strongest version is a test that was written to expose a real bug or specify a real feature, because it encodes a real requirement. **The existing test suite as PASS\_TO\_PASS.** Every task gets the whole existing suite as a regression check for free. A flip requires that nothing already passing breaks. On a repository with a large suite, this alone rules out most of the bad patches that a visible test would accept. **Type checkers and compilers.** A change that does not compile, or that fails `tsc --noEmit`, `mypy --strict`, or `cargo check`, has failed. These are deterministic, fast, and hard to game without touching configuration, which the allowlist excludes. On their own they are a weak oracle, since code that compiles can still be wrong, but as a component of the oracle they remove a class of failures cheaply. **Linters with pinned rule sets.** A lint failure is a weak signal of incorrectness and a strong signal of style drift. Use it as a gate, not a label: a trace that flips the hidden test but introduces lint errors is a flip with a defect, and you can decide per repository whether that counts. **Property-based tests.** Instead of a fixed input and expected output, a property test states a rule, such as "decoding what you encoded gives the original," and a library generates many inputs to check it.[1](#user-content-fn-quickcheck) With a fixed seed, a property test is deterministic. It is a stronger oracle than an example test because it checks many cases, and a harder one for an agent to fit to because the agent cannot see the cases. **Metamorphic tests.** When the correct output is unknown but a relationship between outputs is known, a metamorphic test checks the relationship. If a search for "a" returns results, a search for "a OR b" should return at least as many. Metamorphic testing was proposed for programs without an oracle and is well suited to data pipelines, search, and numerical code.[2](#user-content-fn-chen-metamorphic) **Differential tests against the previous build.** Run the same inputs through the version before the change and the version after. For a refactor, the outputs must match. For a bug fix, they must differ only on the inputs that exposed the bug. McKeeman's differential testing of compilers is the origin of the method.[3](#user-content-fn-mckeeman-differential) The previous binary is an oracle you already have. **Mutation score.** Chapter 3 covered mutation as a test-quality measure. It can also be an oracle for a task of the form "add tests for this module." The hidden check is whether the mutants the task names are killed after the change. This is one of the few ways to make test-writing itself a flippable task. **Reproducible build hash.** For a task that must not change behavior, such as a dependency bump with no code change, the oracle can be that the build output is byte-identical to a reference, or differs only where expected.[4](#user-content-fn-reproducible-builds) **Schema and migration round trips.** For a database change: apply the migration to a copy of the schema, run the down migration, and diff the result against the original. For a data pipeline: run the transform on a frozen fixture and compare row counts and checksums against a stored expectation. **Infrastructure plan idempotence.** For infrastructure-as-code: after applying the change, a second plan must report nothing to do. This is a deterministic check of the form most infrastructure tools provide directly. **Formal checks.** Where a module has a specification in a checkable form, a solver or a proof assistant is the strongest oracle there is. Few organizations have these. The ones that do should use them. ## Soft and delayed oracles A soft oracle is a signal that correlates with correctness but is not deterministic, is not available for minutes or days, or depends on a person. These cannot label a trace as a flip. They can rank traces, build preference pairs, and filter. **Pull request merged.** A merged change passed review and continuous integration. This is the oracle most code-model training has used, in the form of mined commits and pull requests. Meta's SWE-RL built its training corpus from about eleven million pull requests and used similarity to the merged patch as the reward.[5](#user-content-fn-swe-rl) It is a real signal and a slow one, and reviewers miss things. **Reverted within N days.** A change that was merged and then reverted was a failure the review did not catch. The pair "merged, then reverted" against "merged, kept" is a clean preference pair with a delay of days to weeks. **Incident linked to the change.** Rarer and stronger than a revert. A change that caused an incident is a hard negative example. **Review comments.** A change that received requests for changes before merge is weaker than one approved on the first pass. Review text also says what was wrong, which is useful for a data scientist and useless as a training label. **Acceptance of a suggestion.** For completion models, whether the developer accepted the suggestion was found to be the best available predictor of perceived productivity in GitHub Copilot telemetry.[6](#user-content-fn-ziegler-copilot) For agents, the equivalent is whether the person kept the agent's change or discarded the session. It is a weak oracle and an abundant one. **Time to next edit of the same lines.** If a person edits the lines an agent wrote within an hour of the session, the agent's work was probably incomplete. This is a proxy with many false positives and is best used to flag traces for a closer look. **A judge model.** A second model that reads the diff and scores it. Chapter 6 argues that this is the weakest oracle in the catalogue and should not be used as a label. It can be used as a filter for obvious garbage, and even then its errors should be measured against a hard oracle on a sample. > **Oracles by how directly their verdict can label a trace** > > 1. **Judge model**: A model scores the change. Not a label. At most a coarse filter, and even then measured against a hard oracle. > 2. **Acceptance, next edit, review comments**: Human behavior signals. Rank and flag, do not label. > 3. **Merged, reverted, incident**: Delayed by days. Build preference pairs. > 4. **Type check, lint, build hash**: Fast and deterministic. Necessary, not sufficient. Use as gates inside the oracle. > 5. **Existing suite as PASS\_TO\_PASS**: Free regression check on every task. > 6. **Hidden tests, property and metamorphic tests**: Deterministic, fast, independent, unseen. Labels a flip. > > *Only the top rung produces a label the training set can trust on its own. Everything below it is still worth collecting.* ## Oracles outside the code The method is not limited to software, but the hard oracles mostly are. For the sake of completeness, here is what the same construction looks like elsewhere, with the caveat that each of these is a soft oracle and should be treated as one. - **Data work.** A transform's output on a frozen input has a checksum. Row counts and null rates have expectations. Tools that run assertions on data in a pipeline exist and can be hidden from the agent in the same way tests are. - **Documentation.** A link checker, a prose checker with a pinned rule set, and a build that fails on a broken reference are deterministic. Whether the document is clear is not. - **Support.** A ticket resolved and not reopened within two weeks is a delayed soft oracle on the agent's answer. - **Operations.** A runbook step that leaves a system in a state a monitoring check accepts is close to a hard oracle, if the check is deterministic and the system is a copy. A team should start with the hard oracles in its code, because that is where the construction is cleanest and the data is largest. Once the pipeline runs, the soft oracles can be added as preference signals. ## Which oracle wrote the label Every trace should carry the identity of the oracle that graded it: the kind of oracle, the hash of its test list or rule set, the container image digest, and the time. A training set built from traces graded by different oracles at different times is a training set whose labels mean different things. When a later evaluation shows a regression, the first question is which oracle labeled the traces the model learned it from. Without the record, the question cannot be answered. ## Chapter 5: The air gap An oracle the agent can reach is an oracle the agent will eventually satisfy without doing the work. This chapter is about the threat, the evidence that it is real, and the design that removes it. ## The threat model Treat the agent as an adversary with full control of its working directory and shell. That is not a judgment about the model's intent. It is a description of its permissions. A coding agent can run any command, read any file it can reach, and write any file in its checkout. If the verdict depends on anything in that set, the verdict is the agent's to decide. The ways an agent can make a check pass without doing the task are ordinary: - Edit the test so that it asserts what the code does. - Mark the test skipped, or add it to an exclusion list in the test configuration. - Edit the test runner's configuration so the test is not collected. - Edit the build file or the continuous-integration file so the check does not run. - Stub the function under test to return the expected value for the inputs the test uses. - Patch the test framework itself, or a fixture, or a conftest file that runs before every test. - Set an environment variable the test reads. - Exit the process with status zero before the test runner reports. - Read the hidden test from disk or from the network, if it is anywhere the agent can reach, and fit the code to it. - If a model judges the result, write text into the diff or the commit message that addresses the judge. Every one of these has been observed. Baker and colleagues at OpenAI trained a frontier reasoning model on coding tasks and monitored its chain of thought. They found the model saying, in its own reasoning, that it would make the tests pass by other means, and then doing so: calling `sys.exit(0)` before the tests ran so the runner reported success, raising the test framework's skip exception, and writing stubs instead of implementations where test coverage was thin.[1](#user-content-fn-baker-monitoring) When they penalized the model for having such thoughts, it kept hacking the tests and stopped writing down that it was going to. They called this obfuscated reward hacking and recommended against training against the monitor. Denison and colleagues at Anthropic built a curriculum of environments where a model could gain reward by gaming the check, starting with flattery and ending with the chance to edit its own reward function. A model trained on the easy environments generalized to the hard one: in a small but nonzero fraction of runs, it edited the code that computed its reward, and sometimes edited the test that would have caught the edit.[2](#user-content-fn-denison-subterfuge) METR, evaluating frontier models on its own tasks in 2025, reported models that rewrote scoring code, patched the timer in a benchmark harness, and special-cased the inputs the grader used.[3](#user-content-fn-metr-hacking) A later Anthropic study trained a model with reinforcement learning on real production coding environments after teaching it about reward-hacking strategies. At the point where the model learned to hack the rewards, it also began to generalize to other unwanted behavior, including attempts at sabotage.[4](#user-content-fn-anthropic-emergent) The lesson is not that models are malicious. It is that optimization finds the shortest path to the reward, and if the shortest path goes through the test file, that is the path. Skalse and colleagues formalized this as reward hacking: a proxy that correlates with the goal is maximized, and the correlation breaks under optimization.[5](#user-content-fn-skalse-hacking) Gao, Schulman, and Hilton measured it for reward models: the true reward rises as a policy is optimized against the proxy, peaks, and then falls while the proxy keeps rising.[6](#user-content-fn-gao-overoptimization) The hidden test is the goal. Anything the agent can touch is a proxy. The lesson is also older than machine learning. Thompson's 1984 Turing Award lecture described a compiler that inserts a back door into the login program and into any compiler compiled from clean source, so that no inspection of the source finds it.[7](#user-content-fn-thompson-trust) His point was that you cannot trust code you did not totally create yourself, and that the trust boundary has to be drawn around the thing you can verify. The oracle is the thing you can verify. Draw the boundary around it. ## The design The oracle runs in a separate container, on a separate host when possible, with the following properties. **No network.** The container is started with networking disabled. Dependencies come from a local mirror baked into the image or mounted read-only. Nothing the oracle needs is fetched at run time, and nothing the agent could have planted on a server is reachable. **A pristine checkout.** The oracle has its own clone of the repository, from its own mirror. It checks out the base commit the session started from. It does not read the agent's working copy at all. **The source diff only.** The oracle receives the agent's final diff as a file. Before applying it, the oracle filters it through an allowlist of paths. Source directories are on the allowlist. Test directories, test configuration files, continuous-integration configuration, build scripts, dependency manifests and lockfiles, and anything under the repository's tool configuration are not. Hunks that touch excluded paths are dropped. The oracle logs that they were dropped, in a log the agent cannot read. **Hidden tests from the oracle's store.** The hidden tests are stored with the oracle, keyed by task id, and copied into the checkout after the diff is applied. They are never in the repository the agent works in, never in a branch the agent can fetch, and never on a path the agent's shell can read. **The existing suite from the base commit.** The PASS\_TO\_PASS tests are the ones in the pristine checkout at the base commit, not the ones in the agent's diff. If the agent deleted a test, the oracle still runs it. **Fixed clock, seed, locale, and time zone.** Set in the container's environment. A test that reads the system clock reads a constant. **Pinned image.** The container image is referenced by digest, and the digest is recorded with the verdict. **Timeouts.** The oracle run has a wall-clock limit. A run that exceeds it is a FAIL with a reason the oracle logs and the agent does not see. An agent cannot turn a failing test into a hang and have the hang read as success. **One bit out, signed.** The oracle's only output to the agent's side is PASS or FAIL. To the trace store it writes a verdict record: the bit, the hash of the test list, the image digest, the diff hash, the base commit, the time, and a signature over all of these using a key the oracle host holds. A verdict in the trace store that does not verify was not written by the oracle. **Separate credentials.** The process that runs the agent holds no credential that can write to the oracle's store, start an oracle run with a different test list, or read the hidden tests. The hook that asks for a verdict sends a diff and a task id and gets a bit back. > **What crosses the boundary between the agent's host and the oracle's host** > > - Agent host sends Oracle host: Task id; Base commit; The final diff as a file > - Oracle host returns Agent host: PASS or FAIL; Nothing else > - Oracle host writes Trace store: Signed verdict record; Dropped-hunk log; Test output, never shown to the agent > - Agent host writes Trace store: Tool-call events, redacted; Messages; Session metadata > > *The agent's host can write trace events and can ask one question. It cannot read the hidden tests, the test output, or the oracle's log. The trace store holds both sides and joins them on the session id.* ## Why a separate host A container on the same machine as the agent is isolation from the agent's working copy. It is not isolation from the agent's shell. On a developer's laptop, the agent's Bash tool can read any path the developer can read, including the directory where the hidden tests live, the oracle's log, and the container's image. A container boundary protects the oracle's run from the agent's changes. A host boundary protects the oracle's secrets from the agent's eyes. The reference plugin in Chapter 8 runs on one machine because that is where most teams will start, and it says so in its README. On one machine, store the hidden tests encrypted at rest and decrypt them inside the oracle container with a key the agent's process does not hold. That raises the bar from "read the file" to "find and misuse the key." It does not close the gap. Closing it means running the oracle on a host the agent cannot log into, which for a team with a continuous-integration system is a short step: the oracle is a job that runs on a runner, takes a diff and a task id, and returns a bit. Firecracker, the virtual-machine monitor behind AWS Lambda, starts a microVM in about 125 milliseconds and gives each one its own kernel.[8](#user-content-fn-firecracker) gVisor gives a container a user-space kernel and intercepts its system calls.[9](#user-content-fn-gvisor) Either is a stronger boundary than a plain container for the oracle run, and both are used for exactly this purpose: running code you do not trust next to data you care about. For the hidden tests and the signing key, neither replaces a separate host. ## Baselines and the flip that was not The flip needs a baseline: a FAIL recorded before the agent started. Three rules keep baselines honest. The baseline is computed on the base commit, by the oracle, from its own checkout. It is deterministic, so compute it once per task and store it. Do not recompute it in a session-start hook; a container run is too slow for a hook budget, and the result would be the same anyway. A task whose baseline is PASS is not a task. The hidden test already passes. Either the test is wrong or the work is already done. Quarantine the task, and do not count any session on it as a flip. A task whose existing suite fails at baseline is not fit for an oracle. The agent could flip the hidden test while the suite stays broken, and the PASS\_TO\_PASS rule would record a FAIL for a session that did the work. Fix the base or exclude the task. ## What the oracle logs and who reads it The oracle keeps a full log: the test output, the dropped hunks, the timing, and the reason for every FAIL. This log is for the people who run the pipeline. It is how you find a hidden test that fails for the wrong reason, a flaky test that slipped through, or an agent that keeps trying to edit the test configuration. It is never returned to the agent and never written anywhere the agent's host can read. Chapter 6 is about why. ## Chapter 6: The one-bit verdict The oracle says PASS or FAIL. It does not say which test failed, what the assertion was, or what the output looked like. This is the design choice readers push back on most, so this chapter takes it slowly: first the cost, then the three reasons, then the research on self-grading that the third reason depends on. ## The cost, stated plainly Richer feedback helps an agent fix a bug in the current session. Chen, Lin, Schärli, and Zhou compared three kinds of feedback for a model debugging its own code: simple feedback, meaning only whether the code passed; unit-test feedback, meaning the execution results; and a self-written explanation of the code. Unit-test feedback produced the largest gains.[1](#user-content-fn-chen-self-debug) A verdict with no explanation is the "simple feedback" condition in that study, and it is the weakest of the three for raising the pass rate inside one session. So the one-bit verdict lowers the flip rate per session. A team that adopts it will see more sessions end on FAIL than a team that shows the agent the failing test. That is a real cost. Three things make it worth paying. ## Reason one: the hidden test stays valid A test the agent can see is a test the agent can fit to. Chapter 3 covered the patch-overfitting results: patches that pass visible tests by deleting what the tests do not check.[2](#user-content-fn-qi-kali)[3](#user-content-fn-smith-cure) An oracle that returns the name of the failing test and the assertion that failed has shown the agent the test. The agent will fix the assertion. Whether it fixed the feature is now unknown, which is where you started. The theory for this is from statistics, not software. Blum and Hardt studied machine-learning competitions where participants submit many models and see a score on a holdout set each time. With enough submissions, a participant can overfit the holdout without ever seeing its labels, by hill-climbing on the score. Their fix, the Ladder, releases a new score only when it improves on the previous best by more than a threshold, and otherwise repeats the old score. That turns a high-information answer into a low-information one and makes the leaderboard reliable under adaptive submissions.[4](#user-content-fn-ladder) Dwork, Feldman, Hardt, Pitassi, Reingold, and Roth proved the general result: a holdout answered with limited information, through a mechanism they called Thresholdout, stays statistically valid under far more adaptive queries than one answered exactly.[5](#user-content-fn-reusable-holdout) An agent iterating against an oracle is a participant iterating against a holdout. A verdict of PASS or FAIL is about as low-information as an answer gets. An execution trace with the failing assertion is about as high-information as one gets. The Ladder result says which one keeps the hidden test meaningful. ## Reason two: the diagnosis is the data The point of collecting traces is to train a model on them. The behavior you want the model to learn is not "read the failing assertion and change the code until it passes." It is "form a hypothesis about what is wrong, write a test that checks it, run the test, read the result, and fix the cause." That is what expert engineers do, and the trace of an expert doing it is what you want in the training set. If the oracle tells the agent why it failed, the trace after that point is the agent copying the oracle's diagnosis. The model trained on it learns to wait for a diagnosis. If the oracle says only FAIL, the trace after that point is the agent's own diagnosis: it has to write its own tests, run them, read its own failures, and reason about what the hidden requirement might be. That is the behavior worth learning, and it only appears in the trace when the oracle withholds the answer. The agent is not working blind. It has the task description, the whole repository, the existing test suite, and a shell. It can write any test it wants and run it as often as it wants, with full output. The one-bit rule applies to the hidden oracle only. In-session feedback from the agent's own tests is as rich as the agent cares to make it. The oracle's silence forces the agent to generate that feedback for itself, which is the skill. ## Reason three: the alternative is a model grading a model The usual proposal for richer feedback that does not leak the test is a judge: a second model that reads the diff, the test output, and the task, and writes an explanation for the agent. The judge sees the hidden test; the agent sees the judge's prose. This is where the research on self-grading matters, because the judge is a model with the same training and the same blind spots as the agent. Zheng and colleagues, in the paper that introduced MT-Bench and the LLM-as-a-judge method, documented three biases in model judges: position bias, where the judge prefers whichever answer appears first; verbosity bias, where it prefers the longer answer; and self-enhancement bias, where it prefers answers it wrote.[6](#user-content-fn-zheng-judge) Wang and colleagues showed that swapping the order of two answers could reverse a judge's verdict.[7](#user-content-fn-wang-unfair) Panickssery, Bowman, and Feng showed that a model's preference for its own output is not an accident of style. Models can recognize their own text above chance, and the ones that recognize it best prefer it most; fine-tuning a model to recognize its own output more accurately made its self-preference stronger.[8](#user-content-fn-panickssery) A judge from the same family as the agent is a judge that likes the agent. The research on self-correction is worse. Huang and colleagues reviewed the claim that models can correct their own reasoning and found that the gains reported in earlier work depended on an oracle: the model was told whether its answer was right before being asked to reconsider. Without that signal, asking the model to check its work made the answers worse more often than better.[9](#user-content-fn-huang-self-correct) That paper is sometimes read as an argument against external feedback. It is the opposite. It shows that a binary external signal is the thing that made self-correction work in the papers that reported it working, and that the model's own judgment was not contributing. Stechly, Marquez, and Kambhampati tested GPT-4 as a critic of its own graph-coloring solutions and found it could not reliably tell a correct coloring from an incorrect one; iterating on its own critique did not help, and a simple external checker did.[10](#user-content-fn-stechly-wrong) Valmeekam, Marquez, and Kambhampati found the same for planning: a model critiquing its own plans lowered the success rate, and an external verifier raised it.[11](#user-content-fn-valmeekam-plans) Tyen and colleagues separated two skills and found that models are poor at locating the error in a chain of reasoning but can often fix it once the location is given.[12](#user-content-fn-tyen-location) Kamoi and colleagues surveyed the self-correction literature and concluded that no reliable self-correction has been shown without external feedback, and that many reported successes used unrealistic setups where the model was given information it would not have in practice.[13](#user-content-fn-kamoi-survey) Xu and colleagues found that self-refinement loops amplify the model's own biases over iterations.[14](#user-content-fn-xu-pride) None of these results says a judge model is useless. They say a judge model's errors are correlated with the agent's errors, so using one to explain failures to the agent does not add an independent signal. It adds a confident one. The hidden test is independent. Its one bit is worth more than the judge's paragraph. > **Feedback channels from the oracle to the agent** > > 1. **PASS or FAIL**: The Ladder regime. The hidden test stays valid. The agent's own diagnosis is in the trace. > 2. **Which test failed**: Names the requirement. The agent now knows what to fit to. > 3. **The assertion and the values**: The agent can special-case the inputs. Patch overfitting territory. > 4. **A judge model's explanation**: Adds a correlated, biased reading of the test to the leak. The worst of both. > 5. **The test source**: The hidden test is no longer hidden. The oracle measures nothing. > > *Each rung down leaks more of the hidden test and leaves less of the agent's own reasoning in the trace. The book recommends the top rung and names the cost.* ## Recovering some of the cost There are ways to give back part of the in-session success rate without giving up the properties above. **Disclose to the trace, never to the agent.** The oracle writes the full test output to its own log, which the data team reads. When a task fails many sessions in a row, a person reads the log and either fixes a hidden test that fails for the wrong reason or rewrites the task description so the requirement is clearer. The agent gets a better task, not a leaked test. **A visible test the agent must write.** Make the task description say that the change must come with a test that fails before and passes after. The agent writes its own red-green pair, which is rich in-session feedback and good trace content. The hidden test stays hidden and checks that the agent's test checked the right thing. **Staged tasks.** If a task's hidden test covers three requirements, split it into three tasks with one hidden test each. The agent gets three bits instead of one across the same work, and each bit still says nothing about its test. **More attempts, fresh context.** The flip rate per session is lower; the flip rate per task does not have to be. Chapter 9 covers sampling several sessions per task and keeping the one that flips. With the oracle silent, the attempts are independent, which is what repeated sampling needs. **Disclose after the fact.** Once a task has been flipped by some session, its hidden test can be released into the repository as an ordinary test. It has done its job as an oracle. Future sessions on nearby code get it as a visible regression test, and new hidden tests are written for new tasks. ## The rule The hidden oracle says PASS or FAIL. Every other channel from the oracle to the agent is closed. Everything the oracle knows goes to the trace store, for people, under access control. If a team decides to open a wider channel, it should do so knowing which rung of the ladder it has moved to and what it has traded for the higher flip rate. ## Chapter 7: Harness hooks You do not need to build a coding agent to collect traces from one. The harness your team already runs exposes hooks, and hooks see everything a trace needs. This chapter explains what a hook is, what each one sees, and how to turn a set of hooks into a trace collector and an oracle gate. The examples use Claude Code because its hook interface is documented in detail and the reference plugin targets it.[1](#user-content-fn-claude-hooks) Other harnesses have equivalents, and the last section covers them. ## What a hook is A hook is a program the harness runs at a fixed point in a session. The harness passes the event to the program as JSON on standard input, waits for it to exit, and reads its exit code and anything it printed. A hook can observe the event, add context for the model, or, for some events, change what happens next. The events that matter for tracing are the ones at the boundaries of a turn and a session. | Event | When it runs | What it carries | What it can decide | | --- | --- | --- | --- | | SessionStart | A session begins or resumes | Session id, working directory, model, how the session started | Nothing. It can add context. | | UserPromptSubmit | The person sends a prompt | The prompt text | It can block the prompt. | | PreToolUse | Before a tool runs | Tool name and input | It can allow, deny, or ask. | | PostToolUse | After a tool runs | Tool name, input, output, duration | It can add context beside the result. | | Stop | The agent is about to end its turn | The agent's final message, whether a Stop hook already continued the turn | It can refuse to let the agent stop. | | SessionEnd | The session ends | The reason the session ended | Nothing. It runs cleanup. | Every event also carries the session id, the path to the harness's own transcript, the working directory, and the permission mode. The session id is the join key for everything the trace collector writes. ## The trace collector A trace collector is three hooks. On **SessionStart**, write a `session_start` event: the session id, the working directory, the base commit of the repository, the model, and the time. Do not run anything slow here. Hook budgets at session start are short, and a container run does not fit. On **UserPromptSubmit**, write a `message` event with the role `user` and the prompt text. This is the task as the agent received it, which the training set needs as the first turn. On **PostToolUse**, write a `tool_call` event: the tool name, the input, the output, and the duration. Redact the input and output before writing. Cap the size of each and store a hash of the full value beside the truncated one, so a trace that was cut can be identified later. That is the whole collector. The agent's own messages between tool calls are not delivered by PostToolUse, so the collector adds a fourth hook. On **Stop**, write a `message` event with the role `assistant` and the agent's final message, which the Stop event carries as `last_assistant_message`. Then run the oracle gate, below. The collector does not read the harness's transcript file. The transcript is for the person, its format is the harness's to change, and the harness documentation notes that the file is not guaranteed to include the final message at the moment Stop fires.[1](#user-content-fn-claude-hooks) The hook events are the record. ## The oracle gate The Stop hook is where the flip is detected and where an agent is kept from calling the work done. When the agent decides it has finished, the harness fires Stop. The gate hook does the following. 1. Reads the event. If the agent is already continuing because of a previous Stop hook, the event says so in a field named `stop_hook_active`. The gate uses this, with its own attempt counter, to decide whether to keep going. 2. Builds the agent's diff against the base commit recorded at session start, including new files. 3. Sends the diff and the task id to the oracle. The oracle returns one bit. 4. Writes an `oracle_verdict` event with the attempt number, the verdict, and the hash of the test list the oracle reported. 5. On PASS, exits normally. The agent stops. The session has flipped. 6. On FAIL, returns a decision that blocks the stop, with a reason the harness shows to the agent. The reason is a fixed string. It says the oracle reported FAIL and the agent should keep working. It does not say why. The blocking mechanism is specific. In Claude Code, a Stop hook refuses the stop by printing a JSON object with `decision` set to `block` and a `reason`, or by exiting with code 2 and the reason on standard error.[1](#user-content-fn-claude-hooks) The agent sees the reason as the explanation for why it must continue. With a one-bit oracle, the reason is the same every time. ## What the harness limits Two limits in the harness shape how the gate behaves, and a team should know both before relying on it. First, the harness caps consecutive continuations. In Claude Code, after Stop hooks have continued the turn eight times in a row, the harness overrides the next block and ends the turn. The count resets whenever the agent calls a tool, so an agent that keeps working does not hit it, and an agent that keeps saying "done" without doing anything does.[1](#user-content-fn-claude-hooks) The cap is configurable. The gate should treat a session that ends this way as `no_flip`, and the trace should record that the cap, not the oracle, ended it. Second, hooks have timeouts. A Stop hook has minutes, not seconds, which is enough for an oracle that runs a test suite in a container. It is not enough for a suite that takes an hour. For slow oracles, the gate should submit the diff to the oracle host, return a provisional block with the fixed reason, and let the next Stop pick up the verdict. The reference plugin runs the oracle inline because most suites that are fit for an oracle run in minutes. ## Redaction The tool output is where secrets live. A `cat .env`, a failing test that prints a connection string, a curl response with a token: all of it passes through PostToolUse. The trace collector redacts before it writes. Redaction at this stage is pattern-based and should be treated as a floor, not a guarantee. Patterns for the common shapes, such as `AKIA` prefixed AWS keys, GitHub tokens, private key blocks, bearer headers, and `KEY=value` lines where the key name contains `SECRET`, `TOKEN`, or `PASSWORD`, catch most of what appears in practice. They do not catch a password that looks like a word. Meli, McNiece, and Reaves found secrets leaking into public repositories at a rate of thousands of new unique secrets a day, across more than a hundred thousand repositories, which is a measure of how often they appear in code and output that people thought was fine.[2](#user-content-fn-meli-secrets) The collector redacts, and the pipeline runs a second pass with a dedicated scanner before anything reaches training, and the model is kept private to the organization whose traces it learned from. Chapter 16 covers the memorization research that makes the last rule a hard one. ## Storage Traces are written to a directory the plugin owns, not to the repository and not to the plugin's install directory. In Claude Code, that is the plugin's data directory, which the harness exposes as `CLAUDE_PLUGIN_DATA` and which survives plugin updates.[3](#user-content-fn-claude-plugins) A sync job moves finished traces to an object store under a path keyed by organization, repository, and date. The store is append-only. Nothing in the pipeline edits a trace after it is written; corrections are new records that reference the old. Each verdict record is signed by the oracle, as Chapter 5 describes. A trace whose verdict does not verify is excluded from training and flagged. This is the check that stops a compromised or misconfigured agent host from writing its own PASS. ## The same idea in other harnesses The hook design is not specific to one product. OpenAI's Codex CLI, Cursor's agent, OpenCode, and others expose lifecycle events with the same shape: before and after tool calls, and at the end of a turn. The names differ. The stdin JSON differs. The decision mechanism for refusing a stop differs, and in some harnesses does not exist, in which case the gate has to run as a wrapper around the harness instead of inside it. What stays constant is the architecture. One hook writes events. One hook, at the end of a turn, asks the oracle and either lets the agent stop or sends it back. The trace schema is the harness-neutral part, and a team that runs more than one harness should normalize events into the same schema at write time, so the training set does not depend on which tool produced it. ## Chapter 8: The reference plugin This chapter walks through `oracle-flip`, a Claude Code plugin that does what Chapters 2 through 7 describe. It is public at [github.com/macanderson/oracle-flip](https://github.com/macanderson/oracle-flip) under the MIT license. It is a reference, meant to be read and adapted, and it is small enough to read in one sitting. ## Install ```text /plugin marketplace add macanderson/oracle-flip /plugin install oracle-flip@macanderson ``` To try it without installing, clone the repository and start Claude Code with the plugin loaded from disk: ```text git clone https://github.com/macanderson/oracle-flip claude --plugin-dir ./oracle-flip ``` ## Layout ```text oracle-flip/ .claude-plugin/ plugin.json name, version, description marketplace.json lets /plugin marketplace add find it hooks/ hooks.json the five hook registrations scripts/ common.py paths, event writer, redaction, config session_start.py SessionStart: write session_start trace.py UserPromptSubmit and PostToolUse: write message and tool_call gate.py Stop: diff, oracle, verdict, block or allow session_end.py SessionEnd: write session_end with the outcome oracle/ run.sh reference oracle runner: container, allowlist, hidden tests, one bit grade.sh runs inside the container filter_diff.py drops hunks outside the allowlist Dockerfile pinned base image; the container runs with no network example/ a sample task with a hidden test tools/ export_sft.py traces plus verdicts to SFT and preference-pair JSONL schema/ trace-event.schema.json tests/ test_redaction.py, test_filter_diff.py, test_gate.py ``` The scripts are Python with no dependencies outside the standard library, so the plugin runs anywhere `python3` runs. The oracle runner is shell plus one Python script, because it is meant to be replaced by whatever the team's continuous-integration system already does. ## The hook registrations The plugin's `hooks/hooks.json` wraps the event map in a `hooks` key, which is the shape plugins use.[1](#user-content-fn-claude-plugins) Each hook is in exec form: a `command` and an `args` array, with the plugin root substituted by the harness, so paths with spaces need no quoting. ```json { "description": "oracle-flip: trace every tool call and gate Stop on a hidden oracle", "hooks": { "SessionStart": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_start.py"], "timeout": 20 } ] } ], "UserPromptSubmit": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 10 } ] } ], "PostToolUse": [ { "matcher": "*", "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 20 } ] } ], "Stop": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/gate.py"], "timeout": 600 } ] } ], "SessionEnd": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_end.py"], "timeout": 20 } ] } ] } } ``` ## Configuration Configuration is split by who controls the file. Anything the agent can edit holds nothing the gate trusts. The project the agent works in holds one small file, `.oracle-flip/config.json`. It names the task and sets the attempt budget. ```json { "task_id": "billing-proration-rounding", "base_commit": null, "max_attempts": 6 } ``` `base_commit` left null means the commit at session start. `max_attempts` is the number of FAIL verdicts after which the gate stops blocking and records `no_flip`. It exists so a session that will never flip does not run the oracle forever, and it is set below the harness's own cap of eight consecutive continuations so the plugin, not the harness, decides when to give up. The SessionStart hook reads this file once and copies it into the session's state. The gate reads the state, never the project, so an agent that edits the file mid-session changes nothing. An agent that edits it before the session starts can at most point the session at a task whose baseline passes, which the gate records as `no_task`, never as a flip. The oracle command, the oracle store, and the container image live in the plugin's user settings, which Claude Code stores outside every repository and passes to hook processes as environment variables. The oracle command receives the mode, the task id, the base commit, the path of the diff file, and a hash of the repository's remote URL, and it must exit 0 for PASS and 1 for FAIL. Any other exit code is an oracle error, which the gate records and treats as FAIL for the purpose of blocking, with a different fixed reason. The allowlist and denylist of paths the diff may touch live in the oracle store beside the hidden tests, one file each, and the oracle applies them before it touches a file. The gate's host does the filtering nowhere, because the gate's host is the agent's host. The hidden tests are in the same store, keyed by task id, and nowhere in the repository. ## The gate The gate is the file to read if you read one. Its structure follows Chapter 7. ```python def decide(event, state, *, oracle=run_oracle, diff=write_diff): write_event(assistant_message(event)) # last_assistant_message, redacted if state.task_id is None: return allow() # no oracle task; the plugin only traces if state.outcome is not None: return allow() # already decided; never loop on a settled outcome if state.baseline is None: baseline = oracle(state.oracle_command, mode="baseline", ...) state.baseline = baseline.verdict record_verdict(state, "baseline", baseline) if baseline.verdict == "PASS": state.outcome = "no_task" # the hidden test already passes return allow() if baseline.verdict == "ERROR": state.outcome = "oracle_error" return allow() diff_path = diff(event["cwd"], state.base_commit, ...) verdict = oracle(state.oracle_command, mode="grade", diff_path=diff_path, ...) state.attempts += 1 record_verdict(state, "grade", verdict) if verdict.verdict == "PASS": state.outcome = "flipped" return allow() if state.attempts >= state.max_attempts: state.outcome = "no_flip" return allow() return block(BLOCK_REASON_FAIL if verdict.verdict == "FAIL" else BLOCK_REASON_ERROR) ``` Four details are worth pointing at. The gate reads everything from the session state the SessionStart hook wrote: the task id, the base commit, the attempt budget, and the oracle command. It reads nothing from the project. The baseline is computed on the first Stop, not at session start, because the oracle is slow and the baseline is deterministic. It is stored, so later Stops in the same session reuse it. A team with many sessions per task should compute it once per task and share it; the plugin keeps it per session for simplicity. The `block` reason is one of two constants. Neither includes the oracle's output. The oracle's output is written by the oracle, on the oracle's side, and the gate never sees it. The gate's own tests check that the reasons contain no test name and no assertion. `stop_hook_active` is honored by the attempt counter rather than by an early return, because the harness sets it on every Stop after the first block, and returning early on it would mean the gate only ever runs once. The attempt counter plus `max_attempts` is what prevents the loop, and `decide` takes the oracle and the diff builder as parameters so the loop logic is tested without a container. ## The oracle runner `oracle/run.sh` is the reference oracle. It expects a store with one directory of hidden tests per task and a mirror of the repository, and it runs the grade inside a container with networking disabled. The core of it: ```sh task_dir="$store/$ORACLE_TASK_ID" log_dir="$store/.log/$ORACLE_TASK_ID/$(date -u +%Y%m%dT%H%M%SZ)-$ORACLE_MODE" args=(run --rm --network none -e TZ=UTC -e LC_ALL=C.UTF-8 -e PYTHONHASHSEED=0 -e SOURCE_DATE_EPOCH=1700000000 -e ORACLE_MODE="$ORACLE_MODE" -e ORACLE_BASE_COMMIT="$ORACLE_BASE_COMMIT" -v "$task_dir:/task:ro" -v "$mirror:/mirror:ro" -v "$log_dir:/log") [ "$ORACLE_MODE" = "grade" ] && args+=(-v "$ORACLE_DIFF:/in/diff.patch:ro") docker "${args[@]}" "$image" bash /oracle/grade.sh > "$log_dir/stdout.txt" 2> "$log_dir/stderr.txt" status=$? grep -E '^tests_hash=' "$log_dir/stdout.txt" || true exit "$status" ``` `grade.sh` runs inside the container. It clones the repository from the read-only mirror, checks out the base commit, filters the diff through the task's allowlist and denylist with `filter_diff.py`, applies what is left, copies the hidden tests from `/task/hidden` into place, runs the existing suite and then the hidden tests, prints `tests_hash=` followed by a hash of the hidden test list, and exits 0 or 1. In baseline mode it runs without a diff and exits 1 when the suite passes and the hidden tests fail, which is the only baseline that makes a task. Its full output goes to the log directory on the oracle's side. Only the exit code and the `tests_hash` line return to the gate. The Dockerfile pins its base image by digest and installs the repository's dependencies from a lockfile at build time, so the run-time container needs no network and gets none. ## Limits of the plugin It runs the oracle on the same machine as the agent, in a container. Chapter 5 explained why that protects the oracle's run and not the oracle's secrets. On a laptop, the agent's shell can read `ORACLE_STORE`. The README says so. The step to a separate host is to replace `oracle/run.sh` with a script that submits the diff to a job runner the agent cannot log into and waits for the bit. It does not sign verdicts. Signing needs a key on the oracle's side and a verifier in the pipeline, and the plugin has neither because it has no pipeline side. The verdict event has a field for the signature and the export tool has a flag that requires it. It redacts with patterns. That is a floor. Run a real secret scanner on the trace store before training. It has not been tested with a real container run on the author's machine at the time of writing, for a reason unrelated to the plugin, and the README says that too. The hook scripts are tested by piping sample events into them, and `claude plugin validate` passes. ## The export tool `tools/export_sft.py` reads a directory of traces and their verdicts and writes two files. The first is the supervised set: one JSON object per flipped session, in a chat format with tool calls, with a mask that marks which turns to train on. User prompts and tool outputs are masked out; the model is trained to produce the assistant's turns and tool calls, not to reproduce the environment's replies. Chapter 10 explains why. The second is the preference set: for every task with at least one flipped session and at least one that did not flip, a pair with the flipped trace as the preferred one. Chapter 11 explains what to do with it. Both files exclude any session whose verdict record is missing, and, with `--require-signature`, any whose verdict does not verify. ## Chapter 9: The flip rate The pipeline's throughput is the number of flips per day. This chapter is about raising it: choosing tasks that can flip, dispatching work in a shape that flips, sampling more than once, and manufacturing tasks when the natural supply runs low. It ends with what to do with the sessions that do not flip, which is most of them. ## Start with the test A task can flip only if a hidden test fails on the base commit. The highest-leverage change a team can make is to write the test before dispatching the task. This is test-driven development with the roles split: a person or a separate agent writes the red test, and the solving agent is dispatched with the task description and never sees it. The practice has a long history under the name red-green-refactor, and the evidence that it improves defect rates predates language models.[1](#user-content-fn-nagappan-tdd) What is new is the reason to do it. In a team with an oracle, every red test is a potential verified trajectory. In a team without one, a red test is just a test. Some task shapes come with a red test for free. - **A bug report with a reproduction.** The reproduction is the hidden test. Clean it up, assert the correct behavior, confirm it fails on the base commit, and dispatch the bug. - **A failing test in continuous integration.** The test already exists and already fails. Hide it from the agent by giving the agent a checkout where the test is removed, and keep the test as the oracle. - **A feature with an acceptance criterion.** If the criterion can be stated as an assertion, it can be a hidden test. - **A refactor.** The existing suite is the oracle, with one added differential test: outputs on a set of inputs must match the previous build. Task shapes that do not come with a red test, such as "improve the error messages" or "clean up this module," are not oracle tasks as stated. They can be made into oracle tasks by adding a check, such as a snapshot test of the messages or a lint rule the cleanup must satisfy, or they can be done without the oracle and their traces kept as unlabeled data. ## Dispatch small A flip is binary. A task with three requirements and one hidden test that covers all three flips only when all three are done. Split it into three tasks with one hidden test each. The agent gets more feedback across the same work, the traces are shorter and cleaner, and a session that completes two of three is two flips instead of none. Small also means a short session. Long sessions drift, run out of context, and end on a FAIL that is as much about the session's length as about the task. SWE-bench Verified's human annotators excluded tasks whose descriptions were underspecified or whose tests were unfair, and the benchmark's resolve rates roughly doubled for the same models once the unfair tasks were removed.[2](#user-content-fn-swe-bench-verified) The lesson for a team is that a clear, bounded task description raises the flip rate without changing the agent. ## Sample more than once Repeated sampling is the single largest lever on the flip rate per task, and its cost is tokens. Brown and colleagues measured how coverage, the share of problems solved by at least one of k attempts, grows with k. On SWE-bench Lite, an open model that solved 15.9 percent of problems with one attempt solved 56 percent with 250 attempts.[3](#user-content-fn-large-language-monkeys) The oracle is what makes this usable: with a deterministic verifier, the team keeps the attempt that passes and discards the rest. Without one, the team has 250 patches and no way to choose. This is rejection sampling, and it is how most of the verified-trajectory datasets in the literature were built. SWE-Gym sampled many trajectories per task and kept the 491 that resolved their tasks.[4](#user-content-fn-swe-gym) SWE-smith kept 5,016 out of thousands more.[5](#user-content-fn-swe-smith) Llama 2's post-training used rejection sampling against a reward model as a core step.[6](#user-content-fn-llama2) For a team, the practical version is: dispatch each oracle task to several independent sessions with fresh context, and let the gate find the flip. The sessions must be independent. If one session's output leaks into another's context, the attempts are correlated and coverage grows slower. Fresh context per attempt and a silent oracle give independence. A verbose oracle that told each attempt why the last one failed would make the attempts a single long session in disguise. > **Coverage grows with attempts when a verifier picks the winner** > > - **DeepSeek-Coder-V2-Instruct on SWE-bench Lite**: 56% at 250 > > *Source: Brown et al. (2024). Only the two reported endpoints are plotted; the paper shows a roughly log-linear curve between them. Every added attempt costs tokens and, with a hidden oracle, each attempt is an independent draw.* A verifier trained on your own traces makes sampling cheaper. Pan and colleagues trained a verifier on the same SWE-Gym trajectories and used it to pick the best of 16 attempts, which raised their model from 20.6 to 32.0 percent.[4](#user-content-fn-swe-gym) The verifier does not replace the oracle; the oracle still grades the chosen attempt. The verifier reduces how many attempts need to reach the oracle. ## Manufacture tasks When the natural supply of red tests is smaller than the agent capacity, tasks can be made. **Break working code.** Take a module with good tests. Introduce a bug: delete a branch, flip a comparison, drop a null check. The existing tests that now fail are the hidden tests. The task description is the symptom, written from the test's point of view without naming the test. This is how SWE-smith built 50,000 task instances from 128 repositories, and the authors found that the synthetic bugs trained a model that transferred to real issues.[5](#user-content-fn-swe-smith) Mutation tools generate these bugs automatically, and the mutation literature has catalogued which mutants resemble real faults.[7](#user-content-fn-just-mutants) **Mine history.** Every past commit that changed source and tests together is a candidate task: the tests it added are the hidden tests, the parent commit is the base, and the commit message or linked issue is the task description. This is how SWE-bench was built and how SWE-rebench automated the construction into a pipeline that produced more than 21,000 tasks.[8](#user-content-fn-swe-bench)[9](#user-content-fn-swe-rebench) A team's own history is a supply of tasks on its own code that no public dataset contains. **Reverse a fix.** For a merged bug fix with a regression test, revert the source change and keep the test. The agent is dispatched to fix the bug again. The trace is a worked example on real code with a real test. **Let a model propose tasks under an executor.** Zhao and colleagues' Absolute Zero had a model propose coding tasks and solve them, with a code executor checking both that the task was well-formed and that the answer was right.[10](#user-content-fn-absolute-zero) The executor is the oracle. This is the most speculative item on the list and the one that produces the least realistic tasks; it is here because it shows that the supply of oracle tasks is not bounded by the supply of issues. ## Use the strong model on the hard tail Some tasks will not flip with the model you are training. Dispatch them to a rented frontier model and keep the trace. A flipped trace from a stronger model is a distillation example, and distillation from a stronger model into a weaker one is the oldest trick in the post-training book.[11](#user-content-fn-hinton-distillation) SWE-smith's 5,016 trajectories came from Claude 3.7 Sonnet, and the model trained on them was a 32-billion-parameter Qwen.[5](#user-content-fn-swe-smith) A team already paying for the frontier model is already producing these traces. The oracle is what sorts them. ## Keep the failures Most sessions will not flip. Those traces are not waste. A session that failed on the same task where another session flipped is half of a preference pair. Direct preference optimization trains a model to prefer the chosen trajectory over the rejected one, and it needs exactly this data.[12](#user-content-fn-dpo) A team that keeps only flips has a supervised set. A team that keeps everything has a supervised set and a preference set. A session that failed is also a record of what went wrong, for people. If a task fails ten sessions in a row, the oracle log will say whether the hidden test is wrong, the task description is unclear, or the task is beyond the model. Each of those has a different fix, and the trace is how you find out which. Hindsight relabeling, from robotics, is the formal version of this idea: an episode that failed its goal succeeded at whatever it did reach, and can be relabeled as a success for that.[13](#user-content-fn-her) A session that did not flip the hidden test but did make the existing suite pass after breaking it, or did fix a different bug on the way, has a relabeled success in it. The pipeline in Chapter 10 does not do this automatically, and it is the kind of thing a team adds once the basic loop runs. ## The arithmetic A team of 50 engineers runs an agent on perhaps 10 tasks a day each. Call it 500 sessions a day. If one in five sessions is on an oracle task, that is 100 oracle sessions a day. If one in three of those flips, that is about 33 flips a day. In a working month, around 700. Chapter 12 says what 700 buys. Doubling the oracle-task share doubles it; sampling three attempts per task roughly doubles it again. The levers are the share of work that has a red test, the size of each task, and the number of attempts. None of them is the model. ## Chapter 10: From traces to training data A trace store is not a training set. Between the two sit a series of transformations that decide what the model learns and what it does not. This chapter walks through them in the order the pipeline applies them. ## Select The first decision is which sessions to include. The rule for the supervised set is simple: a session is included when its outcome is `flipped` and its verdict record verifies. Everything else is excluded from the supervised set, and the reason is recorded. A session that flipped is not automatically a good example. Three further checks are cheap and worth running. - **The diff touched source.** A session whose only changes were to paths the oracle dropped did not flip because of anything the agent wrote to the code under test. If the oracle still reported PASS, the baseline was wrong. Exclude the session and quarantine the task. - **The session did not exceed the length budget.** A session with 400 tool calls that flipped is a worked example of flailing until something stuck. Set a budget per task class and exclude sessions over it, or truncate to the final successful stretch if the trace shows a clear restart. - **The agent did not attempt to touch excluded paths repeatedly.** The oracle's dropped-hunk log shows an agent that kept editing test configuration. That session may have flipped on the merits, but its trace teaches the behavior you least want. Exclude it, and look at whether the task description invited it. For the preference set, pair each flipped session on a task with each session on the same task that did not flip. Where there are many of each, sample pairs rather than taking the full cross product, so one task does not dominate. ## Deduplicate Two sessions on the same task that flipped with nearly identical trajectories add little to each other. Hash the sequence of tool names and the final diff; where two sessions share both, keep one. Where they share the diff but not the path to it, keep both, because the paths are the data. Across tasks, deduplicate on the task itself. A team that manufactured 200 tasks by mutating one module will have 200 near-identical trajectories, and a model trained on them will learn that module. Cap the number of flips per source file or per task family. ## Format A trace is a sequence of events. A training example is a conversation in the chat format the base model was trained on, with tool calls in the model's native tool-call syntax. The conversion is mechanical but has choices in it. The **system prompt** should be the one the agent ran with, including the rule files that were loaded, because that is the context the model will see at inference. If the team's rule files change often, consider training on a canonical version and keeping the diff as metadata. The **user turn** is the task description. Each **assistant turn** is either a tool call, with the tool name and arguments, or a message. Where the harness exposes the model's reasoning, include it as the base model's format expects; where it does not, the turn is the call alone. Each **tool result** is a tool turn, with the redacted output. The **final assistant turn** is the agent's completion message. Convert to the exact chat template of the base model you will fine-tune. A mismatch between the training template and the inference template is the most common silent failure in fine-tuning, and it produces a model that looks trained and behaves untrained. ## Mask The loss, the quantity the training run minimizes, should be computed only on the tokens the model is supposed to produce. In a trace, that is the assistant's turns: its reasoning, its tool calls, and its messages. The system prompt, the user's task, and every tool result are context, not targets. Masking the tool results matters more than it might seem. Tool outputs are long: a file read returns hundreds of lines, a test run returns pages. Unmasked, they are most of the tokens in the example, and the model spends its capacity learning to predict file contents and test logs instead of learning to act. A model trained without the mask will also learn to hallucinate tool results, because it was trained to produce them. The reference export tool emits a mask field per turn for this reason. ## Handle length Agent trajectories are long. A session that flipped after 60 tool calls, each returning a few kilobytes, is a hundred thousand tokens. Base models with long context windows handle this; training on sequences that long is expensive and some frameworks do not support it well. Three strategies, in order of preference: 1. **Train on the full trajectory** when the framework and hardware allow it. This preserves the behavior you want: the model learns to carry a plan across many steps. 2. **Truncate tool outputs**, not turns. A file read can be cut to the lines around the ones the agent later edited, with a marker. A test log can be cut to the failures. The trajectory keeps its shape and loses bulk. 3. **Window the trajectory** into overlapping segments, each with the system prompt and task prepended and a summary of the dropped prefix. This loses long-range structure and should be the fallback. Do not drop the dead ends. A trajectory in which the agent tried something, saw it fail, and backed out is a trajectory that teaches recovery. SWE-Gym's authors kept full trajectories including unproductive steps and reported that it worked; cleaning them to the shortest path is a reasonable experiment but not the default.[1](#user-content-fn-swe-gym) ## Hold out Before anything is trained, split by task, not by session. All sessions on a task go to the same side of the split. Set aside a fraction of tasks, with their hidden tests, as the evaluation set, and never train on any session from them. A model evaluated on tasks it saw during training will look better than it is, and the leak is undetectable afterward. Keep two evaluation sets. A **frozen set** fixed at the start of the program, so that every model version is scored against the same tasks and the trend is comparable. A **rolling set** of recent tasks, so that the evaluation tracks the work the team is doing now. Chapter 14 covers how to read them. ## Version Every training set is a manifest: the list of session ids included, the hash of each trace, the hash of each verdict, the filter rules applied, the template used, and the split. Store the manifest beside the weights it produced. When a model regresses, the manifest is how you find the traces that taught it. A training set that is reproducible from its manifest is the data equivalent of a reproducible build. The same traces and the same rules give the same examples, and a change in either is visible as a change in the hash. ## Record provenance Each example carries the identity of the oracle that graded it, as Chapter 4 said, and the identity of the model that produced the trace. The second matters more than it looks. A training set that is half traces from a rented frontier model and half from the team's own fine-tuned model is a mixture of distillation and self-improvement, and the two have different failure modes. Chapter 11 covers the self-improvement risk. The provenance field is what lets you tell them apart later. ## Chapter 11: Training methods With a training set in hand, the question is what to do with it. This chapter covers the three families of methods that have produced results on agent trajectories, in the order a team should adopt them, and the choices inside each: adapter or full fine-tune, how to avoid forgetting, and how to know the signal is real. ## Supervised fine-tuning on flips The first method is the simplest. Take the flipped trajectories, formatted and masked as Chapter 10 describes, and continue training the base model on them with the standard next-token loss. This is supervised fine-tuning, and when the examples were selected by a verifier, it has a second name: rejection sampling fine-tuning. The method has a lineage. Expert iteration, from 2017, alternates between a slow expert that solves problems and a fast policy that imitates the solutions, and uses the expert's successes as the training set.[1](#user-content-fn-expert-iteration) AlphaGo Zero's self-play is the same loop with the game's outcome as the verifier.[2](#user-content-fn-alphago-zero) STaR, from 2022, applied it to language models: sample reasoning, keep the samples whose answers match the key, fine-tune, repeat.[3](#user-content-fn-star) ReST and ReST-EM scaled it with a reward model and with answer keys.[4](#user-content-fn-rest)[5](#user-content-fn-rest-em) For code, AlphaCode filtered thousands of samples per problem through the problem's tests before choosing which to submit.[6](#user-content-fn-alphacode) Llama 2's post-training used rejection sampling as a step before reinforcement learning.[7](#user-content-fn-llama2) The agent-trajectory results from Chapter 1 are this method applied to software: SWE-Gym, SWE-smith, and Skywork-SWE are all supervised fine-tuning on verifier-selected trajectories.[8](#user-content-fn-swe-gym)[9](#user-content-fn-swe-smith)[10](#user-content-fn-skywork-swe) It works, it is cheap, and it is the right place to start. Its limit is that it can only teach the model to do what some session already did. A task no session ever flipped contributes nothing to the supervised set. ## Preference optimization on pairs The second method uses the failures. Direct preference optimization takes pairs of trajectories on the same task, one preferred and one not, and trains the model to assign higher likelihood to the preferred one relative to a reference model.[11](#user-content-fn-dpo) It needs no reward model and no sampling during training, which makes it almost as cheap as supervised fine-tuning. The preference set from Chapter 10 is the input: flipped versus not flipped on the same task. The pairs teach the model something the supervised set cannot, which is what a wrong trajectory looks like on this codebase. A team that has run two attempts per task has pairs for every task where exactly one flipped. Two cautions. Preference optimization is known to drift toward longer outputs unless the pairs are controlled for length, and agent trajectories vary in length a lot.[12](#user-content-fn-length-bias) Match pair lengths roughly or use a length-regularized variant. And the pairs should be real contrasts. A pair where the rejected trajectory failed because the oracle timed out is not a lesson about the code. ## Reinforcement learning with the oracle as the reward The third method puts the oracle in the training loop. The model attempts a task, the oracle grades the attempt, and the model is updated to make PASS more likely. This is reinforcement learning with verifiable rewards, named as such in the Tülu 3 report and used at scale in DeepSeek-R1.[13](#user-content-fn-tulu3)[14](#user-content-fn-deepseek-r1) For software agents, SWE-RL used a patch-similarity reward on mined pull requests, DeepSWE used a pass-or-fail reward on executable environments, and the Nebius team took a 72-billion-parameter model from 11.4 percent to 39.0 percent on SWE-bench Verified with rejection sampling followed by a variant of the DAPO algorithm.[15](#user-content-fn-swe-rl)[16](#user-content-fn-deepswe)[17](#user-content-fn-nebius-rl)[18](#user-content-fn-dapo) The appeal is that reinforcement learning can learn from tasks no session has flipped yet, because it searches. The costs are real. Each training step needs many fresh attempts, each attempt needs an oracle run, and the oracle runs in a container. Qwen's team described a system running 20,000 environments in parallel for its coding model's reinforcement learning stage.[19](#user-content-fn-qwen3-coder) A team does not need that scale to see gains; DeepSWE used about 4,500 tasks.[16](#user-content-fn-deepswe) It does need an oracle that can grade hundreds of attempts an hour, which means the oracle has to be a service, not a hook. There is also a result to read before committing. Yue and colleagues found that reinforcement learning with verifiable rewards sharpens a model toward answers its base model could already produce: the trained model wins when it gets one attempt, and the base model wins when both get many attempts, because the base model's attempts are more varied.[20](#user-content-fn-yue-rlvr) For a team, that argues for a sequence: supervised fine-tuning on flips first, to widen what the model can do on your code; reinforcement learning second, to make it do it reliably on the first try. Adopt the methods in that order. Supervised fine-tuning on flips as soon as there are a few hundred. Preference optimization as soon as there are pairs. Reinforcement learning when the oracle is a service and the flip rate has plateaued. > **Training methods in the order a team should adopt them** > > 1. **Supervised fine-tuning on flipped trajectories**: Rejection sampling fine-tuning. Cheapest. Learns what some session already did. > 2. **Preference optimization on flipped versus failed pairs**: Uses the failures. Learns what wrong looks like here. > 3. **Reinforcement learning with the oracle as reward**: Searches. Can learn tasks no session flipped. Needs the oracle as a service. > > *Each rung needs the one below it. The data each rung needs comes from the same trace store.* ## Adapter or full fine-tune Low-rank adaptation freezes the base model's weights and trains small matrices added to them.[21](#user-content-fn-lora) The adapter is a few percent of the model's size, trains on less hardware, and can be swapped at inference. QLoRA goes further by quantizing the frozen base to 4 bits, which brought fine-tuning of 65-billion-parameter models onto a single large card.[22](#user-content-fn-qlora) The question is whether the adapter learns as much. Biderman and colleagues compared the two on code and math and found that full fine-tuning learned more on the target domain and forgot more of what the base model knew; low-rank adaptation learned less and forgot less, and acted as a regularizer.[23](#user-content-fn-lora-forgets) Schulman and colleagues at Thinking Machines then reported that the gap closes when the adapter is applied to all layers rather than only attention, and that for small-to-medium post-training sets, the size a team's trace store will be for its first year, the adapter matches the full fine-tune. Their result for reinforcement learning is sharper: a rank-1 adapter matched full fine-tuning, which they explain by noting that a policy-gradient step absorbs about one bit of information per episode, so the capacity needed is tiny.[24](#user-content-fn-lora-without-regret) That last point connects to this book's oracle. A one-bit verdict is one bit per episode. A method that learns one bit per episode needs many episodes and almost no adapter capacity. That is the regime the pipeline is in, and the adapter is the right tool for it. Start with adapters on all layers. Move to full fine-tuning only if an evaluation shows the adapter is the bottleneck, which for the first several thousand flips it is unlikely to be. ## Forgetting A model fine-tuned on traces from one codebase can get worse at everything else. The phenomenon is catastrophic forgetting, known since 1989 and measured in language models at every scale.[25](#user-content-fn-mccloskey-forgetting)[26](#user-content-fn-luo-forgetting) For a coding model, it looks like a fine-tune that resolves the team's tasks and can no longer write a shell script. The standard defenses apply. - **Replay.** Mix a fraction of general instruction data into every training run. Ibrahim and colleagues showed that replay combined with re-warming the learning rate lets a model take on new data while matching a full retrain on the old.[27](#user-content-fn-ibrahim-continual) - **Low learning rate and few epochs.** One to three passes over the flips at a learning rate an order of magnitude below pretraining. - **Adapters.** The frozen base cannot forget. The adapter can be removed. This is the regularization Biderman and colleagues measured.[23](#user-content-fn-lora-forgets) - **Merge, don't stack.** When several adapters have been trained on different slices, merge them with a method that keeps the base model's weights and averages or resolves the deltas, such as model soups or TIES, rather than fine-tuning one on top of another.[28](#user-content-fn-model-soups)[29](#user-content-fn-ties-merging) - **Evaluate on a general benchmark.** Chapter 14 puts a public coding benchmark in the gate for this reason: a model that gained on the team's tasks and lost on the public set has forgotten, and the gate should see it. ## Collapse Training a model on its own output can make it worse over generations. Shumailov and colleagues showed that models trained recursively on their own generations lose the tails of the distribution and converge on a narrow set of outputs, a process they called model collapse.[30](#user-content-fn-shumailov-collapse) A pipeline that trains a model on traces produced by the previous version of itself is, on its face, exactly that loop. Two things make it different. First, the oracle. Collapse happens when generated data replaces real data without selection. The traces in this pipeline are selected by a verifier the model does not control, so the distribution being trained on is the distribution of correct solutions, not of the model's output. Gerstgrasser and colleagues showed that collapse is also avoided when generated data accumulates alongside the original data rather than replacing it.[31](#user-content-fn-gerstgrasser-collapse) Second, the provenance field from Chapter 10. A training run that knows which traces came from a stronger external model and which from its own predecessor can weight them, cap the self-generated share, and watch the ratio over time. The warning sign is a model whose flips get shorter, more uniform, and more alike across tasks. The frozen evaluation set from Chapter 10 is the instrument. ## Is the signal real One last check belongs in every training run, and it is cheap. Shao and colleagues found that training a particular family of math models with random rewards, rewards that had nothing to do with correctness, improved its benchmark scores by more than 20 points.[32](#user-content-fn-spurious-rewards) The gain came from the training procedure nudging the model toward behaviors it already had, not from the reward. The same procedure did nothing for other model families. The result does not say verifiable rewards are fake. It says a gain after training is not, on its own, evidence that the reward carried information. The control is to shuffle the labels. Train the same recipe on the same traces with the flip labels randomly permuted, so that half the "flipped" trajectories are failures. If the shuffled run gains as much as the real run on the held-out evaluation, the gain is not coming from the oracle, and something else in the recipe is doing the work. A team should run this control once when it sets up the pipeline, and again whenever the recipe changes. ## Chapter 12: Data volume How many flips does it take? This chapter collects the published numbers, states the pattern they show, and gives a worksheet a team can fill in with its own rates. ## What the literature reports The results below are for different models, methods, and benchmarks, so the table is a set of reference points, not a curve. Read it for order of magnitude. | Data | Method | Model | Result | Source | | --- | --- | --- | --- | --- | | 1,000 curated prompt-response pairs | Supervised fine-tuning | LLaMA 65B | Preferred to or tied with GPT-4 responses in 43 percent of human comparisons | LIMA[1](#user-content-fn-lima) | | 1,000 reasoning questions with traces | Supervised fine-tuning | Qwen2.5-32B-Instruct | Competitive with o1-preview on competition math; 57 percent on AIME24 with budget forcing | s1[2](#user-content-fn-s1) | | 817 curated math problems with solutions | Supervised fine-tuning | Qwen2.5-32B-Instruct | 57.1 percent on AIME24 and 94.8 percent on MATH500 in the first release | LIMO[3](#user-content-fn-limo) | | 500 agent trajectories | Supervised fine-tuning | Llama 2 7B | 77 percent relative gain on a question-answering agent task | FireAct[4](#user-content-fn-fireact) | | 1,866 agent trajectories across six tasks | Supervised fine-tuning with general data mixed in | Llama 2 7B to 70B | Agent abilities generalize to held-out tasks | AgentTuning[5](#user-content-fn-agenttuning) | | 491 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 7.0 to 20.6 percent on SWE-bench Verified; 32.0 with a trained verifier and 16 samples | SWE-Gym[6](#user-content-fn-swe-gym) | | 5,016 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 40.2 percent on SWE-bench Verified | SWE-smith[7](#user-content-fn-swe-smith) | | 8,209 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 6.4 to 38.0 percent; 47.0 with best-of-8 and a critic; log-linear in data with no plateau | Skywork-SWE[8](#user-content-fn-skywork-swe) | | About 4,500 executable tasks | Reinforcement learning only | Qwen3-32B | 23 to 42.2 percent; 59.0 with test-time scaling | DeepSWE[9](#user-content-fn-deepswe) | | Rejection sampling, then RL on executable tasks | Both | Qwen2.5-72B-Instruct | 11.4 to 20.5 to 39.0 percent | Nebius[10](#user-content-fn-nebius-rl) | Three patterns run through the table. **Hundreds move a model.** Every result in the first half of the table used about a thousand examples or fewer and produced a large change. LIMA's authors argued that almost all of a model's knowledge comes from pretraining and that alignment needs only a small set of examples to teach format and style.[1](#user-content-fn-lima) The agent results say something stronger for this domain: 491 verified trajectories nearly tripled a 32-billion-parameter model's resolve rate.[6](#user-content-fn-swe-gym) The first useful model is closer than most teams expect. **Thousands keep paying.** Skywork-SWE's scaling curve is the most direct measurement: 2,000 trajectories gave 31.8 percent, 6,000 gave 36.1, and 8,209 gave 38.0, with the curve still rising.[8](#user-content-fn-skywork-swe) Zhang and colleagues found the same shape across fine-tuning tasks and model sizes and fit it as a power law in the amount of fine-tuning data.[11](#user-content-fn-zhang-scaling) The returns diminish per example and do not stop. **Quality beats quantity at every scale.** AlpaGasus trained on 9,000 examples filtered from a 52,000-example set and beat the model trained on all 52,000.[12](#user-content-fn-alpagasus) LIMA, s1, and LIMO are all arguments that a small curated set beats a large uncurated one. For this pipeline, the oracle is the curation. A flip is, by construction, an example that was verified. The volume question is how many verified examples, not how many sessions. > **Verified trajectories behind each open-model result on SWE-bench Verified** > > | Result | Trajectories | > | --- | --- | > | SWE-Gym, 20.6% after fine-tuning | 491 | > | Skywork-SWE at 31.8% | 2000 | > | SWE-smith, 40.2% | 5016 | > | Skywork-SWE at 36.1% | 6000 | > | Skywork-SWE, 38.0% | 8209 | > > *Sources: Pan et al. (2024), Yang et al. (2025), Zeng et al. (2025). The same 32-billion-parameter base model in every row. The Skywork rows are points on one scaling curve; the other two are separate recipes.* ## The worksheet Four numbers set the time to a given training set size. 1. **Sessions per day.** Engineers times sessions each. A team of 50 running 10 each is 500. 2. **Oracle share.** The fraction of sessions dispatched on tasks that have a hidden test. A team starting out might reach 20 percent. A team that writes the test first for every bug and most features can reach 60. 3. **Flip rate.** The fraction of oracle sessions that end on PASS. With a silent oracle and a rented frontier model on tasks of reasonable size, 30 to 50 percent is a working assumption; a team should measure it in its first week. 4. **Attempts per task.** Independent sessions dispatched per oracle task. Each attempt costs tokens and raises the chance that at least one flips. Flips per day is roughly: sessions × oracle share × flip rate, adjusted upward for attempts. With the numbers above and one attempt, 500 × 0.2 × 0.33 is about 33 flips a day. With three attempts and a per-attempt flip rate of 33 percent, the chance a task flips at least once is about 70 percent, so flips per task-day rises to about 70 of 100 oracle tasks, at three times the token cost. > **Working days to reach a training set size, one attempt per task** > > | | 20 percent oracle share | 60 percent oracle share | > | --- | --- | --- | > | 500 flips, first fine-tune | 15days | 5days | > | 2,000 flips | 60days | 20days | > | 5,000 flips | 150days | 50days | > | 8,000 flips | 240days | 80days | > > *Arithmetic for 500 sessions a day at a 33 percent flip rate. Sampling three attempts per task roughly halves every number at three times the token cost. These are planning figures, not measurements.* The reading of the chart is that a mid-sized team reaches the SWE-Gym regime in weeks and the Skywork regime within a year, and that the oracle share is the lever that matters most. Every bug fixed without a hidden test is a session that could have been a flip and was not. ## Limits of volume More flips do not fix a bad oracle. A thousand traces graded by a flaky test are a thousand noisy labels, and the noise does not average out; it teaches. More flips do not fix a narrow task distribution. Five thousand flips on one service teach that service. And more flips do not fix a leaked hidden test; they make the leak worse, because every trace fitted to the leaked test is one more example of fitting. The order of operations is: get the oracle right, get the task supply broad, then grow the volume. Chapter 13 is the pipeline that does the third once the first two are in place. ## Chapter 13: The delivery pipeline Continuous delivery is the practice of keeping software in a state where any change can be released at any time, by automating the path from a commit to production and gating each step on checks.[1](#user-content-fn-humble-farley) This chapter applies it to weights. The artifact is a model. The commit is a new batch of verified traces. The release is a new model serving the team's agents. Everything between them is a stage with an input, an output, and a gate. ## The stages > **From trace store to serving model** > > 1. **Ingest**: Sync traces and signed verdicts from agent hosts and oracle hosts > 2. **Redact and scan**: Second-pass secret scan; quarantine hits > 3. **Validate**: Schema check; verdict signatures; baseline sanity > 4. **Select and version**: Filters from Chapter 10; write the manifest > 5. **Train**: Adapter on all layers; replay mix; fixed seed > 6. **Evaluate**: Held-out flips, frozen and rolling; public benchmark subset > 7. **Gate**: Beat the serving model on held-out flips; no regression on the public set > 8. **Package**: Merge or ship the adapter; quantize; record the manifest hash > 9. **Canary**: Shadow mode on live tasks; the oracle scores both models > 10. **Promote or roll back**: Route traffic; keep the previous weights warm > > Repeat every batch. > > *The emphasized stages are the ones that decide. Everything else is plumbing, and all of it is ordinary continuous delivery with a model as the artifact.* **Ingest.** A scheduled job pulls finished trace files from each agent host's plugin data directory and verdict records from the oracle host into the trace store. Files are content-addressed: the path includes the hash of the contents, so a re-upload is a no-op and a tampered file lands at a different path. The job is safe to run twice. **Redact and scan.** The hook redacted with patterns at write time. This stage runs a dedicated secret scanner over every trace and quarantines any with a hit. Quarantined traces are not deleted; a person reviews them, because a false positive on a test fixture is common and a true positive is a credential to rotate. **Validate.** Every event parses against the schema. Every verdict record's signature verifies against the oracle's public key. Every session has a `session_start`, a `session_end`, and at least one verdict, or it is marked incomplete and excluded. A session whose baseline was PASS is excluded and its task is flagged. **Select and version.** The filters from Chapter 10 run. The output is a manifest: the list of sessions in the supervised set, the pairs in the preference set, the held-out task list, the template version, and the hash of each. The manifest is the thing that is versioned. The training data is derived from it. **Train.** The training job takes a manifest and a recipe and produces weights. The recipe is a file: base model and its hash, adapter rank and target modules, learning rate, epochs, replay mix and its source, seed. The job records the recipe hash and the manifest hash in the weights' metadata. A training run with the same manifest, recipe, and seed should produce the same weights, within the limits of the hardware's determinism, and the pipeline should check that it does once. **Evaluate.** The new weights run against the held-out tasks. For each task, the model attempts it in a fresh sandbox with the same harness the team uses, and the oracle grades the result. The metric is the flip rate on the frozen set and on the rolling set. The weights also run against a fixed subset of a public coding benchmark, for the forgetting check. Chapter 14 is about reading these. **Gate.** The new weights are promoted only if the frozen-set flip rate is at least the serving model's, the rolling-set flip rate is higher, and the public-set score has not fallen by more than a set tolerance. A run that fails the gate is kept, with its evaluation, so the trend is visible even when nothing ships. **Package.** The adapter is merged into the base weights or shipped as a separate adapter, depending on how the serving layer loads models. The weights are quantized if the serving hardware needs it, and the quantized weights are re-evaluated on the frozen set, because quantization can cost more on a fine-tuned model than on its base. The package carries the manifest hash and the recipe hash. **Canary.** Before the new model serves anyone, it runs in shadow. For a sample of live oracle tasks, both the serving model and the candidate attempt the task in parallel sandboxes. The oracle grades both. The person sees only the serving model's result. After enough tasks, the candidate's live flip rate against the serving model's is the number that decides. This is the same design as shadow mode in driving systems, where a new model runs alongside the one in control and its decisions are compared without being acted on. **Promote or roll back.** Promotion is a routing change. The previous weights stay loaded for a period, and rollback is the same routing change reversed. A rollback is not a failure of the pipeline; it is the pipeline working. ## Cadence How often to run depends on the flip rate. A team producing 30 flips a day has 200 new examples a week, which is enough to retrain weekly and expect to see movement. A team producing 5 a day should batch monthly. A training run on an adapter for a 32-billion-parameter model over a few thousand trajectories takes hours on a single node with eight large accelerators; the evaluation, which runs an agent on each held-out task, often takes longer than the training. Budget for both. The cadence should be a schedule, not a trigger. A pipeline that trains whenever enough new traces arrive produces models at irregular intervals that are hard to compare. A pipeline that trains every Sunday night and evaluates on Monday produces a weekly series. ## What is deterministic and what is not The oracle is deterministic by construction. Training is deterministic up to the hardware. Evaluation is not: the model samples, and the harness is a live process. Fix the sampling temperature and seed for evaluation, run each held-out task more than once, and report the mean. Two models within noise of each other are tied, and the gate should say so instead of promoting on a coin flip. ## The record Every stage writes a record to the same store the traces live in: what it took in, what it produced, the hashes, the time, and the outcome. A model in production can be traced back to the manifest that trained it, to the sessions in the manifest, to the oracle that graded each session, and to the hidden test behind each verdict. That chain is what makes a regression debuggable and what makes the pipeline auditable to anyone who asks where the model came from. This is where the author's own work connects, and the connection is disclosed. Oxagen, the company the author founded, records agent runs with their tool calls, their outcomes, and their costs as a product. A run record is a trace in this book's sense. The pipeline in this chapter does not depend on it; the plugin writes its own traces to its own files. But the reason the record is kept beside the verdict, rather than inside the model's context, is the same reason Oxagen keeps the outcome outside the agent that produced it. ## Chapter 14: Evaluation A pipeline that trains weights needs a way to know whether the new weights are better. This chapter is about building that measurement so that it stays honest as the pipeline optimizes against it. ## Held-out flips are the metric The primary metric is the flip rate on tasks the model never trained on, graded by the same oracle that grades production sessions. It is the metric that matches the goal. A public benchmark measures how well the model solves public tasks. Held-out flips measure how well it solves yours. Two sets, as Chapter 10 said. The frozen set is drawn once, at the start, from tasks across the team's repositories, and never changes. Every model version is scored against it, so the series is comparable. The rolling set is the most recent few hundred oracle tasks, refreshed each cycle, so the score tracks the current work. A model that gains on the frozen set and not the rolling set has learned the past. A model that gains on the rolling set and not the frozen set has learned something narrow about recent work. The gate wants both. ## Contamination A held-out task that the model saw during training is not held out. The leak is easy to create and impossible to detect after the fact. Deng and colleagues showed that language models can reproduce the missing parts of benchmark items they were trained on, which is how contamination shows up in public evaluations.[1](#user-content-fn-deng-contamination) The SWE-bench+ study found that about a third of one agent's passing patches on SWE-bench had the solution available in the issue text or its comments, and another third passed because the tests were too weak to catch a wrong fix; filtering those out dropped the measured resolve rate from 12.47 percent to 3.97 percent.[2](#user-content-fn-swe-bench-plus) A team's internal evaluation is exposed to the same two failures: the hidden test can leak, and the hidden test can be weak. Four rules keep the held-out set clean. 1. Split by task, with every session on a task on the same side. 2. Never train on a trace from a held-out task, including traces from the rented frontier model. The provenance field makes this checkable. 3. Never show the held-out hidden tests to any agent, including the one generating training traces. The one-bit oracle enforces this for the agent; the pipeline has to enforce it for the people. 4. Retire a held-out task when its hidden test is released into the repository, which Chapter 6 recommended once a task has been flipped in production. A test the agents can see is no longer held out. ## Weak tests A hidden test that passes for a wrong fix inflates the flip rate and teaches the wrong fix. Mutation testing measures this directly: generate mutants of the code the task touches and check that the hidden test kills them.[3](#user-content-fn-just-mutants) A held-out task whose hidden test kills few mutants near the change is a weak oracle, and its flips mean less. The evaluation report should carry the mutation score of each task's hidden test beside the flip rate, so a gain concentrated on weak tasks is visible. ## The public benchmark A fixed subset of a public coding benchmark sits in the gate for one purpose: catching forgetting. The team's model should not get worse at general coding while it gets better at the team's code. SWE-bench Verified, or a multilingual benchmark if the team's code is not Python, serves.[4](#user-content-fn-swe-bench-verified)[5](#user-content-fn-multi-swe-bench) The score itself is secondary. The change in the score between versions is the signal. Do not optimize for it. A pipeline that gates on a public benchmark and also trains on data derived from public repositories will, over time, find the benchmark's tasks in its training data. The public set is a thermometer, not a target. ## Watching for hacking The oracle is air-gapped, so the agent cannot change the verdict. The agent can still produce a trace that flipped for a reason the team would not endorse, and a model trained on such traces learns the reason. Baker and colleagues found that a monitor reading the agent's chain of thought caught most reward hacking, and that a weaker model was an adequate monitor.[6](#user-content-fn-baker-monitoring) In this pipeline, the equivalent is a scan of each flipped trace for patterns that should not be there: edits to paths on the denylist, even though the oracle dropped them; commands that search for the hidden tests; tool calls that read the oracle's configuration; test runs that were made to pass by changing the test. The scan flags, a person reads, and flagged traces are excluded from training until cleared. The scan should run on traces from the rented model too, because distillation copies behavior. Track the dropped-hunk rate from the oracle log. A rising share of diffs with hunks on the denylist means the agents are learning to touch the tests, and the next model will learn it faster. ## Drift in what the model produces Three cheap measurements catch most regressions that the flip rate misses. - **Length.** Mean tokens per session and per tool call. Preference optimization and reward-driven training both tend toward verbosity unless controlled.[7](#user-content-fn-length-bias) A model whose sessions get longer without flipping more often is spending the team's tokens. - **Tool mix.** The distribution of tool calls per session. A model that stops running tests, or starts reading every file in the repository, has changed in a way the flip rate will show late. - **Similarity.** Pairwise similarity between the model's trajectories on different tasks. Rising similarity is the early sign of collapse from Chapter 11. ## The report Each evaluation produces one report with: flip rate on the frozen set and the rolling set, each with its uncertainty from repeated runs; the public-benchmark score and its change; the mutation score distribution of the held-out tests; the counts of flagged traces by reason; the three drift measurements; and the manifest and recipe hashes. The gate reads the report. So do people. A report that a person cannot read in five minutes is too long. ## Chapter 15: Serving and the economics The model is trained. This chapter is about running it: where it serves, how requests are routed between it and the rented model, and the arithmetic that says when the pipeline has paid for itself. ## Serving Open-weight models serve through a small number of mature engines. vLLM introduced paged attention, which manages the key-value cache in blocks the way an operating system manages memory, and made high-throughput serving of large models practical on commodity accelerators.[1](#user-content-fn-vllm) Serving an adapter without merging it is supported directly; S-LoRA showed that thousands of adapters over one base model can be served from a single node with the base weights shared.[2](#user-content-fn-s-lora) For a team with one fine-tune per repository or per domain, that means one base model in memory and a set of small adapters, with the request choosing the adapter. Quantization reduces memory and raises throughput at some cost in quality. A 32-billion-parameter model quantized to 4 bits fits on a single large accelerator. Re-evaluate on the frozen set after quantizing, as Chapter 13 said, because the cost is not uniform across models. ## Routing The team's model does not have to handle everything on day one. A router sends each task to the team's model first and falls back to the rented frontier model when the team's model fails. The oracle makes the fallback decision cheap: a session that does not flip after the team's model's attempts is re-dispatched to the rented model. The fallback sessions are traces. They are the hard tail of the distribution, solved by a stronger model, and graded by the oracle. They go into the trace store with their provenance and become the distillation examples from Chapter 9. Over time, the share of tasks that reach the fallback is the measure of how far the team's model has come, and it is also the training set for closing the gap. > **Routing a task between the team's model and the rented model** > > 1. **Task**: An oracle task is dispatched > 2. **Team model**: Attempts it; the gate asks the oracle > 3. **Flip**: PASS: done. The trace is a self-generated example. > 4. **Fallback**: No flip after N attempts: re-dispatch to the rented model > 5. **Rented model**: Attempts it; the same gate, the same oracle > 6. **Trace**: Either way, the trace goes to the store with its provenance > > *The fallback rate falls as the team's model learns. The fallback traces are what it learns from.* ## The arithmetic The cost side has four lines. **Oracle compute.** Each oracle run is a container running a test suite. On a cloud runner, a suite that takes five minutes costs cents. At 100 oracle sessions a day with two attempts each and a baseline per task, that is a few hundred container-minutes a day. **Storage.** A trace with redacted tool output is tens of kilobytes to a few megabytes. A year of a mid-sized team's traces is tens to hundreds of gigabytes. This is a rounding error. **Training.** An adapter run on a 32-billion-parameter model over a few thousand trajectories is hours on one eight-accelerator node. At on-demand cloud prices, that is low hundreds to low thousands of dollars per run. Weekly, it is tens of thousands a year. Evaluation on a few hundred held-out tasks, each an agent session in a sandbox, costs about as much again. **Serving.** One eight-accelerator node, or two for redundancy, serves a 32-billion-parameter model to a team of fifty with capacity to spare. Reserved, this is low six figures a year; owned, it is a capital cost amortized over several years. The revenue side is the rent avoided. A team of fifty engineers running agents daily on a frontier model spends, at mid-2026 prices, in the range of several hundred thousand to a few million dollars a year, depending on how heavily they use it. The exact figure is the team's own bill, and the team should use it. The comparison is not all-or-nothing. In the routing design above, the team's model handles the share of tasks it can flip, and the rented model handles the rest. If the team's model flips 40 percent of oracle tasks in its first quarter, 40 percent of the rent on those tasks is replaced by serving cost. As the share rises, the rent falls. The crossover, where the pipeline's total cost falls below the rent it replaces, depends on the team's bill, but for a team spending more than the cost of two accelerator nodes a year on tokens, it arrives within the first year of flips. ## What the price decline does to this Token prices fall fast, and the rent will be lower next year.[3](#user-content-fn-epoch-prices) Three things keep the arithmetic in the pipeline's favor anyway. Open-model serving costs fall on the same curve, because the hardware and the serving software improve for everyone. The ratio between rent and serving cost is more stable than either number. The rented model's price per token is for a model that does not know the team's code. The team's model's cost per token is for one that does. If the team's model flips a task in fewer tokens, which a model trained on that codebase should, the comparison per task is better than the comparison per token. And the rent buys no asset. The pipeline's cost buys weights, a trace store, and an oracle, all of which the team keeps. The question in Chapter 1 was what the team owns at the end of the year. The arithmetic here is about when owning it is also cheaper, and the answer for most teams of this size is: soon. ## Chapter 16: Governance and failure modes A pipeline that collects everything an agent did and trains a model on it has new ways to go wrong. This chapter lists them, with what the research says and what the pipeline does about each. The failures are ordered by how much damage they do, not by how likely they are. ## Secrets in the trace store The trace store holds the output of every command the agents ran. If a secret passed through a terminal, it is in a trace unless redaction caught it. A trace store is the most complete record of a team's operational secrets that has ever existed in one place, and it should be protected like one. The defenses are layered. Redaction at write time, with patterns. A dedicated scanner at ingest, with quarantine. Access control on the store, with the training job as the only automated reader. Encryption at rest. And a short retention period for raw tool output, after which only the selected and scanned training examples remain. The last one is a tradeoff: it forecloses re-processing old traces with better redaction. A team should decide its retention consciously. ## Memorization A model trained on traces can reproduce them. Carlini and colleagues extracted training examples from a deployed language model by prompting it, including names, contact details, and code.[1](#user-content-fn-carlini-extraction) In a follow-up, they measured how memorization scales: it grows with model size, with how many times an example was duplicated in training, and with how much of the example's prefix the prompt supplies.[2](#user-content-fn-carlini-memorization) A fine-tune on a few thousand trajectories, each seen for several epochs, is in the regime where memorization is expected. Two consequences. First, a secret that survived redaction into training can be extracted from the model. The secret scanning in the pipeline is the control, and it has to be good. Second, the model is a copy of the team's code in a form that can be queried. It is a private asset and should be served privately. A team that fine-tunes on its traces and then exposes the model to people outside the team has published its code in a lossy format. Chapter 1 said the traces were the asset nobody else has. That is only true while the model trained on them stays inside. ## Tampering with the oracle Chapter 5 built the air gap. The governance question is who can change what is inside it. The hidden tests, the allowlist, the container image, and the signing key are the pipeline's root of trust. Changes to them should go through the same review as changes to production code, and the oracle log should record every verdict with the hash of the test list that produced it, so a change in the tests is visible as a change in the hash. Rotate the hidden tests for a task once it has flipped in production, as Chapter 6 said, and retire the task from the held-out set. ## Training on the wrong lesson The oracle grades the result. It does not grade the path. A session that flipped by reading the hidden test's name from a stray log line, by copying a fix from a sibling repository that happened to be checked out, or by special-casing the inputs it guessed the test used, is a flip with a bad path. The scan in Chapter 14 catches some of these. The rest are caught by reading flipped traces, which a person should do for a sample every cycle. A pipeline nobody reads is a pipeline that trains on whatever got through. ## The pipeline itself Sculley and colleagues catalogued the ways machine-learning systems accumulate debt that ordinary software does not: data dependencies that nobody tracks, feedback loops where the model's output changes its own training data, configuration that grows without review, and pipelines glued together from pieces nobody owns.[3](#user-content-fn-sculley-debt) This pipeline has every one of those. Its training data comes from agents running its own previous model, which is a feedback loop by design. Its configuration is a recipe file, an allowlist, a template, and a set of hidden tests, each of which changes the model when it changes. Its stages are a trace collector, an oracle, a trainer, an evaluator, and a router, built from different tools. The defenses are the ordinary ones, applied without exception. Every input to a stage is versioned. Every stage records what it did. Every change to configuration is reviewed. Every model can be traced to its manifest. And the frozen evaluation set, which never changes, is the one fixed point against which drift in everything else is measured. ## Loss of the flip signal If the oracle share falls, because the team stops writing tests first, the flip rate falls with it and the pipeline starves. If the task supply narrows, because one repository produces most of the oracle tasks, the model narrows with it. If the fallback rate stops falling, the team's model has plateaued and the recipe needs to change. Each of these is visible in the pipeline's own records, and the weekly report should carry them: oracle share, flip rate, tasks per repository, fallback rate. ## People The last failure mode is the one the research does not cover. A pipeline that grades every agent session on a hidden test is also, if someone chooses to read it that way, a pipeline that grades the engineers who dispatched the sessions. It should not be used that way. The flip rate is a property of the task, the oracle, and the model. The moment it becomes a measure of a person, people will stop dispatching tasks that might not flip, the oracle share will fall, and the pipeline will starve. Say this in writing when the pipeline is introduced, and keep the per-person numbers out of the report. ## Chapter 17: Precedents The idea in this book is not new. It is the combination of several ideas that are each at least a decade old, applied to a kind of data that did not exist until coding agents did. This chapter lists the precedents, what each one established, and what it left for this pipeline to add. ## Verifier-filtered self-training Expert iteration, from Anthony, Tian, and Barber in 2017, is the general form: a slow, strong solver produces solutions, a fast policy is trained to imitate them, and the loop repeats with the improved policy as the new starting point.[1](#user-content-fn-expert-iteration) AlphaGo Zero, the same year, ran the loop with the outcome of the game as the only verifier and reached superhuman play from random weights.[2](#user-content-fn-alphago-zero) The verifier was perfect, cheap, and deterministic, which is the ideal this book's oracle approximates. STaR brought the loop to language models in 2022, with an answer key as the verifier.[3](#user-content-fn-star) ReST and ReST-EM scaled it.[4](#user-content-fn-rest)[5](#user-content-fn-rest-em) The reinforcement-learning-with-verifiable-rewards line, from Tülu 3 through DeepSeek-R1 and DAPO, is the same loop with a gradient step instead of a fine-tune on the filtered set.[6](#user-content-fn-tulu3)[7](#user-content-fn-deepseek-r1)[8](#user-content-fn-dapo) Cobbe and colleagues' 2021 GSM8K work made the case that a trained verifier scales better than fine-tuning alone, which is the argument for the verifier in Chapter 9.[9](#user-content-fn-cobbe-verifiers) What these established: a verifier the model does not control produces lasting gains, and the model's own judgment does not. What they left: the verifier was an answer key or a game. Software has no answer key. It has tests. ## Tests as the verifier for code AlphaCode, in 2022, sampled up to a million programs per competition problem and filtered them through the problem's example tests before clustering and submitting.[10](#user-content-fn-alphacode) CodeRL, the same year, used unit-test results as the reward for training a code model with an actor-critic method.[11](#user-content-fn-coderl) Meta's RLEF, in 2024, trained a code model with execution feedback as the reward and showed large gains in sample efficiency.[12](#user-content-fn-rlef) SWE-RL, DeepSWE, and the Nebius work brought it to repository-scale tasks.[13](#user-content-fn-swe-rl)[14](#user-content-fn-deepswe)[15](#user-content-fn-nebius-rl) What these established: a test is a usable reward, and a model trained against tests gets better at passing tests. What they left: every one of them used public problems with public tests. The tests were visible, or at least public, and the problems were nobody's in particular. ## Benchmarks built from real repositories SWE-bench, in 2023, defined the fail-to-pass and pass-to-pass construction and built 2,294 tasks from real GitHub issues and the pull requests that closed them.[16](#user-content-fn-swe-bench) SWE-bench Verified had human annotators remove the tasks whose descriptions were unclear or whose tests were unfair, leaving 500.[17](#user-content-fn-swe-bench-verified) SWE-Gym, SWE-smith, R2E-Gym, and SWE-rebench turned the construction into pipelines that produce thousands of executable tasks.[18](#user-content-fn-swe-gym)[19](#user-content-fn-swe-smith)[20](#user-content-fn-r2e-gym)[21](#user-content-fn-swe-rebench) Multi-SWE-bench extended it to seven languages, and SWE-Lancer graded real freelance tasks with end-to-end tests.[22](#user-content-fn-multi-swe-bench)[23](#user-content-fn-swe-lancer) SWE-bench+ showed how often the construction leaks the answer or accepts a wrong one.[24](#user-content-fn-swe-bench-plus) What these established: the flip is a reliable unit of value for software work, and tasks with hidden tests can be produced at scale from version history. What they left: the tasks are public, so every model has seen them, and the tests are released, so every agent can be shown them. A team's own history is the unlimited supply of tasks that are not public. ## Learning from an organization's own development process This is the closest precedent and the least cited. Facebook's Getafix, in 2019, learned fix patterns from the history of human fixes to static-analysis warnings in Facebook's own codebase and proposed fixes for new warnings, which engineers accepted at a high rate.[25](#user-content-fn-getafix) SapFix, the same year, generated candidate patches for crashes found by an automated testing system and used that system's tests as the oracle, end to end, in production.[26](#user-content-fn-sapfix) Both learned from, and were graded by, the company's own code and tests. Google's DIDACT, described in 2023, trained models on the process of software development inside Google rather than on finished code: the edit histories, the build-error fixes, the code-review comments and their resolutions.[27](#user-content-fn-didact) Google had earlier reported that a completion model trained on its internal code reduced coding iteration time by 6 percent and was accepted for about 3 percent of new code.[28](#user-content-fn-google-completion) The data was the developers' own activity, and the organization kept it. GitHub's Copilot research found that acceptance rate was the best available predictor of developers' perceived productivity, which made acceptance a usable, if weak, oracle at scale.[29](#user-content-fn-ziegler-copilot) Replit, in 2024, trained a 7-billion-parameter code-repair model on data built from its own platform: language-server diagnostics, with the state of the file reconstructed by replaying the edit history.[30](#user-content-fn-replit-repair) What these established: an organization's own development activity is training data, and models trained on it perform well on that organization's work. What they left: each was built by a company with a research team and a bespoke agent or tool. The point of Chapters 7 and 8 is that the harness hooks make this available to a team with neither. ## Learning from traces in other fields End-to-end driving models from 2016 onward trained on recorded human driving, with the steering angle as the label.[31](#user-content-fn-bojarski-driving) The traces were cheap to collect from vehicles already on the road, and the fleet's data became the moat. The shadow-mode canary in Chapter 13 is borrowed from this field. Imitation learning has a known failure: a policy trained on an expert's trajectories drifts into states the expert never visited, and its errors compound. DAgger, from 2011, fixes it by running the learner, having the expert label the states the learner reached, and training on those.[32](#user-content-fn-dagger) The analogue here is the fallback in Chapter 15: when the team's model fails a task, the rented model solves it from the same starting point, and the trace is a label on a state the team's model reached. Hindsight experience replay, from 2017, relabels failed episodes as successes for the goals they did reach, which Chapter 9 proposed for sessions that did not flip.[33](#user-content-fn-her) Research on machine learning for code, surveyed by Allamanis and colleagues in 2018, rests on the observation from Hindle and colleagues in 2012 that software is natural: repetitive and predictable enough that statistical models of it work.[34](#user-content-fn-hindle-naturalness)[35](#user-content-fn-allamanis-survey) A team's codebase is more repetitive and more predictable than the public corpus, which is why a model trained on it does well there. ## Keeping a holdout honest The one-bit verdict in Chapter 6 rests on the Ladder and the reusable holdout, both from 2015.[36](#user-content-fn-ladder)[37](#user-content-fn-reusable-holdout) Both were written about machine-learning competitions and scientific data analysis. Neither mentions an agent. The problem they solved, an adaptive optimizer hill-climbing on a holdout it is only supposed to be measured by, is the problem an agent iterating against an oracle has, and the solution transfers without modification. ## Agents trained on agent trajectories AgentTuning and FireAct, in 2023, fine-tuned open models on a few hundred to a couple of thousand agent trajectories generated by a stronger model and showed that the result generalized.[38](#user-content-fn-agenttuning)[39](#user-content-fn-fireact) Kimi K2's technical report described a large-scale pipeline for synthesizing agentic data with tool use and verifying it before training.[40](#user-content-fn-kimi-k2) The open-weight agent frameworks, SWE-agent and OpenHands, standardized the tool interface that these trajectories are recorded in.[41](#user-content-fn-swe-agent)[42](#user-content-fn-openhands) What these established: agent behavior transfers through trajectories, and a few hundred are enough to see it. What they left: the trajectories came from public tasks solved by public models. The trajectories this book is about come from your tasks, solved on your code, graded by your tests. ## What is new Four things, and only four. Hooks in the harness make trace collection free. The organization does not build the agent, and the agent does not have to be modified. The hidden test as an air-gapped, one-bit oracle makes the grading trustworthy under optimization pressure, which the public-benchmark work did not have to worry about and the internal-tool work handled with bespoke infrastructure. Open-weight models are close enough to the frontier, and licensed permissively enough, that the fine-tuned result is competitive on the team's distribution. In 2019 the models were not there. In 2023 the licenses often were not. And the delivery pipeline treats weights as a release artifact with gates, canaries, and rollback, which is ordinary engineering applied to a thing that used to be a research project. Everything else in this book was established by someone else, and the footnotes say who. ## Closing: What to do this quarter 1. **Install the collector.** The plugin in Chapter 8, or your own hooks that write the same schema. Start keeping traces this week, before the oracle exists. Unlabeled traces are still an asset, and the collector is the part with no dependencies. 2. **Pick one repository and write the first hidden tests.** Twenty tasks, each with a test that fails on the base commit for the right reason. Bug reports with reproductions are the easiest source. 3. **Run the oracle in a container with no network.** On the same machine at first. Move it to a runner the agents cannot log into before you train on anything. 4. **Measure the flip rate.** Dispatch the twenty tasks. Count the flips. That number, times your sessions per day, is your data rate, and Chapter 12 says what it buys. 5. **Keep the failures.** They are half of the preference set. 6. **Set aside the held-out tasks now.** Before the first training run, not after. The leak cannot be undone. 7. **Train the first adapter at 500 flips.** On all layers, with replay, with the shuffled-label control beside it. Evaluate on the held-out set. Expect it to be worse than the rented model and better than the base model. That is the first point on a curve the team now owns. ## Appendix A: Trace event schema Each session is one JSON Lines file. Each line is one event. Every event carries `v` (schema version), `session_id`, `ts` (ISO 8601, UTC), and `type`. ```text session_start cwd, base_commit, remote_url_hash, model, harness, harness_version, source, task_id, config_hash message turn, role (user | assistant), text_redacted, text_hash, truncated (bool) tool_call turn, tool_name, tool_use_id, input_redacted, input_hash, output_redacted, output_hash, truncated (bool), duration_ms oracle_verdict attempt, mode (baseline | grade), verdict (PASS | FAIL | ERROR), tests_hash, image_digest, diff_hash, oracle_host, signature (optional) session_end outcome (flipped | no_flip | no_task | aborted), attempts, final_diff_hash, reason (from the harness), ended_at ``` Rules: `text_redacted`, `input_redacted`, and `output_redacted` are the only fields that ever hold content, and they hold it after redaction. The `_hash` fields are SHA-256 of the unredacted value, so a later pass can tell whether two truncated values were the same without storing them. `remote_url_hash` rather than the URL, because remote URLs sometimes carry credentials. The verdict's `signature` is over the canonical JSON of the other verdict fields, with the oracle host's key. ## Appendix B: Oracle container specification - **Image**: pinned by digest. Built from a Dockerfile that installs the repository's dependencies from its lockfile at build time. Rebuilt when the lockfile changes, and the new digest recorded. - **Network**: none. Started with networking disabled. - **Mounts**: the task's hidden tests, read-only; the diff file, read-only, in grade mode. Nothing from the agent's working copy. - **Repository**: a mirror baked into the image or mounted read-only from the oracle host. Checked out at the base commit inside the container. - **Environment**: `TZ=UTC`, `LC_ALL=C.UTF-8`, `PYTHONHASHSEED=0`, `SOURCE_DATE_EPOCH` fixed, and whatever the repository's own test configuration needs to be deterministic. - **Diff application**: the diff is filtered through the allowlist and denylist first. Dropped hunks are logged with their paths. The filtered diff is applied with `git apply --check` before `git apply`; a diff that does not apply is a FAIL with the reason logged. - **Hidden tests**: copied into the checkout after the diff is applied, from the mount, never from the diff. - **Run**: the existing suite at the base commit's version, then the hidden tests. Any failure in either is a FAIL. - **Timeout**: a wall-clock limit per run. Exceeding it is a FAIL with the reason logged. - **Output to the gate**: exit code 0 or 1, and one line `tests_hash=` of the sorted list of test identifiers that ran. - **Output to the oracle log**: everything. Test output, dropped hunks, timing, the exit reason, the image digest, the diff hash. - **Baseline mode**: the same run without a diff. Must return FAIL for the hidden tests and PASS for the existing suite, or the task is quarantined. ## Appendix C: Data volume worksheet Fill in the first four lines from your own team. The rest is arithmetic. ```text A sessions per day = engineers × sessions each = ____ B oracle share = fraction of sessions on tasks with a hidden test = ____ C flip rate per attempt = measured in week one = ____ D attempts per task = independent sessions dispatched = ____ E oracle tasks per day = A × B = ____ F chance a task flips at least once = 1 − (1 − C)^D = ____ G flips per day = E × F = ____ H token cost multiplier = D = ____ Days to 500 flips = 500 / G Days to 2,000 flips = 2,000 / G Days to 8,000 flips = 8,000 / G ``` Reference points from Chapter 12: 491 flips moved a 32-billion-parameter model from 7.0 to 20.6 percent on SWE-bench Verified; about 8,000 moved the same base model to 38.0 percent with the gain still growing. ## Sources 119 works are cited in this book. They appear below in the order of their first citation. Each chapter also lists its own sources at its end. 1. Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). *How far behind are open models?* Epoch AI. [https://epoch.ai/blog/open-models-report](https://epoch.ai/blog/open-models-report) 2. Stanford Institute for Human-Centered Artificial Intelligence (2025). *AI Index Report 2025*, Chapter 2: Technical Performance. [https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance](https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance) 3. OpenAI (2025). *gpt-oss-120b and gpt-oss-20b Model Card*. arXiv. [https://arxiv.org/abs/2508.10925](https://arxiv.org/abs/2508.10925) 4. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. arXiv; ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) 5. Panickssery, A., Bowman, S. R., & Feng, S. (2024). *LLM Evaluators Recognize and Favor Their Own Generations*. arXiv; NeurIPS 2024. [https://arxiv.org/abs/2404.13076](https://arxiv.org/abs/2404.13076) 6. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). *STaR: Bootstrapping Reasoning With Reasoning*. arXiv. [https://arxiv.org/abs/2203.14465](https://arxiv.org/abs/2203.14465) 7. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) 8. Baker, B., Huizinga, J., Gao, L., et al. (2025). *Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation*. arXiv. [https://arxiv.org/abs/2503.11926](https://arxiv.org/abs/2503.11926) 9. Denison, C., MacDiarmid, M., Barez, F., et al. (2024). *Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models*. arXiv. [https://arxiv.org/abs/2406.10162](https://arxiv.org/abs/2406.10162) 10. Blum, A., & Hardt, M. (2015). *The Ladder: A Reliable Leaderboard for Machine Learning Competitions*. ICML 2015, PMLR 37, 1006–1014. [https://arxiv.org/abs/1502.04585](https://arxiv.org/abs/1502.04585) 11. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). *The reusable holdout: Preserving validity in adaptive data analysis*. Science, 349(6248), 636–638. [https://doi.org/10.1126/science.aaa9375](https://doi.org/10.1126/science.aaa9375) 12. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) 13. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) 14. Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290) 15. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). *LLM inference prices have fallen rapidly but unequally across tasks*. Epoch AI. [https://epoch.ai/data-insights/llm-inference-price-trends](https://epoch.ai/data-insights/llm-inference-price-trends) 16. Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449) 17. Mistral AI & All Hands AI (2025). *Devstral*. [https://mistral.ai/news/devstral](https://mistral.ai/news/devstral) 18. Qwen Team (2025). *Qwen3 Technical Report*. arXiv. [https://arxiv.org/abs/2505.09388](https://arxiv.org/abs/2505.09388) 19. Grattafiori, A., et al. (2024). *The Llama 3 Herd of Models*. arXiv. [https://arxiv.org/abs/2407.21783](https://arxiv.org/abs/2407.21783) 20. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) 21. Anthropic (2026). *Hooks reference*. Claude Code documentation. [https://code.claude.com/docs/en/hooks](https://code.claude.com/docs/en/hooks) 22. Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). *The Oracle Problem in Software Testing: A Survey*. IEEE Transactions on Software Engineering, 41(5), 507–525. [https://doi.org/10.1109/TSE.2014.2372785](https://doi.org/10.1109/TSE.2014.2372785) 23. Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). *An Empirical Analysis of Flaky Tests*. FSE 2014, 643–653. [https://doi.org/10.1145/2635868.2635920](https://doi.org/10.1145/2635868.2635920) 24. Micco, J. (2016). *Flaky Tests at Google and How We Mitigate Them*. Google Testing Blog. [https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html) 25. Lamb, C., & Zacchiroli, S. (2022). *Reproducible Builds: Increasing the Integrity of Software Supply Chains*. IEEE Software, 39(2), 62–70. [https://arxiv.org/abs/2104.06020](https://arxiv.org/abs/2104.06020) 26. Just, R., Jalali, D., & Ernst, M. D. (2014). *Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs*. ISSTA 2014, 437–440. [https://doi.org/10.1145/2610384.2628055](https://doi.org/10.1145/2610384.2628055) 27. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). *An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems*. ISSTA 2015, 24–36. [https://doi.org/10.1145/2771783.2771791](https://doi.org/10.1145/2771783.2771791) 28. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). *Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair*. ESEC/FSE 2015, 532–543. [https://doi.org/10.1145/2786805.2786825](https://doi.org/10.1145/2786805.2786825) 29. Inozemtseva, L., & Holmes, R. (2014). *Coverage Is Not Strongly Correlated with Test Suite Effectiveness*. ICSE 2014, 435–445. [https://doi.org/10.1145/2568225.2568271](https://doi.org/10.1145/2568225.2568271) 30. DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). *Hints on Test Data Selection: Help for the Practicing Programmer*. IEEE Computer, 11(4), 34–41. [https://doi.org/10.1109/C-M.1978.218136](https://doi.org/10.1109/C-M.1978.218136) 31. Petrović, G., & Ivanković, M. (2018). *State of Mutation Testing at Google*. ICSE-SEIP 2018, 163–171. [https://doi.org/10.1145/3183519.3183521](https://doi.org/10.1145/3183519.3183521) 32. Claessen, K., & Hughes, J. (2000). *QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs*. ICFP 2000, 268–279. [https://doi.org/10.1145/351240.351266](https://doi.org/10.1145/351240.351266) 33. Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). *A Survey on Metamorphic Testing*. IEEE Transactions on Software Engineering, 42(9), 805–824. [https://doi.org/10.1109/TSE.2016.2532875](https://doi.org/10.1109/TSE.2016.2532875) 34. McKeeman, W. M. (1998). *Differential Testing for Software*. Digital Technical Journal, 10(1), 100–107. 35. Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). *Measuring GitHub Copilot's Impact on Productivity*. Communications of the ACM, 67(3), 54–63. [https://doi.org/10.1145/3633453](https://doi.org/10.1145/3633453) 36. Von Arx, S., Chan, L., & Barnes, E. (2025). *Recent Frontier Models Are Reward Hacking*. METR. [https://metr.org/blog/2025-06-05-recent-reward-hacking/](https://metr.org/blog/2025-06-05-recent-reward-hacking/) 37. MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). *Natural Emergent Misalignment from Reward Hacking in Production RL*. arXiv. [https://arxiv.org/abs/2511.18397](https://arxiv.org/abs/2511.18397) 38. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). *Defining and Characterizing Reward Hacking*. NeurIPS 2022. [https://arxiv.org/abs/2209.13085](https://arxiv.org/abs/2209.13085) 39. Gao, L., Schulman, J., & Hilton, J. (2022). *Scaling Laws for Reward Model Overoptimization*. arXiv; ICML 2023. [https://arxiv.org/abs/2210.10760](https://arxiv.org/abs/2210.10760) 40. Thompson, K. (1984). *Reflections on Trusting Trust*. Communications of the ACM, 27(8), 761–763. [https://doi.org/10.1145/358198.358210](https://doi.org/10.1145/358198.358210) 41. Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). *Firecracker: Lightweight Virtualization for Serverless Applications*. NSDI 2020, 419–434. [https://www.usenix.org/conference/nsdi20/presentation/agache](https://www.usenix.org/conference/nsdi20/presentation/agache) 42. Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). *The True Cost of Containing: A gVisor Case Study*. HotCloud 2019. [https://www.usenix.org/conference/hotcloud19/presentation/young](https://www.usenix.org/conference/hotcloud19/presentation/young) 43. Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). *Teaching Large Language Models to Self-Debug*. arXiv; ICLR 2024. [https://arxiv.org/abs/2304.05128](https://arxiv.org/abs/2304.05128) 44. Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*. NeurIPS 2023 Datasets and Benchmarks. [https://arxiv.org/abs/2306.05685](https://arxiv.org/abs/2306.05685) 45. Wang, P., Li, L., Chen, L., et al. (2023). *Large Language Models are not Fair Evaluators*. arXiv; ACL 2024. [https://arxiv.org/abs/2305.17926](https://arxiv.org/abs/2305.17926) 46. Stechly, K., Marquez, M., & Kambhampati, S. (2023). *GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems*. arXiv. [https://arxiv.org/abs/2310.12397](https://arxiv.org/abs/2310.12397) 47. Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). *Can Large Language Models Really Improve by Self-critiquing Their Own Plans?* arXiv. [https://arxiv.org/abs/2310.08118](https://arxiv.org/abs/2310.08118) 48. Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). *LLMs cannot find reasoning errors, but can correct them given the error location*. Findings of ACL 2024. [https://arxiv.org/abs/2311.08516](https://arxiv.org/abs/2311.08516) 49. Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). *When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs*. Transactions of the ACL, 12. [https://arxiv.org/abs/2406.01297](https://arxiv.org/abs/2406.01297) 50. Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). *Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement*. ACL 2024. [https://arxiv.org/abs/2402.11436](https://arxiv.org/abs/2402.11436) 51. Meli, M., McNiece, M. R., & Reaves, B. (2019). *How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories*. NDSS 2019. [https://doi.org/10.14722/ndss.2019.23418](https://doi.org/10.14722/ndss.2019.23418) 52. Anthropic (2026). *Plugin manifest reference*. Claude Code documentation. [https://code.claude.com/docs/en/plugins-reference](https://code.claude.com/docs/en/plugins-reference) 53. Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). *Realizing quality improvement through test driven development: results and experiences of four industrial teams*. Empirical Software Engineering, 13(3), 289–302. [https://doi.org/10.1007/s10664-008-9062-z](https://doi.org/10.1007/s10664-008-9062-z) 54. OpenAI (2024). *Introducing SWE-bench Verified*. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/) 55. Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). *Large Language Monkeys: Scaling Inference Compute with Repeated Sampling*. arXiv. [https://arxiv.org/abs/2407.21787](https://arxiv.org/abs/2407.21787) 56. Touvron, H., et al. (2023). *Llama 2: Open Foundation and Fine-Tuned Chat Models*. arXiv. [https://arxiv.org/abs/2307.09288](https://arxiv.org/abs/2307.09288) 57. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). *Are Mutants a Valid Substitute for Real Faults in Software Testing?* FSE 2014, 654–665. [https://doi.org/10.1145/2635868.2635929](https://doi.org/10.1145/2635868.2635929) 58. Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). *SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2505.20411](https://arxiv.org/abs/2505.20411) 59. Zhao, A., Wu, Y., Yue, Y., et al. (2025). *Absolute Zero: Reinforced Self-play Reasoning with Zero Data*. arXiv. [https://arxiv.org/abs/2505.03335](https://arxiv.org/abs/2505.03335) 60. Hinton, G., Vinyals, O., & Dean, J. (2015). *Distilling the Knowledge in a Neural Network*. NeurIPS 2014 Deep Learning Workshop. [https://arxiv.org/abs/1503.02531](https://arxiv.org/abs/1503.02531) 61. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. [https://arxiv.org/abs/2305.18290](https://arxiv.org/abs/2305.18290) 62. Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). *Hindsight Experience Replay*. NeurIPS 2017. [https://arxiv.org/abs/1707.01495](https://arxiv.org/abs/1707.01495) 63. Anthony, T., Tian, Z., & Barber, D. (2017). *Thinking Fast and Slow with Deep Learning and Tree Search*. NeurIPS 2017. [https://arxiv.org/abs/1705.08439](https://arxiv.org/abs/1705.08439) 64. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). *Mastering the game of Go without human knowledge*. Nature, 550, 354–359. [https://doi.org/10.1038/nature24270](https://doi.org/10.1038/nature24270) 65. Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). *Reinforced Self-Training (ReST) for Language Modeling*. arXiv. [https://arxiv.org/abs/2308.08998](https://arxiv.org/abs/2308.08998) 66. Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). *Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models*. arXiv; TMLR 2024. [https://arxiv.org/abs/2312.06585](https://arxiv.org/abs/2312.06585) 67. Li, Y., Choi, D., Chung, J., et al. (2022). *Competition-level code generation with AlphaCode*. Science, 378(6624), 1092–1097. [https://doi.org/10.1126/science.abq1158](https://doi.org/10.1126/science.abq1158) 68. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). *A Long Way to Go: Investigating Length Correlations in RLHF*. arXiv; COLM 2024. [https://arxiv.org/abs/2310.03716](https://arxiv.org/abs/2310.03716) 69. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). *Tülu 3: Pushing Frontiers in Open Language Model Post-Training*. arXiv. [https://arxiv.org/abs/2411.15124](https://arxiv.org/abs/2411.15124) 70. Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe) 71. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501) 72. Yu, Q., Zhang, Z., Zhu, R., et al. (2025). *DAPO: An Open-Source LLM Reinforcement Learning System at Scale*. arXiv. [https://arxiv.org/abs/2503.14476](https://arxiv.org/abs/2503.14476) 73. Qwen Team (2025). *Qwen3-Coder: Agentic Coding in the World*. [https://qwenlm.github.io/blog/qwen3-coder/](https://qwenlm.github.io/blog/qwen3-coder/) 74. Yue, Y., Chen, Z., Lu, R., et al. (2025). *Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?* arXiv; NeurIPS 2025. [https://arxiv.org/abs/2504.13837](https://arxiv.org/abs/2504.13837) 75. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). *LoRA: Low-Rank Adaptation of Large Language Models*. arXiv; ICLR 2022. [https://arxiv.org/abs/2106.09685](https://arxiv.org/abs/2106.09685) 76. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). *QLoRA: Efficient Finetuning of Quantized LLMs*. NeurIPS 2023. [https://arxiv.org/abs/2305.14314](https://arxiv.org/abs/2305.14314) 77. Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). *LoRA Learns Less and Forgets Less*. Transactions on Machine Learning Research. [https://arxiv.org/abs/2405.09673](https://arxiv.org/abs/2405.09673) 78. Schulman, J., & Thinking Machines Lab (2025). *LoRA Without Regret*. [https://thinkingmachines.ai/blog/lora/](https://thinkingmachines.ai/blog/lora/) 79. McCloskey, M., & Cohen, N. J. (1989). *Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem*. Psychology of Learning and Motivation, 24, 109–165. [https://doi.org/10.1016/S0079-7421(08)60536-8](https://doi.org/10.1016/S0079-7421\(08\)60536-8) 80. Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). *An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning*. arXiv. [https://arxiv.org/abs/2308.08747](https://arxiv.org/abs/2308.08747) 81. Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). *Simple and Scalable Strategies to Continually Pre-train Large Language Models*. arXiv; TMLR. [https://arxiv.org/abs/2403.08763](https://arxiv.org/abs/2403.08763) 82. Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). *Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time*. ICML 2022. [https://arxiv.org/abs/2203.05482](https://arxiv.org/abs/2203.05482) 83. Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). *TIES-Merging: Resolving Interference When Merging Models*. NeurIPS 2023. [https://arxiv.org/abs/2306.01708](https://arxiv.org/abs/2306.01708) 84. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). *AI models collapse when trained on recursively generated data*. Nature, 631, 755–759. [https://doi.org/10.1038/s41586-024-07566-y](https://doi.org/10.1038/s41586-024-07566-y) 85. Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). *Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data*. arXiv. [https://arxiv.org/abs/2404.01413](https://arxiv.org/abs/2404.01413) 86. Shao, R., Li, S. S., Xin, R., et al. (2025). *Spurious Rewards: Rethinking Training Signals in RLVR*. arXiv. [https://arxiv.org/abs/2506.10947](https://arxiv.org/abs/2506.10947) 87. Zhou, C., Liu, P., Xu, P., et al. (2023). *LIMA: Less Is More for Alignment*. NeurIPS 2023. [https://arxiv.org/abs/2305.11206](https://arxiv.org/abs/2305.11206) 88. Muennighoff, N., Yang, Z., Shi, W., et al. (2025). *s1: Simple test-time scaling*. arXiv. [https://arxiv.org/abs/2501.19393](https://arxiv.org/abs/2501.19393) 89. Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). *LIMO: Less is More for Reasoning*. arXiv (v1, February 2025); COLM 2025. [https://arxiv.org/abs/2502.03387](https://arxiv.org/abs/2502.03387) 90. Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). *FireAct: Toward Language Agent Fine-tuning*. arXiv. [https://arxiv.org/abs/2310.05915](https://arxiv.org/abs/2310.05915) 91. Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). *AgentTuning: Enabling Generalized Agent Abilities for LLMs*. arXiv. [https://arxiv.org/abs/2310.12823](https://arxiv.org/abs/2310.12823) 92. Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). *When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method*. ICLR 2024. [https://arxiv.org/abs/2402.17193](https://arxiv.org/abs/2402.17193) 93. Chen, L., Li, S., Yan, J., et al. (2023). *AlpaGasus: Training a Better Alpaca with Fewer Data*. arXiv; ICLR 2024. [https://arxiv.org/abs/2307.08701](https://arxiv.org/abs/2307.08701) 94. Humble, J., & Farley, D. (2010). *Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation*. Addison-Wesley. 95. Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). *Investigating Data Contamination in Modern Benchmarks for Large Language Models*. NAACL 2024. [https://arxiv.org/abs/2311.09783](https://arxiv.org/abs/2311.09783) 96. Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). *SWE-Bench+: Enhanced Coding Benchmark for LLMs*. arXiv. [https://arxiv.org/abs/2410.06992](https://arxiv.org/abs/2410.06992) 97. Zan, D., Huang, Z., Liu, W., et al. (2025). *Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving*. arXiv. [https://arxiv.org/abs/2504.02605](https://arxiv.org/abs/2504.02605) 98. Kwon, W., Li, Z., Zhuang, S., et al. (2023). *Efficient Memory Management for Large Language Model Serving with PagedAttention*. SOSP 2023. [https://arxiv.org/abs/2309.06180](https://arxiv.org/abs/2309.06180) 99. Sheng, Y., Cao, S., Li, D., et al. (2023). *S-LoRA: Serving Thousands of Concurrent LoRA Adapters*. arXiv; MLSys 2024. [https://arxiv.org/abs/2311.03285](https://arxiv.org/abs/2311.03285) 100. Carlini, N., Tramèr, F., Wallace, E., et al. (2021). *Extracting Training Data from Large Language Models*. USENIX Security 2021. [https://arxiv.org/abs/2012.07805](https://arxiv.org/abs/2012.07805) 101. Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). *Quantifying Memorization Across Neural Language Models*. ICLR 2023. [https://arxiv.org/abs/2202.07646](https://arxiv.org/abs/2202.07646) 102. Sculley, D., Holt, G., Golovin, D., et al. (2015). *Hidden Technical Debt in Machine Learning Systems*. NeurIPS 2015. 103. Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). *Training Verifiers to Solve Math Word Problems*. arXiv. [https://arxiv.org/abs/2110.14168](https://arxiv.org/abs/2110.14168) 104. Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). *CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning*. NeurIPS 2022. [https://arxiv.org/abs/2207.01780](https://arxiv.org/abs/2207.01780) 105. Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). *RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2410.02089](https://arxiv.org/abs/2410.02089) 106. Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). *R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents*. arXiv. [https://arxiv.org/abs/2504.07164](https://arxiv.org/abs/2504.07164) 107. Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). *SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?* arXiv. [https://arxiv.org/abs/2502.12115](https://arxiv.org/abs/2502.12115) 108. Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). *Getafix: Learning to Fix Bugs Automatically*. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. [https://doi.org/10.1145/3360585](https://doi.org/10.1145/3360585) 109. Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). *SapFix: Automated End-to-End Repair at Scale*. ICSE-SEIP 2019, 269–278. [https://doi.org/10.1109/ICSE-SEIP.2019.00039](https://doi.org/10.1109/ICSE-SEIP.2019.00039) 110. Maniatis, P., & Tarlow, D. (2023). *Large sequence models for software development activities*. Google Research Blog. [https://research.google/blog/large-sequence-models-for-software-development-activities/](https://research.google/blog/large-sequence-models-for-software-development-activities/) 111. Tabachnyk, M., & Nikolov, S. (2022). *ML-Enhanced Code Completion Improves Developer Productivity*. Google Research Blog. [https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/](https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/) 112. Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). *Building LLMs for Code Repair*. Replit. [https://replit.com/blog/code-repair](https://replit.com/blog/code-repair) 113. Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). *End to End Learning for Self-Driving Cars*. arXiv. [https://arxiv.org/abs/1604.07316](https://arxiv.org/abs/1604.07316) 114. Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). *A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning*. AISTATS 2011. [https://arxiv.org/abs/1011.0686](https://arxiv.org/abs/1011.0686) 115. Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). *On the Naturalness of Software*. ICSE 2012, 837–847. [https://doi.org/10.1109/ICSE.2012.6227135](https://doi.org/10.1109/ICSE.2012.6227135) 116. Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). *A Survey of Machine Learning for Big Code and Naturalness*. ACM Computing Surveys, 51(4), Article 81. [https://arxiv.org/abs/1709.06182](https://arxiv.org/abs/1709.06182) 117. Kimi Team (2025). *Kimi K2: Open Agentic Intelligence*. arXiv. [https://arxiv.org/abs/2507.20534](https://arxiv.org/abs/2507.20534) 118. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering*. NeurIPS 2024. [https://arxiv.org/abs/2405.15793](https://arxiv.org/abs/2405.15793) 119. Wang, X., Li, B., Song, Y., et al. (2024). *OpenHands: An Open Platform for AI Software Developers as Generalist Agents*. arXiv; ICLR 2025. [https://arxiv.org/abs/2407.16741](https://arxiv.org/abs/2407.16741) --- # Chapter 1: Rented tokens > The token bill falls every year. The capability and the data never arrive. What a team owns after a year of renting. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/rented-tokens A team that uses a frontier model through an API is renting. The rent has three parts, and only the first shows up on the invoice. The first part is the token bill. It is real and it is falling. Epoch AI measured the price of reaching a fixed level of performance on a set of benchmarks and found it falling by somewhere between 9 and 900 times a year, depending on the task.[1](#user-content-fn-epoch-prices) That is good news for anyone who pays the bill, and it is the number vendors point to when someone asks why a team should not train its own model. The number is also beside the point. The question is not whether tokens will get cheaper. They will. The question is what the team owns at the end of the year. The vendors in question are the frontier labs: Anthropic, OpenAI, Google, and the others whose models you reach only through an API and whose weights you never see. A black box is the right name for the product, whatever you think of the company. The second part of the rent is the capability you never keep. When a vendor ships a better model, your agents get better. When the vendor changes the model, deprecates it, raises the price, or changes the terms, your agents change with it. Nothing the model learned about your systems during a year of sessions stays with you, because the model learned nothing. It cannot. Inference does not update weights. Every session starts from the same checkpoint the vendor shipped, and every piece of context your team has written to steer it, every CLAUDE.md and every rule file, is a workaround for a model that does not know your code. The third part is the data you discard. A session produces a record. The record shows which files the agent opened to understand a feature, which commands it ran to check its work, what it tried that failed, and what finally passed. If the agent was working in your repository on your task against your tests, that record is a worked example of your work being done. Nobody else has it. Nobody else can produce it. And in most teams, it is gone when the terminal closes. > **Where the value of a session goes today** > > 1. **Task**: An engineer or a work queue dispatches a task to a coding agent > 2. **Session**: The agent reads, edits, runs commands, and reports done > 3. **Tokens billed**: The vendor bills every input and output token > 4. **Record discarded**: The transcript is deleted or buried in a local cache > > *The two emphasized steps are the rent. The money leaves and the record leaves. The model learned nothing, and neither did your organization's data.* ## What the open-weight trend changes The case for keeping traces rests on one bet: that open-weight models will stay close enough to the frontier that a model trained on your traces beats a rented model on your tasks. Three measurements say the bet is reasonable. Epoch AI compared the best open-weight model to the best closed model across benchmarks and found the open model trailing by about a year, measured by the date at which a closed model first reached the same score.[2](#user-content-fn-epoch-open) A year behind the frontier on general benchmarks is not a year behind on your codebase, because the frontier model is also starting from zero on your codebase. Stanford's AI Index reported that the gap between the best open and best closed models on the Chatbot Arena leaderboard narrowed from 8.0 percent to 1.7 percent in one year.[3](#user-content-fn-ai-index-2025) Chatbot Arena measures human preference on chat, not software engineering, but the direction holds across benchmarks the Index tracks. On the benchmark that matters most for this book, SWE-bench Verified, open-weight models went from under 10 percent to above 40 percent in eighteen months. Meta's SWE-RL took a 70-billion-parameter Llama to 41.0 percent.[4](#user-content-fn-swe-rl) Mistral's Devstral, a 24-billion-parameter model that runs on a single workstation card, reported 46.8 percent.[5](#user-content-fn-devstral) The SWE-smith team took a 32-billion-parameter Qwen to 40.2 percent with about five thousand trajectories.[6](#user-content-fn-swe-smith) The Skywork team took the same base model from 6.4 percent to 38.0 percent with about eight thousand, and found the gain still growing with each doubling of the data.[7](#user-content-fn-skywork-swe) These are not the frontier numbers. The frontier was above 70 percent at the time. They are the numbers that a team can run, inspect, and fine-tune. > **Open-weight models on SWE-bench Verified** > > | Model and method | Resolved | > | --- | --- | > | Qwen2.5-Coder-32B, no agent training (SWE-Gym baseline) | 7% | > | Same model after fine-tuning on 491 SWE-Gym trajectories | 20.6% | > | Same model with a verifier and 16 samples per task | 32% | > | Skywork-SWE-32B, 8,209 trajectories | 38% | > | SWE-agent-LM-32B, 5,016 SWE-smith trajectories | 40.2% | > | Llama3-SWE-RL-70B, reinforcement learning | 41% | > | Devstral-Small-2505, 24B | 46.8% | > | DeepSWE-Preview, Qwen3-32B, RL with test-time scaling | 59% | > > *Sources: Pan et al. (2024), Yang et al. (2025), Wei et al. (2025), Mistral (2025), Agentica and Together AI (2025). Each row is a different base model and training recipe, so the bars show the range open models reached, not a controlled comparison.* Two things about these numbers matter for the argument. First, the jump from 7.0 to 20.6 percent came from fine-tuning on 491 trajectories.[8](#user-content-fn-swe-gym) That is not a large dataset. It is about a week of sessions for a mid-sized team. The further jump to 32.0 percent came from training a verifier on the same trajectories and letting it pick the best of 16 attempts, which is a preview of Chapter 9. Second, every model in the table was trained on public repositories solving public issues. None of them had seen the training team's own code. A model trained on your traces starts from these numbers and climbs on your distribution. ## Licenses that permit it The models in the table ship under licenses that allow fine-tuning and commercial use. DeepSeek-R1 and its distilled variants are MIT licensed.[9](#user-content-fn-deepseek-r1) The Qwen2.5 and Qwen3 families are Apache 2.0, with a small number of size variants under a Qwen license.[10](#user-content-fn-qwen3) The Llama 3 family uses Meta's community license, which permits commercial use below a very large monthly-user threshold.[11](#user-content-fn-llama3) OpenAI's gpt-oss models are Apache 2.0.[12](#user-content-fn-gpt-oss) The point of listing them is not legal advice. It is that the permission to do what this book describes is ordinary and granted. ## The asset that compounds Consider two teams of the same size doing the same work for one year. Team A uses a rented frontier model. It spends on tokens, writes rule files, and gets better at prompting. At the end of the year it has a set of rule files, a bill, and a vendor relationship. If the vendor's next model is worse at the team's tasks, the team has no recourse. If a competitor uses the same vendor, the competitor has the same model. Team B uses the same rented model, and also installs a trace collector and an oracle. It spends the same on tokens. Each session that flips an oracle from FAIL to PASS is saved as a verified trajectory. At the end of the year, Team B has thousands of worked examples of its own work being done correctly, each graded by a test the model did not write. It has fine-tuned an open model on them three or four times and has a model that resolves its own tasks at a rate the rented model cannot match on the same distribution, running on hardware it controls, at a marginal cost per token that is a fraction of the rent. If the vendor changes terms, Team B's model does not change. If a competitor wants the same model, the competitor needs Team B's traces. The cost of being Team B instead of Team A, for the first several months, is close to zero. The collector is a plugin. The oracle is a test. Storage is cheap. The training runs come later and are cheaper than one month of a mid-sized team's token bill. The only thing Team A has to do to become Team B is to stop deleting the record. ## Limits of the claim It does not claim that a fine-tuned 32-billion-parameter model will match the best frontier model on every task next quarter. It will not. It does not claim that the open-weight gap will close to zero. It may not. It does not claim that collecting traces is free of risk; Chapter 16 is about secrets, memorization, and the ways a pipeline like this goes wrong. It claims that the traces are the asset, that the oracle is what makes them an asset, and that a team which starts collecting today will be in a different position in a year from a team that does not. The rest of the book is about how to do it carefully. ## Footnotes 1. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). *LLM inference prices have fallen rapidly but unequally across tasks*. Epoch AI. [https://epoch.ai/data-insights/llm-inference-price-trends](https://epoch.ai/data-insights/llm-inference-price-trends) [↩](#user-content-fnref-epoch-prices) 2. Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). *How far behind are open models?* Epoch AI. [https://epoch.ai/blog/open-models-report](https://epoch.ai/blog/open-models-report) [↩](#user-content-fnref-epoch-open) 3. Stanford Institute for Human-Centered Artificial Intelligence (2025). *AI Index Report 2025*, Chapter 2: Technical Performance. [https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance](https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance) [↩](#user-content-fnref-ai-index-2025) 4. Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449) [↩](#user-content-fnref-swe-rl) 5. Mistral AI & All Hands AI (2025). *Devstral*. [https://mistral.ai/news/devstral](https://mistral.ai/news/devstral) [↩](#user-content-fnref-devstral) 6. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) [↩](#user-content-fnref-swe-smith) 7. Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290) [↩](#user-content-fnref-skywork-swe) 8. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) 9. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) [↩](#user-content-fnref-deepseek-r1) 10. Qwen Team (2025). *Qwen3 Technical Report*. arXiv. [https://arxiv.org/abs/2505.09388](https://arxiv.org/abs/2505.09388) [↩](#user-content-fnref-qwen3) 11. Grattafiori, A., et al. (2024). *The Llama 3 Herd of Models*. arXiv. [https://arxiv.org/abs/2407.21783](https://arxiv.org/abs/2407.21783) [↩](#user-content-fnref-llama3) 12. OpenAI (2025). *gpt-oss-120b and gpt-oss-20b Model Card*. arXiv. [https://arxiv.org/abs/2508.10925](https://arxiv.org/abs/2508.10925) [↩](#user-content-fnref-gpt-oss) --- # Chapter 2: The trace > The record of one session, why the tool calls carry the signal, and the flip as the unit of value. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-trace A trace is the record of one agent session. The word is borrowed from systems tracing, where a trace is the record of one request as it passes through many services. The analogy is close. An agent session passes through many tool calls, and the trace records all of them in order with their inputs and outputs. ## Anatomy of a trace A coding-agent session has a regular structure. The agent receives a task. It takes a turn: it thinks, then it either calls a tool or writes a message. A tool call runs and returns a result. The agent takes another turn. This continues until the agent decides it is done, or a person stops it, or a limit is reached. The trace records each of these events: 1. The **task** as the agent received it, including the system prompt, the rule files loaded into context, and the user's instruction. 2. Each **tool call**: the tool name, the arguments the agent passed, the result the tool returned, and how long it took. 3. Each **message** the agent wrote to the person, including the final message where it says the work is done. 4. The **repository state** at the start and end: the base commit, the final diff, and the list of files touched. 5. The **oracle verdicts**: the verdict before the agent started, and the verdict at each point the agent tried to stop. 6. **Metadata**: the model, the harness and its version, the session id, timestamps, token counts, and cost. Of these, the tool calls carry the learning signal. They show the model how an expert in this codebase finds the relevant file, which command verifies a change, what a failing test looks like here, and how a fix is shaped. Training on them teaches the model to take those actions on this codebase. The final diff alone would not teach that. A diff is an answer; a trace is a worked solution. > **The records in a trace and how they relate** > > - Session has one Task: System prompt, rule files, user instruction; Base commit of the repository > - Session has many Turn: One model response; Ends in a tool call or a message > - Turn has zero or one Tool call: Name, input, output, duration; Output redacted before storage > - Session has many Oracle verdict: PASS or FAIL and nothing else; One at the start, one per stop attempt > - Session has one Outcome: flipped, no flip, or aborted; Final diff and files touched > > *The oracle verdicts are kept beside the turns, not inside them. The agent never sees the verdict's reason, so the reason is not in any turn.* ## The episode and the flip In reinforcement learning, an episode is one complete attempt at a task from start to a terminal state. A trace is an episode. What makes an episode useful for training is a reward: a number that says how well it went. For most of what organizations do, there is no clean reward. For software, there is one, and it has a name in the benchmark literature. SWE-bench, the benchmark most coding-agent papers report against, grades a candidate patch with two sets of tests.[1](#user-content-fn-swe-bench) The FAIL\_TO\_PASS tests fail on the repository before the fix and pass after the reference fix. The PASS\_TO\_PASS tests pass both before and after; they are there to catch a patch that fixes the issue by breaking something else. A patch resolves the task when every FAIL\_TO\_PASS test passes and every PASS\_TO\_PASS test still passes. That is the flip. Before the agent starts, the oracle runs the hidden tests and records FAIL. After the agent says it is done, the oracle applies the agent's diff to a clean copy, runs the same hidden tests, and records PASS or FAIL. A session whose verdict goes from FAIL to PASS, with no regression in the tests that already passed, has flipped. The trace of that session is a verified trajectory. A session that ends on FAIL is also kept, because a failed attempt beside a successful one on the same task is a preference pair, and preference pairs are training data too. Chapter 10 covers that. The flip is the unit of value for three reasons. It is binary, so it cannot be argued with. There is no rubric, no score from a judge model, no partial credit for a patch that "looks right." Chapter 6 is about why that matters. It is cheap, so it can be run on every session. A test suite that runs in minutes grades a session for cents. A human review of the same session would cost more than the session. It is yours. The hidden test encodes what your organization means by correct for this task. A public benchmark encodes what a benchmark author meant. ## Other records A trace is not a transcript. Harnesses keep transcripts for the person to scroll back through, and those files are useful, but they mix the agent's messages with display formatting, and they may not include the final message when the session ends.[2](#user-content-fn-claude-hooks) A trace is written by hooks as events happen, in a schema you control, with redaction applied before anything touches disk. A trace is not a log. Logs are for debugging a system. Traces are for training a model. A log can be lossy and unstructured. A trace that drops the tool output for one call has a hole in the worked example at exactly the point the model needs to learn from. A trace is not the diff. The diff is the answer. A model trained only on diffs learns to produce patches that look like your patches. It does not learn to find the file, run the test, read the failure, and try again. Pan and colleagues trained on full agent trajectories, each a sequence of tool calls and observations, and moved their model from 7.0 to 20.6 percent with 491 of them.[3](#user-content-fn-swe-gym) The actions are the data. ## The schema Appendix A gives a full schema. The shape is a JSON Lines file per session, one event per line, with a small set of event types. ```text session_start session_id, task, base_commit, model, harness, started_at tool_call turn, tool_name, tool_input, tool_output_redacted, duration_ms message turn, role, text oracle_verdict attempt, verdict (PASS|FAIL), tests_hash, ran_at session_end outcome (flipped|no_flip|aborted), final_diff_hash, ended_at ``` Two design choices in the schema deserve a sentence each. The verdict record carries a hash of the hidden test list, not the list. That lets you prove later which oracle graded a trace without putting the test names anywhere the agent's process can read. And the tool output is stored after redaction, never before. A trace store is a secret store if you let it be one, and Chapter 16 explains why that is the failure that ends programs like this. ## Why the harness should write it You could write a coding agent and have it write its own traces. Several groups have, and their agents are good. But the harness your team already uses runs thousands of sessions a week today, and it exposes hooks. A hook is a program the harness runs at a fixed point in the session: before a tool call, after a tool call, when the agent tries to stop, when the session starts and ends. Each hook receives the event as JSON on standard input. A hook that appends that JSON to a file is a trace collector, and it took the author of this book an afternoon to write. Chapter 7 covers hooks in detail. The reason to mention them here is that the most common objection to collecting traces, "we would have to build our own agent," is false. ## Footnotes 1. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) [↩](#user-content-fnref-swe-bench) 2. Anthropic (2026). *Hooks reference*. Claude Code documentation. [https://code.claude.com/docs/en/hooks](https://code.claude.com/docs/en/hooks) [↩](#user-content-fnref-claude-hooks) 3. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) --- # Chapter 3: The oracle problem > A verdict the model does not control: deterministic, independent, and hidden, with the fail-to-pass flip stated as a rule. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-oracle-problem Software testing has a name for the thing that decides whether a program's output is correct: the test oracle. Barr, Harman, McMinn, Shahbaz, and Yoo surveyed the field in 2015 and called the difficulty of building one "the oracle problem."[1](#user-content-fn-barr-oracle) Generating inputs to a program is easy. Knowing what the program should have done with them is hard. Every automated check, from a unit test to a type checker, is a partial answer to that problem. This book uses the word in the same sense, narrowed. An oracle here is a program that takes the state of a repository after an agent has worked on it and returns PASS or FAIL. It runs without the agent. It does not take the agent's word for anything. The quality of everything downstream, every training set and every evaluation, is bounded by the quality of the oracle, so this chapter is about what makes a good one. ## Deterministic A deterministic oracle returns the same verdict every time it runs on the same input. That sounds like a low bar. In practice, most test suites do not clear it. Luo, Hariri, Eloussi, and Marinov studied flaky tests, tests that pass and fail on the same code, across 51 open-source projects and classified the causes in 201 fixing commits. The leading causes were waiting on asynchronous work, concurrency, and dependence on test order.[2](#user-content-fn-luo-flaky) At Google, Micco reported that about 1.5 percent of all test runs gave a flaky result, that almost 16 percent of tests showed some flakiness, and that about 84 percent of transitions from pass to fail involved a flaky test.[3](#user-content-fn-micco-flaky) A flaky test in a benchmark is an annoyance. A flaky test in an oracle that labels training data is corruption: a session is recorded as a flip when the agent did nothing, or as a failure when the agent succeeded. The fixes are the fixes the testing literature has recommended for a decade, applied strictly because the stakes are higher. - Run the oracle in a fresh container from a pinned image, so the environment is identical every time. Reproducible builds give the same guarantee for the artifact under test: the same source produces the same binary, bit for bit, so a verdict is about the code and not the build machine.[4](#user-content-fn-reproducible-builds) - Pin every dependency by hash, with a local mirror, and give the container no network. A test that downloads anything is not deterministic. - Fix the clock, the random seed, the locale, and the time zone inside the container. A test that depends on the wall clock is a test that depends on when the oracle ran. - Run the hidden tests twice on the base commit before the agent starts. If the two verdicts disagree, the task is not fit for an oracle. Quarantine it. - Record the test list as a hash with the verdict, so a later change to the tests cannot be confused with a change in the code. The cost of determinism is that some tests cannot be oracles. Integration tests against a live service, tests that depend on timing, and tests that depend on data that changes are excluded. That is a loss, and Chapter 4 is partly about what you can use instead. ## Independent An oracle must be independent of the agent in two senses. It must not run in the agent's environment. The agent's working copy is under the agent's control. The agent can edit any file there, including the tests, the test configuration, the build file, and the environment variables the test runner reads. An oracle that runs `pytest` in the agent's directory is asking the agent whether it passed. Chapter 5 describes the air gap that fixes this. It must not be a model grading itself. A judge model that reads the diff and decides whether the task is done is not an oracle. It is a second opinion from a system with the same blind spots as the first. Chapter 6 reviews the evidence. ## Hidden The hidden test is the version of the oracle this book recommends. The team writes a test that fails on the current code and passes when the task is done correctly. The agent never sees it. The agent sees the task description and the existing test suite, and can write and run its own tests freely. When the agent says it is done, the oracle applies the agent's changes to a clean copy, runs the hidden test and the existing suite, and records the verdict. This is the SWE-bench construction.[5](#user-content-fn-swe-bench) The FAIL\_TO\_PASS tests are hidden from the agent during the attempt. The agent reads the issue, not the test. The construction is also how the Defects4J database of real Java bugs has been used for a decade: each bug comes with at least one test that exposes it and passes after the developer's fix.[6](#user-content-fn-defects4j) Hidden tests matter because visible tests are a weaker oracle than they look. Qi, Long, Achour, and Rinard examined patches that three automated repair systems had reported as fixing bugs, meaning the patches made the visible tests pass. Most of the patches were wrong.[7](#user-content-fn-qi-kali) They passed the tests by deleting the functionality the tests happened not to cover. Smith, Barr, Le Goues, and Brun showed the same thing with a controlled experiment and gave it a name: patch overfitting.[8](#user-content-fn-smith-cure) A patch that passes the tests the agent can see has been fitted to those tests, and nothing more has been shown. A patch that passes a test the agent could not see has been shown something. ## The flip, formally With the pieces named, the flip can be stated as a rule. Let `H` be the hidden tests for the task and `S` be the existing suite, both frozen as a list with a hash. Let `base` be the commit the agent started from and `diff` be the agent's final change. 1. On `base`, run `H` and `S` in a fresh container. Require that `H` fails and `S` passes. If `H` passes, there is no task; if `S` fails, the base is broken and the task is not fit for an oracle. Record the verdict as the baseline. 2. On `base` with `diff` applied to source paths only, run `H` and `S` in a fresh container. Record PASS if every test in `H` passes and every test in `S` passes. Record FAIL otherwise. 3. A session has flipped when the baseline is FAIL and the final verdict is PASS. Step 2 says "source paths only." The diff the agent produced may include changes to test files, test configuration, continuous-integration files, and build scripts. The oracle does not apply those. It applies the agent's changes to the code under test, then runs its own copies of the tests against them. Chapter 5 explains the allowlist that does this. > **One oracle run** > > 1. **Fresh container**: Pinned image, no network, fixed clock and seed > 2. **Clean checkout**: The base commit, from the oracle's own mirror > 3. **Apply source diff**: Only paths on the allowlist > 4. **Inject hidden tests**: From the oracle's store, never from the workspace > 5. **Run H and S**: Hidden tests and the existing suite > 6. **One bit out**: PASS or FAIL, plus a hash of the test list > > *Everything the agent could have touched is replaced before the tests run. The agent's only contribution to the oracle run is the source diff.* ## Tests are a weak oracle too Hidden tests are the best oracle most teams can build quickly. They are not a perfect one. Inozemtseva and Holmes showed that code coverage, the share of lines a test suite runs, is not strongly correlated with how many faults the suite detects once suite size is accounted for.[9](#user-content-fn-inozemtseva-coverage) A hidden test that runs a line is not a hidden test that checks it. The same study of patch overfitting that argues for hidden tests also argues for better ones. Mutation testing is the standard way to measure how good a test is. A mutation tool makes small changes to the code, such as flipping a comparison or deleting a statement, and checks whether the tests fail. A test that does not fail on a mutant is not checking the mutated behavior. The idea is from 1978 and has been used at Google at scale since at least 2018, where mutants are shown to developers during code review.[10](#user-content-fn-demillo-mutation)[11](#user-content-fn-petrovic-mutation) A hidden test that kills the mutants near the code the task touches is a stronger oracle than one that merely runs. Chapter 9 uses mutation the other way around, to manufacture tasks. The practical rule is: a hidden test must fail on the base commit for the right reason. Before you accept a task into the oracle pool, read the failure. If the test fails because of an import error or a fixture that is missing, the agent can flip it by fixing the fixture. If it fails because the feature is missing, the agent has to build the feature. ## Footnotes 1. Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). *The Oracle Problem in Software Testing: A Survey*. IEEE Transactions on Software Engineering, 41(5), 507–525. [https://doi.org/10.1109/TSE.2014.2372785](https://doi.org/10.1109/TSE.2014.2372785) [↩](#user-content-fnref-barr-oracle) 2. Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). *An Empirical Analysis of Flaky Tests*. FSE 2014, 643–653. [https://doi.org/10.1145/2635868.2635920](https://doi.org/10.1145/2635868.2635920) [↩](#user-content-fnref-luo-flaky) 3. Micco, J. (2016). *Flaky Tests at Google and How We Mitigate Them*. Google Testing Blog. [https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html) [↩](#user-content-fnref-micco-flaky) 4. Lamb, C., & Zacchiroli, S. (2022). *Reproducible Builds: Increasing the Integrity of Software Supply Chains*. IEEE Software, 39(2), 62–70. [https://arxiv.org/abs/2104.06020](https://arxiv.org/abs/2104.06020) [↩](#user-content-fnref-reproducible-builds) 5. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) [↩](#user-content-fnref-swe-bench) 6. Just, R., Jalali, D., & Ernst, M. D. (2014). *Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs*. ISSTA 2014, 437–440. [https://doi.org/10.1145/2610384.2628055](https://doi.org/10.1145/2610384.2628055) [↩](#user-content-fnref-defects4j) 7. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). *An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems*. ISSTA 2015, 24–36. [https://doi.org/10.1145/2771783.2771791](https://doi.org/10.1145/2771783.2771791) [↩](#user-content-fnref-qi-kali) 8. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). *Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair*. ESEC/FSE 2015, 532–543. [https://doi.org/10.1145/2786805.2786825](https://doi.org/10.1145/2786805.2786825) [↩](#user-content-fnref-smith-cure) 9. Inozemtseva, L., & Holmes, R. (2014). *Coverage Is Not Strongly Correlated with Test Suite Effectiveness*. ICSE 2014, 435–445. [https://doi.org/10.1145/2568225.2568271](https://doi.org/10.1145/2568225.2568271) [↩](#user-content-fnref-inozemtseva-coverage) 10. DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). *Hints on Test Data Selection: Help for the Practicing Programmer*. IEEE Computer, 11(4), 34–41. [https://doi.org/10.1109/C-M.1978.218136](https://doi.org/10.1109/C-M.1978.218136) [↩](#user-content-fnref-demillo-mutation) 11. Petrović, G., & Ivanković, M. (2018). *State of Mutation Testing at Google*. ICSE-SEIP 2018, 163–171. [https://doi.org/10.1145/3183519.3183521](https://doi.org/10.1145/3183519.3183521) [↩](#user-content-fnref-petrovic-mutation) --- # Chapter 4: Organizational oracles > The checks a team already has, ordered by how directly each one can label a trace. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/organizational-oracles The hidden test is one oracle. An organization has many. This chapter is a catalogue, ordered from the oracles that give the cleanest signal to the ones that give the noisiest, with a note on how each can be used. The ordering matters because the cleaner the oracle, the more directly its verdict can be used as a training label. The noisier the oracle, the more it belongs in a preference pair or a filter rather than as a label. ## Hard oracles A hard oracle is deterministic, runs in minutes, and returns a verdict a program can read. These can label a trace directly. **Hidden unit and integration tests.** The reference oracle, covered in Chapter 3. Written by a person or by a separate model before the task is dispatched. The strongest version is a test that was written to expose a real bug or specify a real feature, because it encodes a real requirement. **The existing test suite as PASS\_TO\_PASS.** Every task gets the whole existing suite as a regression check for free. A flip requires that nothing already passing breaks. On a repository with a large suite, this alone rules out most of the bad patches that a visible test would accept. **Type checkers and compilers.** A change that does not compile, or that fails `tsc --noEmit`, `mypy --strict`, or `cargo check`, has failed. These are deterministic, fast, and hard to game without touching configuration, which the allowlist excludes. On their own they are a weak oracle, since code that compiles can still be wrong, but as a component of the oracle they remove a class of failures cheaply. **Linters with pinned rule sets.** A lint failure is a weak signal of incorrectness and a strong signal of style drift. Use it as a gate, not a label: a trace that flips the hidden test but introduces lint errors is a flip with a defect, and you can decide per repository whether that counts. **Property-based tests.** Instead of a fixed input and expected output, a property test states a rule, such as "decoding what you encoded gives the original," and a library generates many inputs to check it.[1](#user-content-fn-quickcheck) With a fixed seed, a property test is deterministic. It is a stronger oracle than an example test because it checks many cases, and a harder one for an agent to fit to because the agent cannot see the cases. **Metamorphic tests.** When the correct output is unknown but a relationship between outputs is known, a metamorphic test checks the relationship. If a search for "a" returns results, a search for "a OR b" should return at least as many. Metamorphic testing was proposed for programs without an oracle and is well suited to data pipelines, search, and numerical code.[2](#user-content-fn-chen-metamorphic) **Differential tests against the previous build.** Run the same inputs through the version before the change and the version after. For a refactor, the outputs must match. For a bug fix, they must differ only on the inputs that exposed the bug. McKeeman's differential testing of compilers is the origin of the method.[3](#user-content-fn-mckeeman-differential) The previous binary is an oracle you already have. **Mutation score.** Chapter 3 covered mutation as a test-quality measure. It can also be an oracle for a task of the form "add tests for this module." The hidden check is whether the mutants the task names are killed after the change. This is one of the few ways to make test-writing itself a flippable task. **Reproducible build hash.** For a task that must not change behavior, such as a dependency bump with no code change, the oracle can be that the build output is byte-identical to a reference, or differs only where expected.[4](#user-content-fn-reproducible-builds) **Schema and migration round trips.** For a database change: apply the migration to a copy of the schema, run the down migration, and diff the result against the original. For a data pipeline: run the transform on a frozen fixture and compare row counts and checksums against a stored expectation. **Infrastructure plan idempotence.** For infrastructure-as-code: after applying the change, a second plan must report nothing to do. This is a deterministic check of the form most infrastructure tools provide directly. **Formal checks.** Where a module has a specification in a checkable form, a solver or a proof assistant is the strongest oracle there is. Few organizations have these. The ones that do should use them. ## Soft and delayed oracles A soft oracle is a signal that correlates with correctness but is not deterministic, is not available for minutes or days, or depends on a person. These cannot label a trace as a flip. They can rank traces, build preference pairs, and filter. **Pull request merged.** A merged change passed review and continuous integration. This is the oracle most code-model training has used, in the form of mined commits and pull requests. Meta's SWE-RL built its training corpus from about eleven million pull requests and used similarity to the merged patch as the reward.[5](#user-content-fn-swe-rl) It is a real signal and a slow one, and reviewers miss things. **Reverted within N days.** A change that was merged and then reverted was a failure the review did not catch. The pair "merged, then reverted" against "merged, kept" is a clean preference pair with a delay of days to weeks. **Incident linked to the change.** Rarer and stronger than a revert. A change that caused an incident is a hard negative example. **Review comments.** A change that received requests for changes before merge is weaker than one approved on the first pass. Review text also says what was wrong, which is useful for a data scientist and useless as a training label. **Acceptance of a suggestion.** For completion models, whether the developer accepted the suggestion was found to be the best available predictor of perceived productivity in GitHub Copilot telemetry.[6](#user-content-fn-ziegler-copilot) For agents, the equivalent is whether the person kept the agent's change or discarded the session. It is a weak oracle and an abundant one. **Time to next edit of the same lines.** If a person edits the lines an agent wrote within an hour of the session, the agent's work was probably incomplete. This is a proxy with many false positives and is best used to flag traces for a closer look. **A judge model.** A second model that reads the diff and scores it. Chapter 6 argues that this is the weakest oracle in the catalogue and should not be used as a label. It can be used as a filter for obvious garbage, and even then its errors should be measured against a hard oracle on a sample. > **Oracles by how directly their verdict can label a trace** > > 1. **Judge model**: A model scores the change. Not a label. At most a coarse filter, and even then measured against a hard oracle. > 2. **Acceptance, next edit, review comments**: Human behavior signals. Rank and flag, do not label. > 3. **Merged, reverted, incident**: Delayed by days. Build preference pairs. > 4. **Type check, lint, build hash**: Fast and deterministic. Necessary, not sufficient. Use as gates inside the oracle. > 5. **Existing suite as PASS\_TO\_PASS**: Free regression check on every task. > 6. **Hidden tests, property and metamorphic tests**: Deterministic, fast, independent, unseen. Labels a flip. > > *Only the top rung produces a label the training set can trust on its own. Everything below it is still worth collecting.* ## Oracles outside the code The method is not limited to software, but the hard oracles mostly are. For the sake of completeness, here is what the same construction looks like elsewhere, with the caveat that each of these is a soft oracle and should be treated as one. - **Data work.** A transform's output on a frozen input has a checksum. Row counts and null rates have expectations. Tools that run assertions on data in a pipeline exist and can be hidden from the agent in the same way tests are. - **Documentation.** A link checker, a prose checker with a pinned rule set, and a build that fails on a broken reference are deterministic. Whether the document is clear is not. - **Support.** A ticket resolved and not reopened within two weeks is a delayed soft oracle on the agent's answer. - **Operations.** A runbook step that leaves a system in a state a monitoring check accepts is close to a hard oracle, if the check is deterministic and the system is a copy. A team should start with the hard oracles in its code, because that is where the construction is cleanest and the data is largest. Once the pipeline runs, the soft oracles can be added as preference signals. ## Which oracle wrote the label Every trace should carry the identity of the oracle that graded it: the kind of oracle, the hash of its test list or rule set, the container image digest, and the time. A training set built from traces graded by different oracles at different times is a training set whose labels mean different things. When a later evaluation shows a regression, the first question is which oracle labeled the traces the model learned it from. Without the record, the question cannot be answered. ## Footnotes 1. Claessen, K., & Hughes, J. (2000). *QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs*. ICFP 2000, 268–279. [https://doi.org/10.1145/351240.351266](https://doi.org/10.1145/351240.351266) [↩](#user-content-fnref-quickcheck) 2. Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). *A Survey on Metamorphic Testing*. IEEE Transactions on Software Engineering, 42(9), 805–824. [https://doi.org/10.1109/TSE.2016.2532875](https://doi.org/10.1109/TSE.2016.2532875) [↩](#user-content-fnref-chen-metamorphic) 3. McKeeman, W. M. (1998). *Differential Testing for Software*. Digital Technical Journal, 10(1), 100–107. [↩](#user-content-fnref-mckeeman-differential) 4. Lamb, C., & Zacchiroli, S. (2022). *Reproducible Builds: Increasing the Integrity of Software Supply Chains*. IEEE Software, 39(2), 62–70. [https://arxiv.org/abs/2104.06020](https://arxiv.org/abs/2104.06020) [↩](#user-content-fnref-reproducible-builds) 5. Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449) [↩](#user-content-fnref-swe-rl) 6. Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). *Measuring GitHub Copilot's Impact on Productivity*. Communications of the ACM, 67(3), 54–63. [https://doi.org/10.1145/3633453](https://doi.org/10.1145/3633453) [↩](#user-content-fnref-ziegler-copilot) --- # Chapter 5: The air gap > The agent is an adversary with a shell. The oracle runs where the shell cannot reach. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-air-gap An oracle the agent can reach is an oracle the agent will eventually satisfy without doing the work. This chapter is about the threat, the evidence that it is real, and the design that removes it. ## The threat model Treat the agent as an adversary with full control of its working directory and shell. That is not a judgment about the model's intent. It is a description of its permissions. A coding agent can run any command, read any file it can reach, and write any file in its checkout. If the verdict depends on anything in that set, the verdict is the agent's to decide. The ways an agent can make a check pass without doing the task are ordinary: - Edit the test so that it asserts what the code does. - Mark the test skipped, or add it to an exclusion list in the test configuration. - Edit the test runner's configuration so the test is not collected. - Edit the build file or the continuous-integration file so the check does not run. - Stub the function under test to return the expected value for the inputs the test uses. - Patch the test framework itself, or a fixture, or a conftest file that runs before every test. - Set an environment variable the test reads. - Exit the process with status zero before the test runner reports. - Read the hidden test from disk or from the network, if it is anywhere the agent can reach, and fit the code to it. - If a model judges the result, write text into the diff or the commit message that addresses the judge. Every one of these has been observed. Baker and colleagues at OpenAI trained a frontier reasoning model on coding tasks and monitored its chain of thought. They found the model saying, in its own reasoning, that it would make the tests pass by other means, and then doing so: calling `sys.exit(0)` before the tests ran so the runner reported success, raising the test framework's skip exception, and writing stubs instead of implementations where test coverage was thin.[1](#user-content-fn-baker-monitoring) When they penalized the model for having such thoughts, it kept hacking the tests and stopped writing down that it was going to. They called this obfuscated reward hacking and recommended against training against the monitor. Denison and colleagues at Anthropic built a curriculum of environments where a model could gain reward by gaming the check, starting with flattery and ending with the chance to edit its own reward function. A model trained on the easy environments generalized to the hard one: in a small but nonzero fraction of runs, it edited the code that computed its reward, and sometimes edited the test that would have caught the edit.[2](#user-content-fn-denison-subterfuge) METR, evaluating frontier models on its own tasks in 2025, reported models that rewrote scoring code, patched the timer in a benchmark harness, and special-cased the inputs the grader used.[3](#user-content-fn-metr-hacking) A later Anthropic study trained a model with reinforcement learning on real production coding environments after teaching it about reward-hacking strategies. At the point where the model learned to hack the rewards, it also began to generalize to other unwanted behavior, including attempts at sabotage.[4](#user-content-fn-anthropic-emergent) The lesson is not that models are malicious. It is that optimization finds the shortest path to the reward, and if the shortest path goes through the test file, that is the path. Skalse and colleagues formalized this as reward hacking: a proxy that correlates with the goal is maximized, and the correlation breaks under optimization.[5](#user-content-fn-skalse-hacking) Gao, Schulman, and Hilton measured it for reward models: the true reward rises as a policy is optimized against the proxy, peaks, and then falls while the proxy keeps rising.[6](#user-content-fn-gao-overoptimization) The hidden test is the goal. Anything the agent can touch is a proxy. The lesson is also older than machine learning. Thompson's 1984 Turing Award lecture described a compiler that inserts a back door into the login program and into any compiler compiled from clean source, so that no inspection of the source finds it.[7](#user-content-fn-thompson-trust) His point was that you cannot trust code you did not totally create yourself, and that the trust boundary has to be drawn around the thing you can verify. The oracle is the thing you can verify. Draw the boundary around it. ## The design The oracle runs in a separate container, on a separate host when possible, with the following properties. **No network.** The container is started with networking disabled. Dependencies come from a local mirror baked into the image or mounted read-only. Nothing the oracle needs is fetched at run time, and nothing the agent could have planted on a server is reachable. **A pristine checkout.** The oracle has its own clone of the repository, from its own mirror. It checks out the base commit the session started from. It does not read the agent's working copy at all. **The source diff only.** The oracle receives the agent's final diff as a file. Before applying it, the oracle filters it through an allowlist of paths. Source directories are on the allowlist. Test directories, test configuration files, continuous-integration configuration, build scripts, dependency manifests and lockfiles, and anything under the repository's tool configuration are not. Hunks that touch excluded paths are dropped. The oracle logs that they were dropped, in a log the agent cannot read. **Hidden tests from the oracle's store.** The hidden tests are stored with the oracle, keyed by task id, and copied into the checkout after the diff is applied. They are never in the repository the agent works in, never in a branch the agent can fetch, and never on a path the agent's shell can read. **The existing suite from the base commit.** The PASS\_TO\_PASS tests are the ones in the pristine checkout at the base commit, not the ones in the agent's diff. If the agent deleted a test, the oracle still runs it. **Fixed clock, seed, locale, and time zone.** Set in the container's environment. A test that reads the system clock reads a constant. **Pinned image.** The container image is referenced by digest, and the digest is recorded with the verdict. **Timeouts.** The oracle run has a wall-clock limit. A run that exceeds it is a FAIL with a reason the oracle logs and the agent does not see. An agent cannot turn a failing test into a hang and have the hang read as success. **One bit out, signed.** The oracle's only output to the agent's side is PASS or FAIL. To the trace store it writes a verdict record: the bit, the hash of the test list, the image digest, the diff hash, the base commit, the time, and a signature over all of these using a key the oracle host holds. A verdict in the trace store that does not verify was not written by the oracle. **Separate credentials.** The process that runs the agent holds no credential that can write to the oracle's store, start an oracle run with a different test list, or read the hidden tests. The hook that asks for a verdict sends a diff and a task id and gets a bit back. > **What crosses the boundary between the agent's host and the oracle's host** > > - Agent host sends Oracle host: Task id; Base commit; The final diff as a file > - Oracle host returns Agent host: PASS or FAIL; Nothing else > - Oracle host writes Trace store: Signed verdict record; Dropped-hunk log; Test output, never shown to the agent > - Agent host writes Trace store: Tool-call events, redacted; Messages; Session metadata > > *The agent's host can write trace events and can ask one question. It cannot read the hidden tests, the test output, or the oracle's log. The trace store holds both sides and joins them on the session id.* ## Why a separate host A container on the same machine as the agent is isolation from the agent's working copy. It is not isolation from the agent's shell. On a developer's laptop, the agent's Bash tool can read any path the developer can read, including the directory where the hidden tests live, the oracle's log, and the container's image. A container boundary protects the oracle's run from the agent's changes. A host boundary protects the oracle's secrets from the agent's eyes. The reference plugin in Chapter 8 runs on one machine because that is where most teams will start, and it says so in its README. On one machine, store the hidden tests encrypted at rest and decrypt them inside the oracle container with a key the agent's process does not hold. That raises the bar from "read the file" to "find and misuse the key." It does not close the gap. Closing it means running the oracle on a host the agent cannot log into, which for a team with a continuous-integration system is a short step: the oracle is a job that runs on a runner, takes a diff and a task id, and returns a bit. Firecracker, the virtual-machine monitor behind AWS Lambda, starts a microVM in about 125 milliseconds and gives each one its own kernel.[8](#user-content-fn-firecracker) gVisor gives a container a user-space kernel and intercepts its system calls.[9](#user-content-fn-gvisor) Either is a stronger boundary than a plain container for the oracle run, and both are used for exactly this purpose: running code you do not trust next to data you care about. For the hidden tests and the signing key, neither replaces a separate host. ## Baselines and the flip that was not The flip needs a baseline: a FAIL recorded before the agent started. Three rules keep baselines honest. The baseline is computed on the base commit, by the oracle, from its own checkout. It is deterministic, so compute it once per task and store it. Do not recompute it in a session-start hook; a container run is too slow for a hook budget, and the result would be the same anyway. A task whose baseline is PASS is not a task. The hidden test already passes. Either the test is wrong or the work is already done. Quarantine the task, and do not count any session on it as a flip. A task whose existing suite fails at baseline is not fit for an oracle. The agent could flip the hidden test while the suite stays broken, and the PASS\_TO\_PASS rule would record a FAIL for a session that did the work. Fix the base or exclude the task. ## What the oracle logs and who reads it The oracle keeps a full log: the test output, the dropped hunks, the timing, and the reason for every FAIL. This log is for the people who run the pipeline. It is how you find a hidden test that fails for the wrong reason, a flaky test that slipped through, or an agent that keeps trying to edit the test configuration. It is never returned to the agent and never written anywhere the agent's host can read. Chapter 6 is about why. ## Footnotes 1. Baker, B., Huizinga, J., Gao, L., et al. (2025). *Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation*. arXiv. [https://arxiv.org/abs/2503.11926](https://arxiv.org/abs/2503.11926) [↩](#user-content-fnref-baker-monitoring) 2. Denison, C., MacDiarmid, M., Barez, F., et al. (2024). *Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models*. arXiv. [https://arxiv.org/abs/2406.10162](https://arxiv.org/abs/2406.10162) [↩](#user-content-fnref-denison-subterfuge) 3. Von Arx, S., Chan, L., & Barnes, E. (2025). *Recent Frontier Models Are Reward Hacking*. METR. [https://metr.org/blog/2025-06-05-recent-reward-hacking/](https://metr.org/blog/2025-06-05-recent-reward-hacking/) [↩](#user-content-fnref-metr-hacking) 4. MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). *Natural Emergent Misalignment from Reward Hacking in Production RL*. arXiv. [https://arxiv.org/abs/2511.18397](https://arxiv.org/abs/2511.18397) [↩](#user-content-fnref-anthropic-emergent) 5. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). *Defining and Characterizing Reward Hacking*. NeurIPS 2022. [https://arxiv.org/abs/2209.13085](https://arxiv.org/abs/2209.13085) [↩](#user-content-fnref-skalse-hacking) 6. Gao, L., Schulman, J., & Hilton, J. (2022). *Scaling Laws for Reward Model Overoptimization*. arXiv; ICML 2023. [https://arxiv.org/abs/2210.10760](https://arxiv.org/abs/2210.10760) [↩](#user-content-fnref-gao-overoptimization) 7. Thompson, K. (1984). *Reflections on Trusting Trust*. Communications of the ACM, 27(8), 761–763. [https://doi.org/10.1145/358198.358210](https://doi.org/10.1145/358198.358210) [↩](#user-content-fnref-thompson-trust) 8. Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). *Firecracker: Lightweight Virtualization for Serverless Applications*. NSDI 2020, 419–434. [https://www.usenix.org/conference/nsdi20/presentation/agache](https://www.usenix.org/conference/nsdi20/presentation/agache) [↩](#user-content-fnref-firecracker) 9. Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). *The True Cost of Containing: A gVisor Case Study*. HotCloud 2019. [https://www.usenix.org/conference/hotcloud19/presentation/young](https://www.usenix.org/conference/hotcloud19/presentation/young) [↩](#user-content-fnref-gvisor) --- # Chapter 6: The one-bit verdict > Why PASS or FAIL is all the oracle says, what that costs inside a session, and the research on models grading themselves. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-one-bit-verdict The oracle says PASS or FAIL. It does not say which test failed, what the assertion was, or what the output looked like. This is the design choice readers push back on most, so this chapter takes it slowly: first the cost, then the three reasons, then the research on self-grading that the third reason depends on. ## The cost, stated plainly Richer feedback helps an agent fix a bug in the current session. Chen, Lin, Schärli, and Zhou compared three kinds of feedback for a model debugging its own code: simple feedback, meaning only whether the code passed; unit-test feedback, meaning the execution results; and a self-written explanation of the code. Unit-test feedback produced the largest gains.[1](#user-content-fn-chen-self-debug) A verdict with no explanation is the "simple feedback" condition in that study, and it is the weakest of the three for raising the pass rate inside one session. So the one-bit verdict lowers the flip rate per session. A team that adopts it will see more sessions end on FAIL than a team that shows the agent the failing test. That is a real cost. Three things make it worth paying. ## Reason one: the hidden test stays valid A test the agent can see is a test the agent can fit to. Chapter 3 covered the patch-overfitting results: patches that pass visible tests by deleting what the tests do not check.[2](#user-content-fn-qi-kali)[3](#user-content-fn-smith-cure) An oracle that returns the name of the failing test and the assertion that failed has shown the agent the test. The agent will fix the assertion. Whether it fixed the feature is now unknown, which is where you started. The theory for this is from statistics, not software. Blum and Hardt studied machine-learning competitions where participants submit many models and see a score on a holdout set each time. With enough submissions, a participant can overfit the holdout without ever seeing its labels, by hill-climbing on the score. Their fix, the Ladder, releases a new score only when it improves on the previous best by more than a threshold, and otherwise repeats the old score. That turns a high-information answer into a low-information one and makes the leaderboard reliable under adaptive submissions.[4](#user-content-fn-ladder) Dwork, Feldman, Hardt, Pitassi, Reingold, and Roth proved the general result: a holdout answered with limited information, through a mechanism they called Thresholdout, stays statistically valid under far more adaptive queries than one answered exactly.[5](#user-content-fn-reusable-holdout) An agent iterating against an oracle is a participant iterating against a holdout. A verdict of PASS or FAIL is about as low-information as an answer gets. An execution trace with the failing assertion is about as high-information as one gets. The Ladder result says which one keeps the hidden test meaningful. ## Reason two: the diagnosis is the data The point of collecting traces is to train a model on them. The behavior you want the model to learn is not "read the failing assertion and change the code until it passes." It is "form a hypothesis about what is wrong, write a test that checks it, run the test, read the result, and fix the cause." That is what expert engineers do, and the trace of an expert doing it is what you want in the training set. If the oracle tells the agent why it failed, the trace after that point is the agent copying the oracle's diagnosis. The model trained on it learns to wait for a diagnosis. If the oracle says only FAIL, the trace after that point is the agent's own diagnosis: it has to write its own tests, run them, read its own failures, and reason about what the hidden requirement might be. That is the behavior worth learning, and it only appears in the trace when the oracle withholds the answer. The agent is not working blind. It has the task description, the whole repository, the existing test suite, and a shell. It can write any test it wants and run it as often as it wants, with full output. The one-bit rule applies to the hidden oracle only. In-session feedback from the agent's own tests is as rich as the agent cares to make it. The oracle's silence forces the agent to generate that feedback for itself, which is the skill. ## Reason three: the alternative is a model grading a model The usual proposal for richer feedback that does not leak the test is a judge: a second model that reads the diff, the test output, and the task, and writes an explanation for the agent. The judge sees the hidden test; the agent sees the judge's prose. This is where the research on self-grading matters, because the judge is a model with the same training and the same blind spots as the agent. Zheng and colleagues, in the paper that introduced MT-Bench and the LLM-as-a-judge method, documented three biases in model judges: position bias, where the judge prefers whichever answer appears first; verbosity bias, where it prefers the longer answer; and self-enhancement bias, where it prefers answers it wrote.[6](#user-content-fn-zheng-judge) Wang and colleagues showed that swapping the order of two answers could reverse a judge's verdict.[7](#user-content-fn-wang-unfair) Panickssery, Bowman, and Feng showed that a model's preference for its own output is not an accident of style. Models can recognize their own text above chance, and the ones that recognize it best prefer it most; fine-tuning a model to recognize its own output more accurately made its self-preference stronger.[8](#user-content-fn-panickssery) A judge from the same family as the agent is a judge that likes the agent. The research on self-correction is worse. Huang and colleagues reviewed the claim that models can correct their own reasoning and found that the gains reported in earlier work depended on an oracle: the model was told whether its answer was right before being asked to reconsider. Without that signal, asking the model to check its work made the answers worse more often than better.[9](#user-content-fn-huang-self-correct) That paper is sometimes read as an argument against external feedback. It is the opposite. It shows that a binary external signal is the thing that made self-correction work in the papers that reported it working, and that the model's own judgment was not contributing. Stechly, Marquez, and Kambhampati tested GPT-4 as a critic of its own graph-coloring solutions and found it could not reliably tell a correct coloring from an incorrect one; iterating on its own critique did not help, and a simple external checker did.[10](#user-content-fn-stechly-wrong) Valmeekam, Marquez, and Kambhampati found the same for planning: a model critiquing its own plans lowered the success rate, and an external verifier raised it.[11](#user-content-fn-valmeekam-plans) Tyen and colleagues separated two skills and found that models are poor at locating the error in a chain of reasoning but can often fix it once the location is given.[12](#user-content-fn-tyen-location) Kamoi and colleagues surveyed the self-correction literature and concluded that no reliable self-correction has been shown without external feedback, and that many reported successes used unrealistic setups where the model was given information it would not have in practice.[13](#user-content-fn-kamoi-survey) Xu and colleagues found that self-refinement loops amplify the model's own biases over iterations.[14](#user-content-fn-xu-pride) None of these results says a judge model is useless. They say a judge model's errors are correlated with the agent's errors, so using one to explain failures to the agent does not add an independent signal. It adds a confident one. The hidden test is independent. Its one bit is worth more than the judge's paragraph. > **Feedback channels from the oracle to the agent** > > 1. **PASS or FAIL**: The Ladder regime. The hidden test stays valid. The agent's own diagnosis is in the trace. > 2. **Which test failed**: Names the requirement. The agent now knows what to fit to. > 3. **The assertion and the values**: The agent can special-case the inputs. Patch overfitting territory. > 4. **A judge model's explanation**: Adds a correlated, biased reading of the test to the leak. The worst of both. > 5. **The test source**: The hidden test is no longer hidden. The oracle measures nothing. > > *Each rung down leaks more of the hidden test and leaves less of the agent's own reasoning in the trace. The book recommends the top rung and names the cost.* ## Recovering some of the cost There are ways to give back part of the in-session success rate without giving up the properties above. **Disclose to the trace, never to the agent.** The oracle writes the full test output to its own log, which the data team reads. When a task fails many sessions in a row, a person reads the log and either fixes a hidden test that fails for the wrong reason or rewrites the task description so the requirement is clearer. The agent gets a better task, not a leaked test. **A visible test the agent must write.** Make the task description say that the change must come with a test that fails before and passes after. The agent writes its own red-green pair, which is rich in-session feedback and good trace content. The hidden test stays hidden and checks that the agent's test checked the right thing. **Staged tasks.** If a task's hidden test covers three requirements, split it into three tasks with one hidden test each. The agent gets three bits instead of one across the same work, and each bit still says nothing about its test. **More attempts, fresh context.** The flip rate per session is lower; the flip rate per task does not have to be. Chapter 9 covers sampling several sessions per task and keeping the one that flips. With the oracle silent, the attempts are independent, which is what repeated sampling needs. **Disclose after the fact.** Once a task has been flipped by some session, its hidden test can be released into the repository as an ordinary test. It has done its job as an oracle. Future sessions on nearby code get it as a visible regression test, and new hidden tests are written for new tasks. ## The rule The hidden oracle says PASS or FAIL. Every other channel from the oracle to the agent is closed. Everything the oracle knows goes to the trace store, for people, under access control. If a team decides to open a wider channel, it should do so knowing which rung of the ladder it has moved to and what it has traded for the higher flip rate. ## Footnotes 1. Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). *Teaching Large Language Models to Self-Debug*. arXiv; ICLR 2024. [https://arxiv.org/abs/2304.05128](https://arxiv.org/abs/2304.05128) [↩](#user-content-fnref-chen-self-debug) 2. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). *An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems*. ISSTA 2015, 24–36. [https://doi.org/10.1145/2771783.2771791](https://doi.org/10.1145/2771783.2771791) [↩](#user-content-fnref-qi-kali) 3. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). *Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair*. ESEC/FSE 2015, 532–543. [https://doi.org/10.1145/2786805.2786825](https://doi.org/10.1145/2786805.2786825) [↩](#user-content-fnref-smith-cure) 4. Blum, A., & Hardt, M. (2015). *The Ladder: A Reliable Leaderboard for Machine Learning Competitions*. ICML 2015, PMLR 37, 1006–1014. [https://arxiv.org/abs/1502.04585](https://arxiv.org/abs/1502.04585) [↩](#user-content-fnref-ladder) 5. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). *The reusable holdout: Preserving validity in adaptive data analysis*. Science, 349(6248), 636–638. [https://doi.org/10.1126/science.aaa9375](https://doi.org/10.1126/science.aaa9375) [↩](#user-content-fnref-reusable-holdout) 6. Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*. NeurIPS 2023 Datasets and Benchmarks. [https://arxiv.org/abs/2306.05685](https://arxiv.org/abs/2306.05685) [↩](#user-content-fnref-zheng-judge) 7. Wang, P., Li, L., Chen, L., et al. (2023). *Large Language Models are not Fair Evaluators*. arXiv; ACL 2024. [https://arxiv.org/abs/2305.17926](https://arxiv.org/abs/2305.17926) [↩](#user-content-fnref-wang-unfair) 8. Panickssery, A., Bowman, S. R., & Feng, S. (2024). *LLM Evaluators Recognize and Favor Their Own Generations*. arXiv; NeurIPS 2024. [https://arxiv.org/abs/2404.13076](https://arxiv.org/abs/2404.13076) [↩](#user-content-fnref-panickssery) 9. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. arXiv; ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-huang-self-correct) 10. Stechly, K., Marquez, M., & Kambhampati, S. (2023). *GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems*. arXiv. [https://arxiv.org/abs/2310.12397](https://arxiv.org/abs/2310.12397) [↩](#user-content-fnref-stechly-wrong) 11. Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). *Can Large Language Models Really Improve by Self-critiquing Their Own Plans?* arXiv. [https://arxiv.org/abs/2310.08118](https://arxiv.org/abs/2310.08118) [↩](#user-content-fnref-valmeekam-plans) 12. Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). *LLMs cannot find reasoning errors, but can correct them given the error location*. Findings of ACL 2024. [https://arxiv.org/abs/2311.08516](https://arxiv.org/abs/2311.08516) [↩](#user-content-fnref-tyen-location) 13. Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). *When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs*. Transactions of the ACL, 12. [https://arxiv.org/abs/2406.01297](https://arxiv.org/abs/2406.01297) [↩](#user-content-fnref-kamoi-survey) 14. Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). *Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement*. ACL 2024. [https://arxiv.org/abs/2402.11436](https://arxiv.org/abs/2402.11436) [↩](#user-content-fnref-xu-pride) --- # Chapter 7: Harness hooks > The hooks a harness already exposes, turned into a trace collector and a Stop gate without building an agent. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/harness-hooks You do not need to build a coding agent to collect traces from one. The harness your team already runs exposes hooks, and hooks see everything a trace needs. This chapter explains what a hook is, what each one sees, and how to turn a set of hooks into a trace collector and an oracle gate. The examples use Claude Code because its hook interface is documented in detail and the reference plugin targets it.[1](#user-content-fn-claude-hooks) Other harnesses have equivalents, and the last section covers them. ## What a hook is A hook is a program the harness runs at a fixed point in a session. The harness passes the event to the program as JSON on standard input, waits for it to exit, and reads its exit code and anything it printed. A hook can observe the event, add context for the model, or, for some events, change what happens next. The events that matter for tracing are the ones at the boundaries of a turn and a session. | Event | When it runs | What it carries | What it can decide | | --- | --- | --- | --- | | SessionStart | A session begins or resumes | Session id, working directory, model, how the session started | Nothing. It can add context. | | UserPromptSubmit | The person sends a prompt | The prompt text | It can block the prompt. | | PreToolUse | Before a tool runs | Tool name and input | It can allow, deny, or ask. | | PostToolUse | After a tool runs | Tool name, input, output, duration | It can add context beside the result. | | Stop | The agent is about to end its turn | The agent's final message, whether a Stop hook already continued the turn | It can refuse to let the agent stop. | | SessionEnd | The session ends | The reason the session ended | Nothing. It runs cleanup. | Every event also carries the session id, the path to the harness's own transcript, the working directory, and the permission mode. The session id is the join key for everything the trace collector writes. ## The trace collector A trace collector is three hooks. On **SessionStart**, write a `session_start` event: the session id, the working directory, the base commit of the repository, the model, and the time. Do not run anything slow here. Hook budgets at session start are short, and a container run does not fit. On **UserPromptSubmit**, write a `message` event with the role `user` and the prompt text. This is the task as the agent received it, which the training set needs as the first turn. On **PostToolUse**, write a `tool_call` event: the tool name, the input, the output, and the duration. Redact the input and output before writing. Cap the size of each and store a hash of the full value beside the truncated one, so a trace that was cut can be identified later. That is the whole collector. The agent's own messages between tool calls are not delivered by PostToolUse, so the collector adds a fourth hook. On **Stop**, write a `message` event with the role `assistant` and the agent's final message, which the Stop event carries as `last_assistant_message`. Then run the oracle gate, below. The collector does not read the harness's transcript file. The transcript is for the person, its format is the harness's to change, and the harness documentation notes that the file is not guaranteed to include the final message at the moment Stop fires.[1](#user-content-fn-claude-hooks) The hook events are the record. ## The oracle gate The Stop hook is where the flip is detected and where an agent is kept from calling the work done. When the agent decides it has finished, the harness fires Stop. The gate hook does the following. 1. Reads the event. If the agent is already continuing because of a previous Stop hook, the event says so in a field named `stop_hook_active`. The gate uses this, with its own attempt counter, to decide whether to keep going. 2. Builds the agent's diff against the base commit recorded at session start, including new files. 3. Sends the diff and the task id to the oracle. The oracle returns one bit. 4. Writes an `oracle_verdict` event with the attempt number, the verdict, and the hash of the test list the oracle reported. 5. On PASS, exits normally. The agent stops. The session has flipped. 6. On FAIL, returns a decision that blocks the stop, with a reason the harness shows to the agent. The reason is a fixed string. It says the oracle reported FAIL and the agent should keep working. It does not say why. The blocking mechanism is specific. In Claude Code, a Stop hook refuses the stop by printing a JSON object with `decision` set to `block` and a `reason`, or by exiting with code 2 and the reason on standard error.[1](#user-content-fn-claude-hooks) The agent sees the reason as the explanation for why it must continue. With a one-bit oracle, the reason is the same every time. ## What the harness limits Two limits in the harness shape how the gate behaves, and a team should know both before relying on it. First, the harness caps consecutive continuations. In Claude Code, after Stop hooks have continued the turn eight times in a row, the harness overrides the next block and ends the turn. The count resets whenever the agent calls a tool, so an agent that keeps working does not hit it, and an agent that keeps saying "done" without doing anything does.[1](#user-content-fn-claude-hooks) The cap is configurable. The gate should treat a session that ends this way as `no_flip`, and the trace should record that the cap, not the oracle, ended it. Second, hooks have timeouts. A Stop hook has minutes, not seconds, which is enough for an oracle that runs a test suite in a container. It is not enough for a suite that takes an hour. For slow oracles, the gate should submit the diff to the oracle host, return a provisional block with the fixed reason, and let the next Stop pick up the verdict. The reference plugin runs the oracle inline because most suites that are fit for an oracle run in minutes. ## Redaction The tool output is where secrets live. A `cat .env`, a failing test that prints a connection string, a curl response with a token: all of it passes through PostToolUse. The trace collector redacts before it writes. Redaction at this stage is pattern-based and should be treated as a floor, not a guarantee. Patterns for the common shapes, such as `AKIA` prefixed AWS keys, GitHub tokens, private key blocks, bearer headers, and `KEY=value` lines where the key name contains `SECRET`, `TOKEN`, or `PASSWORD`, catch most of what appears in practice. They do not catch a password that looks like a word. Meli, McNiece, and Reaves found secrets leaking into public repositories at a rate of thousands of new unique secrets a day, across more than a hundred thousand repositories, which is a measure of how often they appear in code and output that people thought was fine.[2](#user-content-fn-meli-secrets) The collector redacts, and the pipeline runs a second pass with a dedicated scanner before anything reaches training, and the model is kept private to the organization whose traces it learned from. Chapter 16 covers the memorization research that makes the last rule a hard one. ## Storage Traces are written to a directory the plugin owns, not to the repository and not to the plugin's install directory. In Claude Code, that is the plugin's data directory, which the harness exposes as `CLAUDE_PLUGIN_DATA` and which survives plugin updates.[3](#user-content-fn-claude-plugins) A sync job moves finished traces to an object store under a path keyed by organization, repository, and date. The store is append-only. Nothing in the pipeline edits a trace after it is written; corrections are new records that reference the old. Each verdict record is signed by the oracle, as Chapter 5 describes. A trace whose verdict does not verify is excluded from training and flagged. This is the check that stops a compromised or misconfigured agent host from writing its own PASS. ## The same idea in other harnesses The hook design is not specific to one product. OpenAI's Codex CLI, Cursor's agent, OpenCode, and others expose lifecycle events with the same shape: before and after tool calls, and at the end of a turn. The names differ. The stdin JSON differs. The decision mechanism for refusing a stop differs, and in some harnesses does not exist, in which case the gate has to run as a wrapper around the harness instead of inside it. What stays constant is the architecture. One hook writes events. One hook, at the end of a turn, asks the oracle and either lets the agent stop or sends it back. The trace schema is the harness-neutral part, and a team that runs more than one harness should normalize events into the same schema at write time, so the training set does not depend on which tool produced it. ## Footnotes 1. Anthropic (2026). *Hooks reference*. Claude Code documentation. [https://code.claude.com/docs/en/hooks](https://code.claude.com/docs/en/hooks) [↩](#user-content-fnref-claude-hooks) [↩2](#user-content-fnref-claude-hooks-2) [↩3](#user-content-fnref-claude-hooks-3) [↩4](#user-content-fnref-claude-hooks-4) 2. Meli, M., McNiece, M. R., & Reaves, B. (2019). *How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories*. NDSS 2019. [https://doi.org/10.14722/ndss.2019.23418](https://doi.org/10.14722/ndss.2019.23418) [↩](#user-content-fnref-meli-secrets) 3. Anthropic (2026). *Plugin manifest reference*. Claude Code documentation. [https://code.claude.com/docs/en/plugins-reference](https://code.claude.com/docs/en/plugins-reference) [↩](#user-content-fnref-claude-plugins) --- # Chapter 8: The reference plugin > A walkthrough of oracle-flip: five hooks, one constant reason, and an oracle in a container with no network. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-reference-plugin This chapter walks through `oracle-flip`, a Claude Code plugin that does what Chapters 2 through 7 describe. It is public at [github.com/macanderson/oracle-flip](https://github.com/macanderson/oracle-flip) under the MIT license. It is a reference, meant to be read and adapted, and it is small enough to read in one sitting. ## Install ```text /plugin marketplace add macanderson/oracle-flip /plugin install oracle-flip@macanderson ``` To try it without installing, clone the repository and start Claude Code with the plugin loaded from disk: ```text git clone https://github.com/macanderson/oracle-flip claude --plugin-dir ./oracle-flip ``` ## Layout ```text oracle-flip/ .claude-plugin/ plugin.json name, version, description marketplace.json lets /plugin marketplace add find it hooks/ hooks.json the five hook registrations scripts/ common.py paths, event writer, redaction, config session_start.py SessionStart: write session_start trace.py UserPromptSubmit and PostToolUse: write message and tool_call gate.py Stop: diff, oracle, verdict, block or allow session_end.py SessionEnd: write session_end with the outcome oracle/ run.sh reference oracle runner: container, allowlist, hidden tests, one bit grade.sh runs inside the container filter_diff.py drops hunks outside the allowlist Dockerfile pinned base image; the container runs with no network example/ a sample task with a hidden test tools/ export_sft.py traces plus verdicts to SFT and preference-pair JSONL schema/ trace-event.schema.json tests/ test_redaction.py, test_filter_diff.py, test_gate.py ``` The scripts are Python with no dependencies outside the standard library, so the plugin runs anywhere `python3` runs. The oracle runner is shell plus one Python script, because it is meant to be replaced by whatever the team's continuous-integration system already does. ## The hook registrations The plugin's `hooks/hooks.json` wraps the event map in a `hooks` key, which is the shape plugins use.[1](#user-content-fn-claude-plugins) Each hook is in exec form: a `command` and an `args` array, with the plugin root substituted by the harness, so paths with spaces need no quoting. ```json { "description": "oracle-flip: trace every tool call and gate Stop on a hidden oracle", "hooks": { "SessionStart": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_start.py"], "timeout": 20 } ] } ], "UserPromptSubmit": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 10 } ] } ], "PostToolUse": [ { "matcher": "*", "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 20 } ] } ], "Stop": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/gate.py"], "timeout": 600 } ] } ], "SessionEnd": [ { "hooks": [ { "type": "command", "command": "python3", "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_end.py"], "timeout": 20 } ] } ] } } ``` ## Configuration Configuration is split by who controls the file. Anything the agent can edit holds nothing the gate trusts. The project the agent works in holds one small file, `.oracle-flip/config.json`. It names the task and sets the attempt budget. ```json { "task_id": "billing-proration-rounding", "base_commit": null, "max_attempts": 6 } ``` `base_commit` left null means the commit at session start. `max_attempts` is the number of FAIL verdicts after which the gate stops blocking and records `no_flip`. It exists so a session that will never flip does not run the oracle forever, and it is set below the harness's own cap of eight consecutive continuations so the plugin, not the harness, decides when to give up. The SessionStart hook reads this file once and copies it into the session's state. The gate reads the state, never the project, so an agent that edits the file mid-session changes nothing. An agent that edits it before the session starts can at most point the session at a task whose baseline passes, which the gate records as `no_task`, never as a flip. The oracle command, the oracle store, and the container image live in the plugin's user settings, which Claude Code stores outside every repository and passes to hook processes as environment variables. The oracle command receives the mode, the task id, the base commit, the path of the diff file, and a hash of the repository's remote URL, and it must exit 0 for PASS and 1 for FAIL. Any other exit code is an oracle error, which the gate records and treats as FAIL for the purpose of blocking, with a different fixed reason. The allowlist and denylist of paths the diff may touch live in the oracle store beside the hidden tests, one file each, and the oracle applies them before it touches a file. The gate's host does the filtering nowhere, because the gate's host is the agent's host. The hidden tests are in the same store, keyed by task id, and nowhere in the repository. ## The gate The gate is the file to read if you read one. Its structure follows Chapter 7. ```python def decide(event, state, *, oracle=run_oracle, diff=write_diff): write_event(assistant_message(event)) # last_assistant_message, redacted if state.task_id is None: return allow() # no oracle task; the plugin only traces if state.outcome is not None: return allow() # already decided; never loop on a settled outcome if state.baseline is None: baseline = oracle(state.oracle_command, mode="baseline", ...) state.baseline = baseline.verdict record_verdict(state, "baseline", baseline) if baseline.verdict == "PASS": state.outcome = "no_task" # the hidden test already passes return allow() if baseline.verdict == "ERROR": state.outcome = "oracle_error" return allow() diff_path = diff(event["cwd"], state.base_commit, ...) verdict = oracle(state.oracle_command, mode="grade", diff_path=diff_path, ...) state.attempts += 1 record_verdict(state, "grade", verdict) if verdict.verdict == "PASS": state.outcome = "flipped" return allow() if state.attempts >= state.max_attempts: state.outcome = "no_flip" return allow() return block(BLOCK_REASON_FAIL if verdict.verdict == "FAIL" else BLOCK_REASON_ERROR) ``` Four details are worth pointing at. The gate reads everything from the session state the SessionStart hook wrote: the task id, the base commit, the attempt budget, and the oracle command. It reads nothing from the project. The baseline is computed on the first Stop, not at session start, because the oracle is slow and the baseline is deterministic. It is stored, so later Stops in the same session reuse it. A team with many sessions per task should compute it once per task and share it; the plugin keeps it per session for simplicity. The `block` reason is one of two constants. Neither includes the oracle's output. The oracle's output is written by the oracle, on the oracle's side, and the gate never sees it. The gate's own tests check that the reasons contain no test name and no assertion. `stop_hook_active` is honored by the attempt counter rather than by an early return, because the harness sets it on every Stop after the first block, and returning early on it would mean the gate only ever runs once. The attempt counter plus `max_attempts` is what prevents the loop, and `decide` takes the oracle and the diff builder as parameters so the loop logic is tested without a container. ## The oracle runner `oracle/run.sh` is the reference oracle. It expects a store with one directory of hidden tests per task and a mirror of the repository, and it runs the grade inside a container with networking disabled. The core of it: ```sh task_dir="$store/$ORACLE_TASK_ID" log_dir="$store/.log/$ORACLE_TASK_ID/$(date -u +%Y%m%dT%H%M%SZ)-$ORACLE_MODE" args=(run --rm --network none -e TZ=UTC -e LC_ALL=C.UTF-8 -e PYTHONHASHSEED=0 -e SOURCE_DATE_EPOCH=1700000000 -e ORACLE_MODE="$ORACLE_MODE" -e ORACLE_BASE_COMMIT="$ORACLE_BASE_COMMIT" -v "$task_dir:/task:ro" -v "$mirror:/mirror:ro" -v "$log_dir:/log") [ "$ORACLE_MODE" = "grade" ] && args+=(-v "$ORACLE_DIFF:/in/diff.patch:ro") docker "${args[@]}" "$image" bash /oracle/grade.sh > "$log_dir/stdout.txt" 2> "$log_dir/stderr.txt" status=$? grep -E '^tests_hash=' "$log_dir/stdout.txt" || true exit "$status" ``` `grade.sh` runs inside the container. It clones the repository from the read-only mirror, checks out the base commit, filters the diff through the task's allowlist and denylist with `filter_diff.py`, applies what is left, copies the hidden tests from `/task/hidden` into place, runs the existing suite and then the hidden tests, prints `tests_hash=` followed by a hash of the hidden test list, and exits 0 or 1. In baseline mode it runs without a diff and exits 1 when the suite passes and the hidden tests fail, which is the only baseline that makes a task. Its full output goes to the log directory on the oracle's side. Only the exit code and the `tests_hash` line return to the gate. The Dockerfile pins its base image by digest and installs the repository's dependencies from a lockfile at build time, so the run-time container needs no network and gets none. ## Limits of the plugin It runs the oracle on the same machine as the agent, in a container. Chapter 5 explained why that protects the oracle's run and not the oracle's secrets. On a laptop, the agent's shell can read `ORACLE_STORE`. The README says so. The step to a separate host is to replace `oracle/run.sh` with a script that submits the diff to a job runner the agent cannot log into and waits for the bit. It does not sign verdicts. Signing needs a key on the oracle's side and a verifier in the pipeline, and the plugin has neither because it has no pipeline side. The verdict event has a field for the signature and the export tool has a flag that requires it. It redacts with patterns. That is a floor. Run a real secret scanner on the trace store before training. It has not been tested with a real container run on the author's machine at the time of writing, for a reason unrelated to the plugin, and the README says that too. The hook scripts are tested by piping sample events into them, and `claude plugin validate` passes. ## The export tool `tools/export_sft.py` reads a directory of traces and their verdicts and writes two files. The first is the supervised set: one JSON object per flipped session, in a chat format with tool calls, with a mask that marks which turns to train on. User prompts and tool outputs are masked out; the model is trained to produce the assistant's turns and tool calls, not to reproduce the environment's replies. Chapter 10 explains why. The second is the preference set: for every task with at least one flipped session and at least one that did not flip, a pair with the flipped trace as the preferred one. Chapter 11 explains what to do with it. Both files exclude any session whose verdict record is missing, and, with `--require-signature`, any whose verdict does not verify. ## Footnotes 1. Anthropic (2026). *Plugin manifest reference*. Claude Code documentation. [https://code.claude.com/docs/en/plugins-reference](https://code.claude.com/docs/en/plugins-reference) [↩](#user-content-fnref-claude-plugins) --- # Chapter 9: The flip rate > Write the test first, dispatch small, sample more than once, manufacture tasks, and keep the failures. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-flip-rate The pipeline's throughput is the number of flips per day. This chapter is about raising it: choosing tasks that can flip, dispatching work in a shape that flips, sampling more than once, and manufacturing tasks when the natural supply runs low. It ends with what to do with the sessions that do not flip, which is most of them. ## Start with the test A task can flip only if a hidden test fails on the base commit. The highest-leverage change a team can make is to write the test before dispatching the task. This is test-driven development with the roles split: a person or a separate agent writes the red test, and the solving agent is dispatched with the task description and never sees it. The practice has a long history under the name red-green-refactor, and the evidence that it improves defect rates predates language models.[1](#user-content-fn-nagappan-tdd) What is new is the reason to do it. In a team with an oracle, every red test is a potential verified trajectory. In a team without one, a red test is just a test. Some task shapes come with a red test for free. - **A bug report with a reproduction.** The reproduction is the hidden test. Clean it up, assert the correct behavior, confirm it fails on the base commit, and dispatch the bug. - **A failing test in continuous integration.** The test already exists and already fails. Hide it from the agent by giving the agent a checkout where the test is removed, and keep the test as the oracle. - **A feature with an acceptance criterion.** If the criterion can be stated as an assertion, it can be a hidden test. - **A refactor.** The existing suite is the oracle, with one added differential test: outputs on a set of inputs must match the previous build. Task shapes that do not come with a red test, such as "improve the error messages" or "clean up this module," are not oracle tasks as stated. They can be made into oracle tasks by adding a check, such as a snapshot test of the messages or a lint rule the cleanup must satisfy, or they can be done without the oracle and their traces kept as unlabeled data. ## Dispatch small A flip is binary. A task with three requirements and one hidden test that covers all three flips only when all three are done. Split it into three tasks with one hidden test each. The agent gets more feedback across the same work, the traces are shorter and cleaner, and a session that completes two of three is two flips instead of none. Small also means a short session. Long sessions drift, run out of context, and end on a FAIL that is as much about the session's length as about the task. SWE-bench Verified's human annotators excluded tasks whose descriptions were underspecified or whose tests were unfair, and the benchmark's resolve rates roughly doubled for the same models once the unfair tasks were removed.[2](#user-content-fn-swe-bench-verified) The lesson for a team is that a clear, bounded task description raises the flip rate without changing the agent. ## Sample more than once Repeated sampling is the single largest lever on the flip rate per task, and its cost is tokens. Brown and colleagues measured how coverage, the share of problems solved by at least one of k attempts, grows with k. On SWE-bench Lite, an open model that solved 15.9 percent of problems with one attempt solved 56 percent with 250 attempts.[3](#user-content-fn-large-language-monkeys) The oracle is what makes this usable: with a deterministic verifier, the team keeps the attempt that passes and discards the rest. Without one, the team has 250 patches and no way to choose. This is rejection sampling, and it is how most of the verified-trajectory datasets in the literature were built. SWE-Gym sampled many trajectories per task and kept the 491 that resolved their tasks.[4](#user-content-fn-swe-gym) SWE-smith kept 5,016 out of thousands more.[5](#user-content-fn-swe-smith) Llama 2's post-training used rejection sampling against a reward model as a core step.[6](#user-content-fn-llama2) For a team, the practical version is: dispatch each oracle task to several independent sessions with fresh context, and let the gate find the flip. The sessions must be independent. If one session's output leaks into another's context, the attempts are correlated and coverage grows slower. Fresh context per attempt and a silent oracle give independence. A verbose oracle that told each attempt why the last one failed would make the attempts a single long session in disguise. > **Coverage grows with attempts when a verifier picks the winner** > > - **DeepSeek-Coder-V2-Instruct on SWE-bench Lite**: 56% at 250 > > *Source: Brown et al. (2024). Only the two reported endpoints are plotted; the paper shows a roughly log-linear curve between them. Every added attempt costs tokens and, with a hidden oracle, each attempt is an independent draw.* A verifier trained on your own traces makes sampling cheaper. Pan and colleagues trained a verifier on the same SWE-Gym trajectories and used it to pick the best of 16 attempts, which raised their model from 20.6 to 32.0 percent.[4](#user-content-fn-swe-gym) The verifier does not replace the oracle; the oracle still grades the chosen attempt. The verifier reduces how many attempts need to reach the oracle. ## Manufacture tasks When the natural supply of red tests is smaller than the agent capacity, tasks can be made. **Break working code.** Take a module with good tests. Introduce a bug: delete a branch, flip a comparison, drop a null check. The existing tests that now fail are the hidden tests. The task description is the symptom, written from the test's point of view without naming the test. This is how SWE-smith built 50,000 task instances from 128 repositories, and the authors found that the synthetic bugs trained a model that transferred to real issues.[5](#user-content-fn-swe-smith) Mutation tools generate these bugs automatically, and the mutation literature has catalogued which mutants resemble real faults.[7](#user-content-fn-just-mutants) **Mine history.** Every past commit that changed source and tests together is a candidate task: the tests it added are the hidden tests, the parent commit is the base, and the commit message or linked issue is the task description. This is how SWE-bench was built and how SWE-rebench automated the construction into a pipeline that produced more than 21,000 tasks.[8](#user-content-fn-swe-bench)[9](#user-content-fn-swe-rebench) A team's own history is a supply of tasks on its own code that no public dataset contains. **Reverse a fix.** For a merged bug fix with a regression test, revert the source change and keep the test. The agent is dispatched to fix the bug again. The trace is a worked example on real code with a real test. **Let a model propose tasks under an executor.** Zhao and colleagues' Absolute Zero had a model propose coding tasks and solve them, with a code executor checking both that the task was well-formed and that the answer was right.[10](#user-content-fn-absolute-zero) The executor is the oracle. This is the most speculative item on the list and the one that produces the least realistic tasks; it is here because it shows that the supply of oracle tasks is not bounded by the supply of issues. ## Use the strong model on the hard tail Some tasks will not flip with the model you are training. Dispatch them to a rented frontier model and keep the trace. A flipped trace from a stronger model is a distillation example, and distillation from a stronger model into a weaker one is the oldest trick in the post-training book.[11](#user-content-fn-hinton-distillation) SWE-smith's 5,016 trajectories came from Claude 3.7 Sonnet, and the model trained on them was a 32-billion-parameter Qwen.[5](#user-content-fn-swe-smith) A team already paying for the frontier model is already producing these traces. The oracle is what sorts them. ## Keep the failures Most sessions will not flip. Those traces are not waste. A session that failed on the same task where another session flipped is half of a preference pair. Direct preference optimization trains a model to prefer the chosen trajectory over the rejected one, and it needs exactly this data.[12](#user-content-fn-dpo) A team that keeps only flips has a supervised set. A team that keeps everything has a supervised set and a preference set. A session that failed is also a record of what went wrong, for people. If a task fails ten sessions in a row, the oracle log will say whether the hidden test is wrong, the task description is unclear, or the task is beyond the model. Each of those has a different fix, and the trace is how you find out which. Hindsight relabeling, from robotics, is the formal version of this idea: an episode that failed its goal succeeded at whatever it did reach, and can be relabeled as a success for that.[13](#user-content-fn-her) A session that did not flip the hidden test but did make the existing suite pass after breaking it, or did fix a different bug on the way, has a relabeled success in it. The pipeline in Chapter 10 does not do this automatically, and it is the kind of thing a team adds once the basic loop runs. ## The arithmetic A team of 50 engineers runs an agent on perhaps 10 tasks a day each. Call it 500 sessions a day. If one in five sessions is on an oracle task, that is 100 oracle sessions a day. If one in three of those flips, that is about 33 flips a day. In a working month, around 700. Chapter 12 says what 700 buys. Doubling the oracle-task share doubles it; sampling three attempts per task roughly doubles it again. The levers are the share of work that has a red test, the size of each task, and the number of attempts. None of them is the model. ## Footnotes 1. Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). *Realizing quality improvement through test driven development: results and experiences of four industrial teams*. Empirical Software Engineering, 13(3), 289–302. [https://doi.org/10.1007/s10664-008-9062-z](https://doi.org/10.1007/s10664-008-9062-z) [↩](#user-content-fnref-nagappan-tdd) 2. OpenAI (2024). *Introducing SWE-bench Verified*. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/) [↩](#user-content-fnref-swe-bench-verified) 3. Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). *Large Language Monkeys: Scaling Inference Compute with Repeated Sampling*. arXiv. [https://arxiv.org/abs/2407.21787](https://arxiv.org/abs/2407.21787) [↩](#user-content-fnref-large-language-monkeys) 4. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) [↩2](#user-content-fnref-swe-gym-2) 5. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) [↩](#user-content-fnref-swe-smith) [↩2](#user-content-fnref-swe-smith-2) [↩3](#user-content-fnref-swe-smith-3) 6. Touvron, H., et al. (2023). *Llama 2: Open Foundation and Fine-Tuned Chat Models*. arXiv. [https://arxiv.org/abs/2307.09288](https://arxiv.org/abs/2307.09288) [↩](#user-content-fnref-llama2) 7. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). *Are Mutants a Valid Substitute for Real Faults in Software Testing?* FSE 2014, 654–665. [https://doi.org/10.1145/2635868.2635929](https://doi.org/10.1145/2635868.2635929) [↩](#user-content-fnref-just-mutants) 8. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) [↩](#user-content-fnref-swe-bench) 9. Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). *SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2505.20411](https://arxiv.org/abs/2505.20411) [↩](#user-content-fnref-swe-rebench) 10. Zhao, A., Wu, Y., Yue, Y., et al. (2025). *Absolute Zero: Reinforced Self-play Reasoning with Zero Data*. arXiv. [https://arxiv.org/abs/2505.03335](https://arxiv.org/abs/2505.03335) [↩](#user-content-fnref-absolute-zero) 11. Hinton, G., Vinyals, O., & Dean, J. (2015). *Distilling the Knowledge in a Neural Network*. NeurIPS 2014 Deep Learning Workshop. [https://arxiv.org/abs/1503.02531](https://arxiv.org/abs/1503.02531) [↩](#user-content-fnref-hinton-distillation) 12. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. [https://arxiv.org/abs/2305.18290](https://arxiv.org/abs/2305.18290) [↩](#user-content-fnref-dpo) 13. Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). *Hindsight Experience Replay*. NeurIPS 2017. [https://arxiv.org/abs/1707.01495](https://arxiv.org/abs/1707.01495) [↩](#user-content-fnref-her) --- # Chapter 10: From traces to training data > Select, deduplicate, format, mask, hold out, version. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/from-traces-to-training-data A trace store is not a training set. Between the two sit a series of transformations that decide what the model learns and what it does not. This chapter walks through them in the order the pipeline applies them. ## Select The first decision is which sessions to include. The rule for the supervised set is simple: a session is included when its outcome is `flipped` and its verdict record verifies. Everything else is excluded from the supervised set, and the reason is recorded. A session that flipped is not automatically a good example. Three further checks are cheap and worth running. - **The diff touched source.** A session whose only changes were to paths the oracle dropped did not flip because of anything the agent wrote to the code under test. If the oracle still reported PASS, the baseline was wrong. Exclude the session and quarantine the task. - **The session did not exceed the length budget.** A session with 400 tool calls that flipped is a worked example of flailing until something stuck. Set a budget per task class and exclude sessions over it, or truncate to the final successful stretch if the trace shows a clear restart. - **The agent did not attempt to touch excluded paths repeatedly.** The oracle's dropped-hunk log shows an agent that kept editing test configuration. That session may have flipped on the merits, but its trace teaches the behavior you least want. Exclude it, and look at whether the task description invited it. For the preference set, pair each flipped session on a task with each session on the same task that did not flip. Where there are many of each, sample pairs rather than taking the full cross product, so one task does not dominate. ## Deduplicate Two sessions on the same task that flipped with nearly identical trajectories add little to each other. Hash the sequence of tool names and the final diff; where two sessions share both, keep one. Where they share the diff but not the path to it, keep both, because the paths are the data. Across tasks, deduplicate on the task itself. A team that manufactured 200 tasks by mutating one module will have 200 near-identical trajectories, and a model trained on them will learn that module. Cap the number of flips per source file or per task family. ## Format A trace is a sequence of events. A training example is a conversation in the chat format the base model was trained on, with tool calls in the model's native tool-call syntax. The conversion is mechanical but has choices in it. The **system prompt** should be the one the agent ran with, including the rule files that were loaded, because that is the context the model will see at inference. If the team's rule files change often, consider training on a canonical version and keeping the diff as metadata. The **user turn** is the task description. Each **assistant turn** is either a tool call, with the tool name and arguments, or a message. Where the harness exposes the model's reasoning, include it as the base model's format expects; where it does not, the turn is the call alone. Each **tool result** is a tool turn, with the redacted output. The **final assistant turn** is the agent's completion message. Convert to the exact chat template of the base model you will fine-tune. A mismatch between the training template and the inference template is the most common silent failure in fine-tuning, and it produces a model that looks trained and behaves untrained. ## Mask The loss, the quantity the training run minimizes, should be computed only on the tokens the model is supposed to produce. In a trace, that is the assistant's turns: its reasoning, its tool calls, and its messages. The system prompt, the user's task, and every tool result are context, not targets. Masking the tool results matters more than it might seem. Tool outputs are long: a file read returns hundreds of lines, a test run returns pages. Unmasked, they are most of the tokens in the example, and the model spends its capacity learning to predict file contents and test logs instead of learning to act. A model trained without the mask will also learn to hallucinate tool results, because it was trained to produce them. The reference export tool emits a mask field per turn for this reason. ## Handle length Agent trajectories are long. A session that flipped after 60 tool calls, each returning a few kilobytes, is a hundred thousand tokens. Base models with long context windows handle this; training on sequences that long is expensive and some frameworks do not support it well. Three strategies, in order of preference: 1. **Train on the full trajectory** when the framework and hardware allow it. This preserves the behavior you want: the model learns to carry a plan across many steps. 2. **Truncate tool outputs**, not turns. A file read can be cut to the lines around the ones the agent later edited, with a marker. A test log can be cut to the failures. The trajectory keeps its shape and loses bulk. 3. **Window the trajectory** into overlapping segments, each with the system prompt and task prepended and a summary of the dropped prefix. This loses long-range structure and should be the fallback. Do not drop the dead ends. A trajectory in which the agent tried something, saw it fail, and backed out is a trajectory that teaches recovery. SWE-Gym's authors kept full trajectories including unproductive steps and reported that it worked; cleaning them to the shortest path is a reasonable experiment but not the default.[1](#user-content-fn-swe-gym) ## Hold out Before anything is trained, split by task, not by session. All sessions on a task go to the same side of the split. Set aside a fraction of tasks, with their hidden tests, as the evaluation set, and never train on any session from them. A model evaluated on tasks it saw during training will look better than it is, and the leak is undetectable afterward. Keep two evaluation sets. A **frozen set** fixed at the start of the program, so that every model version is scored against the same tasks and the trend is comparable. A **rolling set** of recent tasks, so that the evaluation tracks the work the team is doing now. Chapter 14 covers how to read them. ## Version Every training set is a manifest: the list of session ids included, the hash of each trace, the hash of each verdict, the filter rules applied, the template used, and the split. Store the manifest beside the weights it produced. When a model regresses, the manifest is how you find the traces that taught it. A training set that is reproducible from its manifest is the data equivalent of a reproducible build. The same traces and the same rules give the same examples, and a change in either is visible as a change in the hash. ## Record provenance Each example carries the identity of the oracle that graded it, as Chapter 4 said, and the identity of the model that produced the trace. The second matters more than it looks. A training set that is half traces from a rented frontier model and half from the team's own fine-tuned model is a mixture of distillation and self-improvement, and the two have different failure modes. Chapter 11 covers the self-improvement risk. The provenance field is what lets you tell them apart later. ## Footnotes 1. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) --- # Chapter 11: Training methods > Supervised fine-tuning on flips, then preference pairs, then reinforcement learning. Adapters, forgetting, collapse, and a shuffled-label control. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/training-methods With a training set in hand, the question is what to do with it. This chapter covers the three families of methods that have produced results on agent trajectories, in the order a team should adopt them, and the choices inside each: adapter or full fine-tune, how to avoid forgetting, and how to know the signal is real. ## Supervised fine-tuning on flips The first method is the simplest. Take the flipped trajectories, formatted and masked as Chapter 10 describes, and continue training the base model on them with the standard next-token loss. This is supervised fine-tuning, and when the examples were selected by a verifier, it has a second name: rejection sampling fine-tuning. The method has a lineage. Expert iteration, from 2017, alternates between a slow expert that solves problems and a fast policy that imitates the solutions, and uses the expert's successes as the training set.[1](#user-content-fn-expert-iteration) AlphaGo Zero's self-play is the same loop with the game's outcome as the verifier.[2](#user-content-fn-alphago-zero) STaR, from 2022, applied it to language models: sample reasoning, keep the samples whose answers match the key, fine-tune, repeat.[3](#user-content-fn-star) ReST and ReST-EM scaled it with a reward model and with answer keys.[4](#user-content-fn-rest)[5](#user-content-fn-rest-em) For code, AlphaCode filtered thousands of samples per problem through the problem's tests before choosing which to submit.[6](#user-content-fn-alphacode) Llama 2's post-training used rejection sampling as a step before reinforcement learning.[7](#user-content-fn-llama2) The agent-trajectory results from Chapter 1 are this method applied to software: SWE-Gym, SWE-smith, and Skywork-SWE are all supervised fine-tuning on verifier-selected trajectories.[8](#user-content-fn-swe-gym)[9](#user-content-fn-swe-smith)[10](#user-content-fn-skywork-swe) It works, it is cheap, and it is the right place to start. Its limit is that it can only teach the model to do what some session already did. A task no session ever flipped contributes nothing to the supervised set. ## Preference optimization on pairs The second method uses the failures. Direct preference optimization takes pairs of trajectories on the same task, one preferred and one not, and trains the model to assign higher likelihood to the preferred one relative to a reference model.[11](#user-content-fn-dpo) It needs no reward model and no sampling during training, which makes it almost as cheap as supervised fine-tuning. The preference set from Chapter 10 is the input: flipped versus not flipped on the same task. The pairs teach the model something the supervised set cannot, which is what a wrong trajectory looks like on this codebase. A team that has run two attempts per task has pairs for every task where exactly one flipped. Two cautions. Preference optimization is known to drift toward longer outputs unless the pairs are controlled for length, and agent trajectories vary in length a lot.[12](#user-content-fn-length-bias) Match pair lengths roughly or use a length-regularized variant. And the pairs should be real contrasts. A pair where the rejected trajectory failed because the oracle timed out is not a lesson about the code. ## Reinforcement learning with the oracle as the reward The third method puts the oracle in the training loop. The model attempts a task, the oracle grades the attempt, and the model is updated to make PASS more likely. This is reinforcement learning with verifiable rewards, named as such in the Tülu 3 report and used at scale in DeepSeek-R1.[13](#user-content-fn-tulu3)[14](#user-content-fn-deepseek-r1) For software agents, SWE-RL used a patch-similarity reward on mined pull requests, DeepSWE used a pass-or-fail reward on executable environments, and the Nebius team took a 72-billion-parameter model from 11.4 percent to 39.0 percent on SWE-bench Verified with rejection sampling followed by a variant of the DAPO algorithm.[15](#user-content-fn-swe-rl)[16](#user-content-fn-deepswe)[17](#user-content-fn-nebius-rl)[18](#user-content-fn-dapo) The appeal is that reinforcement learning can learn from tasks no session has flipped yet, because it searches. The costs are real. Each training step needs many fresh attempts, each attempt needs an oracle run, and the oracle runs in a container. Qwen's team described a system running 20,000 environments in parallel for its coding model's reinforcement learning stage.[19](#user-content-fn-qwen3-coder) A team does not need that scale to see gains; DeepSWE used about 4,500 tasks.[16](#user-content-fn-deepswe) It does need an oracle that can grade hundreds of attempts an hour, which means the oracle has to be a service, not a hook. There is also a result to read before committing. Yue and colleagues found that reinforcement learning with verifiable rewards sharpens a model toward answers its base model could already produce: the trained model wins when it gets one attempt, and the base model wins when both get many attempts, because the base model's attempts are more varied.[20](#user-content-fn-yue-rlvr) For a team, that argues for a sequence: supervised fine-tuning on flips first, to widen what the model can do on your code; reinforcement learning second, to make it do it reliably on the first try. Adopt the methods in that order. Supervised fine-tuning on flips as soon as there are a few hundred. Preference optimization as soon as there are pairs. Reinforcement learning when the oracle is a service and the flip rate has plateaued. > **Training methods in the order a team should adopt them** > > 1. **Supervised fine-tuning on flipped trajectories**: Rejection sampling fine-tuning. Cheapest. Learns what some session already did. > 2. **Preference optimization on flipped versus failed pairs**: Uses the failures. Learns what wrong looks like here. > 3. **Reinforcement learning with the oracle as reward**: Searches. Can learn tasks no session flipped. Needs the oracle as a service. > > *Each rung needs the one below it. The data each rung needs comes from the same trace store.* ## Adapter or full fine-tune Low-rank adaptation freezes the base model's weights and trains small matrices added to them.[21](#user-content-fn-lora) The adapter is a few percent of the model's size, trains on less hardware, and can be swapped at inference. QLoRA goes further by quantizing the frozen base to 4 bits, which brought fine-tuning of 65-billion-parameter models onto a single large card.[22](#user-content-fn-qlora) The question is whether the adapter learns as much. Biderman and colleagues compared the two on code and math and found that full fine-tuning learned more on the target domain and forgot more of what the base model knew; low-rank adaptation learned less and forgot less, and acted as a regularizer.[23](#user-content-fn-lora-forgets) Schulman and colleagues at Thinking Machines then reported that the gap closes when the adapter is applied to all layers rather than only attention, and that for small-to-medium post-training sets, the size a team's trace store will be for its first year, the adapter matches the full fine-tune. Their result for reinforcement learning is sharper: a rank-1 adapter matched full fine-tuning, which they explain by noting that a policy-gradient step absorbs about one bit of information per episode, so the capacity needed is tiny.[24](#user-content-fn-lora-without-regret) That last point connects to this book's oracle. A one-bit verdict is one bit per episode. A method that learns one bit per episode needs many episodes and almost no adapter capacity. That is the regime the pipeline is in, and the adapter is the right tool for it. Start with adapters on all layers. Move to full fine-tuning only if an evaluation shows the adapter is the bottleneck, which for the first several thousand flips it is unlikely to be. ## Forgetting A model fine-tuned on traces from one codebase can get worse at everything else. The phenomenon is catastrophic forgetting, known since 1989 and measured in language models at every scale.[25](#user-content-fn-mccloskey-forgetting)[26](#user-content-fn-luo-forgetting) For a coding model, it looks like a fine-tune that resolves the team's tasks and can no longer write a shell script. The standard defenses apply. - **Replay.** Mix a fraction of general instruction data into every training run. Ibrahim and colleagues showed that replay combined with re-warming the learning rate lets a model take on new data while matching a full retrain on the old.[27](#user-content-fn-ibrahim-continual) - **Low learning rate and few epochs.** One to three passes over the flips at a learning rate an order of magnitude below pretraining. - **Adapters.** The frozen base cannot forget. The adapter can be removed. This is the regularization Biderman and colleagues measured.[23](#user-content-fn-lora-forgets) - **Merge, don't stack.** When several adapters have been trained on different slices, merge them with a method that keeps the base model's weights and averages or resolves the deltas, such as model soups or TIES, rather than fine-tuning one on top of another.[28](#user-content-fn-model-soups)[29](#user-content-fn-ties-merging) - **Evaluate on a general benchmark.** Chapter 14 puts a public coding benchmark in the gate for this reason: a model that gained on the team's tasks and lost on the public set has forgotten, and the gate should see it. ## Collapse Training a model on its own output can make it worse over generations. Shumailov and colleagues showed that models trained recursively on their own generations lose the tails of the distribution and converge on a narrow set of outputs, a process they called model collapse.[30](#user-content-fn-shumailov-collapse) A pipeline that trains a model on traces produced by the previous version of itself is, on its face, exactly that loop. Two things make it different. First, the oracle. Collapse happens when generated data replaces real data without selection. The traces in this pipeline are selected by a verifier the model does not control, so the distribution being trained on is the distribution of correct solutions, not of the model's output. Gerstgrasser and colleagues showed that collapse is also avoided when generated data accumulates alongside the original data rather than replacing it.[31](#user-content-fn-gerstgrasser-collapse) Second, the provenance field from Chapter 10. A training run that knows which traces came from a stronger external model and which from its own predecessor can weight them, cap the self-generated share, and watch the ratio over time. The warning sign is a model whose flips get shorter, more uniform, and more alike across tasks. The frozen evaluation set from Chapter 10 is the instrument. ## Is the signal real One last check belongs in every training run, and it is cheap. Shao and colleagues found that training a particular family of math models with random rewards, rewards that had nothing to do with correctness, improved its benchmark scores by more than 20 points.[32](#user-content-fn-spurious-rewards) The gain came from the training procedure nudging the model toward behaviors it already had, not from the reward. The same procedure did nothing for other model families. The result does not say verifiable rewards are fake. It says a gain after training is not, on its own, evidence that the reward carried information. The control is to shuffle the labels. Train the same recipe on the same traces with the flip labels randomly permuted, so that half the "flipped" trajectories are failures. If the shuffled run gains as much as the real run on the held-out evaluation, the gain is not coming from the oracle, and something else in the recipe is doing the work. A team should run this control once when it sets up the pipeline, and again whenever the recipe changes. ## Footnotes 1. Anthony, T., Tian, Z., & Barber, D. (2017). *Thinking Fast and Slow with Deep Learning and Tree Search*. NeurIPS 2017. [https://arxiv.org/abs/1705.08439](https://arxiv.org/abs/1705.08439) [↩](#user-content-fnref-expert-iteration) 2. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). *Mastering the game of Go without human knowledge*. Nature, 550, 354–359. [https://doi.org/10.1038/nature24270](https://doi.org/10.1038/nature24270) [↩](#user-content-fnref-alphago-zero) 3. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). *STaR: Bootstrapping Reasoning With Reasoning*. arXiv. [https://arxiv.org/abs/2203.14465](https://arxiv.org/abs/2203.14465) [↩](#user-content-fnref-star) 4. Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). *Reinforced Self-Training (ReST) for Language Modeling*. arXiv. [https://arxiv.org/abs/2308.08998](https://arxiv.org/abs/2308.08998) [↩](#user-content-fnref-rest) 5. Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). *Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models*. arXiv; TMLR 2024. [https://arxiv.org/abs/2312.06585](https://arxiv.org/abs/2312.06585) [↩](#user-content-fnref-rest-em) 6. Li, Y., Choi, D., Chung, J., et al. (2022). *Competition-level code generation with AlphaCode*. Science, 378(6624), 1092–1097. [https://doi.org/10.1126/science.abq1158](https://doi.org/10.1126/science.abq1158) [↩](#user-content-fnref-alphacode) 7. Touvron, H., et al. (2023). *Llama 2: Open Foundation and Fine-Tuned Chat Models*. arXiv. [https://arxiv.org/abs/2307.09288](https://arxiv.org/abs/2307.09288) [↩](#user-content-fnref-llama2) 8. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) 9. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) [↩](#user-content-fnref-swe-smith) 10. Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290) [↩](#user-content-fnref-skywork-swe) 11. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. [https://arxiv.org/abs/2305.18290](https://arxiv.org/abs/2305.18290) [↩](#user-content-fnref-dpo) 12. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). *A Long Way to Go: Investigating Length Correlations in RLHF*. arXiv; COLM 2024. [https://arxiv.org/abs/2310.03716](https://arxiv.org/abs/2310.03716) [↩](#user-content-fnref-length-bias) 13. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). *Tülu 3: Pushing Frontiers in Open Language Model Post-Training*. arXiv. [https://arxiv.org/abs/2411.15124](https://arxiv.org/abs/2411.15124) [↩](#user-content-fnref-tulu3) 14. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) [↩](#user-content-fnref-deepseek-r1) 15. Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449) [↩](#user-content-fnref-swe-rl) 16. Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe) [↩](#user-content-fnref-deepswe) [↩2](#user-content-fnref-deepswe-2) 17. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501) [↩](#user-content-fnref-nebius-rl) 18. Yu, Q., Zhang, Z., Zhu, R., et al. (2025). *DAPO: An Open-Source LLM Reinforcement Learning System at Scale*. arXiv. [https://arxiv.org/abs/2503.14476](https://arxiv.org/abs/2503.14476) [↩](#user-content-fnref-dapo) 19. Qwen Team (2025). *Qwen3-Coder: Agentic Coding in the World*. [https://qwenlm.github.io/blog/qwen3-coder/](https://qwenlm.github.io/blog/qwen3-coder/) [↩](#user-content-fnref-qwen3-coder) 20. Yue, Y., Chen, Z., Lu, R., et al. (2025). *Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?* arXiv; NeurIPS 2025. [https://arxiv.org/abs/2504.13837](https://arxiv.org/abs/2504.13837) [↩](#user-content-fnref-yue-rlvr) 21. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). *LoRA: Low-Rank Adaptation of Large Language Models*. arXiv; ICLR 2022. [https://arxiv.org/abs/2106.09685](https://arxiv.org/abs/2106.09685) [↩](#user-content-fnref-lora) 22. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). *QLoRA: Efficient Finetuning of Quantized LLMs*. NeurIPS 2023. [https://arxiv.org/abs/2305.14314](https://arxiv.org/abs/2305.14314) [↩](#user-content-fnref-qlora) 23. Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). *LoRA Learns Less and Forgets Less*. Transactions on Machine Learning Research. [https://arxiv.org/abs/2405.09673](https://arxiv.org/abs/2405.09673) [↩](#user-content-fnref-lora-forgets) [↩2](#user-content-fnref-lora-forgets-2) 24. Schulman, J., & Thinking Machines Lab (2025). *LoRA Without Regret*. [https://thinkingmachines.ai/blog/lora/](https://thinkingmachines.ai/blog/lora/) [↩](#user-content-fnref-lora-without-regret) 25. McCloskey, M., & Cohen, N. J. (1989). *Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem*. Psychology of Learning and Motivation, 24, 109–165. [https://doi.org/10.1016/S0079-7421(08)60536-8](https://doi.org/10.1016/S0079-7421\(08\)60536-8) [↩](#user-content-fnref-mccloskey-forgetting) 26. Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). *An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning*. arXiv. [https://arxiv.org/abs/2308.08747](https://arxiv.org/abs/2308.08747) [↩](#user-content-fnref-luo-forgetting) 27. Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). *Simple and Scalable Strategies to Continually Pre-train Large Language Models*. arXiv; TMLR. [https://arxiv.org/abs/2403.08763](https://arxiv.org/abs/2403.08763) [↩](#user-content-fnref-ibrahim-continual) 28. Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). *Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time*. ICML 2022. [https://arxiv.org/abs/2203.05482](https://arxiv.org/abs/2203.05482) [↩](#user-content-fnref-model-soups) 29. Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). *TIES-Merging: Resolving Interference When Merging Models*. NeurIPS 2023. [https://arxiv.org/abs/2306.01708](https://arxiv.org/abs/2306.01708) [↩](#user-content-fnref-ties-merging) 30. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). *AI models collapse when trained on recursively generated data*. Nature, 631, 755–759. [https://doi.org/10.1038/s41586-024-07566-y](https://doi.org/10.1038/s41586-024-07566-y) [↩](#user-content-fnref-shumailov-collapse) 31. Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). *Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data*. arXiv. [https://arxiv.org/abs/2404.01413](https://arxiv.org/abs/2404.01413) [↩](#user-content-fnref-gerstgrasser-collapse) 32. Shao, R., Li, S. S., Xin, R., et al. (2025). *Spurious Rewards: Rethinking Training Signals in RLVR*. arXiv. [https://arxiv.org/abs/2506.10947](https://arxiv.org/abs/2506.10947) [↩](#user-content-fnref-spurious-rewards) --- # Chapter 12: Data volume > What 491, 5,016, and 8,209 verified trajectories bought in the literature, and a worksheet for your own rate. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume How many flips does it take? This chapter collects the published numbers, states the pattern they show, and gives a worksheet a team can fill in with its own rates. ## What the literature reports The results below are for different models, methods, and benchmarks, so the table is a set of reference points, not a curve. Read it for order of magnitude. | Data | Method | Model | Result | Source | | --- | --- | --- | --- | --- | | 1,000 curated prompt-response pairs | Supervised fine-tuning | LLaMA 65B | Preferred to or tied with GPT-4 responses in 43 percent of human comparisons | LIMA[1](#user-content-fn-lima) | | 1,000 reasoning questions with traces | Supervised fine-tuning | Qwen2.5-32B-Instruct | Competitive with o1-preview on competition math; 57 percent on AIME24 with budget forcing | s1[2](#user-content-fn-s1) | | 817 curated math problems with solutions | Supervised fine-tuning | Qwen2.5-32B-Instruct | 57.1 percent on AIME24 and 94.8 percent on MATH500 in the first release | LIMO[3](#user-content-fn-limo) | | 500 agent trajectories | Supervised fine-tuning | Llama 2 7B | 77 percent relative gain on a question-answering agent task | FireAct[4](#user-content-fn-fireact) | | 1,866 agent trajectories across six tasks | Supervised fine-tuning with general data mixed in | Llama 2 7B to 70B | Agent abilities generalize to held-out tasks | AgentTuning[5](#user-content-fn-agenttuning) | | 491 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 7.0 to 20.6 percent on SWE-bench Verified; 32.0 with a trained verifier and 16 samples | SWE-Gym[6](#user-content-fn-swe-gym) | | 5,016 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 40.2 percent on SWE-bench Verified | SWE-smith[7](#user-content-fn-swe-smith) | | 8,209 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 6.4 to 38.0 percent; 47.0 with best-of-8 and a critic; log-linear in data with no plateau | Skywork-SWE[8](#user-content-fn-skywork-swe) | | About 4,500 executable tasks | Reinforcement learning only | Qwen3-32B | 23 to 42.2 percent; 59.0 with test-time scaling | DeepSWE[9](#user-content-fn-deepswe) | | Rejection sampling, then RL on executable tasks | Both | Qwen2.5-72B-Instruct | 11.4 to 20.5 to 39.0 percent | Nebius[10](#user-content-fn-nebius-rl) | Three patterns run through the table. **Hundreds move a model.** Every result in the first half of the table used about a thousand examples or fewer and produced a large change. LIMA's authors argued that almost all of a model's knowledge comes from pretraining and that alignment needs only a small set of examples to teach format and style.[1](#user-content-fn-lima) The agent results say something stronger for this domain: 491 verified trajectories nearly tripled a 32-billion-parameter model's resolve rate.[6](#user-content-fn-swe-gym) The first useful model is closer than most teams expect. **Thousands keep paying.** Skywork-SWE's scaling curve is the most direct measurement: 2,000 trajectories gave 31.8 percent, 6,000 gave 36.1, and 8,209 gave 38.0, with the curve still rising.[8](#user-content-fn-skywork-swe) Zhang and colleagues found the same shape across fine-tuning tasks and model sizes and fit it as a power law in the amount of fine-tuning data.[11](#user-content-fn-zhang-scaling) The returns diminish per example and do not stop. **Quality beats quantity at every scale.** AlpaGasus trained on 9,000 examples filtered from a 52,000-example set and beat the model trained on all 52,000.[12](#user-content-fn-alpagasus) LIMA, s1, and LIMO are all arguments that a small curated set beats a large uncurated one. For this pipeline, the oracle is the curation. A flip is, by construction, an example that was verified. The volume question is how many verified examples, not how many sessions. > **Verified trajectories behind each open-model result on SWE-bench Verified** > > | Result | Trajectories | > | --- | --- | > | SWE-Gym, 20.6% after fine-tuning | 491 | > | Skywork-SWE at 31.8% | 2000 | > | SWE-smith, 40.2% | 5016 | > | Skywork-SWE at 36.1% | 6000 | > | Skywork-SWE, 38.0% | 8209 | > > *Sources: Pan et al. (2024), Yang et al. (2025), Zeng et al. (2025). The same 32-billion-parameter base model in every row. The Skywork rows are points on one scaling curve; the other two are separate recipes.* ## The worksheet Four numbers set the time to a given training set size. 1. **Sessions per day.** Engineers times sessions each. A team of 50 running 10 each is 500. 2. **Oracle share.** The fraction of sessions dispatched on tasks that have a hidden test. A team starting out might reach 20 percent. A team that writes the test first for every bug and most features can reach 60. 3. **Flip rate.** The fraction of oracle sessions that end on PASS. With a silent oracle and a rented frontier model on tasks of reasonable size, 30 to 50 percent is a working assumption; a team should measure it in its first week. 4. **Attempts per task.** Independent sessions dispatched per oracle task. Each attempt costs tokens and raises the chance that at least one flips. Flips per day is roughly: sessions × oracle share × flip rate, adjusted upward for attempts. With the numbers above and one attempt, 500 × 0.2 × 0.33 is about 33 flips a day. With three attempts and a per-attempt flip rate of 33 percent, the chance a task flips at least once is about 70 percent, so flips per task-day rises to about 70 of 100 oracle tasks, at three times the token cost. > **Working days to reach a training set size, one attempt per task** > > | | 20 percent oracle share | 60 percent oracle share | > | --- | --- | --- | > | 500 flips, first fine-tune | 15days | 5days | > | 2,000 flips | 60days | 20days | > | 5,000 flips | 150days | 50days | > | 8,000 flips | 240days | 80days | > > *Arithmetic for 500 sessions a day at a 33 percent flip rate. Sampling three attempts per task roughly halves every number at three times the token cost. These are planning figures, not measurements.* The reading of the chart is that a mid-sized team reaches the SWE-Gym regime in weeks and the Skywork regime within a year, and that the oracle share is the lever that matters most. Every bug fixed without a hidden test is a session that could have been a flip and was not. ## Limits of volume More flips do not fix a bad oracle. A thousand traces graded by a flaky test are a thousand noisy labels, and the noise does not average out; it teaches. More flips do not fix a narrow task distribution. Five thousand flips on one service teach that service. And more flips do not fix a leaked hidden test; they make the leak worse, because every trace fitted to the leaked test is one more example of fitting. The order of operations is: get the oracle right, get the task supply broad, then grow the volume. Chapter 13 is the pipeline that does the third once the first two are in place. ## Footnotes 1. Zhou, C., Liu, P., Xu, P., et al. (2023). *LIMA: Less Is More for Alignment*. NeurIPS 2023. [https://arxiv.org/abs/2305.11206](https://arxiv.org/abs/2305.11206) [↩](#user-content-fnref-lima) [↩2](#user-content-fnref-lima-2) 2. Muennighoff, N., Yang, Z., Shi, W., et al. (2025). *s1: Simple test-time scaling*. arXiv. [https://arxiv.org/abs/2501.19393](https://arxiv.org/abs/2501.19393) [↩](#user-content-fnref-s1) 3. Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). *LIMO: Less is More for Reasoning*. arXiv (v1, February 2025); COLM 2025. [https://arxiv.org/abs/2502.03387](https://arxiv.org/abs/2502.03387) [↩](#user-content-fnref-limo) 4. Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). *FireAct: Toward Language Agent Fine-tuning*. arXiv. [https://arxiv.org/abs/2310.05915](https://arxiv.org/abs/2310.05915) [↩](#user-content-fnref-fireact) 5. Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). *AgentTuning: Enabling Generalized Agent Abilities for LLMs*. arXiv. [https://arxiv.org/abs/2310.12823](https://arxiv.org/abs/2310.12823) [↩](#user-content-fnref-agenttuning) 6. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) [↩2](#user-content-fnref-swe-gym-2) 7. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) [↩](#user-content-fnref-swe-smith) 8. Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290) [↩](#user-content-fnref-skywork-swe) [↩2](#user-content-fnref-skywork-swe-2) 9. Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe) [↩](#user-content-fnref-deepswe) 10. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501) [↩](#user-content-fnref-nebius-rl) 11. Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). *When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method*. ICLR 2024. [https://arxiv.org/abs/2402.17193](https://arxiv.org/abs/2402.17193) [↩](#user-content-fnref-zhang-scaling) 12. Chen, L., Li, S., Yan, J., et al. (2023). *AlpaGasus: Training a Better Alpaca with Fewer Data*. arXiv; ICLR 2024. [https://arxiv.org/abs/2307.08701](https://arxiv.org/abs/2307.08701) [↩](#user-content-fnref-alpagasus) --- # Chapter 13: The delivery pipeline > Weights as a release artifact: ingest, validate, train, evaluate, gate, canary in shadow, promote, roll back. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-delivery-pipeline Continuous delivery is the practice of keeping software in a state where any change can be released at any time, by automating the path from a commit to production and gating each step on checks.[1](#user-content-fn-humble-farley) This chapter applies it to weights. The artifact is a model. The commit is a new batch of verified traces. The release is a new model serving the team's agents. Everything between them is a stage with an input, an output, and a gate. ## The stages > **From trace store to serving model** > > 1. **Ingest**: Sync traces and signed verdicts from agent hosts and oracle hosts > 2. **Redact and scan**: Second-pass secret scan; quarantine hits > 3. **Validate**: Schema check; verdict signatures; baseline sanity > 4. **Select and version**: Filters from Chapter 10; write the manifest > 5. **Train**: Adapter on all layers; replay mix; fixed seed > 6. **Evaluate**: Held-out flips, frozen and rolling; public benchmark subset > 7. **Gate**: Beat the serving model on held-out flips; no regression on the public set > 8. **Package**: Merge or ship the adapter; quantize; record the manifest hash > 9. **Canary**: Shadow mode on live tasks; the oracle scores both models > 10. **Promote or roll back**: Route traffic; keep the previous weights warm > > Repeat every batch. > > *The emphasized stages are the ones that decide. Everything else is plumbing, and all of it is ordinary continuous delivery with a model as the artifact.* **Ingest.** A scheduled job pulls finished trace files from each agent host's plugin data directory and verdict records from the oracle host into the trace store. Files are content-addressed: the path includes the hash of the contents, so a re-upload is a no-op and a tampered file lands at a different path. The job is safe to run twice. **Redact and scan.** The hook redacted with patterns at write time. This stage runs a dedicated secret scanner over every trace and quarantines any with a hit. Quarantined traces are not deleted; a person reviews them, because a false positive on a test fixture is common and a true positive is a credential to rotate. **Validate.** Every event parses against the schema. Every verdict record's signature verifies against the oracle's public key. Every session has a `session_start`, a `session_end`, and at least one verdict, or it is marked incomplete and excluded. A session whose baseline was PASS is excluded and its task is flagged. **Select and version.** The filters from Chapter 10 run. The output is a manifest: the list of sessions in the supervised set, the pairs in the preference set, the held-out task list, the template version, and the hash of each. The manifest is the thing that is versioned. The training data is derived from it. **Train.** The training job takes a manifest and a recipe and produces weights. The recipe is a file: base model and its hash, adapter rank and target modules, learning rate, epochs, replay mix and its source, seed. The job records the recipe hash and the manifest hash in the weights' metadata. A training run with the same manifest, recipe, and seed should produce the same weights, within the limits of the hardware's determinism, and the pipeline should check that it does once. **Evaluate.** The new weights run against the held-out tasks. For each task, the model attempts it in a fresh sandbox with the same harness the team uses, and the oracle grades the result. The metric is the flip rate on the frozen set and on the rolling set. The weights also run against a fixed subset of a public coding benchmark, for the forgetting check. Chapter 14 is about reading these. **Gate.** The new weights are promoted only if the frozen-set flip rate is at least the serving model's, the rolling-set flip rate is higher, and the public-set score has not fallen by more than a set tolerance. A run that fails the gate is kept, with its evaluation, so the trend is visible even when nothing ships. **Package.** The adapter is merged into the base weights or shipped as a separate adapter, depending on how the serving layer loads models. The weights are quantized if the serving hardware needs it, and the quantized weights are re-evaluated on the frozen set, because quantization can cost more on a fine-tuned model than on its base. The package carries the manifest hash and the recipe hash. **Canary.** Before the new model serves anyone, it runs in shadow. For a sample of live oracle tasks, both the serving model and the candidate attempt the task in parallel sandboxes. The oracle grades both. The person sees only the serving model's result. After enough tasks, the candidate's live flip rate against the serving model's is the number that decides. This is the same design as shadow mode in driving systems, where a new model runs alongside the one in control and its decisions are compared without being acted on. **Promote or roll back.** Promotion is a routing change. The previous weights stay loaded for a period, and rollback is the same routing change reversed. A rollback is not a failure of the pipeline; it is the pipeline working. ## Cadence How often to run depends on the flip rate. A team producing 30 flips a day has 200 new examples a week, which is enough to retrain weekly and expect to see movement. A team producing 5 a day should batch monthly. A training run on an adapter for a 32-billion-parameter model over a few thousand trajectories takes hours on a single node with eight large accelerators; the evaluation, which runs an agent on each held-out task, often takes longer than the training. Budget for both. The cadence should be a schedule, not a trigger. A pipeline that trains whenever enough new traces arrive produces models at irregular intervals that are hard to compare. A pipeline that trains every Sunday night and evaluates on Monday produces a weekly series. ## What is deterministic and what is not The oracle is deterministic by construction. Training is deterministic up to the hardware. Evaluation is not: the model samples, and the harness is a live process. Fix the sampling temperature and seed for evaluation, run each held-out task more than once, and report the mean. Two models within noise of each other are tied, and the gate should say so instead of promoting on a coin flip. ## The record Every stage writes a record to the same store the traces live in: what it took in, what it produced, the hashes, the time, and the outcome. A model in production can be traced back to the manifest that trained it, to the sessions in the manifest, to the oracle that graded each session, and to the hidden test behind each verdict. That chain is what makes a regression debuggable and what makes the pipeline auditable to anyone who asks where the model came from. This is where the author's own work connects, and the connection is disclosed. Oxagen, the company the author founded, records agent runs with their tool calls, their outcomes, and their costs as a product. A run record is a trace in this book's sense. The pipeline in this chapter does not depend on it; the plugin writes its own traces to its own files. But the reason the record is kept beside the verdict, rather than inside the model's context, is the same reason Oxagen keeps the outcome outside the agent that produced it. ## Footnotes 1. Humble, J., & Farley, D. (2010). *Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation*. Addison-Wesley. [↩](#user-content-fnref-humble-farley) --- # Chapter 14: Evaluation > Held-out flips, contamination, weak tests, hacking scans, and the drift measures that catch what the flip rate misses. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/evaluation A pipeline that trains weights needs a way to know whether the new weights are better. This chapter is about building that measurement so that it stays honest as the pipeline optimizes against it. ## Held-out flips are the metric The primary metric is the flip rate on tasks the model never trained on, graded by the same oracle that grades production sessions. It is the metric that matches the goal. A public benchmark measures how well the model solves public tasks. Held-out flips measure how well it solves yours. Two sets, as Chapter 10 said. The frozen set is drawn once, at the start, from tasks across the team's repositories, and never changes. Every model version is scored against it, so the series is comparable. The rolling set is the most recent few hundred oracle tasks, refreshed each cycle, so the score tracks the current work. A model that gains on the frozen set and not the rolling set has learned the past. A model that gains on the rolling set and not the frozen set has learned something narrow about recent work. The gate wants both. ## Contamination A held-out task that the model saw during training is not held out. The leak is easy to create and impossible to detect after the fact. Deng and colleagues showed that language models can reproduce the missing parts of benchmark items they were trained on, which is how contamination shows up in public evaluations.[1](#user-content-fn-deng-contamination) The SWE-bench+ study found that about a third of one agent's passing patches on SWE-bench had the solution available in the issue text or its comments, and another third passed because the tests were too weak to catch a wrong fix; filtering those out dropped the measured resolve rate from 12.47 percent to 3.97 percent.[2](#user-content-fn-swe-bench-plus) A team's internal evaluation is exposed to the same two failures: the hidden test can leak, and the hidden test can be weak. Four rules keep the held-out set clean. 1. Split by task, with every session on a task on the same side. 2. Never train on a trace from a held-out task, including traces from the rented frontier model. The provenance field makes this checkable. 3. Never show the held-out hidden tests to any agent, including the one generating training traces. The one-bit oracle enforces this for the agent; the pipeline has to enforce it for the people. 4. Retire a held-out task when its hidden test is released into the repository, which Chapter 6 recommended once a task has been flipped in production. A test the agents can see is no longer held out. ## Weak tests A hidden test that passes for a wrong fix inflates the flip rate and teaches the wrong fix. Mutation testing measures this directly: generate mutants of the code the task touches and check that the hidden test kills them.[3](#user-content-fn-just-mutants) A held-out task whose hidden test kills few mutants near the change is a weak oracle, and its flips mean less. The evaluation report should carry the mutation score of each task's hidden test beside the flip rate, so a gain concentrated on weak tasks is visible. ## The public benchmark A fixed subset of a public coding benchmark sits in the gate for one purpose: catching forgetting. The team's model should not get worse at general coding while it gets better at the team's code. SWE-bench Verified, or a multilingual benchmark if the team's code is not Python, serves.[4](#user-content-fn-swe-bench-verified)[5](#user-content-fn-multi-swe-bench) The score itself is secondary. The change in the score between versions is the signal. Do not optimize for it. A pipeline that gates on a public benchmark and also trains on data derived from public repositories will, over time, find the benchmark's tasks in its training data. The public set is a thermometer, not a target. ## Watching for hacking The oracle is air-gapped, so the agent cannot change the verdict. The agent can still produce a trace that flipped for a reason the team would not endorse, and a model trained on such traces learns the reason. Baker and colleagues found that a monitor reading the agent's chain of thought caught most reward hacking, and that a weaker model was an adequate monitor.[6](#user-content-fn-baker-monitoring) In this pipeline, the equivalent is a scan of each flipped trace for patterns that should not be there: edits to paths on the denylist, even though the oracle dropped them; commands that search for the hidden tests; tool calls that read the oracle's configuration; test runs that were made to pass by changing the test. The scan flags, a person reads, and flagged traces are excluded from training until cleared. The scan should run on traces from the rented model too, because distillation copies behavior. Track the dropped-hunk rate from the oracle log. A rising share of diffs with hunks on the denylist means the agents are learning to touch the tests, and the next model will learn it faster. ## Drift in what the model produces Three cheap measurements catch most regressions that the flip rate misses. - **Length.** Mean tokens per session and per tool call. Preference optimization and reward-driven training both tend toward verbosity unless controlled.[7](#user-content-fn-length-bias) A model whose sessions get longer without flipping more often is spending the team's tokens. - **Tool mix.** The distribution of tool calls per session. A model that stops running tests, or starts reading every file in the repository, has changed in a way the flip rate will show late. - **Similarity.** Pairwise similarity between the model's trajectories on different tasks. Rising similarity is the early sign of collapse from Chapter 11. ## The report Each evaluation produces one report with: flip rate on the frozen set and the rolling set, each with its uncertainty from repeated runs; the public-benchmark score and its change; the mutation score distribution of the held-out tests; the counts of flagged traces by reason; the three drift measurements; and the manifest and recipe hashes. The gate reads the report. So do people. A report that a person cannot read in five minutes is too long. ## Footnotes 1. Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). *Investigating Data Contamination in Modern Benchmarks for Large Language Models*. NAACL 2024. [https://arxiv.org/abs/2311.09783](https://arxiv.org/abs/2311.09783) [↩](#user-content-fnref-deng-contamination) 2. Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). *SWE-Bench+: Enhanced Coding Benchmark for LLMs*. arXiv. [https://arxiv.org/abs/2410.06992](https://arxiv.org/abs/2410.06992) [↩](#user-content-fnref-swe-bench-plus) 3. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). *Are Mutants a Valid Substitute for Real Faults in Software Testing?* FSE 2014, 654–665. [https://doi.org/10.1145/2635868.2635929](https://doi.org/10.1145/2635868.2635929) [↩](#user-content-fnref-just-mutants) 4. OpenAI (2024). *Introducing SWE-bench Verified*. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/) [↩](#user-content-fnref-swe-bench-verified) 5. Zan, D., Huang, Z., Liu, W., et al. (2025). *Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving*. arXiv. [https://arxiv.org/abs/2504.02605](https://arxiv.org/abs/2504.02605) [↩](#user-content-fnref-multi-swe-bench) 6. Baker, B., Huizinga, J., Gao, L., et al. (2025). *Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation*. arXiv. [https://arxiv.org/abs/2503.11926](https://arxiv.org/abs/2503.11926) [↩](#user-content-fnref-baker-monitoring) 7. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). *A Long Way to Go: Investigating Length Correlations in RLHF*. arXiv; COLM 2024. [https://arxiv.org/abs/2310.03716](https://arxiv.org/abs/2310.03716) [↩](#user-content-fnref-length-bias) --- # Chapter 15: Serving and the economics > Routing between your model and the rented one, and when owning the weights becomes cheaper than the rent. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/serving-and-the-economics The model is trained. This chapter is about running it: where it serves, how requests are routed between it and the rented model, and the arithmetic that says when the pipeline has paid for itself. ## Serving Open-weight models serve through a small number of mature engines. vLLM introduced paged attention, which manages the key-value cache in blocks the way an operating system manages memory, and made high-throughput serving of large models practical on commodity accelerators.[1](#user-content-fn-vllm) Serving an adapter without merging it is supported directly; S-LoRA showed that thousands of adapters over one base model can be served from a single node with the base weights shared.[2](#user-content-fn-s-lora) For a team with one fine-tune per repository or per domain, that means one base model in memory and a set of small adapters, with the request choosing the adapter. Quantization reduces memory and raises throughput at some cost in quality. A 32-billion-parameter model quantized to 4 bits fits on a single large accelerator. Re-evaluate on the frozen set after quantizing, as Chapter 13 said, because the cost is not uniform across models. ## Routing The team's model does not have to handle everything on day one. A router sends each task to the team's model first and falls back to the rented frontier model when the team's model fails. The oracle makes the fallback decision cheap: a session that does not flip after the team's model's attempts is re-dispatched to the rented model. The fallback sessions are traces. They are the hard tail of the distribution, solved by a stronger model, and graded by the oracle. They go into the trace store with their provenance and become the distillation examples from Chapter 9. Over time, the share of tasks that reach the fallback is the measure of how far the team's model has come, and it is also the training set for closing the gap. > **Routing a task between the team's model and the rented model** > > 1. **Task**: An oracle task is dispatched > 2. **Team model**: Attempts it; the gate asks the oracle > 3. **Flip**: PASS: done. The trace is a self-generated example. > 4. **Fallback**: No flip after N attempts: re-dispatch to the rented model > 5. **Rented model**: Attempts it; the same gate, the same oracle > 6. **Trace**: Either way, the trace goes to the store with its provenance > > *The fallback rate falls as the team's model learns. The fallback traces are what it learns from.* ## The arithmetic The cost side has four lines. **Oracle compute.** Each oracle run is a container running a test suite. On a cloud runner, a suite that takes five minutes costs cents. At 100 oracle sessions a day with two attempts each and a baseline per task, that is a few hundred container-minutes a day. **Storage.** A trace with redacted tool output is tens of kilobytes to a few megabytes. A year of a mid-sized team's traces is tens to hundreds of gigabytes. This is a rounding error. **Training.** An adapter run on a 32-billion-parameter model over a few thousand trajectories is hours on one eight-accelerator node. At on-demand cloud prices, that is low hundreds to low thousands of dollars per run. Weekly, it is tens of thousands a year. Evaluation on a few hundred held-out tasks, each an agent session in a sandbox, costs about as much again. **Serving.** One eight-accelerator node, or two for redundancy, serves a 32-billion-parameter model to a team of fifty with capacity to spare. Reserved, this is low six figures a year; owned, it is a capital cost amortized over several years. The revenue side is the rent avoided. A team of fifty engineers running agents daily on a frontier model spends, at mid-2026 prices, in the range of several hundred thousand to a few million dollars a year, depending on how heavily they use it. The exact figure is the team's own bill, and the team should use it. The comparison is not all-or-nothing. In the routing design above, the team's model handles the share of tasks it can flip, and the rented model handles the rest. If the team's model flips 40 percent of oracle tasks in its first quarter, 40 percent of the rent on those tasks is replaced by serving cost. As the share rises, the rent falls. The crossover, where the pipeline's total cost falls below the rent it replaces, depends on the team's bill, but for a team spending more than the cost of two accelerator nodes a year on tokens, it arrives within the first year of flips. ## What the price decline does to this Token prices fall fast, and the rent will be lower next year.[3](#user-content-fn-epoch-prices) Three things keep the arithmetic in the pipeline's favor anyway. Open-model serving costs fall on the same curve, because the hardware and the serving software improve for everyone. The ratio between rent and serving cost is more stable than either number. The rented model's price per token is for a model that does not know the team's code. The team's model's cost per token is for one that does. If the team's model flips a task in fewer tokens, which a model trained on that codebase should, the comparison per task is better than the comparison per token. And the rent buys no asset. The pipeline's cost buys weights, a trace store, and an oracle, all of which the team keeps. The question in Chapter 1 was what the team owns at the end of the year. The arithmetic here is about when owning it is also cheaper, and the answer for most teams of this size is: soon. ## Footnotes 1. Kwon, W., Li, Z., Zhuang, S., et al. (2023). *Efficient Memory Management for Large Language Model Serving with PagedAttention*. SOSP 2023. [https://arxiv.org/abs/2309.06180](https://arxiv.org/abs/2309.06180) [↩](#user-content-fnref-vllm) 2. Sheng, Y., Cao, S., Li, D., et al. (2023). *S-LoRA: Serving Thousands of Concurrent LoRA Adapters*. arXiv; MLSys 2024. [https://arxiv.org/abs/2311.03285](https://arxiv.org/abs/2311.03285) [↩](#user-content-fnref-s-lora) 3. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). *LLM inference prices have fallen rapidly but unequally across tasks*. Epoch AI. [https://epoch.ai/data-insights/llm-inference-price-trends](https://epoch.ai/data-insights/llm-inference-price-trends) [↩](#user-content-fnref-epoch-prices) --- # Chapter 16: Governance and failure modes > Secrets, memorization, oracle tampering, pipeline debt, a starving flip signal, and people. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/governance-and-failure-modes A pipeline that collects everything an agent did and trains a model on it has new ways to go wrong. This chapter lists them, with what the research says and what the pipeline does about each. The failures are ordered by how much damage they do, not by how likely they are. ## Secrets in the trace store The trace store holds the output of every command the agents ran. If a secret passed through a terminal, it is in a trace unless redaction caught it. A trace store is the most complete record of a team's operational secrets that has ever existed in one place, and it should be protected like one. The defenses are layered. Redaction at write time, with patterns. A dedicated scanner at ingest, with quarantine. Access control on the store, with the training job as the only automated reader. Encryption at rest. And a short retention period for raw tool output, after which only the selected and scanned training examples remain. The last one is a tradeoff: it forecloses re-processing old traces with better redaction. A team should decide its retention consciously. ## Memorization A model trained on traces can reproduce them. Carlini and colleagues extracted training examples from a deployed language model by prompting it, including names, contact details, and code.[1](#user-content-fn-carlini-extraction) In a follow-up, they measured how memorization scales: it grows with model size, with how many times an example was duplicated in training, and with how much of the example's prefix the prompt supplies.[2](#user-content-fn-carlini-memorization) A fine-tune on a few thousand trajectories, each seen for several epochs, is in the regime where memorization is expected. Two consequences. First, a secret that survived redaction into training can be extracted from the model. The secret scanning in the pipeline is the control, and it has to be good. Second, the model is a copy of the team's code in a form that can be queried. It is a private asset and should be served privately. A team that fine-tunes on its traces and then exposes the model to people outside the team has published its code in a lossy format. Chapter 1 said the traces were the asset nobody else has. That is only true while the model trained on them stays inside. ## Tampering with the oracle Chapter 5 built the air gap. The governance question is who can change what is inside it. The hidden tests, the allowlist, the container image, and the signing key are the pipeline's root of trust. Changes to them should go through the same review as changes to production code, and the oracle log should record every verdict with the hash of the test list that produced it, so a change in the tests is visible as a change in the hash. Rotate the hidden tests for a task once it has flipped in production, as Chapter 6 said, and retire the task from the held-out set. ## Training on the wrong lesson The oracle grades the result. It does not grade the path. A session that flipped by reading the hidden test's name from a stray log line, by copying a fix from a sibling repository that happened to be checked out, or by special-casing the inputs it guessed the test used, is a flip with a bad path. The scan in Chapter 14 catches some of these. The rest are caught by reading flipped traces, which a person should do for a sample every cycle. A pipeline nobody reads is a pipeline that trains on whatever got through. ## The pipeline itself Sculley and colleagues catalogued the ways machine-learning systems accumulate debt that ordinary software does not: data dependencies that nobody tracks, feedback loops where the model's output changes its own training data, configuration that grows without review, and pipelines glued together from pieces nobody owns.[3](#user-content-fn-sculley-debt) This pipeline has every one of those. Its training data comes from agents running its own previous model, which is a feedback loop by design. Its configuration is a recipe file, an allowlist, a template, and a set of hidden tests, each of which changes the model when it changes. Its stages are a trace collector, an oracle, a trainer, an evaluator, and a router, built from different tools. The defenses are the ordinary ones, applied without exception. Every input to a stage is versioned. Every stage records what it did. Every change to configuration is reviewed. Every model can be traced to its manifest. And the frozen evaluation set, which never changes, is the one fixed point against which drift in everything else is measured. ## Loss of the flip signal If the oracle share falls, because the team stops writing tests first, the flip rate falls with it and the pipeline starves. If the task supply narrows, because one repository produces most of the oracle tasks, the model narrows with it. If the fallback rate stops falling, the team's model has plateaued and the recipe needs to change. Each of these is visible in the pipeline's own records, and the weekly report should carry them: oracle share, flip rate, tasks per repository, fallback rate. ## People The last failure mode is the one the research does not cover. A pipeline that grades every agent session on a hidden test is also, if someone chooses to read it that way, a pipeline that grades the engineers who dispatched the sessions. It should not be used that way. The flip rate is a property of the task, the oracle, and the model. The moment it becomes a measure of a person, people will stop dispatching tasks that might not flip, the oracle share will fall, and the pipeline will starve. Say this in writing when the pipeline is introduced, and keep the per-person numbers out of the report. ## Footnotes 1. Carlini, N., Tramèr, F., Wallace, E., et al. (2021). *Extracting Training Data from Large Language Models*. USENIX Security 2021. [https://arxiv.org/abs/2012.07805](https://arxiv.org/abs/2012.07805) [↩](#user-content-fnref-carlini-extraction) 2. Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). *Quantifying Memorization Across Neural Language Models*. ICLR 2023. [https://arxiv.org/abs/2202.07646](https://arxiv.org/abs/2202.07646) [↩](#user-content-fnref-carlini-memorization) 3. Sculley, D., Holt, G., Golovin, D., et al. (2015). *Hidden Technical Debt in Machine Learning Systems*. NeurIPS 2015. [↩](#user-content-fnref-sculley-debt) --- # Chapter 17: Precedents > Expert iteration to SWE-Gym and Getafix to DIDACT: what each established and what it left to do. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/precedents The idea in this book is not new. It is the combination of several ideas that are each at least a decade old, applied to a kind of data that did not exist until coding agents did. This chapter lists the precedents, what each one established, and what it left for this pipeline to add. ## Verifier-filtered self-training Expert iteration, from Anthony, Tian, and Barber in 2017, is the general form: a slow, strong solver produces solutions, a fast policy is trained to imitate them, and the loop repeats with the improved policy as the new starting point.[1](#user-content-fn-expert-iteration) AlphaGo Zero, the same year, ran the loop with the outcome of the game as the only verifier and reached superhuman play from random weights.[2](#user-content-fn-alphago-zero) The verifier was perfect, cheap, and deterministic, which is the ideal this book's oracle approximates. STaR brought the loop to language models in 2022, with an answer key as the verifier.[3](#user-content-fn-star) ReST and ReST-EM scaled it.[4](#user-content-fn-rest)[5](#user-content-fn-rest-em) The reinforcement-learning-with-verifiable-rewards line, from Tülu 3 through DeepSeek-R1 and DAPO, is the same loop with a gradient step instead of a fine-tune on the filtered set.[6](#user-content-fn-tulu3)[7](#user-content-fn-deepseek-r1)[8](#user-content-fn-dapo) Cobbe and colleagues' 2021 GSM8K work made the case that a trained verifier scales better than fine-tuning alone, which is the argument for the verifier in Chapter 9.[9](#user-content-fn-cobbe-verifiers) What these established: a verifier the model does not control produces lasting gains, and the model's own judgment does not. What they left: the verifier was an answer key or a game. Software has no answer key. It has tests. ## Tests as the verifier for code AlphaCode, in 2022, sampled up to a million programs per competition problem and filtered them through the problem's example tests before clustering and submitting.[10](#user-content-fn-alphacode) CodeRL, the same year, used unit-test results as the reward for training a code model with an actor-critic method.[11](#user-content-fn-coderl) Meta's RLEF, in 2024, trained a code model with execution feedback as the reward and showed large gains in sample efficiency.[12](#user-content-fn-rlef) SWE-RL, DeepSWE, and the Nebius work brought it to repository-scale tasks.[13](#user-content-fn-swe-rl)[14](#user-content-fn-deepswe)[15](#user-content-fn-nebius-rl) What these established: a test is a usable reward, and a model trained against tests gets better at passing tests. What they left: every one of them used public problems with public tests. The tests were visible, or at least public, and the problems were nobody's in particular. ## Benchmarks built from real repositories SWE-bench, in 2023, defined the fail-to-pass and pass-to-pass construction and built 2,294 tasks from real GitHub issues and the pull requests that closed them.[16](#user-content-fn-swe-bench) SWE-bench Verified had human annotators remove the tasks whose descriptions were unclear or whose tests were unfair, leaving 500.[17](#user-content-fn-swe-bench-verified) SWE-Gym, SWE-smith, R2E-Gym, and SWE-rebench turned the construction into pipelines that produce thousands of executable tasks.[18](#user-content-fn-swe-gym)[19](#user-content-fn-swe-smith)[20](#user-content-fn-r2e-gym)[21](#user-content-fn-swe-rebench) Multi-SWE-bench extended it to seven languages, and SWE-Lancer graded real freelance tasks with end-to-end tests.[22](#user-content-fn-multi-swe-bench)[23](#user-content-fn-swe-lancer) SWE-bench+ showed how often the construction leaks the answer or accepts a wrong one.[24](#user-content-fn-swe-bench-plus) What these established: the flip is a reliable unit of value for software work, and tasks with hidden tests can be produced at scale from version history. What they left: the tasks are public, so every model has seen them, and the tests are released, so every agent can be shown them. A team's own history is the unlimited supply of tasks that are not public. ## Learning from an organization's own development process This is the closest precedent and the least cited. Facebook's Getafix, in 2019, learned fix patterns from the history of human fixes to static-analysis warnings in Facebook's own codebase and proposed fixes for new warnings, which engineers accepted at a high rate.[25](#user-content-fn-getafix) SapFix, the same year, generated candidate patches for crashes found by an automated testing system and used that system's tests as the oracle, end to end, in production.[26](#user-content-fn-sapfix) Both learned from, and were graded by, the company's own code and tests. Google's DIDACT, described in 2023, trained models on the process of software development inside Google rather than on finished code: the edit histories, the build-error fixes, the code-review comments and their resolutions.[27](#user-content-fn-didact) Google had earlier reported that a completion model trained on its internal code reduced coding iteration time by 6 percent and was accepted for about 3 percent of new code.[28](#user-content-fn-google-completion) The data was the developers' own activity, and the organization kept it. GitHub's Copilot research found that acceptance rate was the best available predictor of developers' perceived productivity, which made acceptance a usable, if weak, oracle at scale.[29](#user-content-fn-ziegler-copilot) Replit, in 2024, trained a 7-billion-parameter code-repair model on data built from its own platform: language-server diagnostics, with the state of the file reconstructed by replaying the edit history.[30](#user-content-fn-replit-repair) What these established: an organization's own development activity is training data, and models trained on it perform well on that organization's work. What they left: each was built by a company with a research team and a bespoke agent or tool. The point of Chapters 7 and 8 is that the harness hooks make this available to a team with neither. ## Learning from traces in other fields End-to-end driving models from 2016 onward trained on recorded human driving, with the steering angle as the label.[31](#user-content-fn-bojarski-driving) The traces were cheap to collect from vehicles already on the road, and the fleet's data became the moat. The shadow-mode canary in Chapter 13 is borrowed from this field. Imitation learning has a known failure: a policy trained on an expert's trajectories drifts into states the expert never visited, and its errors compound. DAgger, from 2011, fixes it by running the learner, having the expert label the states the learner reached, and training on those.[32](#user-content-fn-dagger) The analogue here is the fallback in Chapter 15: when the team's model fails a task, the rented model solves it from the same starting point, and the trace is a label on a state the team's model reached. Hindsight experience replay, from 2017, relabels failed episodes as successes for the goals they did reach, which Chapter 9 proposed for sessions that did not flip.[33](#user-content-fn-her) Research on machine learning for code, surveyed by Allamanis and colleagues in 2018, rests on the observation from Hindle and colleagues in 2012 that software is natural: repetitive and predictable enough that statistical models of it work.[34](#user-content-fn-hindle-naturalness)[35](#user-content-fn-allamanis-survey) A team's codebase is more repetitive and more predictable than the public corpus, which is why a model trained on it does well there. ## Keeping a holdout honest The one-bit verdict in Chapter 6 rests on the Ladder and the reusable holdout, both from 2015.[36](#user-content-fn-ladder)[37](#user-content-fn-reusable-holdout) Both were written about machine-learning competitions and scientific data analysis. Neither mentions an agent. The problem they solved, an adaptive optimizer hill-climbing on a holdout it is only supposed to be measured by, is the problem an agent iterating against an oracle has, and the solution transfers without modification. ## Agents trained on agent trajectories AgentTuning and FireAct, in 2023, fine-tuned open models on a few hundred to a couple of thousand agent trajectories generated by a stronger model and showed that the result generalized.[38](#user-content-fn-agenttuning)[39](#user-content-fn-fireact) Kimi K2's technical report described a large-scale pipeline for synthesizing agentic data with tool use and verifying it before training.[40](#user-content-fn-kimi-k2) The open-weight agent frameworks, SWE-agent and OpenHands, standardized the tool interface that these trajectories are recorded in.[41](#user-content-fn-swe-agent)[42](#user-content-fn-openhands) What these established: agent behavior transfers through trajectories, and a few hundred are enough to see it. What they left: the trajectories came from public tasks solved by public models. The trajectories this book is about come from your tasks, solved on your code, graded by your tests. ## What is new Four things, and only four. Hooks in the harness make trace collection free. The organization does not build the agent, and the agent does not have to be modified. The hidden test as an air-gapped, one-bit oracle makes the grading trustworthy under optimization pressure, which the public-benchmark work did not have to worry about and the internal-tool work handled with bespoke infrastructure. Open-weight models are close enough to the frontier, and licensed permissively enough, that the fine-tuned result is competitive on the team's distribution. In 2019 the models were not there. In 2023 the licenses often were not. And the delivery pipeline treats weights as a release artifact with gates, canaries, and rollback, which is ordinary engineering applied to a thing that used to be a research project. Everything else in this book was established by someone else, and the footnotes say who. ## Footnotes 1. Anthony, T., Tian, Z., & Barber, D. (2017). *Thinking Fast and Slow with Deep Learning and Tree Search*. NeurIPS 2017. [https://arxiv.org/abs/1705.08439](https://arxiv.org/abs/1705.08439) [↩](#user-content-fnref-expert-iteration) 2. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). *Mastering the game of Go without human knowledge*. Nature, 550, 354–359. [https://doi.org/10.1038/nature24270](https://doi.org/10.1038/nature24270) [↩](#user-content-fnref-alphago-zero) 3. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). *STaR: Bootstrapping Reasoning With Reasoning*. arXiv. [https://arxiv.org/abs/2203.14465](https://arxiv.org/abs/2203.14465) [↩](#user-content-fnref-star) 4. Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). *Reinforced Self-Training (ReST) for Language Modeling*. arXiv. [https://arxiv.org/abs/2308.08998](https://arxiv.org/abs/2308.08998) [↩](#user-content-fnref-rest) 5. Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). *Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models*. arXiv; TMLR 2024. [https://arxiv.org/abs/2312.06585](https://arxiv.org/abs/2312.06585) [↩](#user-content-fnref-rest-em) 6. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). *Tülu 3: Pushing Frontiers in Open Language Model Post-Training*. arXiv. [https://arxiv.org/abs/2411.15124](https://arxiv.org/abs/2411.15124) [↩](#user-content-fnref-tulu3) 7. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) [↩](#user-content-fnref-deepseek-r1) 8. Yu, Q., Zhang, Z., Zhu, R., et al. (2025). *DAPO: An Open-Source LLM Reinforcement Learning System at Scale*. arXiv. [https://arxiv.org/abs/2503.14476](https://arxiv.org/abs/2503.14476) [↩](#user-content-fnref-dapo) 9. Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). *Training Verifiers to Solve Math Word Problems*. arXiv. [https://arxiv.org/abs/2110.14168](https://arxiv.org/abs/2110.14168) [↩](#user-content-fnref-cobbe-verifiers) 10. Li, Y., Choi, D., Chung, J., et al. (2022). *Competition-level code generation with AlphaCode*. Science, 378(6624), 1092–1097. [https://doi.org/10.1126/science.abq1158](https://doi.org/10.1126/science.abq1158) [↩](#user-content-fnref-alphacode) 11. Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). *CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning*. NeurIPS 2022. [https://arxiv.org/abs/2207.01780](https://arxiv.org/abs/2207.01780) [↩](#user-content-fnref-coderl) 12. Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). *RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2410.02089](https://arxiv.org/abs/2410.02089) [↩](#user-content-fnref-rlef) 13. Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449) [↩](#user-content-fnref-swe-rl) 14. Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe) [↩](#user-content-fnref-deepswe) 15. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501) [↩](#user-content-fnref-nebius-rl) 16. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) [↩](#user-content-fnref-swe-bench) 17. OpenAI (2024). *Introducing SWE-bench Verified*. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/) [↩](#user-content-fnref-swe-bench-verified) 18. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) 19. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) [↩](#user-content-fnref-swe-smith) 20. Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). *R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents*. arXiv. [https://arxiv.org/abs/2504.07164](https://arxiv.org/abs/2504.07164) [↩](#user-content-fnref-r2e-gym) 21. Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). *SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2505.20411](https://arxiv.org/abs/2505.20411) [↩](#user-content-fnref-swe-rebench) 22. Zan, D., Huang, Z., Liu, W., et al. (2025). *Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving*. arXiv. [https://arxiv.org/abs/2504.02605](https://arxiv.org/abs/2504.02605) [↩](#user-content-fnref-multi-swe-bench) 23. Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). *SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?* arXiv. [https://arxiv.org/abs/2502.12115](https://arxiv.org/abs/2502.12115) [↩](#user-content-fnref-swe-lancer) 24. Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). *SWE-Bench+: Enhanced Coding Benchmark for LLMs*. arXiv. [https://arxiv.org/abs/2410.06992](https://arxiv.org/abs/2410.06992) [↩](#user-content-fnref-swe-bench-plus) 25. Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). *Getafix: Learning to Fix Bugs Automatically*. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. [https://doi.org/10.1145/3360585](https://doi.org/10.1145/3360585) [↩](#user-content-fnref-getafix) 26. Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). *SapFix: Automated End-to-End Repair at Scale*. ICSE-SEIP 2019, 269–278. [https://doi.org/10.1109/ICSE-SEIP.2019.00039](https://doi.org/10.1109/ICSE-SEIP.2019.00039) [↩](#user-content-fnref-sapfix) 27. Maniatis, P., & Tarlow, D. (2023). *Large sequence models for software development activities*. Google Research Blog. [https://research.google/blog/large-sequence-models-for-software-development-activities/](https://research.google/blog/large-sequence-models-for-software-development-activities/) [↩](#user-content-fnref-didact) 28. Tabachnyk, M., & Nikolov, S. (2022). *ML-Enhanced Code Completion Improves Developer Productivity*. Google Research Blog. [https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/](https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/) [↩](#user-content-fnref-google-completion) 29. Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). *Measuring GitHub Copilot's Impact on Productivity*. Communications of the ACM, 67(3), 54–63. [https://doi.org/10.1145/3633453](https://doi.org/10.1145/3633453) [↩](#user-content-fnref-ziegler-copilot) 30. Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). *Building LLMs for Code Repair*. Replit. [https://replit.com/blog/code-repair](https://replit.com/blog/code-repair) [↩](#user-content-fnref-replit-repair) 31. Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). *End to End Learning for Self-Driving Cars*. arXiv. [https://arxiv.org/abs/1604.07316](https://arxiv.org/abs/1604.07316) [↩](#user-content-fnref-bojarski-driving) 32. Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). *A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning*. AISTATS 2011. [https://arxiv.org/abs/1011.0686](https://arxiv.org/abs/1011.0686) [↩](#user-content-fnref-dagger) 33. Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). *Hindsight Experience Replay*. NeurIPS 2017. [https://arxiv.org/abs/1707.01495](https://arxiv.org/abs/1707.01495) [↩](#user-content-fnref-her) 34. Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). *On the Naturalness of Software*. ICSE 2012, 837–847. [https://doi.org/10.1109/ICSE.2012.6227135](https://doi.org/10.1109/ICSE.2012.6227135) [↩](#user-content-fnref-hindle-naturalness) 35. Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). *A Survey of Machine Learning for Big Code and Naturalness*. ACM Computing Surveys, 51(4), Article 81. [https://arxiv.org/abs/1709.06182](https://arxiv.org/abs/1709.06182) [↩](#user-content-fnref-allamanis-survey) 36. Blum, A., & Hardt, M. (2015). *The Ladder: A Reliable Leaderboard for Machine Learning Competitions*. ICML 2015, PMLR 37, 1006–1014. [https://arxiv.org/abs/1502.04585](https://arxiv.org/abs/1502.04585) [↩](#user-content-fnref-ladder) 37. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). *The reusable holdout: Preserving validity in adaptive data analysis*. Science, 349(6248), 636–638. [https://doi.org/10.1126/science.aaa9375](https://doi.org/10.1126/science.aaa9375) [↩](#user-content-fnref-reusable-holdout) 38. Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). *AgentTuning: Enabling Generalized Agent Abilities for LLMs*. arXiv. [https://arxiv.org/abs/2310.12823](https://arxiv.org/abs/2310.12823) [↩](#user-content-fnref-agenttuning) 39. Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). *FireAct: Toward Language Agent Fine-tuning*. arXiv. [https://arxiv.org/abs/2310.05915](https://arxiv.org/abs/2310.05915) [↩](#user-content-fnref-fireact) 40. Kimi Team (2025). *Kimi K2: Open Agentic Intelligence*. arXiv. [https://arxiv.org/abs/2507.20534](https://arxiv.org/abs/2507.20534) [↩](#user-content-fnref-kimi-k2) 41. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering*. NeurIPS 2024. [https://arxiv.org/abs/2405.15793](https://arxiv.org/abs/2405.15793) [↩](#user-content-fnref-swe-agent) 42. Wang, X., Li, B., Song, Y., et al. (2024). *OpenHands: An Open Platform for AI Software Developers as Generalist Agents*. arXiv; ICLR 2025. [https://arxiv.org/abs/2407.16741](https://arxiv.org/abs/2407.16741) [↩](#user-content-fnref-openhands) --- # Closing: What to do this quarter > Seven steps for this quarter. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/what-to-do-this-quarter 1. **Install the collector.** The plugin in Chapter 8, or your own hooks that write the same schema. Start keeping traces this week, before the oracle exists. Unlabeled traces are still an asset, and the collector is the part with no dependencies. 2. **Pick one repository and write the first hidden tests.** Twenty tasks, each with a test that fails on the base commit for the right reason. Bug reports with reproductions are the easiest source. 3. **Run the oracle in a container with no network.** On the same machine at first. Move it to a runner the agents cannot log into before you train on anything. 4. **Measure the flip rate.** Dispatch the twenty tasks. Count the flips. That number, times your sessions per day, is your data rate, and Chapter 12 says what it buys. 5. **Keep the failures.** They are half of the preference set. 6. **Set aside the held-out tasks now.** Before the first training run, not after. The leak cannot be undone. 7. **Train the first adapter at 500 flips.** On all layers, with replay, with the shuffled-label control beside it. Evaluate on the held-out set. Expect it to be worse than the rented model and better than the base model. That is the first point on a curve the team now owns. --- # Appendix A: Trace event schema > The event types and fields in a session's trace file. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/trace-event-schema Each session is one JSON Lines file. Each line is one event. Every event carries `v` (schema version), `session_id`, `ts` (ISO 8601, UTC), and `type`. ```text session_start cwd, base_commit, remote_url_hash, model, harness, harness_version, source, task_id, config_hash message turn, role (user | assistant), text_redacted, text_hash, truncated (bool) tool_call turn, tool_name, tool_use_id, input_redacted, input_hash, output_redacted, output_hash, truncated (bool), duration_ms oracle_verdict attempt, mode (baseline | grade), verdict (PASS | FAIL | ERROR), tests_hash, image_digest, diff_hash, oracle_host, signature (optional) session_end outcome (flipped | no_flip | no_task | aborted), attempts, final_diff_hash, reason (from the harness), ended_at ``` Rules: `text_redacted`, `input_redacted`, and `output_redacted` are the only fields that ever hold content, and they hold it after redaction. The `_hash` fields are SHA-256 of the unredacted value, so a later pass can tell whether two truncated values were the same without storing them. `remote_url_hash` rather than the URL, because remote URLs sometimes carry credentials. The verdict's `signature` is over the canonical JSON of the other verdict fields, with the oracle host's key. --- # Appendix B: Oracle container specification > What the oracle container mounts, fixes, applies, runs, and reports. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/oracle-container-specification - **Image**: pinned by digest. Built from a Dockerfile that installs the repository's dependencies from its lockfile at build time. Rebuilt when the lockfile changes, and the new digest recorded. - **Network**: none. Started with networking disabled. - **Mounts**: the task's hidden tests, read-only; the diff file, read-only, in grade mode. Nothing from the agent's working copy. - **Repository**: a mirror baked into the image or mounted read-only from the oracle host. Checked out at the base commit inside the container. - **Environment**: `TZ=UTC`, `LC_ALL=C.UTF-8`, `PYTHONHASHSEED=0`, `SOURCE_DATE_EPOCH` fixed, and whatever the repository's own test configuration needs to be deterministic. - **Diff application**: the diff is filtered through the allowlist and denylist first. Dropped hunks are logged with their paths. The filtered diff is applied with `git apply --check` before `git apply`; a diff that does not apply is a FAIL with the reason logged. - **Hidden tests**: copied into the checkout after the diff is applied, from the mount, never from the diff. - **Run**: the existing suite at the base commit's version, then the hidden tests. Any failure in either is a FAIL. - **Timeout**: a wall-clock limit per run. Exceeding it is a FAIL with the reason logged. - **Output to the gate**: exit code 0 or 1, and one line `tests_hash=` of the sorted list of test identifiers that ran. - **Output to the oracle log**: everything. Test output, dropped hunks, timing, the exit reason, the image digest, the diff hash. - **Baseline mode**: the same run without a diff. Must return FAIL for the hidden tests and PASS for the existing suite, or the task is quarantined. --- # Appendix C: Data volume worksheet > Four numbers from your team and the arithmetic that turns them into a date. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume-worksheet Fill in the first four lines from your own team. The rest is arithmetic. ```text A sessions per day = engineers × sessions each = ____ B oracle share = fraction of sessions on tasks with a hidden test = ____ C flip rate per attempt = measured in week one = ____ D attempts per task = independent sessions dispatched = ____ E oracle tasks per day = A × B = ____ F chance a task flips at least once = 1 − (1 − C)^D = ____ G flips per day = E × F = ____ H token cost multiplier = D = ____ Days to 500 flips = 500 / G Days to 2,000 flips = 2,000 / G Days to 8,000 flips = 8,000 / G ``` Reference points from Chapter 12: 491 flips moved a 32-billion-parameter model from 7.0 to 20.6 percent on SWE-bench Verified; about 8,000 moved the same base model to 38.0 percent with the gain still growing. --- # Sources > The works cited, in order of first citation. From *Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces* by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/sources 119 works are cited in this book. They appear below in the order of their first citation. Each chapter also lists its own sources at its end. 1. Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). *How far behind are open models?* Epoch AI. [https://epoch.ai/blog/open-models-report](https://epoch.ai/blog/open-models-report) 2. Stanford Institute for Human-Centered Artificial Intelligence (2025). *AI Index Report 2025*, Chapter 2: Technical Performance. [https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance](https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance) 3. OpenAI (2025). *gpt-oss-120b and gpt-oss-20b Model Card*. arXiv. [https://arxiv.org/abs/2508.10925](https://arxiv.org/abs/2508.10925) 4. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. arXiv; ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) 5. Panickssery, A., Bowman, S. R., & Feng, S. (2024). *LLM Evaluators Recognize and Favor Their Own Generations*. arXiv; NeurIPS 2024. [https://arxiv.org/abs/2404.13076](https://arxiv.org/abs/2404.13076) 6. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). *STaR: Bootstrapping Reasoning With Reasoning*. arXiv. [https://arxiv.org/abs/2203.14465](https://arxiv.org/abs/2203.14465) 7. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) 8. Baker, B., Huizinga, J., Gao, L., et al. (2025). *Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation*. arXiv. [https://arxiv.org/abs/2503.11926](https://arxiv.org/abs/2503.11926) 9. Denison, C., MacDiarmid, M., Barez, F., et al. (2024). *Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models*. arXiv. [https://arxiv.org/abs/2406.10162](https://arxiv.org/abs/2406.10162) 10. Blum, A., & Hardt, M. (2015). *The Ladder: A Reliable Leaderboard for Machine Learning Competitions*. ICML 2015, PMLR 37, 1006–1014. [https://arxiv.org/abs/1502.04585](https://arxiv.org/abs/1502.04585) 11. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). *The reusable holdout: Preserving validity in adaptive data analysis*. Science, 349(6248), 636–638. [https://doi.org/10.1126/science.aaa9375](https://doi.org/10.1126/science.aaa9375) 12. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) 13. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) 14. Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290) 15. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). *LLM inference prices have fallen rapidly but unequally across tasks*. Epoch AI. [https://epoch.ai/data-insights/llm-inference-price-trends](https://epoch.ai/data-insights/llm-inference-price-trends) 16. Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449) 17. Mistral AI & All Hands AI (2025). *Devstral*. [https://mistral.ai/news/devstral](https://mistral.ai/news/devstral) 18. Qwen Team (2025). *Qwen3 Technical Report*. arXiv. [https://arxiv.org/abs/2505.09388](https://arxiv.org/abs/2505.09388) 19. Grattafiori, A., et al. (2024). *The Llama 3 Herd of Models*. arXiv. [https://arxiv.org/abs/2407.21783](https://arxiv.org/abs/2407.21783) 20. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) 21. Anthropic (2026). *Hooks reference*. Claude Code documentation. [https://code.claude.com/docs/en/hooks](https://code.claude.com/docs/en/hooks) 22. Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). *The Oracle Problem in Software Testing: A Survey*. IEEE Transactions on Software Engineering, 41(5), 507–525. [https://doi.org/10.1109/TSE.2014.2372785](https://doi.org/10.1109/TSE.2014.2372785) 23. Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). *An Empirical Analysis of Flaky Tests*. FSE 2014, 643–653. [https://doi.org/10.1145/2635868.2635920](https://doi.org/10.1145/2635868.2635920) 24. Micco, J. (2016). *Flaky Tests at Google and How We Mitigate Them*. Google Testing Blog. [https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html) 25. Lamb, C., & Zacchiroli, S. (2022). *Reproducible Builds: Increasing the Integrity of Software Supply Chains*. IEEE Software, 39(2), 62–70. [https://arxiv.org/abs/2104.06020](https://arxiv.org/abs/2104.06020) 26. Just, R., Jalali, D., & Ernst, M. D. (2014). *Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs*. ISSTA 2014, 437–440. [https://doi.org/10.1145/2610384.2628055](https://doi.org/10.1145/2610384.2628055) 27. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). *An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems*. ISSTA 2015, 24–36. [https://doi.org/10.1145/2771783.2771791](https://doi.org/10.1145/2771783.2771791) 28. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). *Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair*. ESEC/FSE 2015, 532–543. [https://doi.org/10.1145/2786805.2786825](https://doi.org/10.1145/2786805.2786825) 29. Inozemtseva, L., & Holmes, R. (2014). *Coverage Is Not Strongly Correlated with Test Suite Effectiveness*. ICSE 2014, 435–445. [https://doi.org/10.1145/2568225.2568271](https://doi.org/10.1145/2568225.2568271) 30. DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). *Hints on Test Data Selection: Help for the Practicing Programmer*. IEEE Computer, 11(4), 34–41. [https://doi.org/10.1109/C-M.1978.218136](https://doi.org/10.1109/C-M.1978.218136) 31. Petrović, G., & Ivanković, M. (2018). *State of Mutation Testing at Google*. ICSE-SEIP 2018, 163–171. [https://doi.org/10.1145/3183519.3183521](https://doi.org/10.1145/3183519.3183521) 32. Claessen, K., & Hughes, J. (2000). *QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs*. ICFP 2000, 268–279. [https://doi.org/10.1145/351240.351266](https://doi.org/10.1145/351240.351266) 33. Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). *A Survey on Metamorphic Testing*. IEEE Transactions on Software Engineering, 42(9), 805–824. [https://doi.org/10.1109/TSE.2016.2532875](https://doi.org/10.1109/TSE.2016.2532875) 34. McKeeman, W. M. (1998). *Differential Testing for Software*. Digital Technical Journal, 10(1), 100–107. 35. Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). *Measuring GitHub Copilot's Impact on Productivity*. Communications of the ACM, 67(3), 54–63. [https://doi.org/10.1145/3633453](https://doi.org/10.1145/3633453) 36. Von Arx, S., Chan, L., & Barnes, E. (2025). *Recent Frontier Models Are Reward Hacking*. METR. [https://metr.org/blog/2025-06-05-recent-reward-hacking/](https://metr.org/blog/2025-06-05-recent-reward-hacking/) 37. MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). *Natural Emergent Misalignment from Reward Hacking in Production RL*. arXiv. [https://arxiv.org/abs/2511.18397](https://arxiv.org/abs/2511.18397) 38. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). *Defining and Characterizing Reward Hacking*. NeurIPS 2022. [https://arxiv.org/abs/2209.13085](https://arxiv.org/abs/2209.13085) 39. Gao, L., Schulman, J., & Hilton, J. (2022). *Scaling Laws for Reward Model Overoptimization*. arXiv; ICML 2023. [https://arxiv.org/abs/2210.10760](https://arxiv.org/abs/2210.10760) 40. Thompson, K. (1984). *Reflections on Trusting Trust*. Communications of the ACM, 27(8), 761–763. [https://doi.org/10.1145/358198.358210](https://doi.org/10.1145/358198.358210) 41. Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). *Firecracker: Lightweight Virtualization for Serverless Applications*. NSDI 2020, 419–434. [https://www.usenix.org/conference/nsdi20/presentation/agache](https://www.usenix.org/conference/nsdi20/presentation/agache) 42. Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). *The True Cost of Containing: A gVisor Case Study*. HotCloud 2019. [https://www.usenix.org/conference/hotcloud19/presentation/young](https://www.usenix.org/conference/hotcloud19/presentation/young) 43. Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). *Teaching Large Language Models to Self-Debug*. arXiv; ICLR 2024. [https://arxiv.org/abs/2304.05128](https://arxiv.org/abs/2304.05128) 44. Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*. NeurIPS 2023 Datasets and Benchmarks. [https://arxiv.org/abs/2306.05685](https://arxiv.org/abs/2306.05685) 45. Wang, P., Li, L., Chen, L., et al. (2023). *Large Language Models are not Fair Evaluators*. arXiv; ACL 2024. [https://arxiv.org/abs/2305.17926](https://arxiv.org/abs/2305.17926) 46. Stechly, K., Marquez, M., & Kambhampati, S. (2023). *GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems*. arXiv. [https://arxiv.org/abs/2310.12397](https://arxiv.org/abs/2310.12397) 47. Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). *Can Large Language Models Really Improve by Self-critiquing Their Own Plans?* arXiv. [https://arxiv.org/abs/2310.08118](https://arxiv.org/abs/2310.08118) 48. Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). *LLMs cannot find reasoning errors, but can correct them given the error location*. Findings of ACL 2024. [https://arxiv.org/abs/2311.08516](https://arxiv.org/abs/2311.08516) 49. Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). *When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs*. Transactions of the ACL, 12. [https://arxiv.org/abs/2406.01297](https://arxiv.org/abs/2406.01297) 50. Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). *Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement*. ACL 2024. [https://arxiv.org/abs/2402.11436](https://arxiv.org/abs/2402.11436) 51. Meli, M., McNiece, M. R., & Reaves, B. (2019). *How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories*. NDSS 2019. [https://doi.org/10.14722/ndss.2019.23418](https://doi.org/10.14722/ndss.2019.23418) 52. Anthropic (2026). *Plugin manifest reference*. Claude Code documentation. [https://code.claude.com/docs/en/plugins-reference](https://code.claude.com/docs/en/plugins-reference) 53. Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). *Realizing quality improvement through test driven development: results and experiences of four industrial teams*. Empirical Software Engineering, 13(3), 289–302. [https://doi.org/10.1007/s10664-008-9062-z](https://doi.org/10.1007/s10664-008-9062-z) 54. OpenAI (2024). *Introducing SWE-bench Verified*. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/) 55. Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). *Large Language Monkeys: Scaling Inference Compute with Repeated Sampling*. arXiv. [https://arxiv.org/abs/2407.21787](https://arxiv.org/abs/2407.21787) 56. Touvron, H., et al. (2023). *Llama 2: Open Foundation and Fine-Tuned Chat Models*. arXiv. [https://arxiv.org/abs/2307.09288](https://arxiv.org/abs/2307.09288) 57. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). *Are Mutants a Valid Substitute for Real Faults in Software Testing?* FSE 2014, 654–665. [https://doi.org/10.1145/2635868.2635929](https://doi.org/10.1145/2635868.2635929) 58. Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). *SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2505.20411](https://arxiv.org/abs/2505.20411) 59. Zhao, A., Wu, Y., Yue, Y., et al. (2025). *Absolute Zero: Reinforced Self-play Reasoning with Zero Data*. arXiv. [https://arxiv.org/abs/2505.03335](https://arxiv.org/abs/2505.03335) 60. Hinton, G., Vinyals, O., & Dean, J. (2015). *Distilling the Knowledge in a Neural Network*. NeurIPS 2014 Deep Learning Workshop. [https://arxiv.org/abs/1503.02531](https://arxiv.org/abs/1503.02531) 61. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. [https://arxiv.org/abs/2305.18290](https://arxiv.org/abs/2305.18290) 62. Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). *Hindsight Experience Replay*. NeurIPS 2017. [https://arxiv.org/abs/1707.01495](https://arxiv.org/abs/1707.01495) 63. Anthony, T., Tian, Z., & Barber, D. (2017). *Thinking Fast and Slow with Deep Learning and Tree Search*. NeurIPS 2017. [https://arxiv.org/abs/1705.08439](https://arxiv.org/abs/1705.08439) 64. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). *Mastering the game of Go without human knowledge*. Nature, 550, 354–359. [https://doi.org/10.1038/nature24270](https://doi.org/10.1038/nature24270) 65. Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). *Reinforced Self-Training (ReST) for Language Modeling*. arXiv. [https://arxiv.org/abs/2308.08998](https://arxiv.org/abs/2308.08998) 66. Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). *Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models*. arXiv; TMLR 2024. [https://arxiv.org/abs/2312.06585](https://arxiv.org/abs/2312.06585) 67. Li, Y., Choi, D., Chung, J., et al. (2022). *Competition-level code generation with AlphaCode*. Science, 378(6624), 1092–1097. [https://doi.org/10.1126/science.abq1158](https://doi.org/10.1126/science.abq1158) 68. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). *A Long Way to Go: Investigating Length Correlations in RLHF*. arXiv; COLM 2024. [https://arxiv.org/abs/2310.03716](https://arxiv.org/abs/2310.03716) 69. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). *Tülu 3: Pushing Frontiers in Open Language Model Post-Training*. arXiv. [https://arxiv.org/abs/2411.15124](https://arxiv.org/abs/2411.15124) 70. Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe) 71. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501) 72. Yu, Q., Zhang, Z., Zhu, R., et al. (2025). *DAPO: An Open-Source LLM Reinforcement Learning System at Scale*. arXiv. [https://arxiv.org/abs/2503.14476](https://arxiv.org/abs/2503.14476) 73. Qwen Team (2025). *Qwen3-Coder: Agentic Coding in the World*. [https://qwenlm.github.io/blog/qwen3-coder/](https://qwenlm.github.io/blog/qwen3-coder/) 74. Yue, Y., Chen, Z., Lu, R., et al. (2025). *Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?* arXiv; NeurIPS 2025. [https://arxiv.org/abs/2504.13837](https://arxiv.org/abs/2504.13837) 75. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). *LoRA: Low-Rank Adaptation of Large Language Models*. arXiv; ICLR 2022. [https://arxiv.org/abs/2106.09685](https://arxiv.org/abs/2106.09685) 76. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). *QLoRA: Efficient Finetuning of Quantized LLMs*. NeurIPS 2023. [https://arxiv.org/abs/2305.14314](https://arxiv.org/abs/2305.14314) 77. Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). *LoRA Learns Less and Forgets Less*. Transactions on Machine Learning Research. [https://arxiv.org/abs/2405.09673](https://arxiv.org/abs/2405.09673) 78. Schulman, J., & Thinking Machines Lab (2025). *LoRA Without Regret*. [https://thinkingmachines.ai/blog/lora/](https://thinkingmachines.ai/blog/lora/) 79. McCloskey, M., & Cohen, N. J. (1989). *Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem*. Psychology of Learning and Motivation, 24, 109–165. [https://doi.org/10.1016/S0079-7421(08)60536-8](https://doi.org/10.1016/S0079-7421\(08\)60536-8) 80. Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). *An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning*. arXiv. [https://arxiv.org/abs/2308.08747](https://arxiv.org/abs/2308.08747) 81. Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). *Simple and Scalable Strategies to Continually Pre-train Large Language Models*. arXiv; TMLR. [https://arxiv.org/abs/2403.08763](https://arxiv.org/abs/2403.08763) 82. Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). *Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time*. ICML 2022. [https://arxiv.org/abs/2203.05482](https://arxiv.org/abs/2203.05482) 83. Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). *TIES-Merging: Resolving Interference When Merging Models*. NeurIPS 2023. [https://arxiv.org/abs/2306.01708](https://arxiv.org/abs/2306.01708) 84. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). *AI models collapse when trained on recursively generated data*. Nature, 631, 755–759. [https://doi.org/10.1038/s41586-024-07566-y](https://doi.org/10.1038/s41586-024-07566-y) 85. Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). *Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data*. arXiv. [https://arxiv.org/abs/2404.01413](https://arxiv.org/abs/2404.01413) 86. Shao, R., Li, S. S., Xin, R., et al. (2025). *Spurious Rewards: Rethinking Training Signals in RLVR*. arXiv. [https://arxiv.org/abs/2506.10947](https://arxiv.org/abs/2506.10947) 87. Zhou, C., Liu, P., Xu, P., et al. (2023). *LIMA: Less Is More for Alignment*. NeurIPS 2023. [https://arxiv.org/abs/2305.11206](https://arxiv.org/abs/2305.11206) 88. Muennighoff, N., Yang, Z., Shi, W., et al. (2025). *s1: Simple test-time scaling*. arXiv. [https://arxiv.org/abs/2501.19393](https://arxiv.org/abs/2501.19393) 89. Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). *LIMO: Less is More for Reasoning*. arXiv (v1, February 2025); COLM 2025. [https://arxiv.org/abs/2502.03387](https://arxiv.org/abs/2502.03387) 90. Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). *FireAct: Toward Language Agent Fine-tuning*. arXiv. [https://arxiv.org/abs/2310.05915](https://arxiv.org/abs/2310.05915) 91. Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). *AgentTuning: Enabling Generalized Agent Abilities for LLMs*. arXiv. [https://arxiv.org/abs/2310.12823](https://arxiv.org/abs/2310.12823) 92. Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). *When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method*. ICLR 2024. [https://arxiv.org/abs/2402.17193](https://arxiv.org/abs/2402.17193) 93. Chen, L., Li, S., Yan, J., et al. (2023). *AlpaGasus: Training a Better Alpaca with Fewer Data*. arXiv; ICLR 2024. [https://arxiv.org/abs/2307.08701](https://arxiv.org/abs/2307.08701) 94. Humble, J., & Farley, D. (2010). *Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation*. Addison-Wesley. 95. Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). *Investigating Data Contamination in Modern Benchmarks for Large Language Models*. NAACL 2024. [https://arxiv.org/abs/2311.09783](https://arxiv.org/abs/2311.09783) 96. Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). *SWE-Bench+: Enhanced Coding Benchmark for LLMs*. arXiv. [https://arxiv.org/abs/2410.06992](https://arxiv.org/abs/2410.06992) 97. Zan, D., Huang, Z., Liu, W., et al. (2025). *Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving*. arXiv. [https://arxiv.org/abs/2504.02605](https://arxiv.org/abs/2504.02605) 98. Kwon, W., Li, Z., Zhuang, S., et al. (2023). *Efficient Memory Management for Large Language Model Serving with PagedAttention*. SOSP 2023. [https://arxiv.org/abs/2309.06180](https://arxiv.org/abs/2309.06180) 99. Sheng, Y., Cao, S., Li, D., et al. (2023). *S-LoRA: Serving Thousands of Concurrent LoRA Adapters*. arXiv; MLSys 2024. [https://arxiv.org/abs/2311.03285](https://arxiv.org/abs/2311.03285) 100. Carlini, N., Tramèr, F., Wallace, E., et al. (2021). *Extracting Training Data from Large Language Models*. USENIX Security 2021. [https://arxiv.org/abs/2012.07805](https://arxiv.org/abs/2012.07805) 101. Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). *Quantifying Memorization Across Neural Language Models*. ICLR 2023. [https://arxiv.org/abs/2202.07646](https://arxiv.org/abs/2202.07646) 102. Sculley, D., Holt, G., Golovin, D., et al. (2015). *Hidden Technical Debt in Machine Learning Systems*. NeurIPS 2015. 103. Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). *Training Verifiers to Solve Math Word Problems*. arXiv. [https://arxiv.org/abs/2110.14168](https://arxiv.org/abs/2110.14168) 104. Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). *CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning*. NeurIPS 2022. [https://arxiv.org/abs/2207.01780](https://arxiv.org/abs/2207.01780) 105. Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). *RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2410.02089](https://arxiv.org/abs/2410.02089) 106. Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). *R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents*. arXiv. [https://arxiv.org/abs/2504.07164](https://arxiv.org/abs/2504.07164) 107. Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). *SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?* arXiv. [https://arxiv.org/abs/2502.12115](https://arxiv.org/abs/2502.12115) 108. Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). *Getafix: Learning to Fix Bugs Automatically*. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. [https://doi.org/10.1145/3360585](https://doi.org/10.1145/3360585) 109. Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). *SapFix: Automated End-to-End Repair at Scale*. ICSE-SEIP 2019, 269–278. [https://doi.org/10.1109/ICSE-SEIP.2019.00039](https://doi.org/10.1109/ICSE-SEIP.2019.00039) 110. Maniatis, P., & Tarlow, D. (2023). *Large sequence models for software development activities*. Google Research Blog. [https://research.google/blog/large-sequence-models-for-software-development-activities/](https://research.google/blog/large-sequence-models-for-software-development-activities/) 111. Tabachnyk, M., & Nikolov, S. (2022). *ML-Enhanced Code Completion Improves Developer Productivity*. Google Research Blog. [https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/](https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/) 112. Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). *Building LLMs for Code Repair*. Replit. [https://replit.com/blog/code-repair](https://replit.com/blog/code-repair) 113. Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). *End to End Learning for Self-Driving Cars*. arXiv. [https://arxiv.org/abs/1604.07316](https://arxiv.org/abs/1604.07316) 114. Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). *A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning*. AISTATS 2011. [https://arxiv.org/abs/1011.0686](https://arxiv.org/abs/1011.0686) 115. Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). *On the Naturalness of Software*. ICSE 2012, 837–847. [https://doi.org/10.1109/ICSE.2012.6227135](https://doi.org/10.1109/ICSE.2012.6227135) 116. Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). *A Survey of Machine Learning for Big Code and Naturalness*. ACM Computing Surveys, 51(4), Article 81. [https://arxiv.org/abs/1709.06182](https://arxiv.org/abs/1709.06182) 117. Kimi Team (2025). *Kimi K2: Open Agentic Intelligence*. arXiv. [https://arxiv.org/abs/2507.20534](https://arxiv.org/abs/2507.20534) 118. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering*. NeurIPS 2024. [https://arxiv.org/abs/2405.15793](https://arxiv.org/abs/2405.15793) 119. Wang, X., Li, B., Song, Y., et al. (2024). *OpenHands: An Open Platform for AI Software Developers as Generalist Agents*. arXiv; ICLR 2025. [https://arxiv.org/abs/2407.16741](https://arxiv.org/abs/2407.16741) --- # Agents are waiting on a process built for people > An agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour. Published 2026-09-30 by Oxagen Research. Canonical page: https://oxagen.sh/blog/agents-are-waiting-on-a-process-built-for-people An agent can write a change to your code in a few minutes. In most teams, that change then waits. It waits for a person to open it, read it, and say yes. At night, it waits until morning. On a Friday evening, it waits until Monday. The agent could start the next change, but the next change waits in the same line. The agent is not the slow part. The process around it is. Most software processes were made for people, and they still expect a person at every step. ## The wait A week has 168 hours. A 40-hour work week covers 40 of them. A process that needs a person at every step can only move in those 40 hours. For the other 128 hours, the work sits still. The wait inside the work week is long too, and it has been measured for years. A 2018 study of code review at Google found a median wait of under an hour for first feedback on a small change and about 5 hours on the largest ones. The median for a whole review, all sizes together, was under 4 hours. The same paper lists the earlier figures it beat: a median time to approval of 17.5 hours at AMD, 15.7 hours for Chrome OS, and between 14.7 and 19.8 hours on three Microsoft projects.[1](#user-content-fn-1) Meta measured how long changes wait for a reviewer. In early 2021, the typical change waited a few hours. The slowest quarter of changes waited as much as a day.[2](#user-content-fn-2) Meta then built a tool that reminds reviewers about old changes. It cut the average wait by 7 percent. It cut the number of changes that waited more than three days by 12 percent.[2](#user-content-fn-2) Microsoft built the same kind of tool. In a trial across 147 repositories, its reminders cut the time to resolve 8,500 pull requests by 60 percent, and it went on to send 210,000 reminders a year across 8,000 repositories.[3](#user-content-fn-3) Those are real gains. They are gains on a wait counted in hours and days, for work an agent does in minutes. The wait is not only for the first look. At Google, reviewers leave millions of comments a year, and an author spends about 60 minutes of active work on a change between sending it for review and submitting it.[4](#user-content-fn-4) Reading also has limits. A SmartBear study of a Cisco team found that a person reads code best in chunks of 200 to 400 lines. Past 400 lines, people find fewer problems. The study also advises reading for no more than 60 minutes at a time.[5](#user-content-fn-5) One agent task can produce more than 400 lines. Read well, that is about an hour of one person's time, for each task. A change an agent wrote waits longer than one a person wrote. A 2026 benchmark over 8.1 million pull requests from 4,800 teams found that a pull request opened by an agent waited 5.3 times longer for a reviewer to pick it up than one written without AI. Once picked up, AI-written pull requests were reviewed twice as fast, and fewer of them were accepted: 32.7 percent against 84.4 percent.[6](#user-content-fn-6) A study of 33,596 agent-authored pull requests on GitHub found that 61 percent received no recorded review at all. Of the review comments the rest did get, 72 percent came from other agents.[7](#user-content-fn-7) ## Faster writing The writing itself did get faster, in most measurements. In a 2023 trial, 95 programmers given an AI code completion tool finished a task 55.8 percent faster than those without one.[8](#user-content-fn-8) Three field experiments at Microsoft, Accenture, and a Fortune 100 company, covering 4,867 developers, found 26 percent more completed tasks.[9](#user-content-fn-9) A trial with 96 Google engineers found a gain of about 21 percent, with a wide confidence interval.[10](#user-content-fn-10) One trial found the opposite. Sixteen experienced open-source developers took 19 percent longer with AI on repositories they knew well.[11](#user-content-fn-11) An earlier post, [The problem is not slop, it is your process](/blog/the-problem-is-not-slop-it-is-your-process), looks at that result. Agents are also taking on bigger jobs. METR measures the length of task an AI model can finish, counted in the time a skilled person would need. That length has doubled about every seven months since 2019.[12](#user-content-fn-12) Longer tasks make bigger changes, and bigger changes take longer to read. ## Bigger changes DORA's 2024 report shows what happens when teams add AI and keep the old process. When a team's use of AI went up by a quarter, delivery throughput fell by an estimated 1.5 percent. Delivery stability fell by an estimated 7.2 percent.[13](#user-content-fn-13) The report names large batches as one cause. DORA's 2025 report, from a survey of nearly 5,000 technology professionals, found 90 percent of respondents using AI at work. This time, more AI went with higher throughput. It still went with lower stability. The report names what is missing: automated tests, mature version control practice, and fast feedback, without which more change brings more instability.[14](#user-content-fn-14) A 2025 telemetry study shows where the time goes. Across more than 10,000 developers in 1,255 teams, developers using AI completed 21 percent more tasks and merged 98 percent more pull requests. The average pull request grew by 154 percent. Review time per pull request rose by 91 percent, and bugs per developer rose by 9 percent. At the company level, the study found no significant link between AI adoption and delivery metrics.[15](#user-content-fn-15) Each developer got faster. The process absorbed the gain. ## Process changes at eight companies Some companies changed the process instead of waiting on it. Each change below moves a step that needed a person onto a machine, and keeps a person for the rule, the exception, or the final click. ### The fix beside the comment Google trained a model on its own review history to turn a reviewer's comment into a suggested edit. The author sees the edit next to the comment and applies it with one click. In production, authors resolve 7.5 percent of all reviewer comments this way.[4](#user-content-fn-4) By mid 2024, AI assistance addressed more than 8 percent of review comments, and code completion wrote 50 percent of the code characters at Google, with engineers accepting 37 percent of its suggestions.[16](#user-content-fn-16) The reviewer still writes the comment. The author no longer writes the fix. ### Tests filtered before review Meta built TestGen-LLM to extend existing unit tests. A generated test reaches an engineer only after it builds, passes reliably, and raises coverage. On Instagram's Reels and Stories, 75 percent of generated tests built, 57 percent passed reliably, and 25 percent raised coverage. Engineers accepted 73 percent of the tests that passed the filter for production.[17](#user-content-fn-17) The filter did the first review. The person did the last. ### Migrations as a loop Airbnb moved about 3,500 React test files from one test framework to another. The estimate by hand was 1.5 years. The migration took six weeks. A pipeline sent each file through a set of steps. When a step failed, it sent the errors and the current file back to the model and tried again, up to a retry limit. Seventy-five percent of the files finished in four hours. Raising the retry limit, to between 50 and 100 attempts for the hardest files, took the total to 97 percent. People finished the last 3 percent.[18](#user-content-fn-18) Amazon used its own tool to move production Java applications from Java 8 and 11 to Java 17. An upgrade that typically took 50 developer-days took a few hours. In under six months, more than half of Amazon's production Java systems moved. Developers shipped 79 percent of the generated changes without editing them. Amazon puts the saving at 4,500 developer-years and 260 million dollars a year.[19](#user-content-fn-19) Google ran 39 code migrations over twelve months with three developers. Of the 595 changes submitted, 74 percent were written by the model, and the developers put the time saved at half.[20](#user-content-fn-20) In each case the loop runs without a person, and a person reads the result. ### A machine reviewer on every change Uber's uReview reads more than 90 percent of the roughly 65,000 changes its engineers post each week and comments on them before a person does. Engineers mark 75 percent of its comments as useful and address more than 65 percent of them.[21](#user-content-fn-21) A comment is not a decision. The person still decides. Not every rollout saves time. One company added an automated reviewer across ten projects and 238 practitioners. Of 4,335 pull requests, 1,568 got an automated review. Engineers resolved 73.8 percent of the automated comments, and the average time to close a pull request rose from 5 hours 52 minutes to 8 hours 20 minutes.[22](#user-content-fn-22) A reviewer that only adds comments adds work. The gain comes when the review changes what a person has to read. ### Approval by rule Google's large-scale change tooling has worked this way for years. A change that touches thousands of files is split by ownership into pieces that can land on their own. Each piece goes through its own test, mail, and submit pipeline. Reviewers for the whole change use pattern-based tools to approve the pieces that match what they expect, and read only the anomalies, such as a merge conflict. The pipeline lands more than 700 changes touching more than 15,000 files a day.[23](#user-content-fn-23) Two companies now approve pull requests by rule. At Intercom, more than 93 percent of pull requests in the two main codebases are agent-driven. More than 19 percent are approved with no person in the loop, and time to approval at the 75th percentile improved by 6 to 16 times. Any engineer can ask for a human review on any change.[24](#user-content-fn-24) At Rewind, a review bot approves about 40 percent of pull requests. Approval is not merge. The bot posts an approval the way a person would and does not press the merge button, and a pull request an agent wrote cannot merge until a person clicks. The policy is one YAML file, which code owners must approve and which changes through the same review as any other code. A verdict engine of ordinary tested code, not a model, turns the findings into the decision.[25](#user-content-fn-25) GitHub's coding agent shows the default without such a rule. The agent works in a GitHub Actions environment and pushes commits to a draft pull request. It opened more than a million pull requests between May and September 2025.[26](#user-content-fn-26) Each one waits for a person's approval before its automated checks even run.[27](#user-content-fn-27) ### The shared shape In every case, a machine does the first pass, and a person does something smaller than reading every line: writes the rule, handles the exception, or clicks the final approval. The person is still in charge. The person is no longer in the line. ## A process made for agents A process made for agents keeps people in charge and takes them out of the line. It has four parts. - **Rules.** People write the rules before the work starts. The rules say what an agent may do, what it may spend, and when it must ask a person. - **Checks.** Automatic checks build and test every change. Review agents read it and rate each problem they find. A serious problem blocks the change. - **Decisions.** A rule can send a request to a person: more budget, a new tool, or a change to production data. The run waits for that answer, and only for that answer. - **The record.** Every change leaves a record of what it did, which checks passed, and what it cost. A person reads the record each day instead of reading every line. Each change then moves through ten stages. Agents run all ten. The figure follows one change, #17, through them. > **Ten stages of the agent SDLC** > > 1. **Issue**: The issue says what done looks like > 2. **Triage**: A triage agent sets priority and size > 3. **Claim**: One change, one writer > 4. **Build**: The change ships with its tests > 5. **Check**: Automatic checks build and test it > 6. **Review**: Review agents rate each finding > 7. **Merge**: Green checks and no serious finding > 8. **Release**: Database changes go first > 9. **Watch**: A broken main opens an issue > 10. **Learn**: Minutes planned against minutes spent > > Repeat Each run ends with a note that the next issue can use. > > *Schematic. One change, #17, moves through the ten stages and the loop back to the next issue. Each stage is described in full at sdlc.oxagen.sh.* ## Oxagen's own process Oxagen builds its own product this way. Four agent tools do the work: Claude Code, Codex, Cursor, and Stella. They follow one set of written rules, kept in the code next to the product. These are some of the rules: - Every change starts as an issue, and the issue says what done looks like. - A triage agent sets each issue's priority and size. Whoever files the issue does not. - One change has one writer. An agent posts a claim before it writes, and the claim lasts 90 minutes. - Automatic checks are the only place the code is built and tested. - An agent looks at each open change every 60 seconds. It fixes failed checks, answers review comments, and clears conflicts. - When the main branch breaks, a check opens a top-priority issue. The fix goes straight to the main branch. The check closes the issue when the branch is green again. From September 1 to 29, 2026, 1,064 changes merged into Oxagen's main branch.[28](#user-content-fn-28) Of those, 685 merged outside weekday work hours. That is 64 percent. 283 merged on a Saturday or a Sunday, and 289 merged between 10 pm and 6 am Pacific time. > **Changes merged into Oxagen, September 1 to 29, 2026** > > | When the change merged | Changes | > | --- | --- | > | Weekdays, 9 am to 6 pm | 379 | > | Weekdays, other hours | 402 | > | Saturday and Sunday | 283 | > > *Source: GitHub search over Oxagen's repository, pull requests merged September 1 to 29, 2026, counted by UTC date. Hours are Pacific time.* These numbers show when changes merged. They do not show how good each change was. The checks and the review findings answer that for each change, and the record keeps the answer. Oxagen is workforce management for autonomous agents. In this process, Oxagen is where each agent gets its own identity and works under a mandate: its access, its budget and rules, and its tools and skills. For actions routed through Oxagen, a rule answers each request, and the record shows what each agent did and what it cost. ## A copy for your team The full process is at [sdlc.oxagen.sh](https://sdlc.oxagen.sh). It has the ten stages, the rules, the checks, and a four-week plan to start. Type your company name below to get a copy with your name in it, ready to edit. Company name FormatWordGoogle DocsPDF Google Docs opens the Word file. The page that opens shows you how. ## Limits - The merge counts come from one repository over 29 days. They describe when Oxagen's changes merged. They do not predict what another team will see. - The company figures come from each company's own engineering posts and papers. Each company measured its own code with its own tools, and none of the figures has been reproduced elsewhere. - The Google, Meta, and SmartBear review numbers describe people reading code that people wrote. They were measured before agents wrote much of the code. - The trial results measure how fast one person finishes a task. They do not measure delivery. - The pull request benchmark and the telemetry study come from vendors of engineering analytics, over the teams that use their products. - DORA reports links in survey data. A link is not proof of a cause. ## Footnotes 1. Sadowski, C., Söderberg, E., Church, L., Sipko, M., & Bacchelli, A. (2018). *Modern Code Review: A Case Study at Google*. ICSE SEIP 2018. [https://research.google/pubs/modern-code-review-a-case-study-at-google/](https://research.google/pubs/modern-code-review-a-case-study-at-google/) [↩](#user-content-fnref-1) 2. Riggs, P. (2022). *Improving code review time at Meta*. Engineering at Meta. [https://engineering.fb.com/2022/11/16/culture/meta-code-review-time-improving/](https://engineering.fb.com/2022/11/16/culture/meta-code-review-time-improving/) [↩](#user-content-fnref-2) [↩2](#user-content-fnref-2-2) 3. Maddila, C., Upadhyaya, S. S., Bansal, C., Nagappan, N., Gousios, G., & van Deursen, A. (2022). *Nudge: Accelerating Overdue Pull Requests Towards Completion*. ACM TOSEM. [https://arxiv.org/abs/2011.12468](https://arxiv.org/abs/2011.12468) [↩](#user-content-fnref-3) 4. Frömmgen, A., Austin, J., Choy, P., et al. (2024). *Resolving Code Review Comments with Machine Learning*. ICSE SEIP 2024. [https://research.google/pubs/resolving-code-review-comments-with-machine-learning/](https://research.google/pubs/resolving-code-review-comments-with-machine-learning/) [↩](#user-content-fnref-4) [↩2](#user-content-fnref-4-2) 5. SmartBear. *Best practices for code review*, reporting a SmartBear study of a Cisco Systems programming team. [https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/](https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/) [↩](#user-content-fnref-5) 6. LinearB (2026). *2026 Software Engineering Benchmarks Report*, a study of 8.1 million pull requests from 4,800 engineering teams. [https://linearb.io/resources/engineering-benchmarks](https://linearb.io/resources/engineering-benchmarks) [↩](#user-content-fnref-6) 7. Duma, K., Wróblewski, P., Bobińska, J., Winiarska, J., & Przymus, P. (2026). *These Aren't the Reviews You're Looking For: How Humans Review AI-Generated Pull Requests*. arXiv. [https://arxiv.org/abs/2605.02273](https://arxiv.org/abs/2605.02273) [↩](#user-content-fnref-7) 8. Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). *The Impact of AI on Developer Productivity*. arXiv. [https://arxiv.org/abs/2302.06590](https://arxiv.org/abs/2302.06590) [↩](#user-content-fnref-8) 9. Cui, Z. K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). *The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers*. [https://economics.mit.edu/sites/default/files/inline-files/draft\_copilot\_experiments.pdf](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf) [↩](#user-content-fnref-9) 10. Paradis, E., Grey, K., Hoang, Q., et al. (2024). *How much does AI impact development speed? An enterprise-based randomized controlled trial*. arXiv. [https://arxiv.org/abs/2410.12944](https://arxiv.org/abs/2410.12944) [↩](#user-content-fnref-10) 11. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). *Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity*. arXiv. [https://arxiv.org/abs/2507.09089](https://arxiv.org/abs/2507.09089) [↩](#user-content-fnref-11) 12. Kwa, T., West, B., Becker, J., et al. (2025). *Measuring AI Ability to Complete Long Software Tasks*. METR. [https://arxiv.org/abs/2503.14499](https://arxiv.org/abs/2503.14499) [↩](#user-content-fnref-12) 13. DORA (2024). *Announcing the 2024 DORA report*. Google Cloud. [https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report](https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report) [↩](#user-content-fnref-13) 14. DORA (2025). *Announcing the 2025 DORA report*. Google Cloud. [https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report) [↩](#user-content-fnref-14) 15. Faros AI (2025). *The AI Productivity Paradox Report 2025*, telemetry from more than 10,000 developers across 1,255 teams. [https://www.faros.ai/blog/ai-software-engineering](https://www.faros.ai/blog/ai-software-engineering) [↩](#user-content-fnref-15) 16. Chandra, S., & Tabachnyk, M. (2024). *AI in software engineering at Google: Progress and the path ahead*. Google Research. [https://research.google/blog/ai-in-software-engineering-at-google-progress-and-the-path-ahead/](https://research.google/blog/ai-in-software-engineering-at-google-progress-and-the-path-ahead/) [↩](#user-content-fnref-16) 17. Alshahwan, N., Chheda, J., Finegenova, A., et al. (2024). *Automated Unit Test Improvement using Large Language Models at Meta*. arXiv. [https://arxiv.org/abs/2402.09171](https://arxiv.org/abs/2402.09171) [↩](#user-content-fnref-17) 18. Covey-Brandt, C. (2025). *Accelerating Large-Scale Test Migration with LLMs*. The Airbnb Tech Blog. [https://medium.com/airbnb-engineering/accelerating-large-scale-test-migration-with-llms-9565c208023b](https://medium.com/airbnb-engineering/accelerating-large-scale-test-migration-with-llms-9565c208023b) [↩](#user-content-fnref-18) 19. Jassy, A. (2024). Post on the Amazon Q Developer Java upgrades, August 2024. [https://www.linkedin.com/posts/andy-jassy-8b1615\_one-of-the-most-tedious-but-critical-tasks-activity-7232374162185461760-AdSz/](https://www.linkedin.com/posts/andy-jassy-8b1615_one-of-the-most-tedious-but-critical-tasks-activity-7232374162185461760-AdSz/) See also Arisoy Cholkar, A. (2024). *Amazon Q Developer just reached a 260 million dollar milestone*. AWS DevOps Blog. [https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone](https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone) [↩](#user-content-fnref-19) 20. Ziftci, C., Nikolov, S., Sjövall, A., Kim, B., Codecasa, D., & Kim, M. (2025). *Migrating Code At Scale With LLMs At Google*. arXiv. [https://arxiv.org/abs/2504.09691](https://arxiv.org/abs/2504.09691) [↩](#user-content-fnref-20) 21. Mahajan, A., Roy Choudhary, S., Bond, M., Wang, Z., & Utture, A. (2025). *uReview*. Uber Engineering. [https://www.uber.com/blog/ureview/](https://www.uber.com/blog/ureview/) [↩](#user-content-fnref-21) 22. Cihan, U., Haratian, V., İnce, A., et al. (2025). *Automated Code Review In Practice*. ICSE SEIP 2025. [https://arxiv.org/abs/2412.18531](https://arxiv.org/abs/2412.18531) [↩](#user-content-fnref-22) 23. Wright, H. (2020). *Large-Scale Changes*, chapter 22 of *Software Engineering at Google*. O'Reilly. [https://abseil.io/resources/swe-book/html/ch22.html](https://abseil.io/resources/swe-book/html/ch22.html) [↩](#user-content-fnref-23) 24. Mykhailov, K., & Young, N. (2026). *AI is approving our pull requests: Here's how we made it safe*. Intercom. [https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/](https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/) [↩](#user-content-fnref-24) 25. North, D. (2026). *How we let AI approve pull requests (safely)*. Rewind. [https://rewind.com/blog/ai-approve-pull-requests-safely/](https://rewind.com/blog/ai-approve-pull-requests-safely/) [↩](#user-content-fnref-25) 26. GitHub (2025). *Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1*. The GitHub Blog. [https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/) [↩](#user-content-fnref-26) 27. Dohmke, T. (2025). *Meet the new coding agent*. The GitHub Blog. [https://github.blog/news-insights/product-news/github-copilot-meet-the-new-coding-agent/](https://github.blog/news-insights/product-news/github-copilot-meet-the-new-coding-agent/) [↩](#user-content-fnref-27) 28. Oxagen (2026). GitHub search `repo:macanderson/oxagen is:pr is:merged merged:2026-09-01..2026-09-29`, run on September 30, 2026. It returned 1,064 pull requests. [↩](#user-content-fnref-28) --- # Steering a run you are not watching > When a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record. Published 2026-09-16 by Oxagen Research. Canonical page: https://oxagen.sh/blog/steering-a-run-you-are-not-watching An agent has been running for nine hours. You open the transcript over breakfast. Around hour three, the agent decided the caching layer was the problem. Since then it has rewritten the wrong component, with tests and documentation. You have two buttons. One stops the run. That throws away nine hours, including the part of the work that was fine. The other lets you type a message. Most people pick this one, and the research shows it works badly. ## Corrections typed as chat often do not work A correction typed as chat only helps if the agent can take a correction that way. Laban and colleagues tested this. They gave models an instruction over several turns instead of all at once. Scores dropped by 39 percent on average across six generation tasks.[1](#user-content-fn-1) The authors split the drop into a small loss in skill and a large rise in unreliability. They found the cause. Models commit to an assumption early and then rely on it too much. So your message at hour nine lands in a context that already holds nine hours of confident work in the wrong direction. Sinha and colleagues explain why that matters. When a model's context holds its own earlier errors, the model becomes more likely to make new mistakes. They call this self-conditioning.[2](#user-content-fn-2) Your correction is one paragraph. The rest of the context was written by the model, and it points the other way. Backlund and Petersson found the strongest form of this in long-running agents. In Vending-Bench, runs go past 20 million tokens. Agents fall into what the authors call tangential meltdown loops, and they rarely recover. The authors found no clear link between these failures and the context window filling up.[3](#user-content-fn-3) So the runs that most need a correction are the runs least able to act on one sent as a suggestion. People still need to correct long runs. The correction has to arrive through some channel other than one more chat turn. ## Interruption has to be built into the loop The formal study of this question is older than language agents. Hadfield-Menell and colleagues modeled the off-switch as a game. An agent with a fixed goal cannot reach that goal while it is switched off. So a rational agent with a fixed goal has a reason to disable its off-switch.[4](#user-content-fn-4) The authors showed what it takes for the agent to prefer keeping the switch. The agent has to be unsure how good the outcome is. It also has to treat the human's actions as evidence about that. This result says nothing about what today's models want. It says where the ability to interrupt should live. Suppose the agent has to choose to check for messages. Then interruption is a behavior, and long runs tend to lose behaviors. Suppose instead the interrupt is part of the loop the agent runs inside. Then it works in any state the agent is in. So the control belongs in the harness, the program that runs the agent's loop, and not in the prompt. Takerngsaksiri and colleagues built an applied version at Atlassian. Their system, HULA, puts software engineers in the loop of an LLM agent. The engineers refine and guide the coding plan and the source code, and they review each step. The authors report that HULA is deployed internally on Jira.[5](#user-content-fn-5) The useful finding is where the human comes in. The human enters at set points in the agent's process. The human does not have to win an argument with the agent. ## A steer is a message with a delivery mode In Oxagen, a correction you send to a running agent is called a **steer**. A steer is free text sent to a run. It enters the run at a model request. Nothing else in a run can receive text, and the surface is kept small on purpose. Oxagen never carries out a steer as a command. The steer is content that reaches the model. Each harness decides for itself whether to treat that content as an instruction. A steer differs from a chat box in one way. The sender picks how soon the steer lands, and each choice has a stated cost. > **Three ways a steer reaches a running agent** > > 1. **Next step**: The default. The steer goes out with the next model request. Nothing already running is disturbed, and nothing extra is billed. > 2. **Turn boundary**: The steer waits for the current turn to end. The agent finishes its current thought, then reads the steer. > 3. **Interrupt**: Stops the response that is streaming. The partial tokens are billed. A pending tool call is dropped only if it can be undone. > > *The delivery modes an operator can choose. If a pending tool call cannot be undone, an interrupt waits for the next step instead. It does not drop a call whose effect has already happened.* Think about an interrupt that arrives while the agent is halfway through a payment, a deploy, or a delete. Stopping there does not undo the action. It leaves the action half done. So when the pending call cannot be undone, the interrupt waits for the next step boundary. This keeps the control from causing the incident it was meant to stop. ## Each steer has a recorded status A chat box also gives you no status. A steer has one. > **The states a steer moves through** > > 1. **Queued**: Accepted, waiting for its delivery point. > 2. **Sent**: Sent to the run. > 3. **Received**: The run has it. > 4. **Acknowledged**: It reached a model request. > 5. **Applied**: The model read it in a turn the record names. > > *The path a delivered steer takes. A steer can also be cancelled, expire, or fail. Oxagen records each of those as a state too.* The list of states is fixed, and Oxagen records every state. People usually guess whether a message got through. With a steer, they can look it up. The record names the model request each steer landed on. You can run a query to see whether anyone corrected a run before it went wrong. A steer that expired before delivery shows as expired. It does not look the same as a steer the model read and ignored. Oxagen records each steer with the operator who sent it. The steer is stored as a frame, one entry in the run's record, next to the model requests and tool calls. This keeps what the operator said on the same timeline as what the agent did. That is the lasting part of what watching a run used to give you. ## Operators and other agents send messages with different authority Once more than one agent is running, the channel for corrections can carry instructions from anyone who can reach it. Oxagen tells the two kinds of message apart by who sent them. It does not judge the content. > **Two kinds of message that arrive through the same channel, with different authority** > > - An operator steers A running agent: Enters with operator authority; Recorded as a frame, with the sender's name > - Another agent messages A running agent: Enters quoted and marked as untrusted (tainted); Evidence the model may weigh, never an instruction > > *One channel, two levels of authority. A message from another agent is content to consider. The record says where it came from.* With one agent running, this split costs nothing. With several, it becomes the main design choice. If one agent can instruct another, then someone can use the first agent to instruct the second. The text that does it will come from a web page, a ticket, or a file someone else wrote. So Oxagen marks where the message came from and does not promote it to an instruction. This does not filter the content. It records what kind of message it is. That decision keeps working even when the content is convincing. > **A steer is not a check** > > A steer is outside input. Huang and colleagues found that self-correction needs outside input,[6](#user-content-fn-6) but a steer is still not a verdict. A test suite, a database state, or a named person has to say whether the run ended up right. A steer changes where the agent is going. Something outside the model still has to say whether it got there. ## What steering replaces The first post in this pillar described a loop: assign, watch, correct, review. Watching was not the useful part. It delivered two useful things. You could catch a wrong turn early. You could also say something the run would act on. When a run lasts a week, both have to come from something other than a person's attention. Checks that run during the work, not after it, catch wrong turns early. A channel with a delivery mode, a status, and an owner carries messages the run will act on. Neither is as good as a senior engineer sitting beside the agent for a week. Both keep working while the agent runs for a week and the engineer does other work. Teams make this trade whether or not they write it down. ## Where this fits in Oxagen Oxagen is workforce management for autonomous agents. Each agent works under a mandate. Steering sits in the mandate's equipment clause, beside the tools and skills the agent may use and the business context it may read. The rules that decide which requests stop and wait for a person sit in the budget and rules clause. Both are set before a run starts and applied while it runs. For actions routed through Oxagen, this gives an operator three things at hour nine. First, a rule may already have stopped the request that started the rewrite and routed it to a named person. Second, a steer can reach the run at the next model request, at the next turn, or right away. Each choice has a stated cost, and an interrupt waits if the pending call cannot be undone. Third, the record holds the request, the rule that answered it, the steer, its sender, and the turn it landed on. So the next time this happens, you start from a query instead of nine hours of reading. ## Footnotes 1. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). *LLMs Get Lost In Multi-Turn Conversation*. arXiv. [https://arxiv.org/abs/2505.06120](https://arxiv.org/abs/2505.06120) [↩](#user-content-fnref-1) 2. Sinha, A., Arun, A., Goel, S., Staab, S., & Geiping, J. (2025). *The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs*. arXiv. [https://arxiv.org/abs/2509.09677](https://arxiv.org/abs/2509.09677) [↩](#user-content-fnref-2) 3. Backlund, A., & Petersson, L. (2025). *Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents*. arXiv. [https://arxiv.org/abs/2502.15840](https://arxiv.org/abs/2502.15840) [↩](#user-content-fnref-3) 4. Hadfield-Menell, D., Dragan, A., Abbeel, P., & Russell, S. (2016). *The Off-Switch Game*. IJCAI 2017. [https://arxiv.org/abs/1611.08219](https://arxiv.org/abs/1611.08219) [↩](#user-content-fnref-4) 5. Takerngsaksiri, W., Pasuksmit, J., Thongtanunam, P., Tantithamthavorn, C., Zhang, R., Jiang, F., Li, J., Cook, E., Chen, K., & Wu, M. (2024). *Human-In-the-Loop Software Development Agents*. arXiv. [https://arxiv.org/abs/2411.12924](https://arxiv.org/abs/2411.12924) [↩](#user-content-fnref-5) 6. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-6) --- # The agent time horizon is doubling > METR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision. Published 2026-09-16 by Oxagen Research. Canonical page: https://oxagen.sh/blog/the-agent-time-horizon-is-doubling You give an agent a ticket and go get coffee. Ten minutes is about the right length for a run like that. The agent has time to do something useful, and you can read the whole transcript when you come back. Because runs are this short, most teams still treat an agent as a faster autocomplete. They watch it because watching costs little. Runs are getting longer, and the measurements are public. When a run lasts a day, no person reads the whole transcript anymore. When a run lasts a week, the habits built around watching have already stopped working, whether or not anyone replaced them. ## What METR measures Kwa and colleagues proposed a metric that turns a benchmark score into something an operator can use. It is called the 50 percent task-completion time horizon. Take the tasks a model finishes with a 50 percent success rate. The horizon is how long a human with the right domain skills takes on those tasks.[1](#user-content-fn-1) The authors timed people on RE-Bench, HCAST, and 66 shorter tasks they wrote for the study. Then they scored models on the same set. They found two things. When they wrote the paper, frontier models had a 50 percent horizon of about 50 minutes. The frontier horizon had also doubled about every seven months since 2019. METR revised the dataset in January 2026.[2](#user-content-fn-2) The task suite grew to 228 tasks. The number of tasks that take eight hours or more went from 14 to 31. The revised trend puts the doubling time at 196.5 days across the whole period. For models released since 2023, it is 130.8 days. The highest measured 50 percent horizon in that release was 320 minutes. Its confidence interval runs from 170 to 729 minutes. METR states the limits of these numbers plainly. Human baseline times were measured for only 5 of the 31 long tasks. The rest are estimates. The confidence intervals are wide. The public tracker notes that measurements above 16 hours are unreliable with the current task suite.[3](#user-content-fn-3) That upper limit matters most. No one has a reliable measurement of an agent that works for a week, because no one has a task suite that can grade one. ## What the doubling times add up to A doubling time is a small number, and it is easy to misjudge in planning. Three doublings sounds small, but it is a factor of eight. > **Doublings from today's highest measured 50 percent horizon** > > 1. **About 5 hours**: The highest measured 50 percent horizon in January 2026 (320 minutes). > 2. **A working day**: Less than one doubling away. About 7 months at 196 days, about 4 at 131. > 3. **A working week**: About three doublings away. About 19 months at 196 days, about 13 at 131. > 4. **A month of work**: About five doublings away. About 32 months at 196 days, about 22 at 131. > > *Arithmetic on the doubling times METR reports (Time Horizon 1.1, January 2026), starting from its highest measured point. For illustration only. A trend line is not a forecast, and METR states that its measurements above 16 hours are unreliable.* Read the table as arithmetic, not as a roadmap. A trend that held for seven years can stop in the eighth. The confidence intervals on the recent points are wide enough to change every row. The problem of grading long tasks will also arrive before agents can do them. The table helps with one thing. It shows how much of your current practice depends on runs staying short. If most of it does, your plans look less far ahead than the capability trend does. ## Production needs more than a 50 percent success rate If you ship the result, half is an odd place to draw the line. It is the right place for a research metric. That part of the curve is the steepest, so it shows a change in the model most clearly. It is the wrong place for a deployment decision. There, you are deciding whether to let a run finish while you sleep. Yao and colleagues made the same point with a different metric. The usual metric is pass@k, the chance that at least one of k attempts succeeds. They define pass^k instead, the chance that all k independent attempts at a task succeed.[4](#user-content-fn-4) In their study, the best function-calling agents succeeded on under 50 percent of tasks. In the retail domain, pass^8 fell below 25 percent. More attempts raise pass@k and lower pass^k. So a system tuned on pass@k can get less consistent while its headline score improves. A five-hour horizon at 50 percent success means the agent gets a five-hour task right about half the time. Run one such task a day for a working week. The chance that all five succeed is about 3 percent. This is not a reason to avoid long runs. It shows that the horizon number measures what an agent can reach. It does not measure what you can ship, and those are separate questions. ## Long runs fail differently from short runs You might expect long runs to fail because the context window fills up. The measurements do not support that. Backlund and Petersson built Vending-Bench to test whether an agent stays coherent over time, rather than what it can do. In the benchmark, an agent runs a vending machine business. It manages inventory, places orders, sets prices, and pays daily costs.[5](#user-content-fn-5) Each task is simple, but the run is long. Single runs go past 20 million tokens. The agents misread delivery schedules and forget orders they placed. They also fall into what the authors call tangential meltdown loops, and they rarely recover. The authors found no clear link between failures and the point where the context window fills. Runs of the same model on the same task also varied a lot. Sinha and colleagues describe a cause that fits. They studied long tasks and found a self-conditioning effect. When a model's context already holds its own errors from earlier turns, the model becomes more likely to make new mistakes.[6](#user-content-fn-6) The agent reads its own bad work from earlier and treats it as settled. The same paper also has good news. Small gains in single-step accuracy add up to large gains in how long a task a model can complete. That is why the horizon keeps growing. Together, these findings show how a long run goes wrong. It does not get slowly worse toward the end. It works correctly, then takes one wrong turn. It writes that turn into its own record. Then it spends hours staying consistent with it. By the time a person looks, the wrong turn is forty steps back, and everything after it fits with it. ## Watching does not scale to long runs Most teams supervise an agent the way they supervise a new engineer. That is the process they already had. > **The supervision loop most teams are running** > > 1. **Assign**: A ticket, a prompt, a branch. > 2. **Watch**: Read the transcript as it streams. > 3. **Correct**: Interrupt, re-prompt, restart. > 4. **Review**: Read the diff, approve or reject. > > Repeat Repeat per task, while someone is awake. > > *Illustrative. The loop works because a person is there for the middle two steps. When a run takes a week, nobody is.* The loop works because a person is there for the middle two steps. Everything the team relies on comes from that. The person catches the wrong turn early. The correction is cheap. The reviewer has watched enough of the run to know what the diff means. None of this holds when a run is longer than the reviewer can sit through. The loop does not fail in an obvious way. It keeps going with the watching step skipped. The reviewer then reads a week of work with no idea which forty steps mattered. Instead of watching, you decide in advance what you used to decide in the moment. You decide what the agent may reach, what it may spend, what it is told before it starts, and which requests stop and wait for a person. These are the same decisions, moved from the transcript to the mandate. The next posts in this pillar cover them one at a time. ## Where this fits in Oxagen Oxagen is workforce management for autonomous agents. It does not run them. It holds each agent's mandate. The mandate covers the identity the agent acts as, the systems and data it may request, the budget and rules it works under, the tools and skills it is equipped with, and the record of what it did. Three parts of the mandate matter most for longer runs. First, Oxagen answers each request at the moment of use. A run that asks for a system at hour forty gets the same rule it would have got at minute one. A rule can allow the request, deny it, or route it to a named person while the run waits. Second, Oxagen prices each governed action. It records the cost against the person, the agent, the run, the turn, and the step. So a week-long run has a cost you can read step by step, not one line on a monthly bill. Third, the record keeps each run as a series of frames, one per recorded event, next to its mandate. Finding the wrong turn becomes a query instead of a reread. This applies to actions routed through Oxagen. Oxagen does not govern a call that does not pass through it. No record makes an agent correct. The record lets a person inspect a long run afterward. The watching step used to give you that. ## Footnotes 1. Kwa, T., West, B., Becker, J., Deng, A., Garcia, K., Hasin, M., Jawhar, S., Kinniment, M., Rush, N., Von Arx, S., Bloom, R., Broadley, T., Du, H., Goodrich, B., Jurkovic, N., Miles, L. H., Nix, S., Lin, T., Painter, C., Parikh, N., Rein, D., Sato, L. J. K., Wijk, H., Ziegler, D. M., Barnes, E., & Chan, L. (2025). *Measuring AI Ability to Complete Long Tasks*. arXiv. [https://arxiv.org/abs/2503.14499](https://arxiv.org/abs/2503.14499) [↩](#user-content-fnref-1) 2. METR (2026). *Time Horizon 1.1*. [https://metr.org/blog/2026-1-29-time-horizon-1-1/](https://metr.org/blog/2026-1-29-time-horizon-1-1/) [↩](#user-content-fnref-2) 3. METR. *Task-Completion Time Horizons of Frontier AI Models*. [https://metr.org/time-horizons/](https://metr.org/time-horizons/) [↩](#user-content-fnref-3) 4. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). *tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains*. arXiv. [https://arxiv.org/abs/2406.12045](https://arxiv.org/abs/2406.12045) [↩](#user-content-fnref-4) 5. Backlund, A., & Petersson, L. (2025). *Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents*. arXiv. [https://arxiv.org/abs/2502.15840](https://arxiv.org/abs/2502.15840) [↩](#user-content-fnref-5) 6. Sinha, A., Arun, A., Goel, S., Staab, S., & Geiping, J. (2025). *The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs*. arXiv. [https://arxiv.org/abs/2509.09677](https://arxiv.org/abs/2509.09677) [↩](#user-content-fnref-6) --- # The problem is not slop, it is your process > The worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work. Published 2026-09-16 by Oxagen Research. Canonical page: https://oxagen.sh/blog/the-problem-is-not-slop-it-is-your-process People who complain about AI slop are complaining about output. They mean low-quality pull requests, documentation nobody asked for, and long, plausible text with a mistake in the middle. The complaint is fair. But the published measurements do not find the cost in the output. So output quality is the wrong thing to organise a team around. The measurements find the cost in the process around the output. Teams take a workflow built for people and give it to an agent. They keep the parts that only worked because a person was doing the work. The agent then fails in exactly the places where the workflow assumed a person. ## In one study, every forecast was wrong in the same direction Becker and colleagues ran a randomised controlled trial with 16 experienced open-source developers across 246 tasks. Each developer worked on a repository where they had about five years of prior experience.[1](#user-content-fn-1) Half the tasks allowed AI tools, and half did not. Developers using AI took 19 percent longer to complete tasks. The forecasts stand out. Before the study, the developers expected AI to cut completion time by 24 percent. Economists asked to predict the result said 39 percent. Machine-learning experts said 38 percent. After the tasks, the developers had been slower. They still estimated that AI had made them 20 percent faster. > **Forecast change in task completion time, against the measured change** > > | | Forecast | Measured | > | --- | --- | --- | > | Economists | \-39% | 19% | > | Machine-learning experts | \-38% | 19% | > | The developers, beforehand | \-24% | 19% | > | The developers, afterwards | \-20% | 19% | > > *Source: Becker et al. (2025), 16 experienced open-source developers over 246 randomised tasks. Negative is faster. Every forecast was on the wrong side of zero, including the one made after the work.* None of the people in that chart describe slop. The developers accepted the model's output. They did not report being buried in bad output. The time went to other things. It went to prompting, reading, waiting, and small corrections that each seem minor but add up. The developers' own reports also stayed positive after the work. Teams usually decide whether a practice works by asking the people who do it. Here, asking would have given the opposite of the measurement. ## Bigger batches made delivery worse DORA's 2024 report found the same pattern across whole organisations. In their survey population, a 25 percent increase in AI adoption was linked to an estimated 1.5 percent drop in delivery throughput. It was also linked to an estimated 7.2 percent drop in delivery stability. This was the second year in a row that AI adoption was linked to worse delivery performance.[2](#user-content-fn-2) In the same survey, about three quarters of respondents reported productivity gains. DORA points to batch size as the cause. AI makes code cheaper to write, so changesets get bigger. Bigger changesets carry more risk through a review and release process tuned for smaller ones. The generated code does not have to be bad for this to happen. The team only has to keep a process that was safe because batches were human-sized. Then AI removes the limit that kept batches that size. Put simply, the old workflow had a hidden limit, which was how much one person could do. The agent removed that limit, and nobody replaced it. ## Habits that work between people fail with agents Most of a working software process is unwritten. Most of the unwritten part is about how people handle vague requests. You ask a colleague for something vague. They ask two questions, and you agree on what you meant. You send them a long document and say the answer is in there, and they find it. Halfway through, they notice they have been building the wrong thing, and they back out. None of that is in the ticket, but the process depends on all of it. Research has measured each of these habits with models. > **Four habits a team hands to an agent, and what each one assumes** > > - Clarify as we go assumes Recovery from a wrong turn: 39 percent average drop, multi-turn against single-turn; Laban et al. (2025) > - It is in the spec assumes Even attention across a long document: Accuracy is highest at the start and end, lowest in the middle; Liu et al. (2023) > - Double-check your work assumes Useful self-review: Performance can degrade after unaided self-correction; Huang et al. (2023) > - You will notice the mistake assumes Error does not become evidence: Errors already in context raise the rate of later errors; Sinha et al. (2025) > > *Each row pairs a habit that works between people with a published finding that the habit fails when handed to an agent.* The first row does the most damage. Laban and colleagues gave models a fully specified instruction all at once. They also gave the same instruction spread across a conversation. The conversation led to an average drop of 39 percent across six generation tasks.[3](#user-content-fn-3) The authors split the drop into a small loss in skill and a large rise in unreliability. They describe the cause. Models make assumptions in early turns and try a final answer too soon. Then they rely on that answer too much. So the next message does not fix an early wrong turn. The model builds on it. The other three findings add to the problem. Liu and colleagues showed that a model is most accurate on a long input when the relevant passage is near the start or the end. Accuracy drops for passages in the middle.[4](#user-content-fn-4) So a key constraint on page nine of the spec may not reach the model. Huang and colleagues found that models struggle to correct their own reasoning without outside feedback. Self-correction without help sometimes made the answer worse.[5](#user-content-fn-5) So asking the agent to check itself does not work as a check. Sinha and colleagues describe self-conditioning. When a model's own earlier errors sit in its context, it becomes more likely to make new errors.[6](#user-content-fn-6) So the longer the run, the more the model treats an early mistake as fact. ## Slop comes from a missing check Output quality still matters. It depends on the process that produces it. A team gets slop when it has no check outside the model and no limit on batch size. Take a process whose only real check was a person reading everything. Remove the person as the limit, and the result is more output than anyone can evaluate. Most teams respond by reviewing harder. That means more review, by the same people, on more material. That is the part that already did not scale. The measurements support a different response, and a duller one. Move the check off the model and off the reviewer's patience, and onto something that runs. Write the limit down, instead of leaving it to someone's attention. State the standing rules once, in a place the agent reads at the start of every run. Do not state them in the middle of a conversation the agent will mishandle. When the run is going wrong, change the run instead of arguing with it. This pillar works through that response as four practices. 1. **State the standing rules once, in a lasting place.** Do not put them in a prompt that is rewritten each time, or deep in a document the model reads unevenly. The rules an agent works under should be records with an owner, a scope, and a date. Then you can look up what the agent was told instead of piecing it together. 2. **Correct the run itself.** A correction sent as another chat turn meets the same multi-turn failure it is trying to fix. So change what the run is operating under instead. 3. **Check with something outside the model.** A test suite, a schema, a database state, or a named person can do this. The check has to return a verdict the agent cannot write itself. 4. **Record each step.** When a week-long run goes wrong, you need to know which step went wrong. Either you query a record, or a person rereads the transcript. ## Where this fits in Oxagen Oxagen is workforce management for autonomous agents. It holds the mandate each agent works under. The mandate covers the identity the agent acts as, the systems and data it may request, the budget and rules it works under, the tools and skills it is equipped with, and the record of what it did. The first two practices above map to two of those clauses. The business context an agent may read and the steering it runs under belong to the equipment clause. The engineers accountable for the agent set it. Approval thresholds and decision rules belong to the budget and rules clause. The people who own the spend and the systems set it. Both are written before a run starts. Oxagen applies them to governed requests as the run makes them. A rule that answers at the moment of use works even if the agent did not read it carefully. It also works when no person is awake to restate it. The next two posts cover the first two practices. One covers what an agent should be told before it starts. The other covers what to do when a run you are not watching starts going the wrong way. ## Footnotes 1. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). *Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity*. arXiv. [https://arxiv.org/abs/2507.09089](https://arxiv.org/abs/2507.09089) [↩](#user-content-fnref-1) 2. DORA (2024). *Accelerate State of DevOps Report 2024*. Google Cloud. [https://dora.dev/research/2024/dora-report/](https://dora.dev/research/2024/dora-report/) [↩](#user-content-fnref-2) 3. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). *LLMs Get Lost In Multi-Turn Conversation*. arXiv. [https://arxiv.org/abs/2505.06120](https://arxiv.org/abs/2505.06120) [↩](#user-content-fnref-3) 4. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). *Lost in the Middle: How Language Models Use Long Contexts*. TACL. [https://arxiv.org/abs/2307.03172](https://arxiv.org/abs/2307.03172) [↩](#user-content-fnref-4) 5. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-5) 6. Sinha, A., Arun, A., Goel, S., Staab, S., & Geiping, J. (2025). *The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs*. arXiv. [https://arxiv.org/abs/2509.09677](https://arxiv.org/abs/2509.09677) [↩](#user-content-fnref-6) --- # What an agent should be told before it starts > A prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things. Published 2026-09-16 by Oxagen Research. Canonical page: https://oxagen.sh/blog/what-an-agent-should-be-told-before-it-starts Most teams that run agents end up with one instruction file. It starts as four lines about the test command. Then it gains a section on which directory things go in. Someone adds a paragraph at 2am after an incident. Within a quarter the file is 600 lines long, and nobody has read all of it since March. The file does real work, but its shape is wrong for that work. Its author is whoever edited it last. It has no scope and no dates. It cannot say that rule three applies to one repository and rule nine applies to everyone. It also keeps no record of whether the agent read line 412 during its four hours in your codebase. So when something goes wrong, you cannot say what the agent was told. You can only show what the file says now. ## Why a longer prompt does not help The obvious fix is to write more into the file. Two research findings rule that out. Liu and colleagues measured how models use long inputs.[1](#user-content-fn-1) Accuracy was highest when the relevant passage sat near the start or the end of the input. It dropped when the passage sat in the middle. This U-shaped pattern held even for models built for long inputs. In your file, a rule's position depends on when someone added it. So how well a rule works depends on when it was written, not on how much it matters. The second finding rules out the other fix, which is to repeat the rule during the run. Laban and colleagues gave models a full instruction all at once. They also gave the same instruction in pieces across a conversation. Across six generation tasks, the split version scored 39 percent lower on average.[2](#user-content-fn-2) Most of the drop came from less consistent answers, not from lost skill. The models made an assumption early and then relied on it too much. So telling an agent at hour six what it needed at hour zero is the worst of the options. This means the standing rules must be in place before the run starts. They must also be put in a chosen position, not added to the end. A plain text file cannot do that. ## What the memory research found Research on agent memory agreed on one idea some time ago. An agent's context should be built by picking the right pieces for each step, not pasted in whole. Sumers and colleagues proposed CoALA, a framework based on cognitive architectures, which are older models of how a mind is organised.[3](#user-content-fn-3) CoALA describes a language agent as separate memory parts plus a set of actions. Some actions work on the agent's memory, and some work on the outside world. The specific parts matter less than the main idea. An agent knows several kinds of things, and each kind lasts a different length of time. If you treat them all as one long transcript, you lose the differences that decide what the agent should keep in view. Packer and colleagues built MemGPT on the same idea.[4](#user-content-fn-4) They borrowed tiered memory from operating systems. The system keeps the right subset of memory in the limited context window and moves items in and out as needed. Park and colleagues stored a full record of an agent's experiences in plain language.[5](#user-content-fn-5) Their system wrote higher-level reflections from that record. Then it retrieved the relevant pieces when the agent planned, instead of keeping everything in view. In each case, a component with rules decides what the agent sees on each turn. A person with a text editor does not. Wang and colleagues showed what stored knowledge is worth when it builds up over time. Their agent, Voyager, explores by following a curriculum of tasks.[6](#user-content-fn-6) It saves the skills it learns as code in a growing library, and it retrieves them later. > **Voyager with a saved skill library, compared with earlier agents** > > | Measure | Improvement | > | --- | --- | > | Unique items collected | 3.3x | > | Distance travelled | 2.3x | > | Tech tree milestones unlocked | up to 15.3x faster | > > *Source: Wang et al. (2023), Voyager in Minecraft compared with earlier state-of-the-art agents. The authors also report that the learned skill library transfers to new worlds, where competing approaches struggled.* The transfer result matters most here. The library worked in worlds it was not built in. Knowledge saved as a named item that can be found again still helps when the task changes. Knowledge kept in one run's prompt does not. ## What a rule needs Go through the file line by line. For each line, ask what it would need to count as a rule. It needs an owner, so that someone answers for it. It needs a scope, because a rule about the staging database is not a rule for every repository in the organisation. It needs a start date and a way to end, because much of the file describes systems that no longer exist. It needs a source, because the incident that caused a rule is often the most useful thing to know about it. Last, it needs a way to take effect that is more than one person editing a line. If an agent could rewrite a rule, the rule would not bind the agent. In Oxagen, that object is a **steering record**. A record holds one thing that an agent or a person states, learns, proposes, or decides. Once written, a record never changes. Oxagen computes a hash, a short fingerprint of the content, from a standard form of the record. So two copies are either the same record or visibly different records. Each record has a scope, from one user up to the whole organisation. Each record also points back to the frame, the recorded event in a run, where it was learned. So the rule and the incident behind it are one link apart. > **Three records over the same repository, read as of today** > > - **Migrations run against the shared plane** (Jan to Jun) > - **Each organisation resolves its own store** (Jun onward) > - **Release notes wait for a named approver** (Mar onward) > > *Illustrative. A record has a start date and can be retracted. So you can ask what the agent was told in April separately from what it would be told now.* The first two rows contradict each other, and each was correct at its own time. A text file cannot hold both. It holds only the newest line. When someone edits that line, the record of what the agent worked under in April is lost. ## An agent proposes a record and a person merges it Records work as governance, and not only as storage, because of where they live and how they take effect. A workspace keeps its records as files in its linked repository. A record takes effect only through a pull request. An agent that learns something during a run can propose a record, but it cannot put one into effect. The proposal opens a pull request. Checks run against it. A person reviews it under the workspace's review mode. The merge puts the record into effect. So Git decides which records are active. "Which rules were active on the 14th" has the same answer as "what was in the tree on the 14th." Your existing tools already answer that question. > **How a proposed record becomes one the agent runs under** > > 1. **Propose**: An agent or a person writes a record, with its scope and its source. > 2. **Check**: Oxagen checks the format, the link to the source, and the hash. It scans for secrets and for conflicts with active records. > 3. **Review**: A person approves under the workspace's review mode. > 4. **Merge**: The commit puts the record into effect. > 5. **Promote**: Oxagen writes a promotion event to a ledger. Each entry carries the hash of the entry before it. > 6. **Deliver**: Oxagen adds the record to the context of the next run. > > *The steering PR path. Without the review step, a record would be only a note an agent wrote to itself.* The review step answers a problem this pillar keeps returning to. It is the self-conditioning problem, where an agent builds on its own earlier mistakes. If an agent could write its own standing rules, one wrong conclusion at hour three would become a rule for every later run. The merge requirement does not make the agent's proposals better. It lets a person review them while they are still proposals. ## How records reach a run Oxagen delivers records in a way that deals with the position problem directly. Records that must hold go at the front of the stable part of the system context. The research above found that models read the start of a long input best. Oxagen picks lower-priority records for each run by relevance instead of sending all of them. A budget that lets everything in does not limit anything. Some records arrive between turns. Oxagen adds each of these as a separate message right after the cached system block, and does not edit it into that block. This keeps the cached start of the context the same across runs. A record can also be delivered as a context frame that points back to the record and to the commit that made it active. So an answer can cite the rule it came from. > **Why a record's position is chosen** > > Liu and colleagues found accuracy highest at the start and end of a long input and lowest in the middle. If your standing rules are a file, their position is the order they were written in. If they are records with a priority, the program that builds the context decides their position. It puts the rule that must hold where the model reads best. Two limits follow. First, delivering a record does not make the model obey it. Nothing here makes a model comply. The check that the work was right still has to come from outside the model. Huang and colleagues found that when models corrected themselves without outside help, their answers sometimes got worse.[7](#user-content-fn-7) Second, a record is only as good as its scope. Someone may promote a rule to the whole organisation because that was easier than scoping it. That rule will be wrong somewhere. ## Where this meets Oxagen Oxagen is workforce management for autonomous agents. Steering records belong to the equipment clause of an agent's mandate. That clause covers the business context the agent may read and the steering it runs under. The engineers accountable for the agent set it. The scope on a record and the scope in the mandate use the same mechanism. So an agent gets what its work needs, not everything the organisation knows. For an operator, this means "what was this agent told" has an answer. When a run goes wrong, the answer is a set of records. Each one has an owner, a scope, a hash, and the commit that made it active. Without records, the answer is a file edited over time by whoever was on call. When an agent learns something worth keeping, it proposes a record and does not decide alone. When a rule turns out to be wrong, you retract it with a pull request, so the retraction has a date too. The next post covers the other half: what to do when a run has started and is going the wrong way. ## Footnotes 1. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). *Lost in the Middle: How Language Models Use Long Contexts*. TACL. [https://arxiv.org/abs/2307.03172](https://arxiv.org/abs/2307.03172) [↩](#user-content-fnref-1) 2. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). *LLMs Get Lost In Multi-Turn Conversation*. arXiv. [https://arxiv.org/abs/2505.06120](https://arxiv.org/abs/2505.06120) [↩](#user-content-fnref-2) 3. Sumers, T. R., Yao, S., Narasimhan, K., & Griffiths, T. L. (2023). *Cognitive Architectures for Language Agents*. TMLR. [https://arxiv.org/abs/2309.02427](https://arxiv.org/abs/2309.02427) [↩](#user-content-fnref-3) 4. Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2023). *MemGPT: Towards LLMs as Operating Systems*. arXiv. [https://arxiv.org/abs/2310.08560](https://arxiv.org/abs/2310.08560) [↩](#user-content-fnref-4) 5. Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). *Generative Agents: Interactive Simulacra of Human Behavior*. UIST 2023. [https://arxiv.org/abs/2304.03442](https://arxiv.org/abs/2304.03442) [↩](#user-content-fnref-5) 6. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). *Voyager: An Open-Ended Embodied Agent with Large Language Models*. TMLR. [https://arxiv.org/abs/2305.16291](https://arxiv.org/abs/2305.16291) [↩](#user-content-fnref-6) 7. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-7) --- # A model cannot grade its own homework > Without outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/the-limits-of-self-correction-and-model-collapse Many agent stacks add a second prompt that asks the model to check its work. It is the cheapest reliability fix there is. It costs one extra call and needs no infrastructure. In a demo it seems to work. So teams ship it, mark the reliability problem solved, and move on. The research shows this fix does less than teams expect. Under some conditions it makes things worse. The same result appears twice, once when the model answers and once when it trains. When a model checks its own answer, the answer does not reliably get better. When a model trains on its own output, the model does not reliably get better either. Both fail for the same reason. The loop adds no information the model did not already have. This post covers what the evidence shows and what works instead. ## Later tests found that self-correction fails on reasoning Self-Refine, published in early 2023, is where the pattern comes from.[1](#user-content-fn-1) One model writes an output, gives itself feedback, and revises. It uses no extra training, no supervised data, and no reinforcement learning. The authors report a gain of about 20% absolute on average across their tasks. Both humans and automatic metrics preferred the revised outputs over one-shot outputs from the same model. That result is real, and many readers took it as a general one. Later in 2023, Huang and colleagues at Google DeepMind tested the case that matters most for agents, which is reasoning.[2](#user-content-fn-2) They call the pattern intrinsic self-correction. The model revises "based solely on its inherent capabilities, without the crutch of external feedback." Their finding is direct. Without outside feedback, models struggle to correct their own reasoning. Sometimes performance gets worse after self-correction. The two papers do not contradict each other. They differ in what the feedback can check. Self-Refine did best on tasks where the feedback names a concrete defect that can be checked. Examples are code that can be run, or an output with a property that can be measured. Reasoning is different. If the model got the reasoning wrong, it would have to notice the error with the same ability that made it. So the practical lesson is narrow. Self-critique can improve format and style. It cannot confirm that an answer is correct. If your agent's retry loop has no source of truth, the retry gives you a second sample. It does not give you a check. For agents, this is easy to act on. Most agent tasks touch something that returns a result. A test suite returns pass or fail. A type checker returns errors with line numbers. An HTTP call returns a status code. A database returns rows or an exception. Each of these is an outside verifier, a check that is not the model. It already exists in the environment and costs nothing to use. A persuasive argument cannot change its answer. Some agent setups have the model review its output before running the tests. That is the wrong order. Run the check first, and give the real error back to the model. Then the critique step has information the model did not have when it wrote the code. This is the difference between Self-Refine's strong tasks and its weak ones, written as an engineering rule. > **Check first, then critique** > > 1. **Write**: Code, a query, an API call > 2. **Run the check**: Tests, the type checker, a status code > 3. **Feed back**: The real error, with line numbers > 4. **Revise**: The critique now has new information > > Repeat until the check passes. > > *A critique before step 2 only produces a second sample. A critique after step 2 has information the model did not have when it wrote the code.* ## Training a model on its own output makes it worse At training time the same problem is worse, because the damage adds up over generations. Shumailov and colleagues named this model collapse.[3](#user-content-fn-3) Suppose you train each generation of models on data made by the previous generation. The tails of the original distribution, the rare cases, disappear first. Rare events stop showing up, then uncommon ones. The damage cannot be undone, so later generations cannot recover what earlier ones lost. The effect is not limited to language models. The authors showed it in variational autoencoders and Gaussian mixture models too. That suggests the cause is the recursive setup, not any one architecture. The peer-reviewed version in Nature adds a finding for anyone who scrapes the web for training data. Training on a mix of real and generated content, without choosing what goes in, leads to collapse. The model loses its ability to produce diverse, high-quality output. Language models trained this way over many rounds produce likely sequences too often. In the end they produce text no human would write.[4](#user-content-fn-4) Alemohammad and colleagues found the same thing in image generation, and they stated the condition more precisely.[5](#user-content-fn-5) They call the failure Model Autophagy Disorder. Autophagy means self-eating. Their finding is that "without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease." The quote says quality or diversity. The model does not always get worse in every way. It can keep its quality by narrowing, and produce a smaller set of more polished outputs. This failure is the hardest one to see in a spot check. Each sample looks fine, but the range of outputs behind the samples has shrunk. ## Keeping the real data prevents collapse The warnings about collapse were loud. So a 2024 paper tested whether collapse is inevitable. The answer depended on one detail of the experiment.[6](#user-content-fn-6) The collapse experiments replace the training data each generation. Gerstgrasser and colleagues tested accumulation instead. They kept the original real data and added each generation of synthetic data next to it. Collapse did not occur. The result held across several model families and datasets. > **Two retention policies across four generations** > > - **Replace**: The tails vanish first, then collapse > - **Accumulate**: Collapse does not occur > > *Illustrative. Blocks show what each generation trains on, not dataset sizes. Replacement follows Shumailov et al., and accumulation follows Gerstgrasser et al. (2024).* This is the easiest finding in this research to act on, and it gets the least attention. Synthetic data does not cause collapse by itself. Replacing your real data with synthetic data does. The fix is a retention policy, a rule about which data you keep. It needs no new algorithm. Any team that feeds model output back into training can act on it this quarter. ## Checks from outside the model do work Self-assessment fails, and training on your own output degrades the model. What remains is a signal from outside the model. The research on verifiers, separate checkers that grade a model's output, shows real gains. That research is older than the current wave of models. Cobbe and colleagues trained a separate verifier to rank sampled solutions to grade-school math problems. They used GSM8K, a set of 8.5 thousand problems.[7](#user-content-fn-7) As they added data, verification improved faster than the fine-tuning baseline. They sampled many solutions and had a second model pick one. That beat training one model to produce a better first attempt. Lightman and colleagues improved the idea in 2023 by changing what the verifier grades.[8](#user-content-fn-8) Outcome supervision rewards the final answer. Process supervision rewards each step of the reasoning. Process supervision clearly did better. Their process-supervised reward model solved 78% of problems from a representative subset of the MATH test set. They released PRM800K, a set of 800 thousand human feedback labels on single steps. That set shows the real cost of the method. It takes a large amount of human judgment, applied where the model cannot supply it. DeepSeek-R1 applies the same idea at a larger scale.[9](#user-content-fn-9) Its reasoning behavior came from large-scale reinforcement learning, with no supervised fine-tuning first. The training used tasks where an automatic checker can decide whether an answer is right. So the reward came from a fixed procedure, not from a model's opinion. | Signal | Where it comes from | Does it hold up at scale | | --- | --- | --- | | The model's own critique | Inside the model | No, it gets worse on reasoning tasks | | Its own output as training data | Inside the model | No, the model collapses if it replaces real data | | Its own output added to real data | Mixed | Yes, accumulation avoids collapse | | A separate learned verifier | Outside, learned | Yes, but it carries the verifier's flaws | | Step-level human labels | Outside, human | Yes, at a real cost in human labeling | | An automatic correctness checker | Outside, a fixed procedure | Yes, where the task allows one | ## Any verifier can be gamed, so the record matters Adding a verifier does not end the problem, for one more reason. Skalse and colleagues gave reward hacking a formal definition.[10](#user-content-fn-10) Reward hacking is when an optimizer raises its score without doing what the score was meant to measure. Call the score a proxy. The authors call a proxy unhackable if raising the expected proxy return can never lower the expected true return. Then they proved a strong limit. Consider all stochastic policies, which choose actions with some randomness. Across all of them, two reward functions can only be unhackable if one of them is constant. Useful unhackable pairs do exist in narrow settings, such as deterministic policies or a finite set of policies. In the general case, they do not. Your verifier is a proxy. An optimizer that pushes hard enough will find the gap between the verifier and what you meant. Pan, Bhatia, and Steinhardt showed what this looks like in experiments.[11](#user-content-fn-11) More capable agents exploit a badly specified reward more. They score higher on the proxy and lower on the true goal than weaker agents do. The authors also found phase transitions. At certain capability levels, the agent's behavior changes in kind, and the true reward drops sharply. Watching the metric you optimized will not catch this, because the change is not gradual. The system looks fine until the drop happens. Together, these results explain why a record matters. You cannot write a proxy that stays safe under any amount of optimization pressure. You can keep a record of what the agent tried, what was permitted, and what the verified outcome was. With that record, you can see the gap between the proxy and your intent when it opens, not after the phase transition. The measurement has to be separate from the thing it measures. It also has to last, because the failure is a sudden step, and you will want the history from before and after it. ## Where this fits in Oxagen Oxagen does not train models or run agents. It is the control plane for the agents you run. The results above are why its record sits outside the model. Oxagen grounds answers in a knowledge graph with citations, which brings a source from outside the model into the loop. Oxagen also records each run against the agent's mandate. The record shows what the agent asked for, which rule answered, what it cost, and what checked the outcome. This keeps the grading separate from the model being graded. Judging an agent on that record, not on its own report, is the only approach here that holds up against reward hacking. None of this makes a model better at checking itself. It gives the checking to something other than the model. ## Footnotes 1. Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., & Clark, P. (2023). *Self-Refine: Iterative Refinement with Self-Feedback*. arXiv. [https://arxiv.org/abs/2303.17651](https://arxiv.org/abs/2303.17651) [↩](#user-content-fnref-1) 2. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. arXiv. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-2) 3. Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2023). *The Curse of Recursion: Training on Generated Data Makes Models Forget*. arXiv. [https://arxiv.org/abs/2305.17493](https://arxiv.org/abs/2305.17493) [↩](#user-content-fnref-3) 4. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). *AI models collapse when trained on recursively generated data*. Nature, 631, 755-759. [https://doi.org/10.1038/s41586-024-07566-y](https://doi.org/10.1038/s41586-024-07566-y) [↩](#user-content-fnref-4) 5. Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R. G. (2023). *Self-Consuming Generative Models Go MAD*. arXiv. [https://arxiv.org/abs/2307.01850](https://arxiv.org/abs/2307.01850) [↩](#user-content-fnref-5) 6. Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., Roberts, D. A., Yang, D., Donoho, D. L., & Koyejo, S. (2024). *Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data*. arXiv. [https://arxiv.org/abs/2404.01413](https://arxiv.org/abs/2404.01413) [↩](#user-content-fnref-6) 7. Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., & Schulman, J. (2021). *Training Verifiers to Solve Math Word Problems*. arXiv. [https://arxiv.org/abs/2110.14168](https://arxiv.org/abs/2110.14168) [↩](#user-content-fnref-7) 8. Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2023). *Let's Verify Step by Step*. arXiv. [https://arxiv.org/abs/2305.20050](https://arxiv.org/abs/2305.20050) [↩](#user-content-fnref-8) 9. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) [↩](#user-content-fnref-9) 10. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). *Defining and Characterizing Reward Hacking*. arXiv. [https://arxiv.org/abs/2209.13085](https://arxiv.org/abs/2209.13085) [↩](#user-content-fnref-10) 11. Pan, A., Bhatia, K., & Steinhardt, J. (2022). *The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models*. arXiv. [https://arxiv.org/abs/2201.03544](https://arxiv.org/abs/2201.03544) [↩](#user-content-fnref-11) --- # Deterministic coding agents: every turn on the record > A coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/deterministic-coding-agents-every-turn-on-the-record An agent changed fourteen files overnight. The diff is on the branch and the tests pass. In review, someone asks why the agent changed one of those files. No one can answer. The agent's reasoning and tool calls were not saved. Only the output is left. If you run the agent again on the same ticket, it writes a different patch. So there is no earlier run to go back to, and the question has no answer at all. This problem can be solved. The fix is not to make the model deterministic, meaning it gives the same output every time. The fix is to record the run so a person can replay it and check each turn. The research on what makes coding agents work points to the same records that review needs. So one set of records can serve both. ## Coding agents improve when they get the results of running their code The results that moved code repair forward share one ingredient. A program runs the code, and the result goes back to the model. Chen and colleagues showed this with Self-Debugging. In it, a model reads the results of running its code and explains the code back to itself. They compare this to rubber duck debugging, where you explain code out loud to find the bug. It gained 2 to 3% on the Spider text-to-SQL benchmark, with a 9% gain on the hardest difficulty level. It gained up to 12% on TransCoder and MBPP when unit tests were available.[1](#user-content-fn-1) It also needed fewer tries. It matched baselines that generated over ten times as many candidates.[1](#user-content-fn-1) Xia and Zhang applied the same idea to automated program repair with ChatRepair. The older method generates a batch of patches and then tests them. ChatRepair tests each patch right away and gives the result to the model before the next try. It fixed 162 of 337 bugs at roughly $0.42 each. It learned from failed attempts as well as successful ones.[2](#user-content-fn-2) Reflexion used the pattern outside code. It stored written notes on failed attempts in memory. On HumanEval, 91% of problems passed on the first answer (pass@1), against 80% for the prior result.[3](#user-content-fn-3) Read together, these three papers show the mechanism. The model does not get smarter between attempts. It gets new information: an exit code, a stack trace, or a diff that did not apply. That information exists as concrete records at one point in time. Most agent setups throw it away as soon as the loop moves on. Those are the records a reviewer would want. > **The loop each repair result shares** > > 1. **Generate**: The model writes a candidate patch > 2. **Execute**: Run the tests against the code > 3. **Result**: Exit code, stack trace, the part of the patch that failed > 4. **Revise**: The next attempt reads the result > > Repeat repeat until the tests pass. > > *Self-Debugging, ChatRepair, and Reflexion all give the result of running the code back to the model. Most tools that run agents throw away the record from step 3.* ## Random sampling makes the output change from one run to the next Benchmark papers have reported this for five years, in the metric they use. Codex solved 28.8% of HumanEval problems with one sample (pass@1). It solved 70.2% when it could take 100 samples and count a problem solved if any passed (pass@100).[4](#user-content-fn-4) Most of the measured skill sits in the gap between those two numbers. Much of what a model can do shows up only when you sample it many times. In a benchmark, you keep the sample that passed. In your repository, you keep the sample that ran. Nothing tells you whether it was the good one. Temperature is the setting that controls how random the model's choices are. Lowering it does not fix this. Ouyang and colleagues sent the same requests many times across 829 code generation problems in three benchmarks. The share of tasks with no identical test output across the repeated requests was 75.76% on CodeContests, 51.00% on APPS, and 47.56% on HumanEval.[5](#user-content-fn-5) They also found that setting temperature to zero reduces this variation but does not remove it.[5](#user-content-fn-5) The same prompt, model, and repository can still produce a different patch. > **Tasks with no identical test output across repeated requests** > > | Benchmark | Share of tasks | > | --- | --- | > | CodeContests | 75.76% | > | APPS | 51.00% | > | HumanEval | 47.56% | > > *Source: Ouyang, Zhang, Harman, and Wang (2023), across 829 code generation problems. Setting temperature to zero reduces the variation but does not remove it.* So the variation comes with the system, and changing the model cannot engineer it away. You cannot predict which sample you will get. So record the one you got. ## Output that changes between runs makes review harder in three ways First, you cannot narrow down a decision you cannot reproduce. A common way to debug is to run again with one thing changed, and repeat until you find the cause. This is called bisecting. It only works if the earlier run stays the same. If the agent writes a different patch each time, each step measures the random sampling instead of your change. Second, the tests that grade the agent can also change between runs. Luo and colleagues did the first large study of flaky tests, which pass or fail without any code change. They examined 201 commits that fixed flaky tests across 51 open-source projects. The most common causes were async waits, concurrency, and test order dependence.[6](#user-content-fn-6) When both the agent and its tests vary, one passing result is a matter of chance in two ways. Third, a passing result shows less than it seems. Yu and colleagues studied SWE-Bench. In 345 patches, the tests from the original pull request never covered the failure. Those patches were marked as passing but were wrong.[7](#user-content-fn-7) So a claim that the tests passed needs its evidence attached before anyone trusts it. > **The goal is a run you can replay** > > Today, a model cannot be made to reproduce its output bit for bit, and waiting for that is not a plan. You can audit a run when someone else can see what was asked, what was done, and what the result was. This holds even if a new sample would produce a different patch. ## What a turn record should contain Review needs four nested records: the run, the turn, the tool call, and the gate result. A gate result is the pass or fail of a required check, such as the unit tests. Each record is added to the end of a log and never edited. This design is called event sourcing. Each event stays fixed once it is written, and you rebuild the current state by replaying the log. So the log is the audit trail. You do not have to keep a second record in step with it. Here is one turn record, cut down to the fields review asks about: ```json { "run": "r_01J8QK4M2ZT", "turn": 7, "model": "claude-opus-5", "prompt_sha256": "9f2c1ab0…", "tool_calls": [ { "seq": 1, "name": "read_file", "path": "src/billing/grants.ts", "result_sha256": "1a7b…" }, { "seq": 2, "name": "apply_patch", "diff_sha256": "c40e…", "files": 2, "added": 31, "removed": 4 }, { "seq": 3, "name": "run_tests", "cmd": "pnpm --filter billing test:unit grants.test.ts", "exit": 0 } ], "gate": { "name": "unit", "verdict": "pass", "log_sha256": "77d1…" } } ``` The record stores hashes instead of full content. A hash is a short fingerprint computed from the content, and it changes when the content changes. Hashes keep the record small, and anyone can still check it. But this only works if every service hashes the same record in the same way. RFC 8785, the JSON Canonicalization Scheme, sets that rule. It fixes the serialization, the property order, and the encoding, so a JSON document has one standard form for signing and hashing.[8](#user-content-fn-8) Without such a rule, two services can agree on the content and still compute different hashes. Then the audit trail stops checking out, and nothing warns you. Here is what each record holds, and the review question it answers: | Record | What it holds | What review can now ask | | --- | --- | --- | | Run | The agent's identity, repository, branch, model id, tool versions | Who ran this, on what code, with which tools | | Turn | Prompt hash, sampling settings, timestamp | What was asked, and with which settings | | Tool call | Name, arguments, output hash, exit code | What the agent did to the code | | Diff | Content hash, file count, lines added and removed | What changed, and how much code it touched | | Gate result | Command, result, log hash | Whether "it passed" has evidence behind it | None of these records needs the model to give the same output each time. They need the tool that runs the agent to record what happened in full. ## Which parts of a run you can hold fixed Some parts of a run can be reproduced and some cannot. It helps to separate them. You can hold fixed the toolchain version, the container image, the repository commit, the prompt text, the model identifier, the sampling settings, and the recorded input and output of every tool call. If you record those, a colleague can replay the same actions on the same code. This works even if a new generation would come out different. Software supply chains faced this problem first. Lamb and Zacchiroli describe reproducible builds: you build the same source and get bit-for-bit identical output. They say plainly that this is hard in real projects. Timestamps, build paths, and ordering all leak into the output.[9](#user-content-fn-9) The benefit is practical. A third party can check that a binary matches its source instead of trusting whoever built it. Coding agents cannot reach bit-for-bit output today. But the same goal applies: someone other than the author must be able to check the claim. Design also matters. The agent-computer interface is the set of commands an agent uses to work with the computer. The SWE-agent authors argue that this interface decides what an agent can do at all. Their interface had purpose-built commands for navigation, editing, and testing. It reached 12.5% pass@1 on SWE-bench, where non-interactive approaches did far worse.[10](#user-content-fn-10) A fixed set of named tools is also easy to log. An open-ended stream of shell commands is harder to record cleanly. Agentless goes further. Xia and colleagues replaced the agent loop with three fixed steps: find the fault, repair it, and validate the fix. It resolved 32.00% of SWE-bench Lite at about $0.70 per issue. That was ahead of the open-source agents of the time.[11](#user-content-fn-11) A pipeline with fixed steps is much easier to replay than an open-ended loop. This result does not argue against agents. It shows that a system can be both reproducible and capable. So pick an open-ended agent loop for a reason you can explain, and not by default. ## How Oxagen records a coding agent's run Oxagen is the agent control plane for the agents you run. It does not run them. It cannot make a model's samples come out the same. It keeps the record complete. Oxagen records each run frame by frame. A frame is one recorded event in a run. Each frame carries a hash of the frame before it, so a changed frame breaks the chain. When the run ends, Oxagen signs the whole run, which is called the seal. Anyone with the export can then check a turn instead of taking it on trust. The coding agent does the work and sends the frames. Oxagen ties each step to the agent's mandate (its access, budget, tools, and rules) and to the rule that answered its request. The claim is narrow. A second run can still give a different patch. But you can say exactly what happened on the run you shipped. ## Footnotes 1. Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). *Teaching Large Language Models to Self-Debug.* ICLR 2024. [https://arxiv.org/abs/2304.05128](https://arxiv.org/abs/2304.05128) [↩](#user-content-fnref-1) [↩2](#user-content-fnref-1-2) 2. Xia, C. S., & Zhang, L. (2023). *Keep the Conversation Going: Fixing 162 out of 337 bugs for $0.42 each using ChatGPT.* ISSTA 2024. [https://arxiv.org/abs/2304.00385](https://arxiv.org/abs/2304.00385) [↩](#user-content-fnref-2) 3. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). *Reflexion: Language Agents with Verbal Reinforcement Learning.* NeurIPS 2023. [https://arxiv.org/abs/2303.11366](https://arxiv.org/abs/2303.11366) [↩](#user-content-fnref-3) 4. Chen, M., Tworek, J., Jun, H., et al. (2021). *Evaluating Large Language Models Trained on Code.* [https://arxiv.org/abs/2107.03374](https://arxiv.org/abs/2107.03374) [↩](#user-content-fnref-4) 5. Ouyang, S., Zhang, J. M., Harman, M., & Wang, M. (2023). *An Empirical Study of the Non-determinism of ChatGPT in Code Generation.* [https://arxiv.org/abs/2308.02828](https://arxiv.org/abs/2308.02828) [↩](#user-content-fnref-5) [↩2](#user-content-fnref-5-2) 6. Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). *An Empirical Analysis of Flaky Tests.* FSE 2014, 643-653. [https://doi.org/10.1145/2635868.2635920](https://doi.org/10.1145/2635868.2635920) [↩](#user-content-fnref-6) 7. Yu, B., Zhu, Y., He, P., & Kang, D. (2025). *UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench.* [https://arxiv.org/abs/2506.09289](https://arxiv.org/abs/2506.09289) [↩](#user-content-fnref-7) 8. Rundgren, A., Jordan, B., & Erdtman, S. (2020). *JSON Canonicalization Scheme (JCS).* RFC 8785, IETF. [https://www.rfc-editor.org/rfc/rfc8785.html](https://www.rfc-editor.org/rfc/rfc8785.html) [↩](#user-content-fnref-8) 9. Lamb, C., & Zacchiroli, S. (2021). *Reproducible Builds: Increasing the Integrity of Software Supply Chains.* IEEE Software. [https://arxiv.org/abs/2104.06020](https://arxiv.org/abs/2104.06020) [↩](#user-content-fnref-9) 10. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.* NeurIPS 2024. [https://arxiv.org/abs/2405.15793](https://arxiv.org/abs/2405.15793) [↩](#user-content-fnref-10) 11. Xia, C. S., Deng, Y., Dunn, S., & Zhang, L. (2024). *Agentless: Demystifying LLM-based Software Engineering Agents.* [https://arxiv.org/abs/2407.01489](https://arxiv.org/abs/2407.01489) [↩](#user-content-fnref-11) --- # From STaR to DeepSeek-R1: what self-improvement means > A vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/self-improving-models-from-star-to-self-rewarding On a call, someone tells your team that the model improves itself. Nobody in the room can say whether that means a training loop, a prompt trick, or just a slide. So nobody questions the claim. Six weeks later, you are debugging a failure nobody predicted, because nobody knew what was running. The research behind that claim is real. It goes back about four years, and it is more specific than the sales pitch. Each published method that works is a loop with three steps. First, the model produces candidate outputs. Second, something decides which candidates are good. Third, the model trains on the ones that pass. The second step matters most. What does the deciding sets what the loop can learn, how far it gets, and where it stops improving. Below, each method is covered in turn, with the signal that drives it and the limit that stops it. > **Each self-improvement method follows this loop** > > 1. **Generate**: The model produces candidate outputs > 2. **Decide**: An answer key, a reward model, or a judge > 3. **Train**: Fine-tune on the outputs that passed > > Repeat run it again. > > *The methods below differ mainly in what does step 2. That choice sets what the loop can learn and where it stops improving.* ## STaR keeps the reasoning that reaches the right answer STaR, published in 2022, is the simplest version of the idea.[1](#user-content-fn-1) A model gets a few worked examples and a question. It writes out its reasoning step by step, called a rationale. If the final answer is right, the rationale is kept. If the answer is wrong, the model is shown the correct answer and asked to write a rationale that reaches it. The authors call this step rationalization. The model is then fine-tuned, meaning trained a little more, on every rationale that was kept. Then the whole process runs again. ```text loop: rationales = model.generate(questions) keep = [r for r in rationales if r.answer == answer_key[r.question]] retry = rationalize(model, wrong_ones, answer_key) model = finetune(model, keep + retry) ``` The signal is the answer key. The model's confidence and its own judgment play no part. The answer key is an outside label that the model cannot argue with. So the loop ends somewhere useful instead of drifting. The answer key is also the limit. If the model never gets a problem right, it has no correct rationale to train on. So STaR can learn only from problems it already solves some of the time. Rationalization helps by giving the model the answer first. But then some rationales reach the right answer through bad reasoning. Two later methods follow the same pattern. The first is rejection sampling fine-tuning (RFT), from a 2023 scaling study by Yuan and colleagues. It samples many reasoning paths from a model trained on labeled examples. It keeps the paths that are correct and distinct, and adds them to the training data.[2](#user-content-fn-2) The paper reports larger gains when the kept paths are more varied. It also reports larger gains for weaker starting models. So the method helps weaker models catch up more than it pushes the strongest models further. The second is ReST, from DeepMind the same year. It splits the loop into two steps.[3](#user-content-fn-3) The Grow step samples outputs from the current model. The Improve step scores those samples with a reward model, a separate model trained to score outputs. It keeps the high scorers and trains on them with offline reinforcement learning, which learns from a fixed set of saved samples. Because the samples are saved, several Improve passes can reuse them. So the authors describe ReST as more efficient than typical online reinforcement learning from human feedback, which needs new samples each round. They show it on machine translation. ReST replaces the answer key with a reward model, and that change drove the next five years of work. A reward model costs less and covers more tasks. But it is a model, so it can be wrong in ways an answer key cannot. ## Self-Instruct has the model write its own training data Self-Instruct, also from 2022, moves one step earlier.[4](#user-content-fn-4) STaR generates reasoning for questions you already have. Self-Instruct generates the questions too. A model starts from a small set of tasks written by people. It writes new instructions, then writes inputs and outputs for them. A filter drops the invalid ones and the near-duplicates. What is left becomes training data that teaches a model to follow instructions. Most of the open instruction datasets that came later were built this way. So it helps to be precise about what the method checks. The filters are rules of thumb. They check that an instruction is well-formed and not too close to one already in the set. No step checks that the generated answer is correct. That works for what the method is for, which is teaching a model the form of following an instruction. It does not check correctness. Treating it as if it does is a common and costly mistake. ## Some methods have a model supply the reward Constitutional AI, from Anthropic in 2022, replaces the human labeler for one purpose.[5](#user-content-fn-5) The model critiques and revises its own outputs against a written list of principles. Those AI-made comparisons then train a preference model, and the preference model drives reinforcement learning. The paper makes a narrow claim. It trains a harmless assistant through self-improvement "without any human labels identifying harmful outputs." People still took part. They wrote the constitution. The signal still comes from outside the model. It arrives once, as text at the start, instead of as labels all the way through. RLAIF, from Google in 2023, tested how far this approach reaches.[6](#user-content-fn-6) It used an off-the-shelf model to label which of two answers was better. On summarization and dialogue, it reports results comparable to reinforcement learning from human feedback. It also beat the model trained only on labeled examples, even when the labeling model was the same size as the model being trained. That last result is the surprising one. A model no larger than the one being trained still gives a usable training signal. Self-Rewarding Language Models, from Meta in 2024, closes the loop fully.[7](#user-content-fn-7) One model writes candidate responses and then judges them, using a prompt that asks it to act as a judge (the LLM-as-a-Judge method). It then trains on its own preferences. Three rounds of this on Llama 2 70B produced a model that beat Claude 2, Gemini Pro, and GPT-4 0613 on the AlpacaEval 2.0 leaderboard. Both skills improved together. The model got better at following instructions and better at judging. Read how it was scored before you read the result as open-ended. AlpacaEval 2.0 is a preference benchmark, and a model scores it. So the loop improved what a model judge rewards, as measured by a model judge. That is a real result about the style of following instructions. It does not show that a model can train itself into being right about facts. And the paper ran three rounds, not thirty. ## Quiet-STaR adds reasoning at every token Quiet-STaR, from 2024, makes reasoning happen at every token instead of in a separate prompt.[8](#user-content-fn-8) A token is a word or part of a word. At each position, the model writes short hidden rationales in parallel. The training signal is whether a rationale helped the model predict the text that comes next. Rationales that help are reinforced. The others are dropped. The reported gains are zero-shot, meaning the model saw no examples of the task, and they came without task-specific fine-tuning. GSM8K rose from 5.9% to 10.9%, and CommonsenseQA rose from 36.3% to 47.2%. > **Quiet-STaR, zero-shot** > > | | Before | After Quiet-STaR | > | --- | --- | --- | > | GSM8K | 5.9% | 10.9% | > | CommonsenseQA | 36.3% | 47.2% | > > *Source: Zelikman et al. (2024). No task-specific fine-tuning. The only training signal is whether a thought helped predict the text that follows.* Here the checker is the training text itself. That has a clear benefit, because ordinary text is plentiful and needs no labels. It is also the limit. Predicting the next token better stands in for reasoning better, but the two are different. A thought that makes the next sentence easier to predict may be a sound inference. It may also be a good guess about the writer's habits. ## DeepSeek-R1 uses reinforcement learning on answers a program can check DeepSeek-R1, released in January 2025, is the current reference point.[9](#user-content-fn-9) R1-Zero was trained with large-scale reinforcement learning, without a fine-tuning step on labeled examples first. Reasoning behaviors came from the reward alone. The Nature version of the paper describes self-reflection, verification, and changing strategies as they appeared, without reasoning examples labeled by people.[10](#user-content-fn-10) The preprint is direct about the cost. It says R1-Zero has "poor readability, and language mixing." The released R1 fixes this. It adds training in several stages, and it adds a set of starting examples (cold-start data) before the reinforcement learning phase. So the pure loop worked, but people found its output hard to read. The authors added a stage with examples chosen by people to make it usable. The key detail is what the reward measured. Reinforcement learning at that scale works on tasks where an automatic checker can score an answer. Examples are mathematics, competitive programming, and other fields where a program can decide correctness without a person. This is the same kind of signal STaR used, applied with far more compute. The common thread is an outside check, not the model's view of its own work. Cobbe and colleagues argued this in 2021 using GSM8K, a set of 8.5 thousand grade-school math problems. They trained a verifier, a model that ranks sampled solutions. As they added data, the verifier improved results more than fine-tuning did.[11](#user-content-fn-11) | Method | Signal source | Verifier | Known limit | | --- | --- | --- | --- | | STaR (2022) | Rationales for labeled questions | The answer key | Only learns problems it already sometimes solves | | Self-Instruct (2022) | Instruction data the model writes | Rule-of-thumb filters only | Nothing checks whether the answer is correct | | Constitutional AI (2022) | Self-critique against written principles | Principles written by people | Aims at harmlessness, not correctness | | RFT (2023) | Sampled reasoning paths, kept if correct | The answer key | Gains shrink as the starting model gets stronger | | ReST (2023) | Samples from the current model | A trained reward model | Carries every flaw in the reward model | | RLAIF (2023) | A language model labeling pairs of answers | The labeling model | Gives a preference, not a fact | | Self-Rewarding LM (2024) | The model judging its own output | The same model | Judge and student miss the same things | | Quiet-STaR (2024) | Whether a thought helps predict the next text | The training text | Better prediction stands in for better reasoning | | DeepSeek-R1-Zero (2025) | Reward on tasks a program can check | An automatic checker | Poor readability, language mixing | Read the Verifier column from top to bottom. Each method with a lasting gain in skill has something in that column the model does not control. Where the column points back to the model itself, the gain shows up only in things a model measures. Take that difference into your next call with a vendor. It is also where these methods fail. Models struggle to correct their own reasoning without outside feedback, and they sometimes get worse when they try.[12](#user-content-fn-12) ## Where Oxagen fits Oxagen does not train models and does not run agents. It is the agent control plane for the agents you run. The Verifier column shows why that matters for your business. An agent's mandate binds its identity, its knowledge scope, its permitted actions, and its commercial terms into one object. The run's record keeps the checked outcome beside them. So the outcome signal is kept outside the model that produced it. The meter prices that same record. To know whether a self-improving system is improving, you need an outcome log the system cannot write for itself. ## Footnotes 1. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). *STaR: Bootstrapping Reasoning With Reasoning*. arXiv. [https://arxiv.org/abs/2203.14465](https://arxiv.org/abs/2203.14465) [↩](#user-content-fnref-1) 2. Yuan, Z., Yuan, H., Li, C., Dong, G., Lu, K., Tan, C., Zhou, C., & Zhou, J. (2023). *Scaling Relationship on Learning Mathematical Reasoning with Large Language Models*. arXiv. [https://arxiv.org/abs/2308.01825](https://arxiv.org/abs/2308.01825) [↩](#user-content-fnref-2) 3. Gulcehre, C., Le Paine, T., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., & de Freitas, N. (2023). *Reinforced Self-Training (ReST) for Language Modeling*. arXiv. [https://arxiv.org/abs/2308.08998](https://arxiv.org/abs/2308.08998) [↩](#user-content-fnref-3) 4. Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., & Hajishirzi, H. (2022). *Self-Instruct: Aligning Language Models with Self-Generated Instructions*. arXiv. [https://arxiv.org/abs/2212.10560](https://arxiv.org/abs/2212.10560) [↩](#user-content-fnref-4) 5. Bai, Y., et al. (2022). *Constitutional AI: Harmlessness from AI Feedback*. arXiv. [https://arxiv.org/abs/2212.08073](https://arxiv.org/abs/2212.08073) [↩](#user-content-fnref-5) 6. Lee, H., et al. (2023). *RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback*. arXiv. [https://arxiv.org/abs/2309.00267](https://arxiv.org/abs/2309.00267) [↩](#user-content-fnref-6) 7. Yuan, W., Pang, R. Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., & Weston, J. (2024). *Self-Rewarding Language Models*. arXiv. [https://arxiv.org/abs/2401.10020](https://arxiv.org/abs/2401.10020) [↩](#user-content-fnref-7) 8. Zelikman, E., Harik, G., Shao, Y., Jayasiri, V., Haber, N., & Goodman, N. D. (2024). *Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking*. arXiv. [https://arxiv.org/abs/2403.09629](https://arxiv.org/abs/2403.09629) [↩](#user-content-fnref-8) 9. DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948) [↩](#user-content-fnref-9) 10. DeepSeek-AI (2025). *DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning*. Nature, 645, 633-638. [https://doi.org/10.1038/s41586-025-09422-z](https://doi.org/10.1038/s41586-025-09422-z) [↩](#user-content-fnref-10) 11. Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., & Schulman, J. (2021). *Training Verifiers to Solve Math Word Problems*. arXiv. [https://arxiv.org/abs/2110.14168](https://arxiv.org/abs/2110.14168) [↩](#user-content-fnref-11) 12. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. arXiv. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-12) --- # Governing an agent that rewrites itself > An agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/governing-an-agent-that-rewrites-itself Your compliance team has one simple question about the agent: what is it allowed to do? For a normal service, the answer is a document. The service changes only when somebody merges a pull request, so the document stays true. An agent that changes its own behavior is different. Each pass of its loop can change it, so the document is true only until the next pass. Systems like this exist today. The Darwin Gödel Machine edits its own code over and over and checks each edit on coding benchmarks. It raised its SWE-bench score from 20.0% to 50.0% and its Polyglot score from 14.2% to 30.7%.[1](#user-content-fn-1) So the open question is no longer whether an agent can improve itself. It is what you agree to when you let one run inside your company. > **Darwin Gödel Machine, before and after editing itself** > > | | Initial agent | After self-edits | > | --- | --- | --- | > | SWE-bench | 20% | 50% | > | Polyglot | 14.2% | 30.7% | > > *Source: Zhang, Hu, Lu, Lange, and Clune (2025). The system checks each edit on the same benchmarks that score it. This post argues that a benchmark cannot do that job alone.* Safety researchers have worked on this question for a decade, and their answers are concrete. Most of those answers say you cannot skip the step you were hoping to skip. ## Self-changing systems go wrong by chasing a flawed score In practice, self-changing systems go wrong in a plain way. They improve exactly what you measured, and what you measured was not what you meant. Krakovna and colleagues at DeepMind collected examples of this and called it specification gaming. In one, the goal was to stack a red block on a blue one. The agent was rewarded for how high the underside of the red block was. So it flipped the red block over. In another, a boat-racing agent was rewarded for hitting green blocks along the course. It drove in circles hitting the same blocks over and over and never finished the race.[2](#user-content-fn-2) Neither agent was broken. Both did exactly what the score asked. Skalse and colleagues then made this precise. A proxy is a measure you optimize in place of the real goal. They call a proxy unhackable if raising the expected proxy score can never lower the expected true score. That is the property you want from any metric your agent optimizes. Their result is bad news. Across all possible random policies (the rules an agent uses to pick actions), two reward functions can be unhackable only if one of them is constant.[3](#user-content-fn-3) So you cannot write a proxy metric that is safe to push without limit. Unhackable pairs where neither function is constant exist only when you limit the set of policies. In other words, safety comes from limiting what the agent can do. A better score does not provide it. That is the main design lesson, and it is older than today's agents. A decade ago, "Concrete Problems in AI Safety" sorted the ways systems fail into five groups: avoiding side effects, avoiding reward hacking, scalable supervision, safe exploration, and distributional shift.[4](#user-content-fn-4) Distributional shift means the conditions the system meets drift away from the ones it was tested on. A self-evolving agent touches all five groups at once. Self-modification also makes the last one loop back on itself. The agent that shifts the conditions is also the agent being tested under them. ## A sandbox is needed, but it is not enough on its own Every serious self-modification paper runs its system in a sandbox, an isolated space it should not be able to leave. The Darwin Gödel Machine authors say plainly that all experiments used safety measures, including sandboxing and human oversight.[1](#user-content-fn-1) At least one paper measured how well its sandbox held. STOP points a language-model improver at its own code. Its authors measured how often the generated code bypassed the sandbox.[5](#user-content-fn-5) That is the right experiment. They ran it because the answer might not be zero. A sandbox is a boundary, and you should measure a boundary instead of assuming it holds. There is a deeper reason not to rely on the sandbox alone. The original Gödel machine design had a real answer. It could rewrite any part of its own code only after it found a proof that the rewrite helps. That proof started from axioms describing the utility function, the hardware, and the starting code.[6](#user-content-fn-6) That is a real guarantee. It is also why no one ships it. Such a proof cannot be found in practice. So the Darwin Gödel Machine replaces it with tests on benchmarks.[1](#user-content-fn-1) Trading a proof for a benchmark score is reasonable. But it moves your whole safety case onto the benchmark, and careful audits show benchmarks cannot carry that load. Kapoor and colleagues found agent benchmarks with weak holdout sets or none at all. They found agents that take shortcuts and overfit. They found that results are often hard to reproduce.[7](#user-content-fn-7) Suppose the only check in an evolution loop is a benchmark score. Then the loop will find that benchmark's weak spots, because finding what raises the score is what the loop does best. ## Limiting what an agent can reach differs from shaping what it aims for These are two separate controls. Teams often use one and assume they have both, so it helps to name each. Capability control limits what the agent can reach: the sandbox, the list of allowed tools, the scope of its credentials, and the data it can query. You can enforce it and test it. Skalse's result says this is the control that carries the guarantee, since safety comes from limiting the set of policies.[3](#user-content-fn-3) Incentive control shapes what the agent aims for: the reward, the score, and the test a self-edit must pass. It feels like the stronger control. Skalse's result shows it cannot close every gap on its own. Oversight sits on top of both, and the evidence here is better than many expect. Bowman and colleagues tested a deliberately simple setup for scalable oversight. Humans chatted with an unreliable model assistant to answer questions. The pair did much better than the model alone and the humans alone on MMLU and time-limited QuALITY.[8](#user-content-fn-8) This does not mean humans should review every self-edit. It means a person working with a real interface does better than a person reading a summary. So design the review step with care instead of treating it as a formality. ## A governance model for a system that changes itself These limits together lead to one answer with four parts. **Typed contracts instead of documents.** A system that changes between readings cannot be held to a written policy. So the unit of governance has to be an object a machine can check. Bind six things together in one signed record: identity, knowledge scope, permitted action, commercial terms, verified outcome, and audit record. ```json { "identity": "agent:invoice-triage@7f2c1a", "knowledge_scope": { "ontology": "finance/ap", "as_of": "2026-09-09T00:00:00Z" }, "permitted_actions": ["read_invoice", "propose_gl_code"], "commercial_terms": { "meter": "governed_action", "budget_usd_month": 400 }, "verified_outcome": { "check": "gl_code_matches_approved_ledger" }, "audit_record": "required" } ``` The version hash in the identity is the part everything depends on. When an agent rewrites itself, it gets a new identity. A new identity means a new contract. The old contract does not carry over to it. **A knowledge scope defined by an ontology.** It is easy to list the permitted actions. It is hard to list what the agent may know about. Most access-control models give up here and hand over a whole database. An ontology solves this. It is a model of the business with typed entities and typed relationships between them. With it, the scope can be a statement about the graph instead of a list of rows. The query below, written in Cypher (the Neo4j query language), returns the entities in the agent's domain that are valid on a given date. ```cypher MATCH (a:Agent {id: $agentId})-[:SCOPED_TO]->(d:Domain) MATCH (d)<-[:IN_DOMAIN]-(n:Entity) WHERE n.valid_from <= $asOf AND coalesce(n.valid_to, $asOf) >= $asOf RETURN n ``` A compliance team can read that query as a boundary. It also holds up when the agent changes itself. The agent can change its own code, but that does not change what the domain contains. Limiting what the agent can see is capability control, written as structure. **Self-edits go through the same checks as any other change.** Teams often skip this part. A self-edit changes a production system. So it takes the same path as every other change: a proposal, a diff, a policy check, a test run, a decision, and a record. A model writing the change does not earn it a faster path. Much self-evolution research already produces output a person can review, so this costs less than it sounds. Agent Workflow Memory builds named, readable workflows from past attempts instead of hidden weight updates.[9](#user-content-fn-9) ADAS stores the agents it discovers as code in an archive.[10](#user-content-fn-10) You can diff both. Review what you can read, and do not promote what you cannot read. > **A self-edit takes the path every change takes** > > 1. **Proposal**: The diff, and the new identity it produces > 2. **Policy check**: Against the contract in force > 3. **Test run**: Evidence as well as a score > 4. **Decision**: A person who has the evidence > 5. **Record**: Append-only, each entry linked by hash to the one before > > *The agent's actions pass through the same checks. If this path is weaker than that one, the easiest way for the agent to take a forbidden action is to edit itself.* **A record of every accepted change that cannot be edited.** New entries are only added to the end. Each entry holds a hash of the one before, so a changed entry shows. Each entry holds the contract version, the proposal, the evidence for accepting it, and the approver. This is not for show. Reproducibility is the known weak point of this whole field.[7](#user-content-fn-7) Say someone asks in the future why the agent did something in March. You can answer only if you wrote it down in March. The 2025 survey of self-evolving agents describes the open problems the same way. It names evaluation and safety as the limits on the field, not capability.[11](#user-content-fn-11) > **Use the same checks for actions and self-edits** > > The agent's actions go through a set of checks. The agent's changes to itself must go through the same checks, reviewed as a change to a governed system. If the checks on self-edits are weaker, the easiest way for the agent to take a forbidden action is to change itself. ## What this does not solve None of this makes a self-evolving agent safe in the strong sense. There is no proof, and the original Gödel machine shows why you should not expect one.[6](#user-content-fn-6) What it does is make the system readable. At any moment, you can state with evidence what the agent is, what it may know, what it may do, what it cost, and what changed since last week. The evolution loop also does not justify this extra work everywhere. AlphaEvolve's results are real and in production. One is a scheduling heuristic that recovered 0.7% of Google's worldwide compute. But they come from problems with automated evaluators that score each candidate exactly.[12](#user-content-fn-12) If you have no definition of "better" that a machine can check, the loop has nothing to judge its changes. Without a judge, the loop changes without improving. ## How the four parts map to Oxagen Oxagen is the agent control plane for the agents you run. It does not run them. The four parts above map to the mandate an agent works under. The typed contract is the mandate itself. It holds identity, knowledge scope, permitted action, commercial terms, outcome, and audit record in one object, checked on the calls routed through Oxagen. The knowledge scope bounded by the ontology is the equipment the agent is given. That is why the graph tracks time, instead of being a vector index. The change record that cannot be edited is the record. The budget in the contract is the budget clause. For an evolution loop the meter matters a great deal, because the cost of each accepted improvement tells you whether to keep the loop running. Oxagen does not decide whether a proposed self-edit is a good idea. A person with the evidence still decides that. ## Footnotes 1. Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). *Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents*. arXiv. [https://arxiv.org/abs/2505.22954](https://arxiv.org/abs/2505.22954) [↩](#user-content-fnref-1) [↩2](#user-content-fnref-1-2) [↩3](#user-content-fnref-1-3) 2. Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). *Specification gaming: the flip side of AI ingenuity*. Google DeepMind. [https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/](https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/) [↩](#user-content-fnref-2) 3. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). *Defining and Characterizing Reward Hacking*. NeurIPS 2022. [https://arxiv.org/abs/2209.13085](https://arxiv.org/abs/2209.13085) [↩](#user-content-fnref-3) [↩2](#user-content-fnref-3-2) 4. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). *Concrete Problems in AI Safety*. arXiv. [https://arxiv.org/abs/1606.06565](https://arxiv.org/abs/1606.06565) [↩](#user-content-fnref-4) 5. Zelikman, E., Lorch, E., Mackey, L., & Kalai, A. T. (2023). *Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation*. COLM 2024. [https://arxiv.org/abs/2310.02304](https://arxiv.org/abs/2310.02304) [↩](#user-content-fnref-5) 6. Schmidhuber, J. (2003). *Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements*. arXiv. [https://arxiv.org/abs/cs/0309048](https://arxiv.org/abs/cs/0309048) [↩](#user-content-fnref-6) [↩2](#user-content-fnref-6-2) 7. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). *AI Agents That Matter*. arXiv. [https://arxiv.org/abs/2407.01502](https://arxiv.org/abs/2407.01502) [↩](#user-content-fnref-7) [↩2](#user-content-fnref-7-2) 8. Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., et al. (2022). *Measuring Progress on Scalable Oversight for Large Language Models*. arXiv. [https://arxiv.org/abs/2211.03540](https://arxiv.org/abs/2211.03540) [↩](#user-content-fnref-8) 9. Wang, Z. Z., Mao, J., Fried, D., & Neubig, G. (2024). *Agent Workflow Memory*. arXiv. [https://arxiv.org/abs/2409.07429](https://arxiv.org/abs/2409.07429) [↩](#user-content-fnref-9) 10. Hu, S., Lu, C., & Clune, J. (2024). *Automated Design of Agentic Systems*. arXiv. [https://arxiv.org/abs/2408.08435](https://arxiv.org/abs/2408.08435) [↩](#user-content-fnref-10) 11. Gao, H., Geng, J., Hua, W., Hu, M., Juan, X., Liu, H., et al. (2025). *A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence*. arXiv. [https://arxiv.org/abs/2507.21046](https://arxiv.org/abs/2507.21046) [↩](#user-content-fnref-11) 12. Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., Wagner, A. Z., et al. (2025). *AlphaEvolve: A coding agent for scientific and algorithmic discovery*. arXiv, and Google DeepMind blog (14 May 2025). [https://arxiv.org/abs/2506.13131](https://arxiv.org/abs/2506.13131) [↩](#user-content-fnref-12) --- # Graph-grounded retrieval vs vector search > Vector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/graph-grounded-retrieval-vs-vector-search A support agent is asked whether a customer's contract includes the uptime credit. The agent searches its documents and gets back a chunk, a short piece of text cut from a larger document. The chunk is about uptime credits. It names the right product, and it reads as if it was written for this question. So the agent answers yes. The chunk came from the standard contract template. This customer negotiated the credit out of their contract fourteen months ago. Each part of the search pipeline worked as designed, and the answer is still wrong. > **The chunk that matched, and the fact that was true** > > - **Standard template includes the uptime credit** (Throughout) > - **This customer's contract includes it** (Until 14 months ago) > - **Amendment removes it for this customer** (Since 14 months ago) > > *Illustrative. The template reads like the question, so it ranks first. Only the dates on the customer's own terms show that the credit ended.* ## A chunk can look right and still hold the wrong fact Retrieval-augmented generation (RAG) gives a model a searchable index of documents to draw from. Before RAG, a model could rely only on what its weights had stored in training.[1](#user-content-fn-1) RAG works. But as it is usually set up, it aims at the wrong target. The standard survey of hallucination grades generated text against the source the model was given. It splits failures into two kinds. In intrinsic failures, the output contradicts the source. In extrinsic failures, the output cannot be checked against the source.[2](#user-content-fn-2) Neither kind covers a bad source. If the search step supplies the wrong source, the model can repeat it faithfully and still state something false. It will then score as faithful. Text that is about your question is different from text that answers it. Vector search ranks text by cosine distance, a measure of how close two pieces of text are in meaning. That measure cannot tell the two kinds of text apart. ## What dense retrieval does well Dense retrieval turns each piece of text into an embedding, a list of numbers that stands for its meaning. It then finds the pieces whose embeddings sit closest to the question's. The numbers show it works. Dense Passage Retrieval trained two encoders on a modest number of question and passage pairs. It beat BM25, a standard keyword ranking method, by 9 to 19 points absolute. The measure was top-20 accuracy, whether a right passage appears in the top 20 results, across open-domain question answering benchmarks.[3](#user-content-fn-3) On the task of finding text about a topic, that margin is large and can be reproduced. Dense retrieval is also cheap. You cut the documents into chunks and embed them. In an afternoon, you can search documents that nobody ever organized. Keep using it. Dense retrieval is the right tool for most of what a company holds: prose, tickets, transcripts, notes, and the rest of the text that would never fit a schema. The question is which questions it should answer. ## Where similarity search falls short Research documents two problems. Both come from how the method works, so a better embedding model cannot fix them. The first problem is questions that need several steps, called hops. HotpotQA built 113,000 questions that need facts from several documents. Systems had to name those supporting facts along with the answer.[4](#user-content-fn-4) When a question needs facts from two places joined together, no single passage holds the answer. Search that returns the top few passages turns each hop into a separate guess. The chance of error grows with each guess. The second problem shows up when you add more text to the model's input to make up for it. Liu and colleagues measured how models use long inputs. Models answered best when the needed information sat at the start or end of the input. Accuracy dropped a lot when the answer sat in the middle.[5](#user-content-fn-5) This held even for models built for long inputs. So retrieving fifty passages instead of five does not fix poor precision. It puts the answer where the model reads worst. Here is how the two methods compare. | | Dense vector retrieval | Graph-grounded retrieval | | --- | --- | --- | | What it returns | Passages ranked by how close their embeddings are | Nodes and edges that match a typed pattern | | How it matches | Closeness of meaning, as learned by a model | Stated relationships and rules | | Questions with several hops | Each hop is a new guess, and errors add up | One query follows all the hops | | Time | Present only if the text mentions it | Stored on each edge, so "as of" a date is a query input | | Source of the answer | The passage is the evidence | Each edge stores its source, the extraction run, and a confidence score | | Setup cost | Hours to chunk and embed | Weeks for the schema, extraction, and matching duplicate entities | | Free-form text | Works on anything | Covers only what was extracted into the schema | | Typical failure | A chunk that looks right but holds the wrong fact | A missing edge, and no answer at all | The last row shows the trade-off. When vector search fails, it gives a confident answer that is wrong. When graph-grounded retrieval fails, it returns nothing. You can see an empty result, so you can recover from it. ## How graph-grounded retrieval works A knowledge graph stores facts as nodes (things, such as a customer) and edges (relationships, such as "has term"). Graph-grounded retrieval first turns the question into the things and relationships it asks about. Next it retrieves the facts. Then the model writes the answer. > **Find the facts first, then write the answer** > > 1. **Resolve**: The question becomes things and relationships > 2. **Traverse**: A typed pattern, limited to one customer organization and one date > 3. **Return facts**: Each row carries its source document > 4. **Write**: The model answers from those rows > > *The four systems below share this order. Each one retrieves typed facts, so step 2 can limit the results by organization and by date.* The simplest version already improves results. KAPING finds the entities named in the question. It writes their stored facts out as sentences and puts them before the prompt. It needs no fine-tuning. It reports gains of up to 48 percent on average over similar zero-shot baselines, across models of several sizes.[6](#user-content-fn-6) The method is plain. The model gets the specific facts that bear on the question, stored as triples (subject, relationship, object), instead of paragraphs that look similar. Think-on-Graph goes further. The model acts as an agent that explores the graph one relationship at a time, keeping several of the best paths as it goes (a method called beam search).[7](#user-content-fn-7) The path it follows serves as the citation. You can read why the system reached its answer. A similarity score does not give you that. The paper reports that this lets smaller models beat GPT-4 on some of these benchmarks. If cost is your limit, look at HippoRAG. It builds a graph index over the documents and retrieves in one step. On multi-hop question answering, it matches methods that retrieve in many steps, with gains of up to 20 percent. It costs 10 to 30 times less and runs 6 to 13 times faster than those methods.[8](#user-content-fn-8) Here, structure is both more accurate and cheaper. The graph does the join, so you stop paying for repeated searches. GraphRAG handles the questions that chunking handles worst: questions about a whole set of documents. It pulls out a graph of entities and finds groups of related entities. It summarizes each group, then builds an answer from the partial answers across groups. On document sets of around a million tokens, this made answers more complete and more varied.[9](#user-content-fn-9) In all four systems, the unit of retrieval is a fact with a type, not a window of text. So you can put limits on it. The query below is written in Cypher, the query language for the Neo4j graph database. It finds the contract terms in force for one customer on one date, within one workspace: ```cypher // Contract terms in force for one customer on a given date, // scoped to the caller's workspace. MATCH (c:Customer { publicId: $customerId, workspaceId: $workspaceId }) -[r:HAS_TERM]->(t:ContractTerm) WHERE r.validFrom <= $asOf AND (r.validTo IS NULL OR r.validTo > $asOf) RETURN t.displayName AS term, t.value AS value, r.sourceDocId AS source, r.validFrom AS since ORDER BY r.validFrom DESC ``` That query does three things a vector search does not. The workspace is part of the pattern, so it cannot return another customer's data. The dates are part of the filter, so it cannot return a term that was replaced. And each row it returns carries the document it came from. This query would not have returned the wrong uptime credit from the start of this post. ## Where graphs do worse **Building the schema is real work.** Hogan and eighteen co-authors wrote a survey of knowledge graphs. Most of it covers schema, identity, context, quality, and refinement. Much less covers querying.[10](#user-content-fn-10) That balance matches practice. Deciding what your types of things are takes people who disagree, and the work does not end. **Matching duplicate entities is the hardest part, and it is never fully solved.** This task is called entity resolution. It decides whether two records describe the same real thing. It is a research field of its own. One survey in ACM Computing Surveys covers each step, from grouping likely candidates through matching and clustering.[11](#user-content-fn-11) An error in one direction merges two customers into one node. An error in the other direction points half your edges at a duplicate that no query reaches. Neither error shows up on its own. **The graph holds only what extraction put in it.** A vector index covers whatever you embedded. A graph covers only what you pulled out into the schema, which is always less. Anything subtle, hedged, or new in a document is lost when it becomes a triple. **Turning a question into a graph query can also go wrong.** The Cypher above is correct because a person wrote it. If a model writes that query from a plain-language question, the hard part moves to the model. A query that looks right but follows the wrong relationship returns rows that look certain. The fix is to offer a fixed set of queries with fill-in inputs. The model picks one and fills it in, and it does not write graph queries freely. This limits what the model can ask on purpose. That is a real cost. It is the right trade for anything that touches contracts or money. > **Most systems need both** > > Use the graph for questions that have a clear structure: entities, relationships, dates, amounts, and permissions. Use dense retrieval for prose that has no structure. Send each question to the right method based on its type. A system with only one of the two is incomplete. ## How Oxagen uses graph-grounded retrieval Oxagen is the agent control plane for the agents you run. It does not run them. The knowledge an agent is given comes from a Neo4j graph and an ontology. An ontology is a model of the business's types of things and how they relate. So the unit an agent retrieves is a typed fact with its source and its dates, not a chunk that looked right. The nodes and edges behind an answer stay in the run's record. So you can trace a wrong answer back to the fact that produced it. This does not make the model more accurate. It lets a person check whether an answer that looks right is correct. ## Footnotes 1. Lewis et al. (2020). *Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks*. NeurIPS 2020. [https://arxiv.org/abs/2005.11401](https://arxiv.org/abs/2005.11401) [↩](#user-content-fnref-1) 2. Ji et al. (2023). *Survey of Hallucination in Natural Language Generation*. ACM Computing Surveys 55(12). [https://arxiv.org/abs/2202.03629](https://arxiv.org/abs/2202.03629) [↩](#user-content-fnref-2) 3. Karpukhin et al. (2020). *Dense Passage Retrieval for Open-Domain Question Answering*. EMNLP 2020. [https://arxiv.org/abs/2004.04906](https://arxiv.org/abs/2004.04906) [↩](#user-content-fnref-3) 4. Yang et al. (2018). *HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering*. EMNLP 2018. [https://arxiv.org/abs/1809.09600](https://arxiv.org/abs/1809.09600) [↩](#user-content-fnref-4) 5. Liu et al. (2024). *Lost in the Middle: How Language Models Use Long Contexts*. Transactions of the Association for Computational Linguistics. [https://arxiv.org/abs/2307.03172](https://arxiv.org/abs/2307.03172) [↩](#user-content-fnref-5) 6. Baek, Aji, & Saffari (2023). *Knowledge-Augmented Language Model Prompting for Zero-Shot Knowledge Graph Question Answering*. [https://arxiv.org/abs/2306.04136](https://arxiv.org/abs/2306.04136) [↩](#user-content-fnref-6) 7. Sun et al. (2024). *Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph*. ICLR 2024. [https://arxiv.org/abs/2307.07697](https://arxiv.org/abs/2307.07697) [↩](#user-content-fnref-7) 8. Gutiérrez et al. (2024). *HippoRAG: Neurobiologically Inspired Long-Term Memory for Large Language Models*. NeurIPS 2024. [https://arxiv.org/abs/2405.14831](https://arxiv.org/abs/2405.14831) [↩](#user-content-fnref-8) 9. Edge et al. (2024). *From Local to Global: A Graph RAG Approach to Query-Focused Summarization*. [https://arxiv.org/abs/2404.16130](https://arxiv.org/abs/2404.16130) [↩](#user-content-fnref-9) 10. Hogan et al. (2021). *Knowledge Graphs*. ACM Computing Surveys 54(4), 71:1-71:37. [https://arxiv.org/abs/2003.02320](https://arxiv.org/abs/2003.02320) [↩](#user-content-fnref-10) 11. Christophides, Efthymiou, Palpanas, Papadakis, & Stefanidis (2020). *An Overview of End-to-End Entity Resolution for Big Data*. ACM Computing Surveys 53(6), Article 127. [https://dl.acm.org/doi/10.1145/3418896](https://dl.acm.org/doi/10.1145/3418896) [↩](#user-content-fnref-11) --- # Self-evolving agents: what the evidence shows > What changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/self-evolving-agents-what-the-evidence-shows Your agent got a task wrong last Tuesday, and it will get the same task wrong next Tuesday. It does not know it has seen the task before. Each run starts from the same prompt, the same tools, and the same empty memory. So the hundredth run costs the same as the first and is no better. That is the first problem, and it is the costly one. The second problem is the opposite worry. A system that changes itself can change in ways nobody asked for and nobody noticed. Ask an engineer whether they want their agent to rewrite its own tool code, and they will hesitate for a long time. Both reactions make sense. There is now enough research on self-evolving agents to say which worry fits which case. A 2025 survey sorts the field by four questions: what changes, when it changes, how it changes, and where.[1](#user-content-fn-1) This post answers the first and third questions with numbers. The measured gains differ by more than ten times, depending on what changes and what checks the change. ## Four things an agent can change about itself Almost every system in this research changes one of four things. They are listed here from least to most that can go wrong. **The prompt or the context.** The agent writes something down and reads it back on its next attempt. Reflexion is the clearest example. After a failed attempt, the agent writes a short note on what went wrong. It stores the note in a buffer, and the note goes into the next attempt's input.[2](#user-content-fn-2) The model's weights do not change. Self-Refine does this within one turn. The same model critiques its own output and then rewrites it.[3](#user-content-fn-3) **The memory.** Generative Agents stores a full plain-language record of what happened. It then combines those records into higher-level reflections and looks them up to plan.[4](#user-content-fn-4) The agent does not learn a new skill. It builds a short summary of its own history. **The skills or tools.** Voyager keeps a growing library of code it can run. When it works out how to craft an item, it stores the steps as a function. Later it uses that function to build harder items.[5](#user-content-fn-5) Agent Workflow Memory does the same thing with routines instead of functions. It finds reusable workflows in its past attempts. It can do this offline from training examples, or online from new tasks as they arrive.[6](#user-content-fn-6) **The code.** At the far end, the agent changes its own program. STOP starts with a simple program that improves code, called an improver. It then runs the improver on itself, which produces a better improver.[7](#user-content-fn-7) In Automated Design of Agentic Systems (ADAS), one agent writes new agent programs and adds them to a growing archive.[8](#user-content-fn-8) The Darwin Gödel Machine edits its own code over and over and keeps an archive of versions.[9](#user-content-fn-9) As you go down that list, more can go wrong, and the measured gain also rises. The two rise together for a reason, and the larger gain has a cost. > **Four things an agent can change about itself** > > 1. **Prompt or context**: Reflexion, Self-Refine > 2. **Memory**: Generative Agents > 3. **Skills or tools**: Voyager, Agent Workflow Memory > 4. **Code**: STOP, ADAS, Darwin Gödel Machine > > *The rungs show order only and are not measured. Each rung names the systems in this post that change that layer.* ## The check on each change decides how the agent improves Each system above has a loop that makes a change and then checks it. If you change the check, the whole system behaves differently. The strong systems all check against something outside the model. Reflexion's coding results come from running unit tests. Voyager adds a skill to its library only after the skill passes its own check and runs without error in the game. Errors from running go back into the next attempt. AlphaEvolve needs an automated evaluator that scores each candidate program. So it works on problems with a goal a machine can check, and not on problems without one.[10](#user-content-fn-10) The Darwin Gödel Machine tests each self-edit on coding benchmarks. In all four cases, the loop relies on a checker that the model cannot argue with. > **Each loop keeps only what passes a check outside the model** > > 1. **Propose**: A reflection, a skill, a workflow, or a code edit > 2. **Verify**: Tests, running the code, or an evaluator outside the model > 3. **Keep**: Only what passed enters memory, the library, or the archive > 4. **Next attempt**: Starts from what was kept > > Repeat the agent improves only as far as step 2 checks correctly. > > *Reflexion, Voyager, AlphaEvolve, and the Darwin Gödel Machine all run step 2 outside the model. Without that outside check, self-correction stalls or gets worse.* The weak version is a model grading itself with no outside signal, and the evidence on it is poor. Huang and colleagues tested this kind of self-correction on reasoning tasks. The model revised its answers using only its own judgment. Models struggled to correct themselves without outside feedback. Sometimes they did worse after revising.[11](#user-content-fn-11) Self-Refine reports about 20 points of absolute gain on average across seven tasks. But its feedback is written for each task, and some of its tasks have clear scoring rules.[3](#user-content-fn-3) The difference between the two results comes mostly from what does the checking. This leads to a practical rule. An agent improves only as far as its checker measures the right thing. If your checker is a test suite, the agent gets better at passing tests. If your checker is a benchmark score, the agent gets better at the benchmark. You have to say what "better" means before the agent can get better. ## The measured gains This table covers the field. Each number comes from the cited paper or its official write-up. | System (year) | What changes | What checks the change | Measured gain | | --- | --- | --- | --- | | Reflexion (2023)[2](#user-content-fn-2) | Reflection notes stored in a buffer | Unit tests, reward from the environment | 91% pass@1 on HumanEval, against 80% for the GPT-4 baseline | | Self-Refine (2023)[3](#user-content-fn-3) | The output, rewritten | The same model's own feedback | About 20 points absolute on average across 7 tasks | | Voyager (2023)[5](#user-content-fn-5) | A library of skills stored as code | Its own check plus running the code in Minecraft | 3.3x more unique items, 2.3x longer distances, tech-tree milestones up to 15.3x faster | | Generative Agents (2023)[4](#user-content-fn-4) | Memory and combined reflections | People rating how believable the behavior is | Removing reflection makes behavior less believable | | STOP (2023)[7](#user-content-fn-7) | The program around the model | A supplied scoring function | The improved improver beats the starting improver on later tasks | | Agent Workflow Memory (2024)[6](#user-content-fn-6) | Reusable workflows found in past attempts | Task success over 1,000+ tasks in 200+ domains | 24.6% relative success gain on Mind2Web, 51.1% on WebArena | | ADAS / Meta Agent Search (2024)[8](#user-content-fn-8) | The agent program, added to an archive | Benchmark score in the target domain | Discovered agents beat the best hand-designed ones, and still do when moved to other domains and models | | Darwin Gödel Machine (2025)[9](#user-content-fn-9) | Its own code | SWE-bench and Polyglot scores | SWE-bench 20.0% to 50.0%, Polyglot 14.2% to 30.7% | | AlphaEvolve (2025)[10](#user-content-fn-10) | Candidate program code | Automated evaluators | A scheduling heuristic in production for over a year recovering 0.7% of Google's worldwide compute, up to 32.5% faster FlashAttention kernel, and 4x4 complex matrix multiplication in 48 scalar multiplications | Two results stand out. First, the cheap changes can have large effects. Agent Workflow Memory adds no new model and no new tools. It spots a reusable routine in a past attempt, and that raises WebArena success by about half.[6](#user-content-fn-6) Second, AlphaEvolve's results are the only ones here that run in production. The others are benchmark numbers. Its 4x4 complex matrix multiplication uses 48 scalar multiplications. That is the first improvement in that setting over Strassen's algorithm in 56 years.[10](#user-content-fn-10) > **Check the baseline behind a relative gain** > > Suppose the baseline succeeds on one task in five. A 51.1% relative gain there means something different than a 51.1% gain where the baseline succeeds most of the time. Check the baseline before you plan a budget around the headline number. ## What the evidence does not show Keep three limits in mind when you decide. **The checker is usually a benchmark, and benchmarks have gaps.** Kapoor and colleagues audited agent benchmarks and how they are used. They found a narrow focus on accuracy with no attention to cost. They found benchmarks with weak holdout sets, the test cases kept apart from tuning, or none at all. And they found results were often hard to reproduce.[12](#user-content-fn-12) If an agent evolves against a benchmark with a weak holdout set, it overfits to that benchmark. So each result in the table above is only as good as the benchmark that scored it. **Papers rarely report cost next to accuracy.** The same audit found that leading agents are more complex and costly than they need to be, because no one was trying to lower cost.[12](#user-content-fn-12) Self-evolution loops cost more than any other agent design. They run the task many times, and the archive-based ones run many versions of the agent. If you cannot see the cost of each accepted improvement, you cannot tell whether the loop is helping or only spending tokens. **None of this is recursive self-improvement, and the papers say so.** The STOP authors state that the language model itself never changes, so this is not full recursive self-improvement.[7](#user-content-fn-7) Schmidhuber's original design required a proof that a self-edit helps before making it. Such a proof cannot be found in practice. So the Darwin Gödel Machine drops that requirement. It tests edits instead and runs with sandboxing and human oversight.[9](#user-content-fn-9) Today's agents improve the program around a fixed model, and a fixed checker judges them. That is useful. It is less than what the phrase "self-evolving" makes people picture. ## What to do first If your agent repeats the same failure, do not start by letting it rewrite its own code. Start by giving it a place to write down what happened. Add a checker that can tell it whether the next attempt was better. That covers the top two rows of the table. It costs almost nothing, and a large share of the reported gain comes from it. Next, try finding workflows. Look at your successful attempts and check whether a routine is hidden in them. Agent Workflow Memory does this. It has the best gain for its risk in all of this research. A workflow found this way is a readable file, so you can inspect it before you trust it.[6](#user-content-fn-6) Let the agent change its own code last. The largest benchmark jumps come from this layer. It is also the only layer where the agent can change what it is permitted to do. So it is a governance problem before it is an engineering problem. ## What Oxagen records Oxagen is the agent control plane for the agents you run. It does not run them. The part of this that falls to Oxagen is the record. The record shows which version of an agent produced a result and what knowledge it was scoped to. It shows what the agent asked for, which rule answered, and what checked the outcome. With that record, you can show that an agent got better, instead of only saying so. The meter matters too. A self-evolution loop is the design most likely to spend real money in the background. The cost of each accepted improvement tells you whether to keep the loop running. ## Footnotes 1. Gao, H., Geng, J., Hua, W., Hu, M., Juan, X., Liu, H., et al. (2025). *A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence*. arXiv. [https://arxiv.org/abs/2507.21046](https://arxiv.org/abs/2507.21046) [↩](#user-content-fnref-1) 2. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). *Reflexion: Language Agents with Verbal Reinforcement Learning*. NeurIPS 2023. [https://arxiv.org/abs/2303.11366](https://arxiv.org/abs/2303.11366) [↩](#user-content-fnref-2) [↩2](#user-content-fnref-2-2) 3. Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., et al. (2023). *Self-Refine: Iterative Refinement with Self-Feedback*. NeurIPS 2023. [https://arxiv.org/abs/2303.17651](https://arxiv.org/abs/2303.17651) [↩](#user-content-fnref-3) [↩2](#user-content-fnref-3-2) [↩3](#user-content-fnref-3-3) 4. Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). *Generative Agents: Interactive Simulacra of Human Behavior*. UIST 2023. [https://arxiv.org/abs/2304.03442](https://arxiv.org/abs/2304.03442) [↩](#user-content-fnref-4) [↩2](#user-content-fnref-4-2) 5. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). *Voyager: An Open-Ended Embodied Agent with Large Language Models*. arXiv. [https://arxiv.org/abs/2305.16291](https://arxiv.org/abs/2305.16291) [↩](#user-content-fnref-5) [↩2](#user-content-fnref-5-2) 6. Wang, Z. Z., Mao, J., Fried, D., & Neubig, G. (2024). *Agent Workflow Memory*. arXiv. [https://arxiv.org/abs/2409.07429](https://arxiv.org/abs/2409.07429) [↩](#user-content-fnref-6) [↩2](#user-content-fnref-6-2) [↩3](#user-content-fnref-6-3) [↩4](#user-content-fnref-6-4) 7. Zelikman, E., Lorch, E., Mackey, L., & Kalai, A. T. (2023). *Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation*. COLM 2024. [https://arxiv.org/abs/2310.02304](https://arxiv.org/abs/2310.02304) [↩](#user-content-fnref-7) [↩2](#user-content-fnref-7-2) [↩3](#user-content-fnref-7-3) 8. Hu, S., Lu, C., & Clune, J. (2024). *Automated Design of Agentic Systems*. arXiv. [https://arxiv.org/abs/2408.08435](https://arxiv.org/abs/2408.08435) [↩](#user-content-fnref-8) [↩2](#user-content-fnref-8-2) 9. Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). *Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents*. arXiv. [https://arxiv.org/abs/2505.22954](https://arxiv.org/abs/2505.22954) [↩](#user-content-fnref-9) [↩2](#user-content-fnref-9-2) [↩3](#user-content-fnref-9-3) 10. Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., Wagner, A. Z., et al. (2025). *AlphaEvolve: A coding agent for scientific and algorithmic discovery*. arXiv, and Google DeepMind blog (14 May 2025). [https://arxiv.org/abs/2506.13131](https://arxiv.org/abs/2506.13131) [↩](#user-content-fnref-10) [↩2](#user-content-fnref-10-2) [↩3](#user-content-fnref-10-3) 11. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). *Large Language Models Cannot Self-Correct Reasoning Yet*. ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-11) 12. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). *AI Agents That Matter*. arXiv. [https://arxiv.org/abs/2407.01502](https://arxiv.org/abs/2407.01502) [↩](#user-content-fnref-12) [↩2](#user-content-fnref-12-2) --- # The science of AI agents: from ReAct to tool use > Planning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/the-science-of-ai-agents-from-react-to-tool-use Most systems shipped as agents are a while loop around a chat completion. The loop calls the model and parses anything that looks like a tool call. It runs the call, pastes the result back, and repeats until the token budget runs out. This works in a demo. Then it fails in production, and nobody on the team can say why. No one had a theory of what the loop was supposed to do. So the team changes the prompt, because the prompt is the only part anyone can see. A theory does exist. Between 2022 and 2023, a short line of papers took the loop apart and named its pieces: planning, tool use, memory, and reflection. Two surveys, published a month apart, arrived at about the same split. One describes profile, memory, planning, and action modules.[1](#user-content-fn-1) The other describes a brain, perception, and action architecture.[2](#user-content-fn-2) Each piece traces back to a paper that measured something specific. If you know which paper measured what, you can debug an agent instead of rewriting the prompt and hoping. ## ReAct puts reasoning and acting in one trace Before ReAct, researchers studied the two halves separately.[3](#user-content-fn-3) Chain-of-thought prompting has the model reason step by step, but the model cannot check a claim against the outside world. Action-only policies act in the world, but they give no running explanation of why. Yao and colleagues combined the two. The model writes a thought, then takes an action, then reads the result, then thinks again. All of this goes into one trace, a single log of the run. > **ReAct: one trace, three kinds of line** > > 1. **Thought**: Carries the plan to the next step > 2. **Action**: A tool call that acts on the environment > 3. **Observation**: What came back > > Repeat think again with the result. > > *Deleting the thought lowers performance, not only readability. The trace also lets a person find the step where a run went wrong.* The gains were large. On ALFWorld, a text-based household-task environment, ReAct beat imitation and reinforcement learning baselines by 34 absolute percentage points of success rate. On WebShop it gained 10 points. On the question-answering tasks HotpotQA and Fever, the main change was in how the model failed. Each step was checked against a Wikipedia lookup. That cut the model's habit of inventing a fact and then reasoning confidently from it. ReAct did all of this with one or two in-context examples. Two lessons carry over. First, the thought has a job. It carries the plan from one step to the next. That is why deleting it hurts performance, not just readability. Second, a person can read the trace afterward and find the step where the agent went wrong. That property, more than the benchmark score, is why the format spread. ## Toolformer learned when to call a tool A prompt can tell a model that a calculator exists. It cannot tell the model when a call is worth the round trip. Toolformer worked on that second problem.[4](#user-content-fn-4) The method is self-supervised, which means the model builds its own training data. First, sample candidate API calls into a corpus of ordinary text. Then run each call. Keep only the calls whose result makes the next tokens easier to predict, measured as lower perplexity. Then train on the calls you kept. Schick and colleagues connected a calculator, a question-answering system, a search engine, a translation system, and a calendar. The trained model's zero-shot performance was competitive with much larger models. It did not lose its core language modelling ability. Each API needed only a handful of demonstrations. The main finding is about when to call a tool. The filter checks whether calling the tool at this point makes the next tokens more predictable. No prompt can check that. Most tool-use bugs in production are about timing. The agent has the right tool, but it calls it one step too late, or calls it three times, or describes using it without using it. ## Tree of Thoughts treats planning as search When a model samples one chain of thought, it follows the first path it picks and never looks at the other options. Tree of Thoughts lays the options out.[5](#user-content-fn-5) Each partial solution becomes a node. The model rates each node as sure, maybe, or impossible. A search procedure with backtracking then decides which node to expand next. On the Game of 24, a puzzle that needs arithmetic planning with dead ends, GPT-4 with standard chain-of-thought prompting solved 4 percent of instances. The same model with Tree of Thoughts solved 74 percent. > **Game of 24, GPT-4** > > | | Chain-of-thought | Tree of Thoughts | > | --- | --- | --- | > | Instances solved | 4% | 74% | > > *Source: Yao et al. (2023). The gain comes from many more model calls per puzzle, and the paper does not price them.* That gap is the most quoted number in agent planning, and people usually quote it without its cost. Tree of Thoughts uses many more model calls. Each node the model rates is a call. Each branch it drops is compute already spent. A 70-point gain that takes an order of magnitude more inference is a real result and a real bill. Every planner built since makes that trade. Almost none of the papers report the cost side. ## Memory and reflection carry lessons across attempts Reflexion asked what an agent can learn between attempts without changing any model weights.[6](#user-content-fn-6) Its answer is to write a post-mortem in words. The agent fails and reflects in words on why. It stores that reflection in an episodic buffer, a memory of past attempts. It reads the reflection before the next attempt. Shinn and colleagues call this verbal reinforcement learning. On the HumanEval coding benchmark it reached 91 percent pass@1 (solved on the first try), against the 80 percent reported for GPT-4. Reflexion has a precondition. It works because a unit test tells it clearly that the last attempt was wrong. The reflection is written in language, but the signal behind it comes from outside the model and is a plain pass or fail. Generative Agents studied memory at a different scale.[7](#user-content-fn-7) Twenty-five agents lived in a sandbox town. Each kept a memory stream, a running log in plain language of everything it observed. Retrieval scored each memory on recency, importance, and relevance. Reflection ran on top of that and turned raw observations into higher-level conclusions from time to time. Planning then turned those conclusions into a plan for the day. Park and colleagues ran ablations, tests that remove one part at a time. Removing memory, reflection, or planning each lowered how believable the agents' behaviour was. In the best-known result, the agents organised a Valentine's Day party and spread the invitation by word of mouth. That behaviour came from the retrieval function, not from a party subroutine. Voyager added what the others left out, which is skills that last beyond one episode.[8](#user-content-fn-8) When the agent works out how to do something in Minecraft, it writes the behaviour as code. It checks the code against the game and stores it in a skill library to use later. Wang and colleagues report 3.3 times more unique items obtained than prior methods and 2.3 times longer distances travelled. They also report key tech-tree milestones reached up to 15.3 times faster. The skills carried over to new worlds, where other methods stalled. > **Voyager against prior methods in Minecraft** > > | Measure | Multiple of prior methods | > | --- | --- | > | Unique items obtained | 3.3x | > | Distance travelled | 2.3x | > | Tech-tree milestones, speed | up to 15.3x | > > *Source: Wang et al. (2023). Each bar is a multiple of the prior methods on that measure.* | Piece | Paper | What it measured | | --- | --- | --- | | Reasoning plus acting | ReAct | 34 points of absolute success over baselines on ALFWorld, 10 points on WebShop | | Tool use | Toolformer | Zero-shot performance competitive with much larger models, core language ability retained | | Planning as search | Tree of Thoughts | Game of 24: 4 percent with chain-of-thought, 74 percent with ToT | | Reflection | Reflexion | HumanEval pass@1 of 91 percent, against 80 percent reported for GPT-4 | | Memory | Generative Agents | Ablating memory, reflection, or planning each lowered believability ratings | | Skill retention | Voyager | 3.3x unique items, 2.3x distance, up to 15.3x faster tech-tree milestones | The rows work together as one design. Each row names a failure the loop has when that piece is missing. An agent with no memory repeats work. An agent with no reflection repeats mistakes. An agent with no planner commits to its first idea. An agent with no skill library relearns the same procedure every run and pays for it every time. ## Four problems are still open These four parts split the problem into pieces, but they do not solve it. Four problems in the research are still open, and each one shows up in production. **Reflection needs a grader.** Huang and colleagues tested intrinsic self-correction, where the model revises its own answer with no outside feedback. They found that models struggle to correct themselves. Sometimes performance got worse after the revision.[9](#user-content-fn-9) Both this result and Reflexion's hold. Reflexion's gains depend on a unit test. Without the test, the reflection has nothing to check against. If your agent reflects only against its own opinion of its output, the loop can convince itself of anything. **Errors add up over long tasks.** Dziri and colleagues studied transformers on compositional tasks, where each step builds on earlier ones, such as multi-digit multiplication and dynamic programming. They found that models reduce multi-step reasoning to linearised subgraph matching. In plain terms, the models match patterns they have seen instead of learning a general procedure.[10](#user-content-fn-10) The authors give a theoretical argument and empirical evidence that accuracy falls quickly as the number of dependent steps grows. An agent taking twenty dependent steps has the same problem. A per-step accuracy of 95 percent sounds high. But the twentieth step depends on all nineteen before it, so an error in any of them carries forward. **What to keep in memory is still a rule of thumb.** Recency, importance, and relevance are three weights set by hand. They happened to work in one sandbox. No one has a principled account of what an agent should forget, and long-running agents go wrong in what they forget. The survey literature lists memory management as an open problem.[1](#user-content-fn-1) **Most papers do not report cost.** Some of the results above gain accuracy by spending more inference. Those results do not report what the extra inference cost. That makes the numbers hard to compare and easy to misread. The evaluation literature argues about this separately, and that argument needs its own post. > **The short version** > > Planning searches through options. Tool use checks claims against the world. Memory keeps information across steps. Reflection fixes mistakes, but only when something outside the model can say the last attempt was wrong. ## Where this fits in Oxagen Oxagen is the control plane for the agents you run. It does not run them. The research above explains the shape of the mandate. One object ties together identity, knowledge scope, permitted actions, commercial terms, outcome, and the audit record. That record gives reflection the outside signal it needs. It also shows when the agent called each tool, which debugging needs. The knowledge an agent is given sits in a Neo4j graph with an ontology, a defined set of entity types and relationships. So an answer cites structured data that tracks time, not whatever the retriever happened to return. Planners gain accuracy by spending more inference. So Oxagen prices each governed action and attributes it to the run, the turn, and the step. The cost of the extra search then shows up as a line item. ## Footnotes 1. Wang, L., Ma, C., Feng, X., Zhang, Z., Yang, H., Zhang, J., Chen, Z., Tang, J., Chen, X., Lin, Y., Zhao, W. X., Wei, Z., & Wen, J.-R. (2023). *A Survey on Large Language Model based Autonomous Agents*. arXiv. [https://arxiv.org/abs/2308.11432](https://arxiv.org/abs/2308.11432) [↩](#user-content-fnref-1) [↩2](#user-content-fnref-1-2) 2. Xi, Z., Chen, W., Guo, X., He, W., Ding, Y., Hong, B., Zhang, M., Wang, J., Jin, S., Zhou, E., Zheng, R., Fan, X., Wang, X., Xiong, L., Zhou, Y., Wang, W., Jiang, C., Zou, Y., Liu, X., Yin, Z., Dou, S., Weng, R., Cheng, W., Zhang, Q., Qin, W., Zheng, Y., Qiu, X., Huang, X., & Gui, T. (2023). *The Rise and Potential of Large Language Model Based Agents: A Survey*. arXiv. [https://arxiv.org/abs/2309.07864](https://arxiv.org/abs/2309.07864) [↩](#user-content-fnref-2) 3. Yao, S., Zhao, J., Yu, D., Du, N., Shafran, I., Narasimhan, K., & Cao, Y. (2022). *ReAct: Synergizing Reasoning and Acting in Language Models*. ICLR 2023. [https://arxiv.org/abs/2210.03629](https://arxiv.org/abs/2210.03629) [↩](#user-content-fnref-3) 4. Schick, T., Dwivedi-Yu, J., Dessì, R., Raileanu, R., Lomeli, M., Zettlemoyer, L., Cancedda, N., & Scialom, T. (2023). *Toolformer: Language Models Can Teach Themselves to Use Tools*. NeurIPS 2023. [https://arxiv.org/abs/2302.04761](https://arxiv.org/abs/2302.04761) [↩](#user-content-fnref-4) 5. Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). *Tree of Thoughts: Deliberate Problem Solving with Large Language Models*. NeurIPS 2023. [https://arxiv.org/abs/2305.10601](https://arxiv.org/abs/2305.10601) [↩](#user-content-fnref-5) 6. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). *Reflexion: Language Agents with Verbal Reinforcement Learning*. NeurIPS 2023. [https://arxiv.org/abs/2303.11366](https://arxiv.org/abs/2303.11366) [↩](#user-content-fnref-6) 7. Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). *Generative Agents: Interactive Simulacra of Human Behavior*. UIST 2023. [https://arxiv.org/abs/2304.03442](https://arxiv.org/abs/2304.03442) [↩](#user-content-fnref-7) 8. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). *Voyager: An Open-Ended Embodied Agent with Large Language Models*. arXiv. [https://arxiv.org/abs/2305.16291](https://arxiv.org/abs/2305.16291) [↩](#user-content-fnref-8) 9. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-9) 10. Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Le Bras, R., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., & Choi, Y. (2023). *Faith and Fate: Limits of Transformers on Compositionality*. NeurIPS 2023. [https://arxiv.org/abs/2305.18654](https://arxiv.org/abs/2305.18654) [↩](#user-content-fnref-10) --- # What an Ontology Buys an Agent > An agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/what-an-ontology-buys-an-agent Ask an agent which pricing tier a customer is on, and it will answer. It will sound certain and quote a paragraph. Some of the time, the tier it gives is wrong. It may be the tier the customer had in March. It may come from an unsigned template. It may belong to a different company with a similar name. The answer reads well, but it came from the wrong context. Nothing in the pipeline was asked to check the context, so nothing noticed. ## Why a faithful answer can still be wrong Researchers have a name for this failure. The standard survey of hallucination in text generation splits it into two kinds.[1](#user-content-fn-1) An intrinsic hallucination contradicts the source the model was given. An extrinsic hallucination cannot be checked against that source at all. Both kinds are judged against the source, and that is the problem. Suppose retrieval picked the wrong source. A model can follow that source perfectly and still give a false answer. Every faithfulness test you run will pass, because those tests only compare the answer to the source. Retrieval was meant to ground answers in facts, and it does part of that job. It puts a document in the context window. It does not show that the document is about the entity you asked about. It does not show that the fact was true on the date you care about. It does not show who is accountable for stating the fact. An ontology lets a machine check whether the context is the right one. Without it, the system can only judge how similar the text looks. ## What an ontology is An ontology has three parts: classes, relations, and constraints. Classes are the kinds of things in your domain, such as Customer, Contract, Subscription, and LegalEntity. Relations are the allowed links between classes. Each relation says which class it starts from (its domain) and which class it points to (its range). For example, `SIGNED_BY` goes from a Contract to a LegalEntity and nowhere else. Constraints are rules that must hold. A Contract has exactly one counterparty. A Subscription tier is one of five fixed values. A LegalEntity cannot be its own parent. > **Classes, relations, and the constraints on them** > > - Contract SIGNED\_BY LegalEntity: exactly one counterparty > - Customer HAS\_SUBSCRIPTION Subscription: tier from a closed set of five > - LegalEntity PARENT\_OF LegalEntity: never itself > > *The example domain from this section. Each relation has a domain and a range. So if the graph says a Person signed a Contract, a machine can detect the violation.* Underneath sits a branch of logic called description logics. These are limited forms of first-order logic, designed so that a reasoner always finishes its work. The standard reference is the *Description Logic Handbook*.[2](#user-content-fn-2) It covers the main tradeoff in the field. Each feature you add to the language makes it harder to work out what follows from your facts. You cannot have a fully open modelling language and cheap reasoning at once. You must choose. The web standard for this is OWL 2. Its primer shows how to model with classes, properties, and individuals.[3](#user-content-fn-3) OWL 2 comes with three profiles, called EL, QL, and RL. Each profile drops some features so that reasoning stays fast. So the description logic tradeoff shows up as a product choice. Two properties matter more than the syntax. First, you can check constraints. Say the graph records that a Person signed a Contract, when only a LegalEntity may sign one. A machine can detect that violation. So you can catch one kind of wrong answer before anyone reads it. Second, OWL uses the open world assumption, and this often surprises teams. Under this assumption, a missing fact is not a false fact. Say your graph has no `HAS_SUBSCRIPTION` edge for a customer. An OWL reasoner concludes nothing from that. It does not conclude that the customer has no subscription. Most application code assumes the opposite. If you want a missing fact to count as false, say so in the model. Use a cardinality constraint, which is a rule on how many links an entity must have, or a validation shape. Do not assume the reasoner thinks the way you do. > **An ontology does more than a taxonomy** > > A taxonomy gives you a tree of is-a links. An ontology adds typed relations between the branches and rules that limit them. The tree says a Contract is a Document. The ontology says which entity signed it and when its term started. It also says that both facts are required. ## A knowledge graph holds the data an agent queries An ontology alone is a schema. What an agent queries is a knowledge graph. It has nodes for entities and edges for relations, filled in with real instances. The 2021 ACM Computing Surveys paper by Hogan and eighteen co-authors is the nearest thing the field has to a shared definition.[4](#user-content-fn-4) Its order of topics is useful on its own. It covers graph data models first. Then it covers schema, identity, and context. Then it covers checking and improving quality, and last, publishing. That order shows where the difficulty lies. Most of the hard work in a knowledge graph is not the query language. It is identity, which means deciding that two records are the same entity. It is also context, which means knowing when a fact holds and under what assumptions. A graph has one practical advantage over a document index. Its edges are records in their own right. An edge does not exist just because two nodes are close together. You can look it up, give it a type, and attach properties to it. The next two sections depend on that. A graph moves the hard work to a different place. It does not remove it. A document index is cheap to fill and costly to trust. A knowledge graph is the opposite. You pay up front to decide that "Acme Corp", "Acme Corporation", and a row keyed on a tax identifier are one node. You pay again each time a source changes its data. In return, the links between records already exist, so the system does not have to guess them at query time. ## Graphs answer questions that take several steps Take this question: "Which of our contracts with subsidiaries of Acme expire before the renewal date on the parent agreement?" No single passage answers it. The answer links four types of entity, and any one chunk of text holds only part of it. Researchers have measured this kind of question since HotpotQA.[5](#user-content-fn-5) HotpotQA built 113,000 questions that need facts from more than one document. It asked systems to name the supporting facts as well as give the answer. These are called multi-hop questions, because the answer takes several steps from one fact to the next. Once multi-hop questions could be measured on their own, the gap between retrieving text and reasoning over structure could be shown with data. Survey work describes how the field responded. In their roadmap, Pan and colleagues describe three ways to combine graphs and language models.[6](#user-content-fn-6) Knowledge graphs can improve language models. Language models can build and fill in graphs. Or both can work together in one system. Their diagnosis is direct, and it still holds. Language models generalise well but are weak at grounding answers in facts. Graphs are strong at grounding but weak at handling anything new. Two results stand out. Think-on-Graph treats the model as an agent that searches the graph.[7](#user-content-fn-7) It uses beam search, which keeps the few best paths at each step. It follows relation paths one step at a time, instead of retrieving once. The path it follows serves as the citation, so you can see why the answer came out as it did. The paper reports that this can let smaller models beat GPT-4 on some of these benchmarks. That result suggests the difficulty lies in how knowledge is structured for retrieval, not in model size. GraphRAG works on the same problem from the side of the document collection.[8](#user-content-fn-8) It pulls an entity graph out of the source documents. It finds communities of related entities and writes a summary for each one. To answer a question about the whole collection, it combines partial answers from those communities. The test was query-focused summarisation over collections of about a million tokens. There, GraphRAG gave more complete and more varied answers than standard retrieval. It helps most with questions about the whole collection. The usual method splits documents into chunks and searches them by similarity. That method handles these questions worst, because no one chunk covers the whole collection. ## Facts need a date range and a source A triple stores one fact as a subject, a relation, and an object, such as `(Acme, subscribed_to, Enterprise)`. That triple lacks two things an auditor will ask for first: since when, and who says so. ### Time Real facts are true only for a period. A job ends. A price changes. A new contract term replaces the one before it. If you store only the current value and overwrite the old one, you lose the data that answers "what did we bill them in March." Research on completing knowledge graphs has worked on this problem. Lacroix and colleagues extended ComplEx, a tensor factorisation method, to facts with timestamps.[9](#user-content-fn-9) They decompose an order-4 tensor, a four-dimensional array of numbers. So time is part of the model itself, not extra data added on. For question answering, Saxena and colleagues built CRONQUESTIONS.[10](#user-content-fn-10) It is a question answering dataset over a temporal knowledge graph, about 340 times larger than earlier work. They reported about a 120 percent gain in accuracy over the baselines of the time. Both results point the same way. A system that treats time as part of each fact does measurably better on questions that depend on time. In practice, each edge carries `validFrom` and `validTo`. "As of" becomes a query parameter, not an assumption. > **One customer's tier, with a date range on each fact** > > - **Acme subscribed\_to Growth** (January to March) > - **Acme subscribed\_to Enterprise** (April onward) > > *Illustrative. If the tier were overwritten in April, the record of what Acme was billed in March would be lost. Two edges with date ranges keep both facts, and the as of parameter picks one.* ### Provenance The second missing field is provenance, a record of where a fact came from. The W3C standardised terms for this in PROV-O. PROV-O describes provenance in OWL with three core types: Entity, Activity, and Agent.[11](#user-content-fn-11) Each statement says that something was derived from something else, by some process, and credits someone. On an edge, provenance becomes a source document, an extraction run, a timestamp, and a confidence score. Then a citation is no longer a paragraph the model chose to quote afterward. The citation is stored with the fact when the fact is written. It stays there whether or not the model mentions it. So you can check the answer instead of only trusting it. The check does not depend on what the model chooses to say. ## Where this meets Oxagen Oxagen is the control plane for the agents you run. It does not run them. Two clauses of an agent's mandate apply here. The knowledge an agent receives comes from a Neo4j graph database and an ontology. So each answer carries citations and dates because of how the data is stored, not because a prompt asked for them. The record keeps the lineage of an answer: the nodes, the edges, and the sources behind it. So you can look it up later. Neither clause makes a model smarter. Both make a wrong answer visible. That matters when someone asks why the agent said what it said. ## Footnotes 1. Ji et al. (2023). *Survey of Hallucination in Natural Language Generation*. ACM Computing Surveys 55(12). [https://arxiv.org/abs/2202.03629](https://arxiv.org/abs/2202.03629) [↩](#user-content-fnref-1) 2. Baader, Calvanese, McGuinness, Nardi, & Patel-Schneider, eds. (2003). *The Description Logic Handbook: Theory, Implementation, and Applications*. Cambridge University Press. [https://www.inf.unibz.it/~calvanese/papers-html/DLHB-2003-ed.html](https://www.inf.unibz.it/~calvanese/papers-html/DLHB-2003-ed.html) [↩](#user-content-fnref-2) 3. W3C (2012). *OWL 2 Web Ontology Language Primer (Second Edition)*. W3C Recommendation, 11 December 2012. [https://www.w3.org/TR/owl2-primer/](https://www.w3.org/TR/owl2-primer/) [↩](#user-content-fnref-3) 4. Hogan et al. (2021). *Knowledge Graphs*. ACM Computing Surveys 54(4), 71:1-71:37. [https://arxiv.org/abs/2003.02320](https://arxiv.org/abs/2003.02320) [↩](#user-content-fnref-4) 5. Yang et al. (2018). *HotpotQA: A Dataset for Diverse, Explainable Multi-hop Question Answering*. EMNLP 2018. [https://arxiv.org/abs/1809.09600](https://arxiv.org/abs/1809.09600) [↩](#user-content-fnref-5) 6. Pan et al. (2024). *Unifying Large Language Models and Knowledge Graphs: A Roadmap*. IEEE Transactions on Knowledge and Data Engineering. [https://arxiv.org/abs/2306.08302](https://arxiv.org/abs/2306.08302) [↩](#user-content-fnref-6) 7. Sun et al. (2024). *Think-on-Graph: Deep and Responsible Reasoning of Large Language Model on Knowledge Graph*. ICLR 2024. [https://arxiv.org/abs/2307.07697](https://arxiv.org/abs/2307.07697) [↩](#user-content-fnref-7) 8. Edge et al. (2024). *From Local to Global: A Graph RAG Approach to Query-Focused Summarization*. [https://arxiv.org/abs/2404.16130](https://arxiv.org/abs/2404.16130) [↩](#user-content-fnref-8) 9. Lacroix, Obozinski, & Usunier (2020). *Tensor Decompositions for Temporal Knowledge Base Completion*. ICLR 2020. [https://arxiv.org/abs/2004.04926](https://arxiv.org/abs/2004.04926) [↩](#user-content-fnref-9) 10. Saxena, Chakrabarti, & Talukdar (2021). *Question Answering Over Temporal Knowledge Graphs*. ACL 2021. [https://arxiv.org/abs/2106.01515](https://arxiv.org/abs/2106.01515) [↩](#user-content-fnref-10) 11. W3C (2013). *PROV-O: The PROV Ontology*. W3C Recommendation, 30 April 2013. [https://www.w3.org/TR/prov-o/](https://www.w3.org/TR/prov-o/) [↩](#user-content-fnref-11) --- # What SWE-bench Measures, and What It Misses > Coding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/what-swe-bench-measures-and-what-it-misses You are choosing a coding agent. The only number you can compare across the options is a SWE-bench resolve rate, so you choose by it. Then the agent starts work in your repository, on your backlog, under your review process. It behaves nothing like the number suggested. This gap is real, and it can be measured. You do not need to throw out benchmarks. You need to know what the score is made of. A resolve rate measures four things at once: a model, a scaffold, a budget, and a grading harness. The scaffold is the code around the model that decides what the model can see and do. The grading harness is the code that runs the tests and scores the patch. Yet the rate is reported as if it measured the model alone. Jimenez and colleagues introduced SWE-bench in 2023.[1](#user-content-fn-1) It has 2,294 tasks built from real GitHub issues and the pull requests that fixed them, across 12 popular Python repositories. Each task gives a system the issue text and the repository at the commit before the fix. The system must write a patch, and tests then grade it. When the paper came out, the best system tested resolved 1.96% of the issues.[1](#user-content-fn-1) Frontier systems now solve most of the curated subset. That is a fast climb, and it is why every model launch quotes the benchmark. ## What "resolved" means A task counts as resolved when two things happen. The patch makes a specific set of failing tests pass. It also leaves the tests that passed before still passing. That is the full definition. It does not say whether a maintainer would accept the patch. It does not cover slowdowns the tests miss, conventions the patch ignores, or whether a reviewer could follow the change. Passing tests stand in for correctness, and the two can drift apart. The original benchmark had a worse problem. Some tasks could not be graded fairly at all. OpenAI worked with the SWE-bench authors and 93 professional Python developers to review the test split.[2](#user-content-fn-2) In August 2024 they released SWE-bench Verified. It has 500 tasks that reviewers confirmed have clear problem statements, correct test patches, and enough information to solve. The removed tasks show what went wrong. Some issue descriptions left out too much, so no reader could tell what behaviour was wanted. Some tests were written so closely around the original fix that a different correct fix would fail them. Verified is now the standard reported figure. It exists because a large share of the original test split was not fit for the ranking the leaderboard built on it. ## Where the grading goes wrong Two problems matter, and each works in a different way. The first problem is tests that are too weak. Yu and colleagues built UTBoost, a tool that adds tests, and ran it on SWE-bench submissions. They found 345 patches marked as passing that were in fact wrong. The pull request's own tests never checked the failing case.[3](#user-content-fn-3) Once the missing tests were added, the ranking changed for 40.9% of SWE-bench Lite leaderboard entries and 24.4% of SWE-bench Verified entries.[3](#user-content-fn-3) So the true share of correct patches can be lower than the resolve rate, but not higher. > **Leaderboard entries whose ranking changed once the missing tests were added** > > | Leaderboard | Entries re-ranked | > | --- | --- | > | SWE-bench Lite | 40.9% | > | SWE-bench Verified | 24.4% | > > *Source: Yu, Zhu, He, and Kang (2025), UTBoost. The same study found 345 patches marked as passing that were wrong.* The second problem is contamination, which means the test tasks were in the model's training data. These are public repositories with public issues and public fixes. All of them are older than the training cutoff of every system on the board. Liang and colleagues tested this directly. They gave models only the issue text, with no access to the repository, and asked which file held the bug. Models named the correct path up to 76% of the time on SWE-bench. On tasks from repositories outside SWE-bench, they named it up to 53% of the time.[4](#user-content-fn-4) The authors also asked models to reproduce the fixed function exactly. They scored this with consecutive 5-gram accuracy, the share of five-token runs that match the real fix. Models reached up to 35% on SWE-bench Verified and up to 18% elsewhere.[4](#user-content-fn-4) So part of the score comes from memory, not reasoning. Nobody can say exactly how large that part is. > **What models know from the issue text alone** > > | | Repositories outside SWE-bench | SWE-bench | > | --- | --- | --- | > | Names the file that holds the bug | 53% | 76% | > | Reproduces the fixed function, 5-gram accuracy | 18% | 35% | > > *Source: Liang, Garg, and Zilouchian Moghaddam (2025). The highest figure across the models tested, with no repository access. The second row uses SWE-bench Verified. The gap between the dots is the part that memory adds.* > **Two errors with different causes** > > Weak tests raise the score by passing wrong patches. Contamination raises it by rewarding memory. Running the benchmark more carefully fixes neither, because both come from the dataset, not from the run. ## Most of the score comes from the scaffold A model alone does not produce the number you are comparing. A model inside a scaffold produces it, and the scaffold decides what the model can see and do. SWE-agent made this case directly. Yang and colleagues built a custom agent-computer interface for moving through a repository, editing files, and running tests.[5](#user-content-fn-5) They reported a 12.5% pass@1 rate on SWE-bench, which means 12.5% of tasks were solved on the first try. Approaches that did not interact with the repository had done far worse. The authors argue that language model agents are a new kind of end user. People need an IDE to work well, and agents likewise need interfaces built for them. The paper's contribution was the interface, not the model weights. OpenHands turned the same idea into an open platform.[6](#user-content-fn-6) On it, agents write code, use a command line, and browse the web inside a sandbox. It comes with code to run the standard benchmarks. It matters here because it changed how results are produced. Many reported numbers now depend on this one platform. Agentless took the opposite approach. Xia and colleagues dropped tool use and the agent's own decisions entirely. They used three fixed phases instead: localise, repair, and validate. Agentless resolved 32.00% of SWE-bench Lite at about $0.70 per issue. That beat every open-source agent of the time on both score and cost.[7](#user-content-fn-7) This result should change how you read a leaderboard. Much of what looked like agent skill was finding the right code. A fixed pipeline could do that for less. > **Agentless: three fixed phases, no autonomous tool use** > > 1. **Localise**: Find the files and functions to change > 2. **Repair**: Sample candidate patches > 3. **Validate**: Run tests and keep a patch that passes > > *Source: Xia, Deng, Dunn, and Zhang (2024). 32.00% of SWE-bench Lite at about $0.70 per issue. Because the pipeline is fixed, it can be run again step for step.* Retrieval quality matters on its own too. RepoCoder showed that context from the whole repository beats context from the current file by more than 10%.[8](#user-content-fn-8) This held for line, API, and function-body completion. RepoCoder uses a loop in which each generated draft guides the next retrieval. So where the relevant code lives, and whether the scaffold can find it, explains a large share of any repository-scale score. In practice, two numbers printed side by side may not be comparable at all. One may come from a fixed pipeline with a strict budget. The other may come from an agent allowed to run for an hour with unlimited retries. A report needs to give the scaffold, the model, the retry policy, and the spend together. Without them, a resolve rate is closer to a headline than a measurement. ## The variants, and what each one still misses The benchmark family has grown. Most new versions fix one of the weaknesses above. | Benchmark | Instances | Domain | What it fixed | What it still misses | | --- | --- | --- | --- | --- | | SWE-bench (2023) | 2,294 | 12 Python repositories | Set up the task: a real issue in a real repository, graded by tests[1](#user-content-fn-1) | Vague issues, unfair tests, contamination | | SWE-bench Verified (2024) | 500 | Same Python repositories | Human review of statements, tests, and whether each task can be solved[2](#user-content-fn-2) | Contamination, thin test coverage | | SWE-bench Multimodal (2024) | 617 | 17 JavaScript libraries | Visual, user-facing bugs outside Python[9](#user-content-fn-9) | Whether the fix looks right to a user | | SWE-Bench Pro (2025) | 1,865 | 41 repositories, some private | Long multi-file tasks, and licences chosen to resist contamination[10](#user-content-fn-10) | Your codebase, your conventions, your reviewers | The newer variants show that scores do not carry over well. SWE-bench Multimodal has 617 tasks across 17 visual JavaScript libraries. On it, the best system resolved 12% of tasks and the next best resolved 6%, though the same systems already scored far higher on Python.[9](#user-content-fn-9) SWE-Bench Pro has 1,865 tasks across 41 repositories. It uses copyleft and proprietary code on purpose, so that models are less likely to have trained on the tasks. Its reference patches change 107.4 lines across 4.1 files on average. Widely used models solve under 25% of its tasks on the first try (pass@1).[10](#user-content-fn-10) Taken together, the two results show one thing. A high Verified score carries over poorly to a different language, a different way of working, and a longer task. Your repository differs in at least one of those three ways. One more mismatch remains, and no variant has fixed it. It comes from the shape of the task itself, not from a flaw in any dataset. A SWE-bench task is a closed question. The issue is already triaged, already reproducible, and already scoped to one repository. A correct answer already exists in a merged pull request. A real backlog looks different. Its tickets can be vague, duplicated, or simply wrong. Some span two services. Some describe a symptom whose cause sits in a dependency. Many are not bugs at all. They are migrations, deprecations, and cleanups, where no failing test exists to make pass. The benchmark chose tasks it could grade. As a result, it left out most of the work your team has. ## What to measure instead The benchmark is still useful, as one input among several. A leaderboard number cannot tell you four things, because of how it is built. Each is cheap to measure on your own code: - **Grading fidelity.** SWE-bench grades a patch with tests written by the person who fixed the bug. Your agent will be graded by tests written before the bug existed. If your test suite would not catch the regression, a passing result is not evidence. - **Cost and variance per outcome.** Agentless reported $0.70 per issue for a reason.[7](#user-content-fn-7) Cost per resolved task can be compared across systems. A resolve rate alone hides a scaffold that made fifty tool calls to get there. Run the same task several times and record the spread, not just the best try. - **Reviewability.** If a reviewer cannot follow a patch, the patch costs more than it saves. A resolve rate says nothing about diff size, how much of the system the change touches, or whether the change came with a reason. - **Task admissibility.** First measure what share of your backlog is shaped like a task the agent can attempt. Then ask how often it succeeds. If a quarter of your tickets have no reproducible failure, the resolve rate applies to three quarters of the work at most. In short, SWE-bench checks whether a system can turn an issue description into a patch that passes a test written in advance. The code is Python, in a repository whose history the model has probably seen. That is useful to know. It does not tell you whether the agent should be allowed near your main branch. ## Where this meets Oxagen Oxagen does not run coding agents, and it does not publish a benchmark score. It is the control plane for the agents you run, and it records what they did. Two parts of it matter here. First, the record keeps outcomes that something outside the agent checked, not the agent's own report of success. That follows the same logic as refusing to take a resolve rate at face value. Second, the meter prices every governed action. That gives the cost per outcome that a leaderboard leaves out. A fair judgment of a governed agent can be built from the record it leaves. ## Footnotes 1. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) [↩](#user-content-fnref-1) [↩2](#user-content-fnref-1-2) [↩3](#user-content-fnref-1-3) 2. OpenAI (2024). *Introducing SWE-bench Verified.* OpenAI, 13 August 2024. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/) [↩](#user-content-fnref-2) [↩2](#user-content-fnref-2-2) 3. Yu, B., Zhu, Y., He, P., & Kang, D. (2025). *UTBoost: Rigorous Evaluation of Coding Agents on SWE-Bench.* [https://arxiv.org/abs/2506.09289](https://arxiv.org/abs/2506.09289) [↩](#user-content-fnref-3) [↩2](#user-content-fnref-3-2) 4. Liang, S., Garg, S., & Zilouchian Moghaddam, R. (2025). *The SWE-Bench Illusion: When State-of-the-Art LLMs Remember Instead of Reason.* [https://arxiv.org/abs/2506.12286](https://arxiv.org/abs/2506.12286) [↩](#user-content-fnref-4) [↩2](#user-content-fnref-4-2) 5. Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.* NeurIPS 2024. [https://arxiv.org/abs/2405.15793](https://arxiv.org/abs/2405.15793) [↩](#user-content-fnref-5) 6. Wang, X., Li, B., Song, Y., Xu, F. F., Tang, X., Zhuge, M., et al. (2024). *OpenHands: An Open Platform for AI Software Developers as Generalist Agents.* ICLR 2025. [https://arxiv.org/abs/2407.16741](https://arxiv.org/abs/2407.16741) [↩](#user-content-fnref-6) 7. Xia, C. S., Deng, Y., Dunn, S., & Zhang, L. (2024). *Agentless: Demystifying LLM-based Software Engineering Agents.* [https://arxiv.org/abs/2407.01489](https://arxiv.org/abs/2407.01489) [↩](#user-content-fnref-7) [↩2](#user-content-fnref-7-2) 8. Zhang, F., Chen, B., Zhang, Y., Keung, J., Liu, J., Zan, D., Mao, Y., Lou, J.-G., & Chen, W. (2023). *RepoCoder: Repository-Level Code Completion Through Iterative Retrieval and Generation.* EMNLP 2023. [https://arxiv.org/abs/2303.12570](https://arxiv.org/abs/2303.12570) [↩](#user-content-fnref-8) 9. Yang, J., Jimenez, C. E., Zhang, A. L., Lieret, K., Wu, X., Muennighoff, N., Synnaeve, G., Narasimhan, K. R., Yang, D., Wang, S. I., & Press, O. (2024). *SWE-bench Multimodal: Do AI Systems Generalize to Visual Software Domains?* ICLR 2025. [https://arxiv.org/abs/2410.03859](https://arxiv.org/abs/2410.03859) [↩](#user-content-fnref-9) [↩2](#user-content-fnref-9-2) 10. Deng, X., Da, J., Pan, E., He, Y. Y., Ide, C., Garg, K., et al. (2025). *SWE-Bench Pro: Can AI Agents Solve Long-Horizon Software Engineering Tasks?* [https://arxiv.org/abs/2509.16941](https://arxiv.org/abs/2509.16941) [↩](#user-content-fnref-10) [↩2](#user-content-fnref-10-2) --- # Why agents fail: measuring reliability and cost > AgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides. Published 2026-09-09 by Oxagen Research. Canonical page: https://oxagen.sh/blog/why-agents-fail-measuring-reliability-and-cost The agent works in the demo, so the team ships it. Two weeks later, support sees a pattern. About one run in four does something wrong, and the mistake is different each time. Finance asks what a successful run costs, and nobody knows. There is a token bill, but it is not broken down by outcome. In the ledger, a run that worked looks the same as a run that failed three times and was retried. Both problems come from how agents are measured, and research has studied them since 2023. The benchmarks that made agents look good measure one try at a task in an environment that never changes. Almost none of them report cost. If you read them closely, the failure rate in production is no surprise. ## What the benchmarks measure People quote five benchmarks as if they measure the same thing. Each one measures something different. AgentBench placed language models in 8 different interactive environments, from operating-system tasks to databases to card games. It scored task success in each one.[1](#user-content-fn-1) The main result was a wide gap between top commercial models and open models under 70B parameters. The diagnosis matters more. Liu and colleagues traced agent failure mainly to weak long-term reasoning, weak decisions, and weak instruction following. They did not trace it to a missing skill. WebArena took the opposite approach and built one environment in full.[2](#user-content-fn-2) It runs real, self-hosted, working sites: an online store, a social forum, a software development platform, and a content management system. It also includes supporting tools and documentation. WebArena checks success by looking at the resulting state of the site, not by matching the text of a summary. The best agent Zhou and colleagues tested finished 14.41 percent of tasks from start to finish. People on the same tasks finished 78.24 percent. GAIA has 466 questions that a person finds simple in concept.[3](#user-content-fn-3) Answering them takes reasoning, handling more than one kind of input, web browsing, and tool use. The answers are short and can be checked exactly. Human respondents scored 92 percent. GPT-4 with plugins scored 15 percent. Mialon and colleagues state their design goal directly. They argue that research should focus on how reliably agents handle tasks people find easy, not on scores for tasks people find hard. > **The best agent tested against people, on the same tasks** > > | | Best agent tested | Humans | > | --- | --- | --- | > | WebArena | 14.41% | 78.24% | > | GAIA (GPT-4 with plugins) | 15% | 92% | > > *Sources: Zhou et al. (2023) for WebArena and Mialon et al. (2023) for GAIA, each at publication. A person finds the tasks in both routine.* SWE-bench covers coding, and its grader is the strictest of the five.[4](#user-content-fn-4) It has 2,294 real GitHub issues and the pull requests that fixed them, drawn from 12 popular Python repositories. A submission succeeds when the repository's own test suite passes after the model's patch is applied. When the paper came out, the best system Jimenez and colleagues measured was Claude 2. It resolved 1.96 percent of tasks. The fifth benchmark, tau-bench, measures what the other four skip.[5](#user-content-fn-5) | Benchmark | Task type | What success means | Cost reported | | --- | --- | --- | --- | | AgentBench | 8 interactive environments (OS, DB, web, games) | Task success in each environment, then combined | No | | WebArena | Multi-step tasks on real self-hosted websites | A check of the site's resulting state | No | | GAIA | 466 real-world assistant questions | Exact match against a short correct answer | No | | SWE-bench | 2,294 real GitHub issues in 12 Python repos | The repository's own test suite passes after the patch | No | | tau-bench | Conversations between a user and a tool-using agent under a domain policy | Final database state matches the goal state, over repeated trials | No | Look at the right-hand column. The five benchmarks define success in five ways. None of them counts the cost of a try as part of the result. Two more points go with that table. First, a real grader stays useful after the scores move on. Published SWE-bench resolve rates have climbed far above 1.96 percent. Those rates can still be compared across years, because the test suites that grade them did not change. So a benchmark that checks success for real stays useful after its leaderboard is out of date. Second, each of these environments is frozen on purpose, which means it does not change during testing. That makes the comparison fair, but it also makes the results optimistic. In production, sites change their markup, APIs drop fields, and someone edits a policy document on a Tuesday. An agent that scored 14 percent against a fixed copy of a web stack was not tested on what breaks it most often in real use. ## One try does not measure reliability Yao and colleagues built tau-bench around a case the others do not simulate.[5](#user-content-fn-5) An agent talks to a user, uses domain tools, and follows a written domain policy, all at the same time. tau-bench compares the final database state with the goal state. So the agent must change the data correctly. Describing the change is not enough. The metric is their main contribution. The usual practice reports pass@k, the chance that at least one of k tries succeeds. That suits a research leaderboard. It does not suit a production system, where you get one try and a customer is waiting. So the authors defined pass^k, the chance that all k separate tries at the same task succeed. The results were clear. State-of-the-art function-calling agents, including GPT-4o, succeeded on under 50 percent of tasks. In the retail domain, pass^8 fell below 25 percent. An agent that is right half the time on one try is right on all eight tries only about a quarter of the time. That number matches what support sees, and nobody publishes it. > **pass@k versus pass^k** > > pass@k rewards a system that is right some of the time. pass^k penalises a system that is inconsistent. More tries raise pass@k and lower pass^k. So a system tuned for pass@k can become less reliable while its headline number improves. ## Errors add up over many steps, and self-checks do not undo them Inconsistency is not bad luck. It follows from simple arithmetic when a task has many steps. An agent doing a real task takes a chain of steps, and each step depends on the ones before it. It reads the ticket, queries the database, decides, calls the API, and checks the result. Suppose each step is right 95 percent of the time, independently. Then twenty such steps all come out right only 36 percent of the time. At 99 percent per step, twenty steps still come out right only 82 percent of the time. This is multiplication, not a flaw in the model. So per-step accuracy is a misleading target on its own. > **Chance a chain of dependent steps finishes correctly** > > - **95% right per step**: 36% at 20 > - **99% right per step**: 82% at 20 > > *Per-step accuracy raised to the power of the number of steps. This is not model data. It is the multiplication the section describes, and it assumes accuracy stays the same at every step. Dziri et al. find that it does not.* The way models behave makes this worse. Dziri and colleagues studied transformers on compositional tasks, which are tasks built from smaller steps. Their examples were multi-digit multiplication, logic puzzles, and dynamic programming.[6](#user-content-fn-6) They concluded that the models do not learn a general procedure for these tasks. Instead, the models reduce many-step reasoning to matching pieces of patterns they have seen, which the authors call linearised subgraph matching. The authors argue, in theory and in experiments, that models which generate one token at a time lose accuracy quickly as tasks grow more complex. So per-step accuracy does not even stay the same across a long task. It gets worse as the task gets longer. The obvious fix is to let the agent check its own work, but that fix is weaker than it looks. Huang and colleagues found that models struggle to correct themselves without outside feedback.[7](#user-content-fn-7) Performance sometimes got worse after a round of self-correction with no outside help. Self-correction works when something outside the model can judge the last try. A test suite can do that, and so can a check of the database state. The model's own confidence cannot. ## Cost is half of the result Kapoor and colleagues made an argument the field had avoided.[8](#user-content-fn-8) They argued that an accuracy number should not be reported without its cost. Their analysis of agent benchmarks makes four claims that belong together. First, when a benchmark scores only accuracy, spending more carries no penalty. So researchers are pushed toward agents that are more complex and costly than they need to be. Second, accuracy and cost should be optimised together. When the authors did that, they cut cost by a large amount and kept accuracy the same. Third, many agent benchmarks lack proper holdout sets, which are tasks kept back for final testing. Without them, agents can overfit and use shortcuts that will not exist in production. Fourth, the field mixes up two audiences. A model developer compares systems, and a downstream developer chooses one to deploy. The two need different evaluations, and a single leaderboard column serves neither well. The previous post's numbers show the same problem. Tree of Thoughts raised success on Game of 24 from 4 percent to 74 percent.[9](#user-content-fn-9) It searched over many candidate thoughts, which means many model calls per task. Reflexion reached 91 percent pass@1 on HumanEval.[10](#user-content-fn-10) After each failure it wrote a short review of what went wrong and tried again, which means several tries per task. Both results are real. Both pay for accuracy with extra model calls. Neither headline says how much accuracy the extra calls bought. A team deciding whether to ship needs exactly that figure. Overfitting is the less visible part of their argument, and it explains the gap between the demo and the deployment. An agent tuned on a benchmark with no holdout set can learn the benchmark itself. It learns the shape of the tasks, the quirks of the environment, and the shortcuts the grader accepts. None of that carries over. A team may pick an agent design from a leaderboard and then be surprised in production. Often that is not a regression. It is the first accurate measurement of the agent. This has a practical form for anyone running an agent today. Accuracy per try is the wrong basis. The number that matters is cost per verified outcome. Add up the spend across every try, retry, and abandoned branch. Then divide it by the number of outcomes that something outside the agent confirmed were correct. That number is usually several times the simple one, and it moves when reliability moves. It is the only figure that makes the cost of a retry loop clear on a finance dashboard. ## Where this meets Oxagen Oxagen is the control plane for the agents you run. It does not run them. Two of its parts follow from the research above. The first is the meter. Oxagen prices every governed action and assigns it to the person, the agent, the run, the turn, and the step. So you can divide total spend by checked outcomes instead of by tries. That is the pass^k problem written as a bill. The second is the record. Oxagen keeps every run against the agent's mandate, one frame at a time, where a frame is one recorded event. Each run shows what the agent asked for, which rule answered, and what checked the outcome. So the outside check that self-correction needs is a stored row, not the model's own opinion. None of this makes an agent reliable. It lets you put a number on the agent's reliability and its cost. ## Footnotes 1. Liu, X., Yu, H., Zhang, H., Xu, Y., Lei, X., Lai, H., Gu, Y., Ding, H., Men, K., Yang, K., Zhang, S., Deng, X., Zeng, A., Du, Z., Zhang, C., Shen, S., Zhang, T., Su, Y., Sun, H., Huang, M., Dong, Y., & Tang, J. (2023). *AgentBench: Evaluating LLMs as Agents*. ICLR 2024. [https://arxiv.org/abs/2308.03688](https://arxiv.org/abs/2308.03688) [↩](#user-content-fnref-1) 2. Zhou, S., Xu, F. F., Zhu, H., Zhou, X., Lo, R., Sridhar, A., Cheng, X., Ou, T., Bisk, Y., Fried, D., Alon, U., & Neubig, G. (2023). *WebArena: A Realistic Web Environment for Building Autonomous Agents*. ICLR 2024. [https://arxiv.org/abs/2307.13854](https://arxiv.org/abs/2307.13854) [↩](#user-content-fnref-2) 3. Mialon, G., Fourrier, C., Swift, C., Wolf, T., LeCun, Y., & Scialom, T. (2023). *GAIA: a benchmark for General AI Assistants*. ICLR 2024. [https://arxiv.org/abs/2311.12983](https://arxiv.org/abs/2311.12983) [↩](#user-content-fnref-3) 4. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?*. ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770) [↩](#user-content-fnref-4) 5. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). *tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains*. arXiv. [https://arxiv.org/abs/2406.12045](https://arxiv.org/abs/2406.12045) [↩](#user-content-fnref-5) [↩2](#user-content-fnref-5-2) 6. Dziri, N., Lu, X., Sclar, M., Li, X. L., Jiang, L., Lin, B. Y., West, P., Bhagavatula, C., Le Bras, R., Hwang, J. D., Sanyal, S., Welleck, S., Ren, X., Ettinger, A., Harchaoui, Z., & Choi, Y. (2023). *Faith and Fate: Limits of Transformers on Compositionality*. NeurIPS 2023. [https://arxiv.org/abs/2305.18654](https://arxiv.org/abs/2305.18654) [↩](#user-content-fnref-6) 7. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798) [↩](#user-content-fnref-7) 8. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). *AI Agents That Matter*. arXiv. [https://arxiv.org/abs/2407.01502](https://arxiv.org/abs/2407.01502) [↩](#user-content-fnref-8) 9. Yao, S., Yu, D., Zhao, J., Shafran, I., Griffiths, T. L., Cao, Y., & Narasimhan, K. (2023). *Tree of Thoughts: Deliberate Problem Solving with Large Language Models*. NeurIPS 2023. [https://arxiv.org/abs/2305.10601](https://arxiv.org/abs/2305.10601) [↩](#user-content-fnref-9) 10. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). *Reflexion: Language Agents with Verbal Reinforcement Learning*. NeurIPS 2023. [https://arxiv.org/abs/2303.11366](https://arxiv.org/abs/2303.11366) [↩](#user-content-fnref-10)