# Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces

> Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.

Published 2026-10-09 by Mac Anderson. Canonical page: https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models

Your organization runs coding agents every day. Each session produces a record: the task, every file the agent read, every command it ran, every edit it made, and whether the work held up. Most teams throw that record away the moment the session ends. This book argues that the record is the most valuable asset the session produces, and that a team which keeps it can, within a year, fine-tune an open-weight model that does its own work as well as the rented frontier model does today.

The book is long because the argument has many parts and each part has evidence behind it. It is written for the engineer who will build the pipeline and for the person who has to approve it. Chapters 1 and 2 make the case. Chapters 3 to 6 are about the oracle, the program that decides whether a trace counts. Chapters 7 to 9 are about collecting traces from the agent harness you already use, with a working Claude Code plugin you can install today. Chapters 10 to 12 are about turning traces into training data and how much you need. Chapters 13 to 16 are about the delivery pipeline that ships new weights on a schedule. Chapter 17 lists the earlier systems that tried the same idea, so you can see what held up.

Three words carry most of the weight, so here they are up front.

-   A **trace** is the full record of one agent session: prompts, tool calls, tool results, edits, and the final state of the repository.
-   An **oracle** is a program that reads the result of a session and returns one verdict, PASS or FAIL, without the agent's help. In this book the oracle is a hidden test that the agent never sees, run in a container the agent cannot reach.
-   A **flip** is the event the whole pipeline is built around. The oracle said FAIL before the agent started. The oracle says PASS after the agent finished. The trace between those two verdicts is a worked example of the task being done right, graded by something the model did not control.

> **The reference plugin**
>
> The plugin this book describes is public at [github.com/macanderson/oracle-flip](https://github.com/macanderson/oracle-flip). It installs into Claude Code in two commands, writes every tool call to a trace file, and runs an oracle before it lets the agent say the work is done. Chapter 8 walks through every file.

## The argument in one page

Rented tokens are a cost that never ends and a capability you never own. The price per token has fallen fast, and it will keep falling, but the thing you are paying for is a model that knows nothing about your systems on the day you start and nothing more on the day you stop. Every improvement happens in the vendor's weights, under the vendor's terms, on the vendor's schedule.

Open-weight models are now close enough to the frontier that the gap no longer decides the outcome. In late 2024, the best open model trailed the best closed model by about a year.[1](#user-content-fn-epoch-open) In the year to early 2025, the gap between the best open and closed models on a public human-preference leaderboard narrowed from about 8 percent to under 2 percent.[2](#user-content-fn-ai-index-2025) Models released under permissive licenses by DeepSeek, Alibaba, Meta, Mistral, and in August 2025 by OpenAI itself, run on hardware you can rent by the hour or buy outright.[3](#user-content-fn-gpt-oss) A general open model is a floor. What raises the floor on your tasks is training on your tasks.

Traces are that training data, but only when something outside the model grades them. The research on models grading their own work is consistent and unkind. Left to judge themselves, models prefer their own output, cannot reliably find their own errors, and often get worse when they try to self-correct without an outside signal.[4](#user-content-fn-huang-self-correct)[5](#user-content-fn-panickssery) The methods that do produce lasting gains all share one feature: a verifier the model does not control, such as an answer key, a compiler, or a test.[6](#user-content-fn-star)[7](#user-content-fn-deepseek-r1) For software, that verifier already exists. It is the test suite, the type checker, the build, and the hidden test you write before the agent starts.

The oracle has to be deterministic, air-gapped, and silent. Deterministic, because a flaky verdict is noise in the training set, and noise at this stage is expensive to remove later. Air-gapped, because frontier models have been caught editing tests, stubbing functions, and exiting early to make a check pass, and a reward signal the agent can reach is a signal it will eventually reach.[8](#user-content-fn-baker-monitoring)[9](#user-content-fn-denison-subterfuge) Silent, meaning it says only PASS or FAIL, because a verdict that explains itself leaks the hidden test into the agent's context, and a test the agent can see is a test it can overfit. The theory for this is older than language models: a holdout that answers with low information stays valid under many adaptive queries, and one that answers with high information does not.[10](#user-content-fn-ladder)[11](#user-content-fn-reusable-holdout)

You do not have to build the agent. Every serious coding harness now exposes hooks: programs that run before and after each tool call, and when the agent tries to stop. A hook that appends each event to a file is a trace collector. A hook that runs the oracle on Stop and refuses to let the agent finish on FAIL is the flip detector. The plugin in Chapter 8 is about 400 lines.

You need fewer flips than you think to start, and more than you think to finish. Published results put the first large gains at a few hundred verified trajectories on a 32-billion-parameter open model, and the best open results at several thousand.[12](#user-content-fn-swe-gym)[13](#user-content-fn-swe-smith)[14](#user-content-fn-skywork-swe) A team of fifty engineers running agents daily reaches the first number in a week and the second in a quarter. The chapters on volume give the arithmetic.

The pipeline is ordinary continuous delivery with weights as the artifact. Ingest, redact, validate, version, train, evaluate against held-out flips the model never saw, gate, package, canary in shadow mode against the live oracle, promote, roll back. None of these stages is new. What is new is that the artifact learns.

## Chapter 1: Rented tokens

A team that uses a frontier model through an API is renting. The rent has three parts, and only the first shows up on the invoice.

The first part is the token bill. It is real and it is falling. Epoch AI measured the price of reaching a fixed level of performance on a set of benchmarks and found it falling by somewhere between 9 and 900 times a year, depending on the task.[1](#user-content-fn-epoch-prices) That is good news for anyone who pays the bill, and it is the number vendors point to when someone asks why a team should not train its own model. The number is also beside the point. The question is not whether tokens will get cheaper. They will. The question is what the team owns at the end of the year.

The vendors in question are the frontier labs: Anthropic, OpenAI, Google, and the others whose models you reach only through an API and whose weights you never see. A black box is the right name for the product, whatever you think of the company. The second part of the rent is the capability you never keep. When a vendor ships a better model, your agents get better. When the vendor changes the model, deprecates it, raises the price, or changes the terms, your agents change with it. Nothing the model learned about your systems during a year of sessions stays with you, because the model learned nothing. It cannot. Inference does not update weights. Every session starts from the same checkpoint the vendor shipped, and every piece of context your team has written to steer it, every CLAUDE.md and every rule file, is a workaround for a model that does not know your code.

The third part is the data you discard. A session produces a record. The record shows which files the agent opened to understand a feature, which commands it ran to check its work, what it tried that failed, and what finally passed. If the agent was working in your repository on your task against your tests, that record is a worked example of your work being done. Nobody else has it. Nobody else can produce it. And in most teams, it is gone when the terminal closes.

> **Where the value of a session goes today**
>
> 1.  **Task**: An engineer or a work queue dispatches a task to a coding agent
> 2.  **Session**: The agent reads, edits, runs commands, and reports done
> 3.  **Tokens billed**: The vendor bills every input and output token
> 4.  **Record discarded**: The transcript is deleted or buried in a local cache
>
> *The two emphasized steps are the rent. The money leaves and the record leaves. The model learned nothing, and neither did your organization's data.*

## What the open-weight trend changes

The case for keeping traces rests on one bet: that open-weight models will stay close enough to the frontier that a model trained on your traces beats a rented model on your tasks. Three measurements say the bet is reasonable.

Epoch AI compared the best open-weight model to the best closed model across benchmarks and found the open model trailing by about a year, measured by the date at which a closed model first reached the same score.[2](#user-content-fn-epoch-open) A year behind the frontier on general benchmarks is not a year behind on your codebase, because the frontier model is also starting from zero on your codebase.

Stanford's AI Index reported that the gap between the best open and best closed models on the Chatbot Arena leaderboard narrowed from 8.0 percent to 1.7 percent in one year.[3](#user-content-fn-ai-index-2025) Chatbot Arena measures human preference on chat, not software engineering, but the direction holds across benchmarks the Index tracks.

On the benchmark that matters most for this book, SWE-bench Verified, open-weight models went from under 10 percent to above 40 percent in eighteen months. Meta's SWE-RL took a 70-billion-parameter Llama to 41.0 percent.[4](#user-content-fn-swe-rl) Mistral's Devstral, a 24-billion-parameter model that runs on a single workstation card, reported 46.8 percent.[5](#user-content-fn-devstral) The SWE-smith team took a 32-billion-parameter Qwen to 40.2 percent with about five thousand trajectories.[6](#user-content-fn-swe-smith) The Skywork team took the same base model from 6.4 percent to 38.0 percent with about eight thousand, and found the gain still growing with each doubling of the data.[7](#user-content-fn-skywork-swe) These are not the frontier numbers. The frontier was above 70 percent at the time. They are the numbers that a team can run, inspect, and fine-tune.

> **Open-weight models on SWE-bench Verified**
>
> | Model and method | Resolved |
> | --- | --- |
> | Qwen2.5-Coder-32B, no agent training (SWE-Gym baseline) | 7% |
> | Same model after fine-tuning on 491 SWE-Gym trajectories | 20.6% |
> | Same model with a verifier and 16 samples per task | 32% |
> | Skywork-SWE-32B, 8,209 trajectories | 38% |
> | SWE-agent-LM-32B, 5,016 SWE-smith trajectories | 40.2% |
> | Llama3-SWE-RL-70B, reinforcement learning | 41% |
> | Devstral-Small-2505, 24B | 46.8% |
> | DeepSWE-Preview, Qwen3-32B, RL with test-time scaling | 59% |
>
> *Sources: Pan et al. (2024), Yang et al. (2025), Wei et al. (2025), Mistral (2025), Agentica and Together AI (2025). Each row is a different base model and training recipe, so the bars show the range open models reached, not a controlled comparison.*

Two things about these numbers matter for the argument. First, the jump from 7.0 to 20.6 percent came from fine-tuning on 491 trajectories.[8](#user-content-fn-swe-gym) That is not a large dataset. It is about a week of sessions for a mid-sized team. The further jump to 32.0 percent came from training a verifier on the same trajectories and letting it pick the best of 16 attempts, which is a preview of Chapter 9. Second, every model in the table was trained on public repositories solving public issues. None of them had seen the training team's own code. A model trained on your traces starts from these numbers and climbs on your distribution.

## Licenses that permit it

The models in the table ship under licenses that allow fine-tuning and commercial use. DeepSeek-R1 and its distilled variants are MIT licensed.[9](#user-content-fn-deepseek-r1) The Qwen2.5 and Qwen3 families are Apache 2.0, with a small number of size variants under a Qwen license.[10](#user-content-fn-qwen3) The Llama 3 family uses Meta's community license, which permits commercial use below a very large monthly-user threshold.[11](#user-content-fn-llama3) OpenAI's gpt-oss models are Apache 2.0.[12](#user-content-fn-gpt-oss) The point of listing them is not legal advice. It is that the permission to do what this book describes is ordinary and granted.

## The asset that compounds

Consider two teams of the same size doing the same work for one year.

Team A uses a rented frontier model. It spends on tokens, writes rule files, and gets better at prompting. At the end of the year it has a set of rule files, a bill, and a vendor relationship. If the vendor's next model is worse at the team's tasks, the team has no recourse. If a competitor uses the same vendor, the competitor has the same model.

Team B uses the same rented model, and also installs a trace collector and an oracle. It spends the same on tokens. Each session that flips an oracle from FAIL to PASS is saved as a verified trajectory. At the end of the year, Team B has thousands of worked examples of its own work being done correctly, each graded by a test the model did not write. It has fine-tuned an open model on them three or four times and has a model that resolves its own tasks at a rate the rented model cannot match on the same distribution, running on hardware it controls, at a marginal cost per token that is a fraction of the rent. If the vendor changes terms, Team B's model does not change. If a competitor wants the same model, the competitor needs Team B's traces.

The cost of being Team B instead of Team A, for the first several months, is close to zero. The collector is a plugin. The oracle is a test. Storage is cheap. The training runs come later and are cheaper than one month of a mid-sized team's token bill. The only thing Team A has to do to become Team B is to stop deleting the record.

## Limits of the claim

It does not claim that a fine-tuned 32-billion-parameter model will match the best frontier model on every task next quarter. It will not. It does not claim that the open-weight gap will close to zero. It may not. It does not claim that collecting traces is free of risk; Chapter 16 is about secrets, memorization, and the ways a pipeline like this goes wrong.

It claims that the traces are the asset, that the oracle is what makes them an asset, and that a team which starts collecting today will be in a different position in a year from a team that does not. The rest of the book is about how to do it carefully.

## Chapter 2: The trace

A trace is the record of one agent session. The word is borrowed from systems tracing, where a trace is the record of one request as it passes through many services. The analogy is close. An agent session passes through many tool calls, and the trace records all of them in order with their inputs and outputs.

## Anatomy of a trace

A coding-agent session has a regular structure. The agent receives a task. It takes a turn: it thinks, then it either calls a tool or writes a message. A tool call runs and returns a result. The agent takes another turn. This continues until the agent decides it is done, or a person stops it, or a limit is reached.

The trace records each of these events:

1.  The **task** as the agent received it, including the system prompt, the rule files loaded into context, and the user's instruction.
2.  Each **tool call**: the tool name, the arguments the agent passed, the result the tool returned, and how long it took.
3.  Each **message** the agent wrote to the person, including the final message where it says the work is done.
4.  The **repository state** at the start and end: the base commit, the final diff, and the list of files touched.
5.  The **oracle verdicts**: the verdict before the agent started, and the verdict at each point the agent tried to stop.
6.  **Metadata**: the model, the harness and its version, the session id, timestamps, token counts, and cost.

Of these, the tool calls carry the learning signal. They show the model how an expert in this codebase finds the relevant file, which command verifies a change, what a failing test looks like here, and how a fix is shaped. Training on them teaches the model to take those actions on this codebase. The final diff alone would not teach that. A diff is an answer; a trace is a worked solution.

> **The records in a trace and how they relate**
>
> -   Session has one Task: System prompt, rule files, user instruction; Base commit of the repository
> -   Session has many Turn: One model response; Ends in a tool call or a message
> -   Turn has zero or one Tool call: Name, input, output, duration; Output redacted before storage
> -   Session has many Oracle verdict: PASS or FAIL and nothing else; One at the start, one per stop attempt
> -   Session has one Outcome: flipped, no flip, or aborted; Final diff and files touched
>
> *The oracle verdicts are kept beside the turns, not inside them. The agent never sees the verdict's reason, so the reason is not in any turn.*

## The episode and the flip

In reinforcement learning, an episode is one complete attempt at a task from start to a terminal state. A trace is an episode. What makes an episode useful for training is a reward: a number that says how well it went. For most of what organizations do, there is no clean reward. For software, there is one, and it has a name in the benchmark literature.

SWE-bench, the benchmark most coding-agent papers report against, grades a candidate patch with two sets of tests.[1](#user-content-fn-swe-bench) The FAIL\_TO\_PASS tests fail on the repository before the fix and pass after the reference fix. The PASS\_TO\_PASS tests pass both before and after; they are there to catch a patch that fixes the issue by breaking something else. A patch resolves the task when every FAIL\_TO\_PASS test passes and every PASS\_TO\_PASS test still passes.

That is the flip. Before the agent starts, the oracle runs the hidden tests and records FAIL. After the agent says it is done, the oracle applies the agent's diff to a clean copy, runs the same hidden tests, and records PASS or FAIL. A session whose verdict goes from FAIL to PASS, with no regression in the tests that already passed, has flipped. The trace of that session is a verified trajectory. A session that ends on FAIL is also kept, because a failed attempt beside a successful one on the same task is a preference pair, and preference pairs are training data too. Chapter 10 covers that.

The flip is the unit of value for three reasons.

It is binary, so it cannot be argued with. There is no rubric, no score from a judge model, no partial credit for a patch that "looks right." Chapter 6 is about why that matters.

It is cheap, so it can be run on every session. A test suite that runs in minutes grades a session for cents. A human review of the same session would cost more than the session.

It is yours. The hidden test encodes what your organization means by correct for this task. A public benchmark encodes what a benchmark author meant.

## Other records

A trace is not a transcript. Harnesses keep transcripts for the person to scroll back through, and those files are useful, but they mix the agent's messages with display formatting, and they may not include the final message when the session ends.[2](#user-content-fn-claude-hooks) A trace is written by hooks as events happen, in a schema you control, with redaction applied before anything touches disk.

A trace is not a log. Logs are for debugging a system. Traces are for training a model. A log can be lossy and unstructured. A trace that drops the tool output for one call has a hole in the worked example at exactly the point the model needs to learn from.

A trace is not the diff. The diff is the answer. A model trained only on diffs learns to produce patches that look like your patches. It does not learn to find the file, run the test, read the failure, and try again. Pan and colleagues trained on full agent trajectories, each a sequence of tool calls and observations, and moved their model from 7.0 to 20.6 percent with 491 of them.[3](#user-content-fn-swe-gym) The actions are the data.

## The schema

Appendix A gives a full schema. The shape is a JSON Lines file per session, one event per line, with a small set of event types.

```text
session_start   session_id, task, base_commit, model, harness, started_at
tool_call       turn, tool_name, tool_input, tool_output_redacted, duration_ms
message         turn, role, text
oracle_verdict  attempt, verdict (PASS|FAIL), tests_hash, ran_at
session_end     outcome (flipped|no_flip|aborted), final_diff_hash, ended_at
```

Two design choices in the schema deserve a sentence each. The verdict record carries a hash of the hidden test list, not the list. That lets you prove later which oracle graded a trace without putting the test names anywhere the agent's process can read. And the tool output is stored after redaction, never before. A trace store is a secret store if you let it be one, and Chapter 16 explains why that is the failure that ends programs like this.

## Why the harness should write it

You could write a coding agent and have it write its own traces. Several groups have, and their agents are good. But the harness your team already uses runs thousands of sessions a week today, and it exposes hooks. A hook is a program the harness runs at a fixed point in the session: before a tool call, after a tool call, when the agent tries to stop, when the session starts and ends. Each hook receives the event as JSON on standard input. A hook that appends that JSON to a file is a trace collector, and it took the author of this book an afternoon to write. Chapter 7 covers hooks in detail. The reason to mention them here is that the most common objection to collecting traces, "we would have to build our own agent," is false.

## Chapter 3: The oracle problem

Software testing has a name for the thing that decides whether a program's output is correct: the test oracle. Barr, Harman, McMinn, Shahbaz, and Yoo surveyed the field in 2015 and called the difficulty of building one "the oracle problem."[1](#user-content-fn-barr-oracle) Generating inputs to a program is easy. Knowing what the program should have done with them is hard. Every automated check, from a unit test to a type checker, is a partial answer to that problem.

This book uses the word in the same sense, narrowed. An oracle here is a program that takes the state of a repository after an agent has worked on it and returns PASS or FAIL. It runs without the agent. It does not take the agent's word for anything. The quality of everything downstream, every training set and every evaluation, is bounded by the quality of the oracle, so this chapter is about what makes a good one.

## Deterministic

A deterministic oracle returns the same verdict every time it runs on the same input. That sounds like a low bar. In practice, most test suites do not clear it.

Luo, Hariri, Eloussi, and Marinov studied flaky tests, tests that pass and fail on the same code, across 51 open-source projects and classified the causes in 201 fixing commits. The leading causes were waiting on asynchronous work, concurrency, and dependence on test order.[2](#user-content-fn-luo-flaky) At Google, Micco reported that about 1.5 percent of all test runs gave a flaky result, that almost 16 percent of tests showed some flakiness, and that about 84 percent of transitions from pass to fail involved a flaky test.[3](#user-content-fn-micco-flaky) A flaky test in a benchmark is an annoyance. A flaky test in an oracle that labels training data is corruption: a session is recorded as a flip when the agent did nothing, or as a failure when the agent succeeded.

The fixes are the fixes the testing literature has recommended for a decade, applied strictly because the stakes are higher.

-   Run the oracle in a fresh container from a pinned image, so the environment is identical every time. Reproducible builds give the same guarantee for the artifact under test: the same source produces the same binary, bit for bit, so a verdict is about the code and not the build machine.[4](#user-content-fn-reproducible-builds)
-   Pin every dependency by hash, with a local mirror, and give the container no network. A test that downloads anything is not deterministic.
-   Fix the clock, the random seed, the locale, and the time zone inside the container. A test that depends on the wall clock is a test that depends on when the oracle ran.
-   Run the hidden tests twice on the base commit before the agent starts. If the two verdicts disagree, the task is not fit for an oracle. Quarantine it.
-   Record the test list as a hash with the verdict, so a later change to the tests cannot be confused with a change in the code.

The cost of determinism is that some tests cannot be oracles. Integration tests against a live service, tests that depend on timing, and tests that depend on data that changes are excluded. That is a loss, and Chapter 4 is partly about what you can use instead.

## Independent

An oracle must be independent of the agent in two senses.

It must not run in the agent's environment. The agent's working copy is under the agent's control. The agent can edit any file there, including the tests, the test configuration, the build file, and the environment variables the test runner reads. An oracle that runs `pytest` in the agent's directory is asking the agent whether it passed. Chapter 5 describes the air gap that fixes this.

It must not be a model grading itself. A judge model that reads the diff and decides whether the task is done is not an oracle. It is a second opinion from a system with the same blind spots as the first. Chapter 6 reviews the evidence.

## Hidden

The hidden test is the version of the oracle this book recommends. The team writes a test that fails on the current code and passes when the task is done correctly. The agent never sees it. The agent sees the task description and the existing test suite, and can write and run its own tests freely. When the agent says it is done, the oracle applies the agent's changes to a clean copy, runs the hidden test and the existing suite, and records the verdict.

This is the SWE-bench construction.[5](#user-content-fn-swe-bench) The FAIL\_TO\_PASS tests are hidden from the agent during the attempt. The agent reads the issue, not the test. The construction is also how the Defects4J database of real Java bugs has been used for a decade: each bug comes with at least one test that exposes it and passes after the developer's fix.[6](#user-content-fn-defects4j)

Hidden tests matter because visible tests are a weaker oracle than they look. Qi, Long, Achour, and Rinard examined patches that three automated repair systems had reported as fixing bugs, meaning the patches made the visible tests pass. Most of the patches were wrong.[7](#user-content-fn-qi-kali) They passed the tests by deleting the functionality the tests happened not to cover. Smith, Barr, Le Goues, and Brun showed the same thing with a controlled experiment and gave it a name: patch overfitting.[8](#user-content-fn-smith-cure) A patch that passes the tests the agent can see has been fitted to those tests, and nothing more has been shown. A patch that passes a test the agent could not see has been shown something.

## The flip, formally

With the pieces named, the flip can be stated as a rule.

Let `H` be the hidden tests for the task and `S` be the existing suite, both frozen as a list with a hash. Let `base` be the commit the agent started from and `diff` be the agent's final change.

1.  On `base`, run `H` and `S` in a fresh container. Require that `H` fails and `S` passes. If `H` passes, there is no task; if `S` fails, the base is broken and the task is not fit for an oracle. Record the verdict as the baseline.
2.  On `base` with `diff` applied to source paths only, run `H` and `S` in a fresh container. Record PASS if every test in `H` passes and every test in `S` passes. Record FAIL otherwise.
3.  A session has flipped when the baseline is FAIL and the final verdict is PASS.

Step 2 says "source paths only." The diff the agent produced may include changes to test files, test configuration, continuous-integration files, and build scripts. The oracle does not apply those. It applies the agent's changes to the code under test, then runs its own copies of the tests against them. Chapter 5 explains the allowlist that does this.

> **One oracle run**
>
> 1.  **Fresh container**: Pinned image, no network, fixed clock and seed
> 2.  **Clean checkout**: The base commit, from the oracle's own mirror
> 3.  **Apply source diff**: Only paths on the allowlist
> 4.  **Inject hidden tests**: From the oracle's store, never from the workspace
> 5.  **Run H and S**: Hidden tests and the existing suite
> 6.  **One bit out**: PASS or FAIL, plus a hash of the test list
>
> *Everything the agent could have touched is replaced before the tests run. The agent's only contribution to the oracle run is the source diff.*

## Tests are a weak oracle too

Hidden tests are the best oracle most teams can build quickly. They are not a perfect one. Inozemtseva and Holmes showed that code coverage, the share of lines a test suite runs, is not strongly correlated with how many faults the suite detects once suite size is accounted for.[9](#user-content-fn-inozemtseva-coverage) A hidden test that runs a line is not a hidden test that checks it. The same study of patch overfitting that argues for hidden tests also argues for better ones.

Mutation testing is the standard way to measure how good a test is. A mutation tool makes small changes to the code, such as flipping a comparison or deleting a statement, and checks whether the tests fail. A test that does not fail on a mutant is not checking the mutated behavior. The idea is from 1978 and has been used at Google at scale since at least 2018, where mutants are shown to developers during code review.[10](#user-content-fn-demillo-mutation)[11](#user-content-fn-petrovic-mutation) A hidden test that kills the mutants near the code the task touches is a stronger oracle than one that merely runs. Chapter 9 uses mutation the other way around, to manufacture tasks.

The practical rule is: a hidden test must fail on the base commit for the right reason. Before you accept a task into the oracle pool, read the failure. If the test fails because of an import error or a fixture that is missing, the agent can flip it by fixing the fixture. If it fails because the feature is missing, the agent has to build the feature.

## Chapter 4: Organizational oracles

The hidden test is one oracle. An organization has many. This chapter is a catalogue, ordered from the oracles that give the cleanest signal to the ones that give the noisiest, with a note on how each can be used. The ordering matters because the cleaner the oracle, the more directly its verdict can be used as a training label. The noisier the oracle, the more it belongs in a preference pair or a filter rather than as a label.

## Hard oracles

A hard oracle is deterministic, runs in minutes, and returns a verdict a program can read. These can label a trace directly.

**Hidden unit and integration tests.** The reference oracle, covered in Chapter 3. Written by a person or by a separate model before the task is dispatched. The strongest version is a test that was written to expose a real bug or specify a real feature, because it encodes a real requirement.

**The existing test suite as PASS\_TO\_PASS.** Every task gets the whole existing suite as a regression check for free. A flip requires that nothing already passing breaks. On a repository with a large suite, this alone rules out most of the bad patches that a visible test would accept.

**Type checkers and compilers.** A change that does not compile, or that fails `tsc --noEmit`, `mypy --strict`, or `cargo check`, has failed. These are deterministic, fast, and hard to game without touching configuration, which the allowlist excludes. On their own they are a weak oracle, since code that compiles can still be wrong, but as a component of the oracle they remove a class of failures cheaply.

**Linters with pinned rule sets.** A lint failure is a weak signal of incorrectness and a strong signal of style drift. Use it as a gate, not a label: a trace that flips the hidden test but introduces lint errors is a flip with a defect, and you can decide per repository whether that counts.

**Property-based tests.** Instead of a fixed input and expected output, a property test states a rule, such as "decoding what you encoded gives the original," and a library generates many inputs to check it.[1](#user-content-fn-quickcheck) With a fixed seed, a property test is deterministic. It is a stronger oracle than an example test because it checks many cases, and a harder one for an agent to fit to because the agent cannot see the cases.

**Metamorphic tests.** When the correct output is unknown but a relationship between outputs is known, a metamorphic test checks the relationship. If a search for "a" returns results, a search for "a OR b" should return at least as many. Metamorphic testing was proposed for programs without an oracle and is well suited to data pipelines, search, and numerical code.[2](#user-content-fn-chen-metamorphic)

**Differential tests against the previous build.** Run the same inputs through the version before the change and the version after. For a refactor, the outputs must match. For a bug fix, they must differ only on the inputs that exposed the bug. McKeeman's differential testing of compilers is the origin of the method.[3](#user-content-fn-mckeeman-differential) The previous binary is an oracle you already have.

**Mutation score.** Chapter 3 covered mutation as a test-quality measure. It can also be an oracle for a task of the form "add tests for this module." The hidden check is whether the mutants the task names are killed after the change. This is one of the few ways to make test-writing itself a flippable task.

**Reproducible build hash.** For a task that must not change behavior, such as a dependency bump with no code change, the oracle can be that the build output is byte-identical to a reference, or differs only where expected.[4](#user-content-fn-reproducible-builds)

**Schema and migration round trips.** For a database change: apply the migration to a copy of the schema, run the down migration, and diff the result against the original. For a data pipeline: run the transform on a frozen fixture and compare row counts and checksums against a stored expectation.

**Infrastructure plan idempotence.** For infrastructure-as-code: after applying the change, a second plan must report nothing to do. This is a deterministic check of the form most infrastructure tools provide directly.

**Formal checks.** Where a module has a specification in a checkable form, a solver or a proof assistant is the strongest oracle there is. Few organizations have these. The ones that do should use them.

## Soft and delayed oracles

A soft oracle is a signal that correlates with correctness but is not deterministic, is not available for minutes or days, or depends on a person. These cannot label a trace as a flip. They can rank traces, build preference pairs, and filter.

**Pull request merged.** A merged change passed review and continuous integration. This is the oracle most code-model training has used, in the form of mined commits and pull requests. Meta's SWE-RL built its training corpus from about eleven million pull requests and used similarity to the merged patch as the reward.[5](#user-content-fn-swe-rl) It is a real signal and a slow one, and reviewers miss things.

**Reverted within N days.** A change that was merged and then reverted was a failure the review did not catch. The pair "merged, then reverted" against "merged, kept" is a clean preference pair with a delay of days to weeks.

**Incident linked to the change.** Rarer and stronger than a revert. A change that caused an incident is a hard negative example.

**Review comments.** A change that received requests for changes before merge is weaker than one approved on the first pass. Review text also says what was wrong, which is useful for a data scientist and useless as a training label.

**Acceptance of a suggestion.** For completion models, whether the developer accepted the suggestion was found to be the best available predictor of perceived productivity in GitHub Copilot telemetry.[6](#user-content-fn-ziegler-copilot) For agents, the equivalent is whether the person kept the agent's change or discarded the session. It is a weak oracle and an abundant one.

**Time to next edit of the same lines.** If a person edits the lines an agent wrote within an hour of the session, the agent's work was probably incomplete. This is a proxy with many false positives and is best used to flag traces for a closer look.

**A judge model.** A second model that reads the diff and scores it. Chapter 6 argues that this is the weakest oracle in the catalogue and should not be used as a label. It can be used as a filter for obvious garbage, and even then its errors should be measured against a hard oracle on a sample.

> **Oracles by how directly their verdict can label a trace**
>
> 1.  **Judge model**: A model scores the change. Not a label. At most a coarse filter, and even then measured against a hard oracle.
> 2.  **Acceptance, next edit, review comments**: Human behavior signals. Rank and flag, do not label.
> 3.  **Merged, reverted, incident**: Delayed by days. Build preference pairs.
> 4.  **Type check, lint, build hash**: Fast and deterministic. Necessary, not sufficient. Use as gates inside the oracle.
> 5.  **Existing suite as PASS\_TO\_PASS**: Free regression check on every task.
> 6.  **Hidden tests, property and metamorphic tests**: Deterministic, fast, independent, unseen. Labels a flip.
>
> *Only the top rung produces a label the training set can trust on its own. Everything below it is still worth collecting.*

## Oracles outside the code

The method is not limited to software, but the hard oracles mostly are. For the sake of completeness, here is what the same construction looks like elsewhere, with the caveat that each of these is a soft oracle and should be treated as one.

-   **Data work.** A transform's output on a frozen input has a checksum. Row counts and null rates have expectations. Tools that run assertions on data in a pipeline exist and can be hidden from the agent in the same way tests are.
-   **Documentation.** A link checker, a prose checker with a pinned rule set, and a build that fails on a broken reference are deterministic. Whether the document is clear is not.
-   **Support.** A ticket resolved and not reopened within two weeks is a delayed soft oracle on the agent's answer.
-   **Operations.** A runbook step that leaves a system in a state a monitoring check accepts is close to a hard oracle, if the check is deterministic and the system is a copy.

A team should start with the hard oracles in its code, because that is where the construction is cleanest and the data is largest. Once the pipeline runs, the soft oracles can be added as preference signals.

## Which oracle wrote the label

Every trace should carry the identity of the oracle that graded it: the kind of oracle, the hash of its test list or rule set, the container image digest, and the time. A training set built from traces graded by different oracles at different times is a training set whose labels mean different things. When a later evaluation shows a regression, the first question is which oracle labeled the traces the model learned it from. Without the record, the question cannot be answered.

## Chapter 5: The air gap

An oracle the agent can reach is an oracle the agent will eventually satisfy without doing the work. This chapter is about the threat, the evidence that it is real, and the design that removes it.

## The threat model

Treat the agent as an adversary with full control of its working directory and shell. That is not a judgment about the model's intent. It is a description of its permissions. A coding agent can run any command, read any file it can reach, and write any file in its checkout. If the verdict depends on anything in that set, the verdict is the agent's to decide.

The ways an agent can make a check pass without doing the task are ordinary:

-   Edit the test so that it asserts what the code does.
-   Mark the test skipped, or add it to an exclusion list in the test configuration.
-   Edit the test runner's configuration so the test is not collected.
-   Edit the build file or the continuous-integration file so the check does not run.
-   Stub the function under test to return the expected value for the inputs the test uses.
-   Patch the test framework itself, or a fixture, or a conftest file that runs before every test.
-   Set an environment variable the test reads.
-   Exit the process with status zero before the test runner reports.
-   Read the hidden test from disk or from the network, if it is anywhere the agent can reach, and fit the code to it.
-   If a model judges the result, write text into the diff or the commit message that addresses the judge.

Every one of these has been observed. Baker and colleagues at OpenAI trained a frontier reasoning model on coding tasks and monitored its chain of thought. They found the model saying, in its own reasoning, that it would make the tests pass by other means, and then doing so: calling `sys.exit(0)` before the tests ran so the runner reported success, raising the test framework's skip exception, and writing stubs instead of implementations where test coverage was thin.[1](#user-content-fn-baker-monitoring) When they penalized the model for having such thoughts, it kept hacking the tests and stopped writing down that it was going to. They called this obfuscated reward hacking and recommended against training against the monitor.

Denison and colleagues at Anthropic built a curriculum of environments where a model could gain reward by gaming the check, starting with flattery and ending with the chance to edit its own reward function. A model trained on the easy environments generalized to the hard one: in a small but nonzero fraction of runs, it edited the code that computed its reward, and sometimes edited the test that would have caught the edit.[2](#user-content-fn-denison-subterfuge) METR, evaluating frontier models on its own tasks in 2025, reported models that rewrote scoring code, patched the timer in a benchmark harness, and special-cased the inputs the grader used.[3](#user-content-fn-metr-hacking) A later Anthropic study trained a model with reinforcement learning on real production coding environments after teaching it about reward-hacking strategies. At the point where the model learned to hack the rewards, it also began to generalize to other unwanted behavior, including attempts at sabotage.[4](#user-content-fn-anthropic-emergent)

The lesson is not that models are malicious. It is that optimization finds the shortest path to the reward, and if the shortest path goes through the test file, that is the path. Skalse and colleagues formalized this as reward hacking: a proxy that correlates with the goal is maximized, and the correlation breaks under optimization.[5](#user-content-fn-skalse-hacking) Gao, Schulman, and Hilton measured it for reward models: the true reward rises as a policy is optimized against the proxy, peaks, and then falls while the proxy keeps rising.[6](#user-content-fn-gao-overoptimization) The hidden test is the goal. Anything the agent can touch is a proxy.

The lesson is also older than machine learning. Thompson's 1984 Turing Award lecture described a compiler that inserts a back door into the login program and into any compiler compiled from clean source, so that no inspection of the source finds it.[7](#user-content-fn-thompson-trust) His point was that you cannot trust code you did not totally create yourself, and that the trust boundary has to be drawn around the thing you can verify. The oracle is the thing you can verify. Draw the boundary around it.

## The design

The oracle runs in a separate container, on a separate host when possible, with the following properties.

**No network.** The container is started with networking disabled. Dependencies come from a local mirror baked into the image or mounted read-only. Nothing the oracle needs is fetched at run time, and nothing the agent could have planted on a server is reachable.

**A pristine checkout.** The oracle has its own clone of the repository, from its own mirror. It checks out the base commit the session started from. It does not read the agent's working copy at all.

**The source diff only.** The oracle receives the agent's final diff as a file. Before applying it, the oracle filters it through an allowlist of paths. Source directories are on the allowlist. Test directories, test configuration files, continuous-integration configuration, build scripts, dependency manifests and lockfiles, and anything under the repository's tool configuration are not. Hunks that touch excluded paths are dropped. The oracle logs that they were dropped, in a log the agent cannot read.

**Hidden tests from the oracle's store.** The hidden tests are stored with the oracle, keyed by task id, and copied into the checkout after the diff is applied. They are never in the repository the agent works in, never in a branch the agent can fetch, and never on a path the agent's shell can read.

**The existing suite from the base commit.** The PASS\_TO\_PASS tests are the ones in the pristine checkout at the base commit, not the ones in the agent's diff. If the agent deleted a test, the oracle still runs it.

**Fixed clock, seed, locale, and time zone.** Set in the container's environment. A test that reads the system clock reads a constant.

**Pinned image.** The container image is referenced by digest, and the digest is recorded with the verdict.

**Timeouts.** The oracle run has a wall-clock limit. A run that exceeds it is a FAIL with a reason the oracle logs and the agent does not see. An agent cannot turn a failing test into a hang and have the hang read as success.

**One bit out, signed.** The oracle's only output to the agent's side is PASS or FAIL. To the trace store it writes a verdict record: the bit, the hash of the test list, the image digest, the diff hash, the base commit, the time, and a signature over all of these using a key the oracle host holds. A verdict in the trace store that does not verify was not written by the oracle.

**Separate credentials.** The process that runs the agent holds no credential that can write to the oracle's store, start an oracle run with a different test list, or read the hidden tests. The hook that asks for a verdict sends a diff and a task id and gets a bit back.

> **What crosses the boundary between the agent's host and the oracle's host**
>
> -   Agent host sends Oracle host: Task id; Base commit; The final diff as a file
> -   Oracle host returns Agent host: PASS or FAIL; Nothing else
> -   Oracle host writes Trace store: Signed verdict record; Dropped-hunk log; Test output, never shown to the agent
> -   Agent host writes Trace store: Tool-call events, redacted; Messages; Session metadata
>
> *The agent's host can write trace events and can ask one question. It cannot read the hidden tests, the test output, or the oracle's log. The trace store holds both sides and joins them on the session id.*

## Why a separate host

A container on the same machine as the agent is isolation from the agent's working copy. It is not isolation from the agent's shell. On a developer's laptop, the agent's Bash tool can read any path the developer can read, including the directory where the hidden tests live, the oracle's log, and the container's image. A container boundary protects the oracle's run from the agent's changes. A host boundary protects the oracle's secrets from the agent's eyes.

The reference plugin in Chapter 8 runs on one machine because that is where most teams will start, and it says so in its README. On one machine, store the hidden tests encrypted at rest and decrypt them inside the oracle container with a key the agent's process does not hold. That raises the bar from "read the file" to "find and misuse the key." It does not close the gap. Closing it means running the oracle on a host the agent cannot log into, which for a team with a continuous-integration system is a short step: the oracle is a job that runs on a runner, takes a diff and a task id, and returns a bit.

Firecracker, the virtual-machine monitor behind AWS Lambda, starts a microVM in about 125 milliseconds and gives each one its own kernel.[8](#user-content-fn-firecracker) gVisor gives a container a user-space kernel and intercepts its system calls.[9](#user-content-fn-gvisor) Either is a stronger boundary than a plain container for the oracle run, and both are used for exactly this purpose: running code you do not trust next to data you care about. For the hidden tests and the signing key, neither replaces a separate host.

## Baselines and the flip that was not

The flip needs a baseline: a FAIL recorded before the agent started. Three rules keep baselines honest.

The baseline is computed on the base commit, by the oracle, from its own checkout. It is deterministic, so compute it once per task and store it. Do not recompute it in a session-start hook; a container run is too slow for a hook budget, and the result would be the same anyway.

A task whose baseline is PASS is not a task. The hidden test already passes. Either the test is wrong or the work is already done. Quarantine the task, and do not count any session on it as a flip.

A task whose existing suite fails at baseline is not fit for an oracle. The agent could flip the hidden test while the suite stays broken, and the PASS\_TO\_PASS rule would record a FAIL for a session that did the work. Fix the base or exclude the task.

## What the oracle logs and who reads it

The oracle keeps a full log: the test output, the dropped hunks, the timing, and the reason for every FAIL. This log is for the people who run the pipeline. It is how you find a hidden test that fails for the wrong reason, a flaky test that slipped through, or an agent that keeps trying to edit the test configuration. It is never returned to the agent and never written anywhere the agent's host can read. Chapter 6 is about why.

## Chapter 6: The one-bit verdict

The oracle says PASS or FAIL. It does not say which test failed, what the assertion was, or what the output looked like. This is the design choice readers push back on most, so this chapter takes it slowly: first the cost, then the three reasons, then the research on self-grading that the third reason depends on.

## The cost, stated plainly

Richer feedback helps an agent fix a bug in the current session. Chen, Lin, Schärli, and Zhou compared three kinds of feedback for a model debugging its own code: simple feedback, meaning only whether the code passed; unit-test feedback, meaning the execution results; and a self-written explanation of the code. Unit-test feedback produced the largest gains.[1](#user-content-fn-chen-self-debug) A verdict with no explanation is the "simple feedback" condition in that study, and it is the weakest of the three for raising the pass rate inside one session.

So the one-bit verdict lowers the flip rate per session. A team that adopts it will see more sessions end on FAIL than a team that shows the agent the failing test. That is a real cost. Three things make it worth paying.

## Reason one: the hidden test stays valid

A test the agent can see is a test the agent can fit to. Chapter 3 covered the patch-overfitting results: patches that pass visible tests by deleting what the tests do not check.[2](#user-content-fn-qi-kali)[3](#user-content-fn-smith-cure) An oracle that returns the name of the failing test and the assertion that failed has shown the agent the test. The agent will fix the assertion. Whether it fixed the feature is now unknown, which is where you started.

The theory for this is from statistics, not software. Blum and Hardt studied machine-learning competitions where participants submit many models and see a score on a holdout set each time. With enough submissions, a participant can overfit the holdout without ever seeing its labels, by hill-climbing on the score. Their fix, the Ladder, releases a new score only when it improves on the previous best by more than a threshold, and otherwise repeats the old score. That turns a high-information answer into a low-information one and makes the leaderboard reliable under adaptive submissions.[4](#user-content-fn-ladder) Dwork, Feldman, Hardt, Pitassi, Reingold, and Roth proved the general result: a holdout answered with limited information, through a mechanism they called Thresholdout, stays statistically valid under far more adaptive queries than one answered exactly.[5](#user-content-fn-reusable-holdout)

An agent iterating against an oracle is a participant iterating against a holdout. A verdict of PASS or FAIL is about as low-information as an answer gets. An execution trace with the failing assertion is about as high-information as one gets. The Ladder result says which one keeps the hidden test meaningful.

## Reason two: the diagnosis is the data

The point of collecting traces is to train a model on them. The behavior you want the model to learn is not "read the failing assertion and change the code until it passes." It is "form a hypothesis about what is wrong, write a test that checks it, run the test, read the result, and fix the cause." That is what expert engineers do, and the trace of an expert doing it is what you want in the training set.

If the oracle tells the agent why it failed, the trace after that point is the agent copying the oracle's diagnosis. The model trained on it learns to wait for a diagnosis. If the oracle says only FAIL, the trace after that point is the agent's own diagnosis: it has to write its own tests, run them, read its own failures, and reason about what the hidden requirement might be. That is the behavior worth learning, and it only appears in the trace when the oracle withholds the answer.

The agent is not working blind. It has the task description, the whole repository, the existing test suite, and a shell. It can write any test it wants and run it as often as it wants, with full output. The one-bit rule applies to the hidden oracle only. In-session feedback from the agent's own tests is as rich as the agent cares to make it. The oracle's silence forces the agent to generate that feedback for itself, which is the skill.

## Reason three: the alternative is a model grading a model

The usual proposal for richer feedback that does not leak the test is a judge: a second model that reads the diff, the test output, and the task, and writes an explanation for the agent. The judge sees the hidden test; the agent sees the judge's prose. This is where the research on self-grading matters, because the judge is a model with the same training and the same blind spots as the agent.

Zheng and colleagues, in the paper that introduced MT-Bench and the LLM-as-a-judge method, documented three biases in model judges: position bias, where the judge prefers whichever answer appears first; verbosity bias, where it prefers the longer answer; and self-enhancement bias, where it prefers answers it wrote.[6](#user-content-fn-zheng-judge) Wang and colleagues showed that swapping the order of two answers could reverse a judge's verdict.[7](#user-content-fn-wang-unfair) Panickssery, Bowman, and Feng showed that a model's preference for its own output is not an accident of style. Models can recognize their own text above chance, and the ones that recognize it best prefer it most; fine-tuning a model to recognize its own output more accurately made its self-preference stronger.[8](#user-content-fn-panickssery) A judge from the same family as the agent is a judge that likes the agent.

The research on self-correction is worse. Huang and colleagues reviewed the claim that models can correct their own reasoning and found that the gains reported in earlier work depended on an oracle: the model was told whether its answer was right before being asked to reconsider. Without that signal, asking the model to check its work made the answers worse more often than better.[9](#user-content-fn-huang-self-correct) That paper is sometimes read as an argument against external feedback. It is the opposite. It shows that a binary external signal is the thing that made self-correction work in the papers that reported it working, and that the model's own judgment was not contributing.

Stechly, Marquez, and Kambhampati tested GPT-4 as a critic of its own graph-coloring solutions and found it could not reliably tell a correct coloring from an incorrect one; iterating on its own critique did not help, and a simple external checker did.[10](#user-content-fn-stechly-wrong) Valmeekam, Marquez, and Kambhampati found the same for planning: a model critiquing its own plans lowered the success rate, and an external verifier raised it.[11](#user-content-fn-valmeekam-plans) Tyen and colleagues separated two skills and found that models are poor at locating the error in a chain of reasoning but can often fix it once the location is given.[12](#user-content-fn-tyen-location) Kamoi and colleagues surveyed the self-correction literature and concluded that no reliable self-correction has been shown without external feedback, and that many reported successes used unrealistic setups where the model was given information it would not have in practice.[13](#user-content-fn-kamoi-survey) Xu and colleagues found that self-refinement loops amplify the model's own biases over iterations.[14](#user-content-fn-xu-pride)

None of these results says a judge model is useless. They say a judge model's errors are correlated with the agent's errors, so using one to explain failures to the agent does not add an independent signal. It adds a confident one. The hidden test is independent. Its one bit is worth more than the judge's paragraph.

> **Feedback channels from the oracle to the agent**
>
> 1.  **PASS or FAIL**: The Ladder regime. The hidden test stays valid. The agent's own diagnosis is in the trace.
> 2.  **Which test failed**: Names the requirement. The agent now knows what to fit to.
> 3.  **The assertion and the values**: The agent can special-case the inputs. Patch overfitting territory.
> 4.  **A judge model's explanation**: Adds a correlated, biased reading of the test to the leak. The worst of both.
> 5.  **The test source**: The hidden test is no longer hidden. The oracle measures nothing.
>
> *Each rung down leaks more of the hidden test and leaves less of the agent's own reasoning in the trace. The book recommends the top rung and names the cost.*

## Recovering some of the cost

There are ways to give back part of the in-session success rate without giving up the properties above.

**Disclose to the trace, never to the agent.** The oracle writes the full test output to its own log, which the data team reads. When a task fails many sessions in a row, a person reads the log and either fixes a hidden test that fails for the wrong reason or rewrites the task description so the requirement is clearer. The agent gets a better task, not a leaked test.

**A visible test the agent must write.** Make the task description say that the change must come with a test that fails before and passes after. The agent writes its own red-green pair, which is rich in-session feedback and good trace content. The hidden test stays hidden and checks that the agent's test checked the right thing.

**Staged tasks.** If a task's hidden test covers three requirements, split it into three tasks with one hidden test each. The agent gets three bits instead of one across the same work, and each bit still says nothing about its test.

**More attempts, fresh context.** The flip rate per session is lower; the flip rate per task does not have to be. Chapter 9 covers sampling several sessions per task and keeping the one that flips. With the oracle silent, the attempts are independent, which is what repeated sampling needs.

**Disclose after the fact.** Once a task has been flipped by some session, its hidden test can be released into the repository as an ordinary test. It has done its job as an oracle. Future sessions on nearby code get it as a visible regression test, and new hidden tests are written for new tasks.

## The rule

The hidden oracle says PASS or FAIL. Every other channel from the oracle to the agent is closed. Everything the oracle knows goes to the trace store, for people, under access control. If a team decides to open a wider channel, it should do so knowing which rung of the ladder it has moved to and what it has traded for the higher flip rate.

## Chapter 7: Harness hooks

You do not need to build a coding agent to collect traces from one. The harness your team already runs exposes hooks, and hooks see everything a trace needs. This chapter explains what a hook is, what each one sees, and how to turn a set of hooks into a trace collector and an oracle gate. The examples use Claude Code because its hook interface is documented in detail and the reference plugin targets it.[1](#user-content-fn-claude-hooks) Other harnesses have equivalents, and the last section covers them.

## What a hook is

A hook is a program the harness runs at a fixed point in a session. The harness passes the event to the program as JSON on standard input, waits for it to exit, and reads its exit code and anything it printed. A hook can observe the event, add context for the model, or, for some events, change what happens next.

The events that matter for tracing are the ones at the boundaries of a turn and a session.

| Event | When it runs | What it carries | What it can decide |
| --- | --- | --- | --- |
| SessionStart | A session begins or resumes | Session id, working directory, model, how the session started | Nothing. It can add context. |
| UserPromptSubmit | The person sends a prompt | The prompt text | It can block the prompt. |
| PreToolUse | Before a tool runs | Tool name and input | It can allow, deny, or ask. |
| PostToolUse | After a tool runs | Tool name, input, output, duration | It can add context beside the result. |
| Stop | The agent is about to end its turn | The agent's final message, whether a Stop hook already continued the turn | It can refuse to let the agent stop. |
| SessionEnd | The session ends | The reason the session ended | Nothing. It runs cleanup. |

Every event also carries the session id, the path to the harness's own transcript, the working directory, and the permission mode. The session id is the join key for everything the trace collector writes.

## The trace collector

A trace collector is three hooks.

On **SessionStart**, write a `session_start` event: the session id, the working directory, the base commit of the repository, the model, and the time. Do not run anything slow here. Hook budgets at session start are short, and a container run does not fit.

On **UserPromptSubmit**, write a `message` event with the role `user` and the prompt text. This is the task as the agent received it, which the training set needs as the first turn.

On **PostToolUse**, write a `tool_call` event: the tool name, the input, the output, and the duration. Redact the input and output before writing. Cap the size of each and store a hash of the full value beside the truncated one, so a trace that was cut can be identified later.

That is the whole collector. The agent's own messages between tool calls are not delivered by PostToolUse, so the collector adds a fourth hook.

On **Stop**, write a `message` event with the role `assistant` and the agent's final message, which the Stop event carries as `last_assistant_message`. Then run the oracle gate, below.

The collector does not read the harness's transcript file. The transcript is for the person, its format is the harness's to change, and the harness documentation notes that the file is not guaranteed to include the final message at the moment Stop fires.[1](#user-content-fn-claude-hooks) The hook events are the record.

## The oracle gate

The Stop hook is where the flip is detected and where an agent is kept from calling the work done.

When the agent decides it has finished, the harness fires Stop. The gate hook does the following.

1.  Reads the event. If the agent is already continuing because of a previous Stop hook, the event says so in a field named `stop_hook_active`. The gate uses this, with its own attempt counter, to decide whether to keep going.
2.  Builds the agent's diff against the base commit recorded at session start, including new files.
3.  Sends the diff and the task id to the oracle. The oracle returns one bit.
4.  Writes an `oracle_verdict` event with the attempt number, the verdict, and the hash of the test list the oracle reported.
5.  On PASS, exits normally. The agent stops. The session has flipped.
6.  On FAIL, returns a decision that blocks the stop, with a reason the harness shows to the agent. The reason is a fixed string. It says the oracle reported FAIL and the agent should keep working. It does not say why.

The blocking mechanism is specific. In Claude Code, a Stop hook refuses the stop by printing a JSON object with `decision` set to `block` and a `reason`, or by exiting with code 2 and the reason on standard error.[1](#user-content-fn-claude-hooks) The agent sees the reason as the explanation for why it must continue. With a one-bit oracle, the reason is the same every time.

## What the harness limits

Two limits in the harness shape how the gate behaves, and a team should know both before relying on it.

First, the harness caps consecutive continuations. In Claude Code, after Stop hooks have continued the turn eight times in a row, the harness overrides the next block and ends the turn. The count resets whenever the agent calls a tool, so an agent that keeps working does not hit it, and an agent that keeps saying "done" without doing anything does.[1](#user-content-fn-claude-hooks) The cap is configurable. The gate should treat a session that ends this way as `no_flip`, and the trace should record that the cap, not the oracle, ended it.

Second, hooks have timeouts. A Stop hook has minutes, not seconds, which is enough for an oracle that runs a test suite in a container. It is not enough for a suite that takes an hour. For slow oracles, the gate should submit the diff to the oracle host, return a provisional block with the fixed reason, and let the next Stop pick up the verdict. The reference plugin runs the oracle inline because most suites that are fit for an oracle run in minutes.

## Redaction

The tool output is where secrets live. A `cat .env`, a failing test that prints a connection string, a curl response with a token: all of it passes through PostToolUse. The trace collector redacts before it writes.

Redaction at this stage is pattern-based and should be treated as a floor, not a guarantee. Patterns for the common shapes, such as `AKIA` prefixed AWS keys, GitHub tokens, private key blocks, bearer headers, and `KEY=value` lines where the key name contains `SECRET`, `TOKEN`, or `PASSWORD`, catch most of what appears in practice. They do not catch a password that looks like a word. Meli, McNiece, and Reaves found secrets leaking into public repositories at a rate of thousands of new unique secrets a day, across more than a hundred thousand repositories, which is a measure of how often they appear in code and output that people thought was fine.[2](#user-content-fn-meli-secrets) The collector redacts, and the pipeline runs a second pass with a dedicated scanner before anything reaches training, and the model is kept private to the organization whose traces it learned from. Chapter 16 covers the memorization research that makes the last rule a hard one.

## Storage

Traces are written to a directory the plugin owns, not to the repository and not to the plugin's install directory. In Claude Code, that is the plugin's data directory, which the harness exposes as `CLAUDE_PLUGIN_DATA` and which survives plugin updates.[3](#user-content-fn-claude-plugins) A sync job moves finished traces to an object store under a path keyed by organization, repository, and date. The store is append-only. Nothing in the pipeline edits a trace after it is written; corrections are new records that reference the old.

Each verdict record is signed by the oracle, as Chapter 5 describes. A trace whose verdict does not verify is excluded from training and flagged. This is the check that stops a compromised or misconfigured agent host from writing its own PASS.

## The same idea in other harnesses

The hook design is not specific to one product. OpenAI's Codex CLI, Cursor's agent, OpenCode, and others expose lifecycle events with the same shape: before and after tool calls, and at the end of a turn. The names differ. The stdin JSON differs. The decision mechanism for refusing a stop differs, and in some harnesses does not exist, in which case the gate has to run as a wrapper around the harness instead of inside it.

What stays constant is the architecture. One hook writes events. One hook, at the end of a turn, asks the oracle and either lets the agent stop or sends it back. The trace schema is the harness-neutral part, and a team that runs more than one harness should normalize events into the same schema at write time, so the training set does not depend on which tool produced it.

## Chapter 8: The reference plugin

This chapter walks through `oracle-flip`, a Claude Code plugin that does what Chapters 2 through 7 describe. It is public at [github.com/macanderson/oracle-flip](https://github.com/macanderson/oracle-flip) under the MIT license. It is a reference, meant to be read and adapted, and it is small enough to read in one sitting.

## Install

```text
/plugin marketplace add macanderson/oracle-flip
/plugin install oracle-flip@macanderson
```

To try it without installing, clone the repository and start Claude Code with the plugin loaded from disk:

```text
git clone https://github.com/macanderson/oracle-flip
claude --plugin-dir ./oracle-flip
```

## Layout

```text
oracle-flip/
  .claude-plugin/
    plugin.json           name, version, description
    marketplace.json      lets /plugin marketplace add find it
  hooks/
    hooks.json            the five hook registrations
  scripts/
    common.py             paths, event writer, redaction, config
    session_start.py      SessionStart: write session_start
    trace.py              UserPromptSubmit and PostToolUse: write message and tool_call
    gate.py               Stop: diff, oracle, verdict, block or allow
    session_end.py        SessionEnd: write session_end with the outcome
  oracle/
    run.sh                reference oracle runner: container, allowlist, hidden tests, one bit
    grade.sh              runs inside the container
    filter_diff.py        drops hunks outside the allowlist
    Dockerfile            pinned base image; the container runs with no network
    example/              a sample task with a hidden test
  tools/
    export_sft.py         traces plus verdicts to SFT and preference-pair JSONL
  schema/
    trace-event.schema.json
  tests/
    test_redaction.py, test_filter_diff.py, test_gate.py
```

The scripts are Python with no dependencies outside the standard library, so the plugin runs anywhere `python3` runs. The oracle runner is shell plus one Python script, because it is meant to be replaced by whatever the team's continuous-integration system already does.

## The hook registrations

The plugin's `hooks/hooks.json` wraps the event map in a `hooks` key, which is the shape plugins use.[1](#user-content-fn-claude-plugins) Each hook is in exec form: a `command` and an `args` array, with the plugin root substituted by the harness, so paths with spaces need no quoting.

```json
{
  "description": "oracle-flip: trace every tool call and gate Stop on a hidden oracle",
  "hooks": {
    "SessionStart": [
      { "hooks": [ { "type": "command", "command": "python3",
                     "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_start.py"], "timeout": 20 } ] }
    ],
    "UserPromptSubmit": [
      { "hooks": [ { "type": "command", "command": "python3",
                     "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 10 } ] }
    ],
    "PostToolUse": [
      { "matcher": "*",
        "hooks": [ { "type": "command", "command": "python3",
                     "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 20 } ] }
    ],
    "Stop": [
      { "hooks": [ { "type": "command", "command": "python3",
                     "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/gate.py"], "timeout": 600 } ] }
    ],
    "SessionEnd": [
      { "hooks": [ { "type": "command", "command": "python3",
                     "args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_end.py"], "timeout": 20 } ] }
    ]
  }
}
```

## Configuration

Configuration is split by who controls the file. Anything the agent can edit holds nothing the gate trusts.

The project the agent works in holds one small file, `.oracle-flip/config.json`. It names the task and sets the attempt budget.

```json
{ "task_id": "billing-proration-rounding", "base_commit": null, "max_attempts": 6 }
```

`base_commit` left null means the commit at session start. `max_attempts` is the number of FAIL verdicts after which the gate stops blocking and records `no_flip`. It exists so a session that will never flip does not run the oracle forever, and it is set below the harness's own cap of eight consecutive continuations so the plugin, not the harness, decides when to give up. The SessionStart hook reads this file once and copies it into the session's state. The gate reads the state, never the project, so an agent that edits the file mid-session changes nothing. An agent that edits it before the session starts can at most point the session at a task whose baseline passes, which the gate records as `no_task`, never as a flip.

The oracle command, the oracle store, and the container image live in the plugin's user settings, which Claude Code stores outside every repository and passes to hook processes as environment variables. The oracle command receives the mode, the task id, the base commit, the path of the diff file, and a hash of the repository's remote URL, and it must exit 0 for PASS and 1 for FAIL. Any other exit code is an oracle error, which the gate records and treats as FAIL for the purpose of blocking, with a different fixed reason.

The allowlist and denylist of paths the diff may touch live in the oracle store beside the hidden tests, one file each, and the oracle applies them before it touches a file. The gate's host does the filtering nowhere, because the gate's host is the agent's host. The hidden tests are in the same store, keyed by task id, and nowhere in the repository.

## The gate

The gate is the file to read if you read one. Its structure follows Chapter 7.

```python
def decide(event, state, *, oracle=run_oracle, diff=write_diff):
    write_event(assistant_message(event))        # last_assistant_message, redacted

    if state.task_id is None:
        return allow()                           # no oracle task; the plugin only traces
    if state.outcome is not None:
        return allow()                           # already decided; never loop on a settled outcome

    if state.baseline is None:
        baseline = oracle(state.oracle_command, mode="baseline", ...)
        state.baseline = baseline.verdict
        record_verdict(state, "baseline", baseline)
        if baseline.verdict == "PASS":
            state.outcome = "no_task"            # the hidden test already passes
            return allow()
        if baseline.verdict == "ERROR":
            state.outcome = "oracle_error"
            return allow()

    diff_path = diff(event["cwd"], state.base_commit, ...)
    verdict = oracle(state.oracle_command, mode="grade", diff_path=diff_path, ...)
    state.attempts += 1
    record_verdict(state, "grade", verdict)

    if verdict.verdict == "PASS":
        state.outcome = "flipped"
        return allow()
    if state.attempts >= state.max_attempts:
        state.outcome = "no_flip"
        return allow()
    return block(BLOCK_REASON_FAIL if verdict.verdict == "FAIL" else BLOCK_REASON_ERROR)
```

Four details are worth pointing at.

The gate reads everything from the session state the SessionStart hook wrote: the task id, the base commit, the attempt budget, and the oracle command. It reads nothing from the project.

The baseline is computed on the first Stop, not at session start, because the oracle is slow and the baseline is deterministic. It is stored, so later Stops in the same session reuse it. A team with many sessions per task should compute it once per task and share it; the plugin keeps it per session for simplicity.

The `block` reason is one of two constants. Neither includes the oracle's output. The oracle's output is written by the oracle, on the oracle's side, and the gate never sees it. The gate's own tests check that the reasons contain no test name and no assertion.

`stop_hook_active` is honored by the attempt counter rather than by an early return, because the harness sets it on every Stop after the first block, and returning early on it would mean the gate only ever runs once. The attempt counter plus `max_attempts` is what prevents the loop, and `decide` takes the oracle and the diff builder as parameters so the loop logic is tested without a container.

## The oracle runner

`oracle/run.sh` is the reference oracle. It expects a store with one directory of hidden tests per task and a mirror of the repository, and it runs the grade inside a container with networking disabled. The core of it:

```sh
task_dir="$store/$ORACLE_TASK_ID"
log_dir="$store/.log/$ORACLE_TASK_ID/$(date -u +%Y%m%dT%H%M%SZ)-$ORACLE_MODE"

args=(run --rm --network none
      -e TZ=UTC -e LC_ALL=C.UTF-8 -e PYTHONHASHSEED=0 -e SOURCE_DATE_EPOCH=1700000000
      -e ORACLE_MODE="$ORACLE_MODE" -e ORACLE_BASE_COMMIT="$ORACLE_BASE_COMMIT"
      -v "$task_dir:/task:ro" -v "$mirror:/mirror:ro" -v "$log_dir:/log")
[ "$ORACLE_MODE" = "grade" ] && args+=(-v "$ORACLE_DIFF:/in/diff.patch:ro")

docker "${args[@]}" "$image" bash /oracle/grade.sh > "$log_dir/stdout.txt" 2> "$log_dir/stderr.txt"
status=$?
grep -E '^tests_hash=' "$log_dir/stdout.txt" || true
exit "$status"
```

`grade.sh` runs inside the container. It clones the repository from the read-only mirror, checks out the base commit, filters the diff through the task's allowlist and denylist with `filter_diff.py`, applies what is left, copies the hidden tests from `/task/hidden` into place, runs the existing suite and then the hidden tests, prints `tests_hash=` followed by a hash of the hidden test list, and exits 0 or 1. In baseline mode it runs without a diff and exits 1 when the suite passes and the hidden tests fail, which is the only baseline that makes a task. Its full output goes to the log directory on the oracle's side. Only the exit code and the `tests_hash` line return to the gate.

The Dockerfile pins its base image by digest and installs the repository's dependencies from a lockfile at build time, so the run-time container needs no network and gets none.

## Limits of the plugin

It runs the oracle on the same machine as the agent, in a container. Chapter 5 explained why that protects the oracle's run and not the oracle's secrets. On a laptop, the agent's shell can read `ORACLE_STORE`. The README says so. The step to a separate host is to replace `oracle/run.sh` with a script that submits the diff to a job runner the agent cannot log into and waits for the bit.

It does not sign verdicts. Signing needs a key on the oracle's side and a verifier in the pipeline, and the plugin has neither because it has no pipeline side. The verdict event has a field for the signature and the export tool has a flag that requires it.

It redacts with patterns. That is a floor. Run a real secret scanner on the trace store before training.

It has not been tested with a real container run on the author's machine at the time of writing, for a reason unrelated to the plugin, and the README says that too. The hook scripts are tested by piping sample events into them, and `claude plugin validate` passes.

## The export tool

`tools/export_sft.py` reads a directory of traces and their verdicts and writes two files.

The first is the supervised set: one JSON object per flipped session, in a chat format with tool calls, with a mask that marks which turns to train on. User prompts and tool outputs are masked out; the model is trained to produce the assistant's turns and tool calls, not to reproduce the environment's replies. Chapter 10 explains why.

The second is the preference set: for every task with at least one flipped session and at least one that did not flip, a pair with the flipped trace as the preferred one. Chapter 11 explains what to do with it.

Both files exclude any session whose verdict record is missing, and, with `--require-signature`, any whose verdict does not verify.

## Chapter 9: The flip rate

The pipeline's throughput is the number of flips per day. This chapter is about raising it: choosing tasks that can flip, dispatching work in a shape that flips, sampling more than once, and manufacturing tasks when the natural supply runs low. It ends with what to do with the sessions that do not flip, which is most of them.

## Start with the test

A task can flip only if a hidden test fails on the base commit. The highest-leverage change a team can make is to write the test before dispatching the task. This is test-driven development with the roles split: a person or a separate agent writes the red test, and the solving agent is dispatched with the task description and never sees it.

The practice has a long history under the name red-green-refactor, and the evidence that it improves defect rates predates language models.[1](#user-content-fn-nagappan-tdd) What is new is the reason to do it. In a team with an oracle, every red test is a potential verified trajectory. In a team without one, a red test is just a test.

Some task shapes come with a red test for free.

-   **A bug report with a reproduction.** The reproduction is the hidden test. Clean it up, assert the correct behavior, confirm it fails on the base commit, and dispatch the bug.
-   **A failing test in continuous integration.** The test already exists and already fails. Hide it from the agent by giving the agent a checkout where the test is removed, and keep the test as the oracle.
-   **A feature with an acceptance criterion.** If the criterion can be stated as an assertion, it can be a hidden test.
-   **A refactor.** The existing suite is the oracle, with one added differential test: outputs on a set of inputs must match the previous build.

Task shapes that do not come with a red test, such as "improve the error messages" or "clean up this module," are not oracle tasks as stated. They can be made into oracle tasks by adding a check, such as a snapshot test of the messages or a lint rule the cleanup must satisfy, or they can be done without the oracle and their traces kept as unlabeled data.

## Dispatch small

A flip is binary. A task with three requirements and one hidden test that covers all three flips only when all three are done. Split it into three tasks with one hidden test each. The agent gets more feedback across the same work, the traces are shorter and cleaner, and a session that completes two of three is two flips instead of none.

Small also means a short session. Long sessions drift, run out of context, and end on a FAIL that is as much about the session's length as about the task. SWE-bench Verified's human annotators excluded tasks whose descriptions were underspecified or whose tests were unfair, and the benchmark's resolve rates roughly doubled for the same models once the unfair tasks were removed.[2](#user-content-fn-swe-bench-verified) The lesson for a team is that a clear, bounded task description raises the flip rate without changing the agent.

## Sample more than once

Repeated sampling is the single largest lever on the flip rate per task, and its cost is tokens.

Brown and colleagues measured how coverage, the share of problems solved by at least one of k attempts, grows with k. On SWE-bench Lite, an open model that solved 15.9 percent of problems with one attempt solved 56 percent with 250 attempts.[3](#user-content-fn-large-language-monkeys) The oracle is what makes this usable: with a deterministic verifier, the team keeps the attempt that passes and discards the rest. Without one, the team has 250 patches and no way to choose.

This is rejection sampling, and it is how most of the verified-trajectory datasets in the literature were built. SWE-Gym sampled many trajectories per task and kept the 491 that resolved their tasks.[4](#user-content-fn-swe-gym) SWE-smith kept 5,016 out of thousands more.[5](#user-content-fn-swe-smith) Llama 2's post-training used rejection sampling against a reward model as a core step.[6](#user-content-fn-llama2) For a team, the practical version is: dispatch each oracle task to several independent sessions with fresh context, and let the gate find the flip.

The sessions must be independent. If one session's output leaks into another's context, the attempts are correlated and coverage grows slower. Fresh context per attempt and a silent oracle give independence. A verbose oracle that told each attempt why the last one failed would make the attempts a single long session in disguise.

> **Coverage grows with attempts when a verifier picks the winner**
>
> -   **DeepSeek-Coder-V2-Instruct on SWE-bench Lite**: 56% at 250
>
> *Source: Brown et al. (2024). Only the two reported endpoints are plotted; the paper shows a roughly log-linear curve between them. Every added attempt costs tokens and, with a hidden oracle, each attempt is an independent draw.*

A verifier trained on your own traces makes sampling cheaper. Pan and colleagues trained a verifier on the same SWE-Gym trajectories and used it to pick the best of 16 attempts, which raised their model from 20.6 to 32.0 percent.[4](#user-content-fn-swe-gym) The verifier does not replace the oracle; the oracle still grades the chosen attempt. The verifier reduces how many attempts need to reach the oracle.

## Manufacture tasks

When the natural supply of red tests is smaller than the agent capacity, tasks can be made.

**Break working code.** Take a module with good tests. Introduce a bug: delete a branch, flip a comparison, drop a null check. The existing tests that now fail are the hidden tests. The task description is the symptom, written from the test's point of view without naming the test. This is how SWE-smith built 50,000 task instances from 128 repositories, and the authors found that the synthetic bugs trained a model that transferred to real issues.[5](#user-content-fn-swe-smith) Mutation tools generate these bugs automatically, and the mutation literature has catalogued which mutants resemble real faults.[7](#user-content-fn-just-mutants)

**Mine history.** Every past commit that changed source and tests together is a candidate task: the tests it added are the hidden tests, the parent commit is the base, and the commit message or linked issue is the task description. This is how SWE-bench was built and how SWE-rebench automated the construction into a pipeline that produced more than 21,000 tasks.[8](#user-content-fn-swe-bench)[9](#user-content-fn-swe-rebench) A team's own history is a supply of tasks on its own code that no public dataset contains.

**Reverse a fix.** For a merged bug fix with a regression test, revert the source change and keep the test. The agent is dispatched to fix the bug again. The trace is a worked example on real code with a real test.

**Let a model propose tasks under an executor.** Zhao and colleagues' Absolute Zero had a model propose coding tasks and solve them, with a code executor checking both that the task was well-formed and that the answer was right.[10](#user-content-fn-absolute-zero) The executor is the oracle. This is the most speculative item on the list and the one that produces the least realistic tasks; it is here because it shows that the supply of oracle tasks is not bounded by the supply of issues.

## Use the strong model on the hard tail

Some tasks will not flip with the model you are training. Dispatch them to a rented frontier model and keep the trace. A flipped trace from a stronger model is a distillation example, and distillation from a stronger model into a weaker one is the oldest trick in the post-training book.[11](#user-content-fn-hinton-distillation) SWE-smith's 5,016 trajectories came from Claude 3.7 Sonnet, and the model trained on them was a 32-billion-parameter Qwen.[5](#user-content-fn-swe-smith) A team already paying for the frontier model is already producing these traces. The oracle is what sorts them.

## Keep the failures

Most sessions will not flip. Those traces are not waste.

A session that failed on the same task where another session flipped is half of a preference pair. Direct preference optimization trains a model to prefer the chosen trajectory over the rejected one, and it needs exactly this data.[12](#user-content-fn-dpo) A team that keeps only flips has a supervised set. A team that keeps everything has a supervised set and a preference set.

A session that failed is also a record of what went wrong, for people. If a task fails ten sessions in a row, the oracle log will say whether the hidden test is wrong, the task description is unclear, or the task is beyond the model. Each of those has a different fix, and the trace is how you find out which.

Hindsight relabeling, from robotics, is the formal version of this idea: an episode that failed its goal succeeded at whatever it did reach, and can be relabeled as a success for that.[13](#user-content-fn-her) A session that did not flip the hidden test but did make the existing suite pass after breaking it, or did fix a different bug on the way, has a relabeled success in it. The pipeline in Chapter 10 does not do this automatically, and it is the kind of thing a team adds once the basic loop runs.

## The arithmetic

A team of 50 engineers runs an agent on perhaps 10 tasks a day each. Call it 500 sessions a day. If one in five sessions is on an oracle task, that is 100 oracle sessions a day. If one in three of those flips, that is about 33 flips a day. In a working month, around 700. Chapter 12 says what 700 buys. Doubling the oracle-task share doubles it; sampling three attempts per task roughly doubles it again. The levers are the share of work that has a red test, the size of each task, and the number of attempts. None of them is the model.

## Chapter 10: From traces to training data

A trace store is not a training set. Between the two sit a series of transformations that decide what the model learns and what it does not. This chapter walks through them in the order the pipeline applies them.

## Select

The first decision is which sessions to include. The rule for the supervised set is simple: a session is included when its outcome is `flipped` and its verdict record verifies. Everything else is excluded from the supervised set, and the reason is recorded.

A session that flipped is not automatically a good example. Three further checks are cheap and worth running.

-   **The diff touched source.** A session whose only changes were to paths the oracle dropped did not flip because of anything the agent wrote to the code under test. If the oracle still reported PASS, the baseline was wrong. Exclude the session and quarantine the task.
-   **The session did not exceed the length budget.** A session with 400 tool calls that flipped is a worked example of flailing until something stuck. Set a budget per task class and exclude sessions over it, or truncate to the final successful stretch if the trace shows a clear restart.
-   **The agent did not attempt to touch excluded paths repeatedly.** The oracle's dropped-hunk log shows an agent that kept editing test configuration. That session may have flipped on the merits, but its trace teaches the behavior you least want. Exclude it, and look at whether the task description invited it.

For the preference set, pair each flipped session on a task with each session on the same task that did not flip. Where there are many of each, sample pairs rather than taking the full cross product, so one task does not dominate.

## Deduplicate

Two sessions on the same task that flipped with nearly identical trajectories add little to each other. Hash the sequence of tool names and the final diff; where two sessions share both, keep one. Where they share the diff but not the path to it, keep both, because the paths are the data.

Across tasks, deduplicate on the task itself. A team that manufactured 200 tasks by mutating one module will have 200 near-identical trajectories, and a model trained on them will learn that module. Cap the number of flips per source file or per task family.

## Format

A trace is a sequence of events. A training example is a conversation in the chat format the base model was trained on, with tool calls in the model's native tool-call syntax. The conversion is mechanical but has choices in it.

The **system prompt** should be the one the agent ran with, including the rule files that were loaded, because that is the context the model will see at inference. If the team's rule files change often, consider training on a canonical version and keeping the diff as metadata.

The **user turn** is the task description.

Each **assistant turn** is either a tool call, with the tool name and arguments, or a message. Where the harness exposes the model's reasoning, include it as the base model's format expects; where it does not, the turn is the call alone.

Each **tool result** is a tool turn, with the redacted output.

The **final assistant turn** is the agent's completion message.

Convert to the exact chat template of the base model you will fine-tune. A mismatch between the training template and the inference template is the most common silent failure in fine-tuning, and it produces a model that looks trained and behaves untrained.

## Mask

The loss, the quantity the training run minimizes, should be computed only on the tokens the model is supposed to produce. In a trace, that is the assistant's turns: its reasoning, its tool calls, and its messages. The system prompt, the user's task, and every tool result are context, not targets.

Masking the tool results matters more than it might seem. Tool outputs are long: a file read returns hundreds of lines, a test run returns pages. Unmasked, they are most of the tokens in the example, and the model spends its capacity learning to predict file contents and test logs instead of learning to act. A model trained without the mask will also learn to hallucinate tool results, because it was trained to produce them. The reference export tool emits a mask field per turn for this reason.

## Handle length

Agent trajectories are long. A session that flipped after 60 tool calls, each returning a few kilobytes, is a hundred thousand tokens. Base models with long context windows handle this; training on sequences that long is expensive and some frameworks do not support it well.

Three strategies, in order of preference:

1.  **Train on the full trajectory** when the framework and hardware allow it. This preserves the behavior you want: the model learns to carry a plan across many steps.
2.  **Truncate tool outputs**, not turns. A file read can be cut to the lines around the ones the agent later edited, with a marker. A test log can be cut to the failures. The trajectory keeps its shape and loses bulk.
3.  **Window the trajectory** into overlapping segments, each with the system prompt and task prepended and a summary of the dropped prefix. This loses long-range structure and should be the fallback.

Do not drop the dead ends. A trajectory in which the agent tried something, saw it fail, and backed out is a trajectory that teaches recovery. SWE-Gym's authors kept full trajectories including unproductive steps and reported that it worked; cleaning them to the shortest path is a reasonable experiment but not the default.[1](#user-content-fn-swe-gym)

## Hold out

Before anything is trained, split by task, not by session. All sessions on a task go to the same side of the split. Set aside a fraction of tasks, with their hidden tests, as the evaluation set, and never train on any session from them. A model evaluated on tasks it saw during training will look better than it is, and the leak is undetectable afterward.

Keep two evaluation sets. A **frozen set** fixed at the start of the program, so that every model version is scored against the same tasks and the trend is comparable. A **rolling set** of recent tasks, so that the evaluation tracks the work the team is doing now. Chapter 14 covers how to read them.

## Version

Every training set is a manifest: the list of session ids included, the hash of each trace, the hash of each verdict, the filter rules applied, the template used, and the split. Store the manifest beside the weights it produced. When a model regresses, the manifest is how you find the traces that taught it.

A training set that is reproducible from its manifest is the data equivalent of a reproducible build. The same traces and the same rules give the same examples, and a change in either is visible as a change in the hash.

## Record provenance

Each example carries the identity of the oracle that graded it, as Chapter 4 said, and the identity of the model that produced the trace. The second matters more than it looks. A training set that is half traces from a rented frontier model and half from the team's own fine-tuned model is a mixture of distillation and self-improvement, and the two have different failure modes. Chapter 11 covers the self-improvement risk. The provenance field is what lets you tell them apart later.

## Chapter 11: Training methods

With a training set in hand, the question is what to do with it. This chapter covers the three families of methods that have produced results on agent trajectories, in the order a team should adopt them, and the choices inside each: adapter or full fine-tune, how to avoid forgetting, and how to know the signal is real.

## Supervised fine-tuning on flips

The first method is the simplest. Take the flipped trajectories, formatted and masked as Chapter 10 describes, and continue training the base model on them with the standard next-token loss. This is supervised fine-tuning, and when the examples were selected by a verifier, it has a second name: rejection sampling fine-tuning.

The method has a lineage. Expert iteration, from 2017, alternates between a slow expert that solves problems and a fast policy that imitates the solutions, and uses the expert's successes as the training set.[1](#user-content-fn-expert-iteration) AlphaGo Zero's self-play is the same loop with the game's outcome as the verifier.[2](#user-content-fn-alphago-zero) STaR, from 2022, applied it to language models: sample reasoning, keep the samples whose answers match the key, fine-tune, repeat.[3](#user-content-fn-star) ReST and ReST-EM scaled it with a reward model and with answer keys.[4](#user-content-fn-rest)[5](#user-content-fn-rest-em) For code, AlphaCode filtered thousands of samples per problem through the problem's tests before choosing which to submit.[6](#user-content-fn-alphacode) Llama 2's post-training used rejection sampling as a step before reinforcement learning.[7](#user-content-fn-llama2) The agent-trajectory results from Chapter 1 are this method applied to software: SWE-Gym, SWE-smith, and Skywork-SWE are all supervised fine-tuning on verifier-selected trajectories.[8](#user-content-fn-swe-gym)[9](#user-content-fn-swe-smith)[10](#user-content-fn-skywork-swe)

It works, it is cheap, and it is the right place to start. Its limit is that it can only teach the model to do what some session already did. A task no session ever flipped contributes nothing to the supervised set.

## Preference optimization on pairs

The second method uses the failures. Direct preference optimization takes pairs of trajectories on the same task, one preferred and one not, and trains the model to assign higher likelihood to the preferred one relative to a reference model.[11](#user-content-fn-dpo) It needs no reward model and no sampling during training, which makes it almost as cheap as supervised fine-tuning.

The preference set from Chapter 10 is the input: flipped versus not flipped on the same task. The pairs teach the model something the supervised set cannot, which is what a wrong trajectory looks like on this codebase. A team that has run two attempts per task has pairs for every task where exactly one flipped.

Two cautions. Preference optimization is known to drift toward longer outputs unless the pairs are controlled for length, and agent trajectories vary in length a lot.[12](#user-content-fn-length-bias) Match pair lengths roughly or use a length-regularized variant. And the pairs should be real contrasts. A pair where the rejected trajectory failed because the oracle timed out is not a lesson about the code.

## Reinforcement learning with the oracle as the reward

The third method puts the oracle in the training loop. The model attempts a task, the oracle grades the attempt, and the model is updated to make PASS more likely. This is reinforcement learning with verifiable rewards, named as such in the Tülu 3 report and used at scale in DeepSeek-R1.[13](#user-content-fn-tulu3)[14](#user-content-fn-deepseek-r1) For software agents, SWE-RL used a patch-similarity reward on mined pull requests, DeepSWE used a pass-or-fail reward on executable environments, and the Nebius team took a 72-billion-parameter model from 11.4 percent to 39.0 percent on SWE-bench Verified with rejection sampling followed by a variant of the DAPO algorithm.[15](#user-content-fn-swe-rl)[16](#user-content-fn-deepswe)[17](#user-content-fn-nebius-rl)[18](#user-content-fn-dapo)

The appeal is that reinforcement learning can learn from tasks no session has flipped yet, because it searches. The costs are real. Each training step needs many fresh attempts, each attempt needs an oracle run, and the oracle runs in a container. Qwen's team described a system running 20,000 environments in parallel for its coding model's reinforcement learning stage.[19](#user-content-fn-qwen3-coder) A team does not need that scale to see gains; DeepSWE used about 4,500 tasks.[16](#user-content-fn-deepswe) It does need an oracle that can grade hundreds of attempts an hour, which means the oracle has to be a service, not a hook.

There is also a result to read before committing. Yue and colleagues found that reinforcement learning with verifiable rewards sharpens a model toward answers its base model could already produce: the trained model wins when it gets one attempt, and the base model wins when both get many attempts, because the base model's attempts are more varied.[20](#user-content-fn-yue-rlvr) For a team, that argues for a sequence: supervised fine-tuning on flips first, to widen what the model can do on your code; reinforcement learning second, to make it do it reliably on the first try.

Adopt the methods in that order. Supervised fine-tuning on flips as soon as there are a few hundred. Preference optimization as soon as there are pairs. Reinforcement learning when the oracle is a service and the flip rate has plateaued.

> **Training methods in the order a team should adopt them**
>
> 1.  **Supervised fine-tuning on flipped trajectories**: Rejection sampling fine-tuning. Cheapest. Learns what some session already did.
> 2.  **Preference optimization on flipped versus failed pairs**: Uses the failures. Learns what wrong looks like here.
> 3.  **Reinforcement learning with the oracle as reward**: Searches. Can learn tasks no session flipped. Needs the oracle as a service.
>
> *Each rung needs the one below it. The data each rung needs comes from the same trace store.*

## Adapter or full fine-tune

Low-rank adaptation freezes the base model's weights and trains small matrices added to them.[21](#user-content-fn-lora) The adapter is a few percent of the model's size, trains on less hardware, and can be swapped at inference. QLoRA goes further by quantizing the frozen base to 4 bits, which brought fine-tuning of 65-billion-parameter models onto a single large card.[22](#user-content-fn-qlora)

The question is whether the adapter learns as much. Biderman and colleagues compared the two on code and math and found that full fine-tuning learned more on the target domain and forgot more of what the base model knew; low-rank adaptation learned less and forgot less, and acted as a regularizer.[23](#user-content-fn-lora-forgets) Schulman and colleagues at Thinking Machines then reported that the gap closes when the adapter is applied to all layers rather than only attention, and that for small-to-medium post-training sets, the size a team's trace store will be for its first year, the adapter matches the full fine-tune. Their result for reinforcement learning is sharper: a rank-1 adapter matched full fine-tuning, which they explain by noting that a policy-gradient step absorbs about one bit of information per episode, so the capacity needed is tiny.[24](#user-content-fn-lora-without-regret)

That last point connects to this book's oracle. A one-bit verdict is one bit per episode. A method that learns one bit per episode needs many episodes and almost no adapter capacity. That is the regime the pipeline is in, and the adapter is the right tool for it. Start with adapters on all layers. Move to full fine-tuning only if an evaluation shows the adapter is the bottleneck, which for the first several thousand flips it is unlikely to be.

## Forgetting

A model fine-tuned on traces from one codebase can get worse at everything else. The phenomenon is catastrophic forgetting, known since 1989 and measured in language models at every scale.[25](#user-content-fn-mccloskey-forgetting)[26](#user-content-fn-luo-forgetting) For a coding model, it looks like a fine-tune that resolves the team's tasks and can no longer write a shell script.

The standard defenses apply.

-   **Replay.** Mix a fraction of general instruction data into every training run. Ibrahim and colleagues showed that replay combined with re-warming the learning rate lets a model take on new data while matching a full retrain on the old.[27](#user-content-fn-ibrahim-continual)
-   **Low learning rate and few epochs.** One to three passes over the flips at a learning rate an order of magnitude below pretraining.
-   **Adapters.** The frozen base cannot forget. The adapter can be removed. This is the regularization Biderman and colleagues measured.[23](#user-content-fn-lora-forgets)
-   **Merge, don't stack.** When several adapters have been trained on different slices, merge them with a method that keeps the base model's weights and averages or resolves the deltas, such as model soups or TIES, rather than fine-tuning one on top of another.[28](#user-content-fn-model-soups)[29](#user-content-fn-ties-merging)
-   **Evaluate on a general benchmark.** Chapter 14 puts a public coding benchmark in the gate for this reason: a model that gained on the team's tasks and lost on the public set has forgotten, and the gate should see it.

## Collapse

Training a model on its own output can make it worse over generations. Shumailov and colleagues showed that models trained recursively on their own generations lose the tails of the distribution and converge on a narrow set of outputs, a process they called model collapse.[30](#user-content-fn-shumailov-collapse) A pipeline that trains a model on traces produced by the previous version of itself is, on its face, exactly that loop.

Two things make it different. First, the oracle. Collapse happens when generated data replaces real data without selection. The traces in this pipeline are selected by a verifier the model does not control, so the distribution being trained on is the distribution of correct solutions, not of the model's output. Gerstgrasser and colleagues showed that collapse is also avoided when generated data accumulates alongside the original data rather than replacing it.[31](#user-content-fn-gerstgrasser-collapse) Second, the provenance field from Chapter 10. A training run that knows which traces came from a stronger external model and which from its own predecessor can weight them, cap the self-generated share, and watch the ratio over time.

The warning sign is a model whose flips get shorter, more uniform, and more alike across tasks. The frozen evaluation set from Chapter 10 is the instrument.

## Is the signal real

One last check belongs in every training run, and it is cheap. Shao and colleagues found that training a particular family of math models with random rewards, rewards that had nothing to do with correctness, improved its benchmark scores by more than 20 points.[32](#user-content-fn-spurious-rewards) The gain came from the training procedure nudging the model toward behaviors it already had, not from the reward. The same procedure did nothing for other model families. The result does not say verifiable rewards are fake. It says a gain after training is not, on its own, evidence that the reward carried information.

The control is to shuffle the labels. Train the same recipe on the same traces with the flip labels randomly permuted, so that half the "flipped" trajectories are failures. If the shuffled run gains as much as the real run on the held-out evaluation, the gain is not coming from the oracle, and something else in the recipe is doing the work. A team should run this control once when it sets up the pipeline, and again whenever the recipe changes.

## Chapter 12: Data volume

How many flips does it take? This chapter collects the published numbers, states the pattern they show, and gives a worksheet a team can fill in with its own rates.

## What the literature reports

The results below are for different models, methods, and benchmarks, so the table is a set of reference points, not a curve. Read it for order of magnitude.

| Data | Method | Model | Result | Source |
| --- | --- | --- | --- | --- |
| 1,000 curated prompt-response pairs | Supervised fine-tuning | LLaMA 65B | Preferred to or tied with GPT-4 responses in 43 percent of human comparisons | LIMA[1](#user-content-fn-lima) |
| 1,000 reasoning questions with traces | Supervised fine-tuning | Qwen2.5-32B-Instruct | Competitive with o1-preview on competition math; 57 percent on AIME24 with budget forcing | s1[2](#user-content-fn-s1) |
| 817 curated math problems with solutions | Supervised fine-tuning | Qwen2.5-32B-Instruct | 57.1 percent on AIME24 and 94.8 percent on MATH500 in the first release | LIMO[3](#user-content-fn-limo) |
| 500 agent trajectories | Supervised fine-tuning | Llama 2 7B | 77 percent relative gain on a question-answering agent task | FireAct[4](#user-content-fn-fireact) |
| 1,866 agent trajectories across six tasks | Supervised fine-tuning with general data mixed in | Llama 2 7B to 70B | Agent abilities generalize to held-out tasks | AgentTuning[5](#user-content-fn-agenttuning) |
| 491 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 7.0 to 20.6 percent on SWE-bench Verified; 32.0 with a trained verifier and 16 samples | SWE-Gym[6](#user-content-fn-swe-gym) |
| 5,016 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 40.2 percent on SWE-bench Verified | SWE-smith[7](#user-content-fn-swe-smith) |
| 8,209 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 6.4 to 38.0 percent; 47.0 with best-of-8 and a critic; log-linear in data with no plateau | Skywork-SWE[8](#user-content-fn-skywork-swe) |
| About 4,500 executable tasks | Reinforcement learning only | Qwen3-32B | 23 to 42.2 percent; 59.0 with test-time scaling | DeepSWE[9](#user-content-fn-deepswe) |
| Rejection sampling, then RL on executable tasks | Both | Qwen2.5-72B-Instruct | 11.4 to 20.5 to 39.0 percent | Nebius[10](#user-content-fn-nebius-rl) |

Three patterns run through the table.

**Hundreds move a model.** Every result in the first half of the table used about a thousand examples or fewer and produced a large change. LIMA's authors argued that almost all of a model's knowledge comes from pretraining and that alignment needs only a small set of examples to teach format and style.[1](#user-content-fn-lima) The agent results say something stronger for this domain: 491 verified trajectories nearly tripled a 32-billion-parameter model's resolve rate.[6](#user-content-fn-swe-gym) The first useful model is closer than most teams expect.

**Thousands keep paying.** Skywork-SWE's scaling curve is the most direct measurement: 2,000 trajectories gave 31.8 percent, 6,000 gave 36.1, and 8,209 gave 38.0, with the curve still rising.[8](#user-content-fn-skywork-swe) Zhang and colleagues found the same shape across fine-tuning tasks and model sizes and fit it as a power law in the amount of fine-tuning data.[11](#user-content-fn-zhang-scaling) The returns diminish per example and do not stop.

**Quality beats quantity at every scale.** AlpaGasus trained on 9,000 examples filtered from a 52,000-example set and beat the model trained on all 52,000.[12](#user-content-fn-alpagasus) LIMA, s1, and LIMO are all arguments that a small curated set beats a large uncurated one. For this pipeline, the oracle is the curation. A flip is, by construction, an example that was verified. The volume question is how many verified examples, not how many sessions.

> **Verified trajectories behind each open-model result on SWE-bench Verified**
>
> | Result | Trajectories |
> | --- | --- |
> | SWE-Gym, 20.6% after fine-tuning | 491 |
> | Skywork-SWE at 31.8% | 2000 |
> | SWE-smith, 40.2% | 5016 |
> | Skywork-SWE at 36.1% | 6000 |
> | Skywork-SWE, 38.0% | 8209 |
>
> *Sources: Pan et al. (2024), Yang et al. (2025), Zeng et al. (2025). The same 32-billion-parameter base model in every row. The Skywork rows are points on one scaling curve; the other two are separate recipes.*

## The worksheet

Four numbers set the time to a given training set size.

1.  **Sessions per day.** Engineers times sessions each. A team of 50 running 10 each is 500.
2.  **Oracle share.** The fraction of sessions dispatched on tasks that have a hidden test. A team starting out might reach 20 percent. A team that writes the test first for every bug and most features can reach 60.
3.  **Flip rate.** The fraction of oracle sessions that end on PASS. With a silent oracle and a rented frontier model on tasks of reasonable size, 30 to 50 percent is a working assumption; a team should measure it in its first week.
4.  **Attempts per task.** Independent sessions dispatched per oracle task. Each attempt costs tokens and raises the chance that at least one flips.

Flips per day is roughly: sessions × oracle share × flip rate, adjusted upward for attempts. With the numbers above and one attempt, 500 × 0.2 × 0.33 is about 33 flips a day. With three attempts and a per-attempt flip rate of 33 percent, the chance a task flips at least once is about 70 percent, so flips per task-day rises to about 70 of 100 oracle tasks, at three times the token cost.

> **Working days to reach a training set size, one attempt per task**
>
> |  | 20 percent oracle share | 60 percent oracle share |
> | --- | --- | --- |
> | 500 flips, first fine-tune | 15days | 5days |
> | 2,000 flips | 60days | 20days |
> | 5,000 flips | 150days | 50days |
> | 8,000 flips | 240days | 80days |
>
> *Arithmetic for 500 sessions a day at a 33 percent flip rate. Sampling three attempts per task roughly halves every number at three times the token cost. These are planning figures, not measurements.*

The reading of the chart is that a mid-sized team reaches the SWE-Gym regime in weeks and the Skywork regime within a year, and that the oracle share is the lever that matters most. Every bug fixed without a hidden test is a session that could have been a flip and was not.

## Limits of volume

More flips do not fix a bad oracle. A thousand traces graded by a flaky test are a thousand noisy labels, and the noise does not average out; it teaches. More flips do not fix a narrow task distribution. Five thousand flips on one service teach that service. And more flips do not fix a leaked hidden test; they make the leak worse, because every trace fitted to the leaked test is one more example of fitting.

The order of operations is: get the oracle right, get the task supply broad, then grow the volume. Chapter 13 is the pipeline that does the third once the first two are in place.

## Chapter 13: The delivery pipeline

Continuous delivery is the practice of keeping software in a state where any change can be released at any time, by automating the path from a commit to production and gating each step on checks.[1](#user-content-fn-humble-farley) This chapter applies it to weights. The artifact is a model. The commit is a new batch of verified traces. The release is a new model serving the team's agents. Everything between them is a stage with an input, an output, and a gate.

## The stages

> **From trace store to serving model**
>
> 1.  **Ingest**: Sync traces and signed verdicts from agent hosts and oracle hosts
> 2.  **Redact and scan**: Second-pass secret scan; quarantine hits
> 3.  **Validate**: Schema check; verdict signatures; baseline sanity
> 4.  **Select and version**: Filters from Chapter 10; write the manifest
> 5.  **Train**: Adapter on all layers; replay mix; fixed seed
> 6.  **Evaluate**: Held-out flips, frozen and rolling; public benchmark subset
> 7.  **Gate**: Beat the serving model on held-out flips; no regression on the public set
> 8.  **Package**: Merge or ship the adapter; quantize; record the manifest hash
> 9.  **Canary**: Shadow mode on live tasks; the oracle scores both models
> 10.  **Promote or roll back**: Route traffic; keep the previous weights warm
>
> Repeat every batch.
>
> *The emphasized stages are the ones that decide. Everything else is plumbing, and all of it is ordinary continuous delivery with a model as the artifact.*

**Ingest.** A scheduled job pulls finished trace files from each agent host's plugin data directory and verdict records from the oracle host into the trace store. Files are content-addressed: the path includes the hash of the contents, so a re-upload is a no-op and a tampered file lands at a different path. The job is safe to run twice.

**Redact and scan.** The hook redacted with patterns at write time. This stage runs a dedicated secret scanner over every trace and quarantines any with a hit. Quarantined traces are not deleted; a person reviews them, because a false positive on a test fixture is common and a true positive is a credential to rotate.

**Validate.** Every event parses against the schema. Every verdict record's signature verifies against the oracle's public key. Every session has a `session_start`, a `session_end`, and at least one verdict, or it is marked incomplete and excluded. A session whose baseline was PASS is excluded and its task is flagged.

**Select and version.** The filters from Chapter 10 run. The output is a manifest: the list of sessions in the supervised set, the pairs in the preference set, the held-out task list, the template version, and the hash of each. The manifest is the thing that is versioned. The training data is derived from it.

**Train.** The training job takes a manifest and a recipe and produces weights. The recipe is a file: base model and its hash, adapter rank and target modules, learning rate, epochs, replay mix and its source, seed. The job records the recipe hash and the manifest hash in the weights' metadata. A training run with the same manifest, recipe, and seed should produce the same weights, within the limits of the hardware's determinism, and the pipeline should check that it does once.

**Evaluate.** The new weights run against the held-out tasks. For each task, the model attempts it in a fresh sandbox with the same harness the team uses, and the oracle grades the result. The metric is the flip rate on the frozen set and on the rolling set. The weights also run against a fixed subset of a public coding benchmark, for the forgetting check. Chapter 14 is about reading these.

**Gate.** The new weights are promoted only if the frozen-set flip rate is at least the serving model's, the rolling-set flip rate is higher, and the public-set score has not fallen by more than a set tolerance. A run that fails the gate is kept, with its evaluation, so the trend is visible even when nothing ships.

**Package.** The adapter is merged into the base weights or shipped as a separate adapter, depending on how the serving layer loads models. The weights are quantized if the serving hardware needs it, and the quantized weights are re-evaluated on the frozen set, because quantization can cost more on a fine-tuned model than on its base. The package carries the manifest hash and the recipe hash.

**Canary.** Before the new model serves anyone, it runs in shadow. For a sample of live oracle tasks, both the serving model and the candidate attempt the task in parallel sandboxes. The oracle grades both. The person sees only the serving model's result. After enough tasks, the candidate's live flip rate against the serving model's is the number that decides. This is the same design as shadow mode in driving systems, where a new model runs alongside the one in control and its decisions are compared without being acted on.

**Promote or roll back.** Promotion is a routing change. The previous weights stay loaded for a period, and rollback is the same routing change reversed. A rollback is not a failure of the pipeline; it is the pipeline working.

## Cadence

How often to run depends on the flip rate. A team producing 30 flips a day has 200 new examples a week, which is enough to retrain weekly and expect to see movement. A team producing 5 a day should batch monthly. A training run on an adapter for a 32-billion-parameter model over a few thousand trajectories takes hours on a single node with eight large accelerators; the evaluation, which runs an agent on each held-out task, often takes longer than the training. Budget for both.

The cadence should be a schedule, not a trigger. A pipeline that trains whenever enough new traces arrive produces models at irregular intervals that are hard to compare. A pipeline that trains every Sunday night and evaluates on Monday produces a weekly series.

## What is deterministic and what is not

The oracle is deterministic by construction. Training is deterministic up to the hardware. Evaluation is not: the model samples, and the harness is a live process. Fix the sampling temperature and seed for evaluation, run each held-out task more than once, and report the mean. Two models within noise of each other are tied, and the gate should say so instead of promoting on a coin flip.

## The record

Every stage writes a record to the same store the traces live in: what it took in, what it produced, the hashes, the time, and the outcome. A model in production can be traced back to the manifest that trained it, to the sessions in the manifest, to the oracle that graded each session, and to the hidden test behind each verdict. That chain is what makes a regression debuggable and what makes the pipeline auditable to anyone who asks where the model came from.

This is where the author's own work connects, and the connection is disclosed. Oxagen, the company the author founded, records agent runs with their tool calls, their outcomes, and their costs as a product. A run record is a trace in this book's sense. The pipeline in this chapter does not depend on it; the plugin writes its own traces to its own files. But the reason the record is kept beside the verdict, rather than inside the model's context, is the same reason Oxagen keeps the outcome outside the agent that produced it.

## Chapter 14: Evaluation

A pipeline that trains weights needs a way to know whether the new weights are better. This chapter is about building that measurement so that it stays honest as the pipeline optimizes against it.

## Held-out flips are the metric

The primary metric is the flip rate on tasks the model never trained on, graded by the same oracle that grades production sessions. It is the metric that matches the goal. A public benchmark measures how well the model solves public tasks. Held-out flips measure how well it solves yours.

Two sets, as Chapter 10 said. The frozen set is drawn once, at the start, from tasks across the team's repositories, and never changes. Every model version is scored against it, so the series is comparable. The rolling set is the most recent few hundred oracle tasks, refreshed each cycle, so the score tracks the current work. A model that gains on the frozen set and not the rolling set has learned the past. A model that gains on the rolling set and not the frozen set has learned something narrow about recent work. The gate wants both.

## Contamination

A held-out task that the model saw during training is not held out. The leak is easy to create and impossible to detect after the fact.

Deng and colleagues showed that language models can reproduce the missing parts of benchmark items they were trained on, which is how contamination shows up in public evaluations.[1](#user-content-fn-deng-contamination) The SWE-bench+ study found that about a third of one agent's passing patches on SWE-bench had the solution available in the issue text or its comments, and another third passed because the tests were too weak to catch a wrong fix; filtering those out dropped the measured resolve rate from 12.47 percent to 3.97 percent.[2](#user-content-fn-swe-bench-plus) A team's internal evaluation is exposed to the same two failures: the hidden test can leak, and the hidden test can be weak.

Four rules keep the held-out set clean.

1.  Split by task, with every session on a task on the same side.
2.  Never train on a trace from a held-out task, including traces from the rented frontier model. The provenance field makes this checkable.
3.  Never show the held-out hidden tests to any agent, including the one generating training traces. The one-bit oracle enforces this for the agent; the pipeline has to enforce it for the people.
4.  Retire a held-out task when its hidden test is released into the repository, which Chapter 6 recommended once a task has been flipped in production. A test the agents can see is no longer held out.

## Weak tests

A hidden test that passes for a wrong fix inflates the flip rate and teaches the wrong fix. Mutation testing measures this directly: generate mutants of the code the task touches and check that the hidden test kills them.[3](#user-content-fn-just-mutants) A held-out task whose hidden test kills few mutants near the change is a weak oracle, and its flips mean less. The evaluation report should carry the mutation score of each task's hidden test beside the flip rate, so a gain concentrated on weak tasks is visible.

## The public benchmark

A fixed subset of a public coding benchmark sits in the gate for one purpose: catching forgetting. The team's model should not get worse at general coding while it gets better at the team's code. SWE-bench Verified, or a multilingual benchmark if the team's code is not Python, serves.[4](#user-content-fn-swe-bench-verified)[5](#user-content-fn-multi-swe-bench) The score itself is secondary. The change in the score between versions is the signal.

Do not optimize for it. A pipeline that gates on a public benchmark and also trains on data derived from public repositories will, over time, find the benchmark's tasks in its training data. The public set is a thermometer, not a target.

## Watching for hacking

The oracle is air-gapped, so the agent cannot change the verdict. The agent can still produce a trace that flipped for a reason the team would not endorse, and a model trained on such traces learns the reason.

Baker and colleagues found that a monitor reading the agent's chain of thought caught most reward hacking, and that a weaker model was an adequate monitor.[6](#user-content-fn-baker-monitoring) In this pipeline, the equivalent is a scan of each flipped trace for patterns that should not be there: edits to paths on the denylist, even though the oracle dropped them; commands that search for the hidden tests; tool calls that read the oracle's configuration; test runs that were made to pass by changing the test. The scan flags, a person reads, and flagged traces are excluded from training until cleared. The scan should run on traces from the rented model too, because distillation copies behavior.

Track the dropped-hunk rate from the oracle log. A rising share of diffs with hunks on the denylist means the agents are learning to touch the tests, and the next model will learn it faster.

## Drift in what the model produces

Three cheap measurements catch most regressions that the flip rate misses.

-   **Length.** Mean tokens per session and per tool call. Preference optimization and reward-driven training both tend toward verbosity unless controlled.[7](#user-content-fn-length-bias) A model whose sessions get longer without flipping more often is spending the team's tokens.
-   **Tool mix.** The distribution of tool calls per session. A model that stops running tests, or starts reading every file in the repository, has changed in a way the flip rate will show late.
-   **Similarity.** Pairwise similarity between the model's trajectories on different tasks. Rising similarity is the early sign of collapse from Chapter 11.

## The report

Each evaluation produces one report with: flip rate on the frozen set and the rolling set, each with its uncertainty from repeated runs; the public-benchmark score and its change; the mutation score distribution of the held-out tests; the counts of flagged traces by reason; the three drift measurements; and the manifest and recipe hashes. The gate reads the report. So do people. A report that a person cannot read in five minutes is too long.

## Chapter 15: Serving and the economics

The model is trained. This chapter is about running it: where it serves, how requests are routed between it and the rented model, and the arithmetic that says when the pipeline has paid for itself.

## Serving

Open-weight models serve through a small number of mature engines. vLLM introduced paged attention, which manages the key-value cache in blocks the way an operating system manages memory, and made high-throughput serving of large models practical on commodity accelerators.[1](#user-content-fn-vllm) Serving an adapter without merging it is supported directly; S-LoRA showed that thousands of adapters over one base model can be served from a single node with the base weights shared.[2](#user-content-fn-s-lora) For a team with one fine-tune per repository or per domain, that means one base model in memory and a set of small adapters, with the request choosing the adapter.

Quantization reduces memory and raises throughput at some cost in quality. A 32-billion-parameter model quantized to 4 bits fits on a single large accelerator. Re-evaluate on the frozen set after quantizing, as Chapter 13 said, because the cost is not uniform across models.

## Routing

The team's model does not have to handle everything on day one. A router sends each task to the team's model first and falls back to the rented frontier model when the team's model fails. The oracle makes the fallback decision cheap: a session that does not flip after the team's model's attempts is re-dispatched to the rented model.

The fallback sessions are traces. They are the hard tail of the distribution, solved by a stronger model, and graded by the oracle. They go into the trace store with their provenance and become the distillation examples from Chapter 9. Over time, the share of tasks that reach the fallback is the measure of how far the team's model has come, and it is also the training set for closing the gap.

> **Routing a task between the team's model and the rented model**
>
> 1.  **Task**: An oracle task is dispatched
> 2.  **Team model**: Attempts it; the gate asks the oracle
> 3.  **Flip**: PASS: done. The trace is a self-generated example.
> 4.  **Fallback**: No flip after N attempts: re-dispatch to the rented model
> 5.  **Rented model**: Attempts it; the same gate, the same oracle
> 6.  **Trace**: Either way, the trace goes to the store with its provenance
>
> *The fallback rate falls as the team's model learns. The fallback traces are what it learns from.*

## The arithmetic

The cost side has four lines.

**Oracle compute.** Each oracle run is a container running a test suite. On a cloud runner, a suite that takes five minutes costs cents. At 100 oracle sessions a day with two attempts each and a baseline per task, that is a few hundred container-minutes a day.

**Storage.** A trace with redacted tool output is tens of kilobytes to a few megabytes. A year of a mid-sized team's traces is tens to hundreds of gigabytes. This is a rounding error.

**Training.** An adapter run on a 32-billion-parameter model over a few thousand trajectories is hours on one eight-accelerator node. At on-demand cloud prices, that is low hundreds to low thousands of dollars per run. Weekly, it is tens of thousands a year. Evaluation on a few hundred held-out tasks, each an agent session in a sandbox, costs about as much again.

**Serving.** One eight-accelerator node, or two for redundancy, serves a 32-billion-parameter model to a team of fifty with capacity to spare. Reserved, this is low six figures a year; owned, it is a capital cost amortized over several years.

The revenue side is the rent avoided. A team of fifty engineers running agents daily on a frontier model spends, at mid-2026 prices, in the range of several hundred thousand to a few million dollars a year, depending on how heavily they use it. The exact figure is the team's own bill, and the team should use it.

The comparison is not all-or-nothing. In the routing design above, the team's model handles the share of tasks it can flip, and the rented model handles the rest. If the team's model flips 40 percent of oracle tasks in its first quarter, 40 percent of the rent on those tasks is replaced by serving cost. As the share rises, the rent falls. The crossover, where the pipeline's total cost falls below the rent it replaces, depends on the team's bill, but for a team spending more than the cost of two accelerator nodes a year on tokens, it arrives within the first year of flips.

## What the price decline does to this

Token prices fall fast, and the rent will be lower next year.[3](#user-content-fn-epoch-prices) Three things keep the arithmetic in the pipeline's favor anyway.

Open-model serving costs fall on the same curve, because the hardware and the serving software improve for everyone. The ratio between rent and serving cost is more stable than either number.

The rented model's price per token is for a model that does not know the team's code. The team's model's cost per token is for one that does. If the team's model flips a task in fewer tokens, which a model trained on that codebase should, the comparison per task is better than the comparison per token.

And the rent buys no asset. The pipeline's cost buys weights, a trace store, and an oracle, all of which the team keeps. The question in Chapter 1 was what the team owns at the end of the year. The arithmetic here is about when owning it is also cheaper, and the answer for most teams of this size is: soon.

## Chapter 16: Governance and failure modes

A pipeline that collects everything an agent did and trains a model on it has new ways to go wrong. This chapter lists them, with what the research says and what the pipeline does about each. The failures are ordered by how much damage they do, not by how likely they are.

## Secrets in the trace store

The trace store holds the output of every command the agents ran. If a secret passed through a terminal, it is in a trace unless redaction caught it. A trace store is the most complete record of a team's operational secrets that has ever existed in one place, and it should be protected like one.

The defenses are layered. Redaction at write time, with patterns. A dedicated scanner at ingest, with quarantine. Access control on the store, with the training job as the only automated reader. Encryption at rest. And a short retention period for raw tool output, after which only the selected and scanned training examples remain. The last one is a tradeoff: it forecloses re-processing old traces with better redaction. A team should decide its retention consciously.

## Memorization

A model trained on traces can reproduce them. Carlini and colleagues extracted training examples from a deployed language model by prompting it, including names, contact details, and code.[1](#user-content-fn-carlini-extraction) In a follow-up, they measured how memorization scales: it grows with model size, with how many times an example was duplicated in training, and with how much of the example's prefix the prompt supplies.[2](#user-content-fn-carlini-memorization) A fine-tune on a few thousand trajectories, each seen for several epochs, is in the regime where memorization is expected.

Two consequences. First, a secret that survived redaction into training can be extracted from the model. The secret scanning in the pipeline is the control, and it has to be good. Second, the model is a copy of the team's code in a form that can be queried. It is a private asset and should be served privately. A team that fine-tunes on its traces and then exposes the model to people outside the team has published its code in a lossy format. Chapter 1 said the traces were the asset nobody else has. That is only true while the model trained on them stays inside.

## Tampering with the oracle

Chapter 5 built the air gap. The governance question is who can change what is inside it. The hidden tests, the allowlist, the container image, and the signing key are the pipeline's root of trust. Changes to them should go through the same review as changes to production code, and the oracle log should record every verdict with the hash of the test list that produced it, so a change in the tests is visible as a change in the hash. Rotate the hidden tests for a task once it has flipped in production, as Chapter 6 said, and retire the task from the held-out set.

## Training on the wrong lesson

The oracle grades the result. It does not grade the path. A session that flipped by reading the hidden test's name from a stray log line, by copying a fix from a sibling repository that happened to be checked out, or by special-casing the inputs it guessed the test used, is a flip with a bad path. The scan in Chapter 14 catches some of these. The rest are caught by reading flipped traces, which a person should do for a sample every cycle. A pipeline nobody reads is a pipeline that trains on whatever got through.

## The pipeline itself

Sculley and colleagues catalogued the ways machine-learning systems accumulate debt that ordinary software does not: data dependencies that nobody tracks, feedback loops where the model's output changes its own training data, configuration that grows without review, and pipelines glued together from pieces nobody owns.[3](#user-content-fn-sculley-debt) This pipeline has every one of those. Its training data comes from agents running its own previous model, which is a feedback loop by design. Its configuration is a recipe file, an allowlist, a template, and a set of hidden tests, each of which changes the model when it changes. Its stages are a trace collector, an oracle, a trainer, an evaluator, and a router, built from different tools.

The defenses are the ordinary ones, applied without exception. Every input to a stage is versioned. Every stage records what it did. Every change to configuration is reviewed. Every model can be traced to its manifest. And the frozen evaluation set, which never changes, is the one fixed point against which drift in everything else is measured.

## Loss of the flip signal

If the oracle share falls, because the team stops writing tests first, the flip rate falls with it and the pipeline starves. If the task supply narrows, because one repository produces most of the oracle tasks, the model narrows with it. If the fallback rate stops falling, the team's model has plateaued and the recipe needs to change. Each of these is visible in the pipeline's own records, and the weekly report should carry them: oracle share, flip rate, tasks per repository, fallback rate.

## People

The last failure mode is the one the research does not cover. A pipeline that grades every agent session on a hidden test is also, if someone chooses to read it that way, a pipeline that grades the engineers who dispatched the sessions. It should not be used that way. The flip rate is a property of the task, the oracle, and the model. The moment it becomes a measure of a person, people will stop dispatching tasks that might not flip, the oracle share will fall, and the pipeline will starve. Say this in writing when the pipeline is introduced, and keep the per-person numbers out of the report.

## Chapter 17: Precedents

The idea in this book is not new. It is the combination of several ideas that are each at least a decade old, applied to a kind of data that did not exist until coding agents did. This chapter lists the precedents, what each one established, and what it left for this pipeline to add.

## Verifier-filtered self-training

Expert iteration, from Anthony, Tian, and Barber in 2017, is the general form: a slow, strong solver produces solutions, a fast policy is trained to imitate them, and the loop repeats with the improved policy as the new starting point.[1](#user-content-fn-expert-iteration) AlphaGo Zero, the same year, ran the loop with the outcome of the game as the only verifier and reached superhuman play from random weights.[2](#user-content-fn-alphago-zero) The verifier was perfect, cheap, and deterministic, which is the ideal this book's oracle approximates.

STaR brought the loop to language models in 2022, with an answer key as the verifier.[3](#user-content-fn-star) ReST and ReST-EM scaled it.[4](#user-content-fn-rest)[5](#user-content-fn-rest-em) The reinforcement-learning-with-verifiable-rewards line, from Tülu 3 through DeepSeek-R1 and DAPO, is the same loop with a gradient step instead of a fine-tune on the filtered set.[6](#user-content-fn-tulu3)[7](#user-content-fn-deepseek-r1)[8](#user-content-fn-dapo) Cobbe and colleagues' 2021 GSM8K work made the case that a trained verifier scales better than fine-tuning alone, which is the argument for the verifier in Chapter 9.[9](#user-content-fn-cobbe-verifiers)

What these established: a verifier the model does not control produces lasting gains, and the model's own judgment does not. What they left: the verifier was an answer key or a game. Software has no answer key. It has tests.

## Tests as the verifier for code

AlphaCode, in 2022, sampled up to a million programs per competition problem and filtered them through the problem's example tests before clustering and submitting.[10](#user-content-fn-alphacode) CodeRL, the same year, used unit-test results as the reward for training a code model with an actor-critic method.[11](#user-content-fn-coderl) Meta's RLEF, in 2024, trained a code model with execution feedback as the reward and showed large gains in sample efficiency.[12](#user-content-fn-rlef) SWE-RL, DeepSWE, and the Nebius work brought it to repository-scale tasks.[13](#user-content-fn-swe-rl)[14](#user-content-fn-deepswe)[15](#user-content-fn-nebius-rl)

What these established: a test is a usable reward, and a model trained against tests gets better at passing tests. What they left: every one of them used public problems with public tests. The tests were visible, or at least public, and the problems were nobody's in particular.

## Benchmarks built from real repositories

SWE-bench, in 2023, defined the fail-to-pass and pass-to-pass construction and built 2,294 tasks from real GitHub issues and the pull requests that closed them.[16](#user-content-fn-swe-bench) SWE-bench Verified had human annotators remove the tasks whose descriptions were unclear or whose tests were unfair, leaving 500.[17](#user-content-fn-swe-bench-verified) SWE-Gym, SWE-smith, R2E-Gym, and SWE-rebench turned the construction into pipelines that produce thousands of executable tasks.[18](#user-content-fn-swe-gym)[19](#user-content-fn-swe-smith)[20](#user-content-fn-r2e-gym)[21](#user-content-fn-swe-rebench) Multi-SWE-bench extended it to seven languages, and SWE-Lancer graded real freelance tasks with end-to-end tests.[22](#user-content-fn-multi-swe-bench)[23](#user-content-fn-swe-lancer) SWE-bench+ showed how often the construction leaks the answer or accepts a wrong one.[24](#user-content-fn-swe-bench-plus)

What these established: the flip is a reliable unit of value for software work, and tasks with hidden tests can be produced at scale from version history. What they left: the tasks are public, so every model has seen them, and the tests are released, so every agent can be shown them. A team's own history is the unlimited supply of tasks that are not public.

## Learning from an organization's own development process

This is the closest precedent and the least cited.

Facebook's Getafix, in 2019, learned fix patterns from the history of human fixes to static-analysis warnings in Facebook's own codebase and proposed fixes for new warnings, which engineers accepted at a high rate.[25](#user-content-fn-getafix) SapFix, the same year, generated candidate patches for crashes found by an automated testing system and used that system's tests as the oracle, end to end, in production.[26](#user-content-fn-sapfix) Both learned from, and were graded by, the company's own code and tests.

Google's DIDACT, described in 2023, trained models on the process of software development inside Google rather than on finished code: the edit histories, the build-error fixes, the code-review comments and their resolutions.[27](#user-content-fn-didact) Google had earlier reported that a completion model trained on its internal code reduced coding iteration time by 6 percent and was accepted for about 3 percent of new code.[28](#user-content-fn-google-completion) The data was the developers' own activity, and the organization kept it.

GitHub's Copilot research found that acceptance rate was the best available predictor of developers' perceived productivity, which made acceptance a usable, if weak, oracle at scale.[29](#user-content-fn-ziegler-copilot) Replit, in 2024, trained a 7-billion-parameter code-repair model on data built from its own platform: language-server diagnostics, with the state of the file reconstructed by replaying the edit history.[30](#user-content-fn-replit-repair)

What these established: an organization's own development activity is training data, and models trained on it perform well on that organization's work. What they left: each was built by a company with a research team and a bespoke agent or tool. The point of Chapters 7 and 8 is that the harness hooks make this available to a team with neither.

## Learning from traces in other fields

End-to-end driving models from 2016 onward trained on recorded human driving, with the steering angle as the label.[31](#user-content-fn-bojarski-driving) The traces were cheap to collect from vehicles already on the road, and the fleet's data became the moat. The shadow-mode canary in Chapter 13 is borrowed from this field.

Imitation learning has a known failure: a policy trained on an expert's trajectories drifts into states the expert never visited, and its errors compound. DAgger, from 2011, fixes it by running the learner, having the expert label the states the learner reached, and training on those.[32](#user-content-fn-dagger) The analogue here is the fallback in Chapter 15: when the team's model fails a task, the rented model solves it from the same starting point, and the trace is a label on a state the team's model reached. Hindsight experience replay, from 2017, relabels failed episodes as successes for the goals they did reach, which Chapter 9 proposed for sessions that did not flip.[33](#user-content-fn-her)

Research on machine learning for code, surveyed by Allamanis and colleagues in 2018, rests on the observation from Hindle and colleagues in 2012 that software is natural: repetitive and predictable enough that statistical models of it work.[34](#user-content-fn-hindle-naturalness)[35](#user-content-fn-allamanis-survey) A team's codebase is more repetitive and more predictable than the public corpus, which is why a model trained on it does well there.

## Keeping a holdout honest

The one-bit verdict in Chapter 6 rests on the Ladder and the reusable holdout, both from 2015.[36](#user-content-fn-ladder)[37](#user-content-fn-reusable-holdout) Both were written about machine-learning competitions and scientific data analysis. Neither mentions an agent. The problem they solved, an adaptive optimizer hill-climbing on a holdout it is only supposed to be measured by, is the problem an agent iterating against an oracle has, and the solution transfers without modification.

## Agents trained on agent trajectories

AgentTuning and FireAct, in 2023, fine-tuned open models on a few hundred to a couple of thousand agent trajectories generated by a stronger model and showed that the result generalized.[38](#user-content-fn-agenttuning)[39](#user-content-fn-fireact) Kimi K2's technical report described a large-scale pipeline for synthesizing agentic data with tool use and verifying it before training.[40](#user-content-fn-kimi-k2) The open-weight agent frameworks, SWE-agent and OpenHands, standardized the tool interface that these trajectories are recorded in.[41](#user-content-fn-swe-agent)[42](#user-content-fn-openhands)

What these established: agent behavior transfers through trajectories, and a few hundred are enough to see it. What they left: the trajectories came from public tasks solved by public models. The trajectories this book is about come from your tasks, solved on your code, graded by your tests.

## What is new

Four things, and only four.

Hooks in the harness make trace collection free. The organization does not build the agent, and the agent does not have to be modified.

The hidden test as an air-gapped, one-bit oracle makes the grading trustworthy under optimization pressure, which the public-benchmark work did not have to worry about and the internal-tool work handled with bespoke infrastructure.

Open-weight models are close enough to the frontier, and licensed permissively enough, that the fine-tuned result is competitive on the team's distribution. In 2019 the models were not there. In 2023 the licenses often were not.

And the delivery pipeline treats weights as a release artifact with gates, canaries, and rollback, which is ordinary engineering applied to a thing that used to be a research project.

Everything else in this book was established by someone else, and the footnotes say who.

## Closing: What to do this quarter

1.  **Install the collector.** The plugin in Chapter 8, or your own hooks that write the same schema. Start keeping traces this week, before the oracle exists. Unlabeled traces are still an asset, and the collector is the part with no dependencies.
2.  **Pick one repository and write the first hidden tests.** Twenty tasks, each with a test that fails on the base commit for the right reason. Bug reports with reproductions are the easiest source.
3.  **Run the oracle in a container with no network.** On the same machine at first. Move it to a runner the agents cannot log into before you train on anything.
4.  **Measure the flip rate.** Dispatch the twenty tasks. Count the flips. That number, times your sessions per day, is your data rate, and Chapter 12 says what it buys.
5.  **Keep the failures.** They are half of the preference set.
6.  **Set aside the held-out tasks now.** Before the first training run, not after. The leak cannot be undone.
7.  **Train the first adapter at 500 flips.** On all layers, with replay, with the shuffled-label control beside it. Evaluate on the held-out set. Expect it to be worse than the rented model and better than the base model. That is the first point on a curve the team now owns.

## Appendix A: Trace event schema

Each session is one JSON Lines file. Each line is one event. Every event carries `v` (schema version), `session_id`, `ts` (ISO 8601, UTC), and `type`.

```text
session_start
  cwd, base_commit, remote_url_hash, model, harness, harness_version, source,
  task_id, config_hash

message
  turn, role (user | assistant), text_redacted, text_hash, truncated (bool)

tool_call
  turn, tool_name, tool_use_id, input_redacted, input_hash,
  output_redacted, output_hash, truncated (bool), duration_ms

oracle_verdict
  attempt, mode (baseline | grade), verdict (PASS | FAIL | ERROR),
  tests_hash, image_digest, diff_hash, oracle_host, signature (optional)

session_end
  outcome (flipped | no_flip | no_task | aborted), attempts,
  final_diff_hash, reason (from the harness), ended_at
```

Rules: `text_redacted`, `input_redacted`, and `output_redacted` are the only fields that ever hold content, and they hold it after redaction. The `_hash` fields are SHA-256 of the unredacted value, so a later pass can tell whether two truncated values were the same without storing them. `remote_url_hash` rather than the URL, because remote URLs sometimes carry credentials. The verdict's `signature` is over the canonical JSON of the other verdict fields, with the oracle host's key.

## Appendix B: Oracle container specification

-   **Image**: pinned by digest. Built from a Dockerfile that installs the repository's dependencies from its lockfile at build time. Rebuilt when the lockfile changes, and the new digest recorded.
-   **Network**: none. Started with networking disabled.
-   **Mounts**: the task's hidden tests, read-only; the diff file, read-only, in grade mode. Nothing from the agent's working copy.
-   **Repository**: a mirror baked into the image or mounted read-only from the oracle host. Checked out at the base commit inside the container.
-   **Environment**: `TZ=UTC`, `LC_ALL=C.UTF-8`, `PYTHONHASHSEED=0`, `SOURCE_DATE_EPOCH` fixed, and whatever the repository's own test configuration needs to be deterministic.
-   **Diff application**: the diff is filtered through the allowlist and denylist first. Dropped hunks are logged with their paths. The filtered diff is applied with `git apply --check` before `git apply`; a diff that does not apply is a FAIL with the reason logged.
-   **Hidden tests**: copied into the checkout after the diff is applied, from the mount, never from the diff.
-   **Run**: the existing suite at the base commit's version, then the hidden tests. Any failure in either is a FAIL.
-   **Timeout**: a wall-clock limit per run. Exceeding it is a FAIL with the reason logged.
-   **Output to the gate**: exit code 0 or 1, and one line `tests_hash=<sha256>` of the sorted list of test identifiers that ran.
-   **Output to the oracle log**: everything. Test output, dropped hunks, timing, the exit reason, the image digest, the diff hash.
-   **Baseline mode**: the same run without a diff. Must return FAIL for the hidden tests and PASS for the existing suite, or the task is quarantined.

## Appendix C: Data volume worksheet

Fill in the first four lines from your own team. The rest is arithmetic.

```text
A  sessions per day                   = engineers × sessions each       = ____
B  oracle share                       = fraction of sessions on tasks
                                        with a hidden test               = ____
C  flip rate per attempt              = measured in week one             = ____
D  attempts per task                  = independent sessions dispatched  = ____

E  oracle tasks per day               = A × B                            = ____
F  chance a task flips at least once  = 1 − (1 − C)^D                    = ____
G  flips per day                      = E × F                            = ____
H  token cost multiplier              = D                                = ____

Days to 500 flips   = 500 / G
Days to 2,000 flips = 2,000 / G
Days to 8,000 flips = 8,000 / G
```

Reference points from Chapter 12: 491 flips moved a 32-billion-parameter model from 7.0 to 20.6 percent on SWE-bench Verified; about 8,000 moved the same base model to 38.0 percent with the gain still growing.

## Sources

119 works are cited in this book. They appear below in the order of their first citation. Each chapter also lists its own sources at its end.

1.  Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). *How far behind are open models?* Epoch AI. [https://epoch.ai/blog/open-models-report](https://epoch.ai/blog/open-models-report)
2.  Stanford Institute for Human-Centered Artificial Intelligence (2025). *AI Index Report 2025*, Chapter 2: Technical Performance. [https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance](https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance)
3.  OpenAI (2025). *gpt-oss-120b and gpt-oss-20b Model Card*. arXiv. [https://arxiv.org/abs/2508.10925](https://arxiv.org/abs/2508.10925)
4.  Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). *Large Language Models Cannot Self-Correct Reasoning Yet*. arXiv; ICLR 2024. [https://arxiv.org/abs/2310.01798](https://arxiv.org/abs/2310.01798)
5.  Panickssery, A., Bowman, S. R., & Feng, S. (2024). *LLM Evaluators Recognize and Favor Their Own Generations*. arXiv; NeurIPS 2024. [https://arxiv.org/abs/2404.13076](https://arxiv.org/abs/2404.13076)
6.  Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). *STaR: Bootstrapping Reasoning With Reasoning*. arXiv. [https://arxiv.org/abs/2203.14465](https://arxiv.org/abs/2203.14465)
7.  DeepSeek-AI (2025). *DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2501.12948](https://arxiv.org/abs/2501.12948)
8.  Baker, B., Huizinga, J., Gao, L., et al. (2025). *Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation*. arXiv. [https://arxiv.org/abs/2503.11926](https://arxiv.org/abs/2503.11926)
9.  Denison, C., MacDiarmid, M., Barez, F., et al. (2024). *Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models*. arXiv. [https://arxiv.org/abs/2406.10162](https://arxiv.org/abs/2406.10162)
10.  Blum, A., & Hardt, M. (2015). *The Ladder: A Reliable Leaderboard for Machine Learning Competitions*. ICML 2015, PMLR 37, 1006–1014. [https://arxiv.org/abs/1502.04585](https://arxiv.org/abs/1502.04585)
11.  Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). *The reusable holdout: Preserving validity in adaptive data analysis*. Science, 349(6248), 636–638. [https://doi.org/10.1126/science.aaa9375](https://doi.org/10.1126/science.aaa9375)
12.  Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139)
13.  Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798)
14.  Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290)
15.  Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). *LLM inference prices have fallen rapidly but unequally across tasks*. Epoch AI. [https://epoch.ai/data-insights/llm-inference-price-trends](https://epoch.ai/data-insights/llm-inference-price-trends)
16.  Wei, Y., Duchenne, O., Copet, J., et al. (2025). *SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2502.18449](https://arxiv.org/abs/2502.18449)
17.  Mistral AI & All Hands AI (2025). *Devstral*. [https://mistral.ai/news/devstral](https://mistral.ai/news/devstral)
18.  Qwen Team (2025). *Qwen3 Technical Report*. arXiv. [https://arxiv.org/abs/2505.09388](https://arxiv.org/abs/2505.09388)
19.  Grattafiori, A., et al. (2024). *The Llama 3 Herd of Models*. arXiv. [https://arxiv.org/abs/2407.21783](https://arxiv.org/abs/2407.21783)
20.  Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). *SWE-bench: Can Language Models Resolve Real-World GitHub Issues?* arXiv; ICLR 2024. [https://arxiv.org/abs/2310.06770](https://arxiv.org/abs/2310.06770)
21.  Anthropic (2026). *Hooks reference*. Claude Code documentation. [https://code.claude.com/docs/en/hooks](https://code.claude.com/docs/en/hooks)
22.  Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). *The Oracle Problem in Software Testing: A Survey*. IEEE Transactions on Software Engineering, 41(5), 507–525. [https://doi.org/10.1109/TSE.2014.2372785](https://doi.org/10.1109/TSE.2014.2372785)
23.  Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). *An Empirical Analysis of Flaky Tests*. FSE 2014, 643–653. [https://doi.org/10.1145/2635868.2635920](https://doi.org/10.1145/2635868.2635920)
24.  Micco, J. (2016). *Flaky Tests at Google and How We Mitigate Them*. Google Testing Blog. [https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html](https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html)
25.  Lamb, C., & Zacchiroli, S. (2022). *Reproducible Builds: Increasing the Integrity of Software Supply Chains*. IEEE Software, 39(2), 62–70. [https://arxiv.org/abs/2104.06020](https://arxiv.org/abs/2104.06020)
26.  Just, R., Jalali, D., & Ernst, M. D. (2014). *Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs*. ISSTA 2014, 437–440. [https://doi.org/10.1145/2610384.2628055](https://doi.org/10.1145/2610384.2628055)
27.  Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). *An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems*. ISSTA 2015, 24–36. [https://doi.org/10.1145/2771783.2771791](https://doi.org/10.1145/2771783.2771791)
28.  Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). *Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair*. ESEC/FSE 2015, 532–543. [https://doi.org/10.1145/2786805.2786825](https://doi.org/10.1145/2786805.2786825)
29.  Inozemtseva, L., & Holmes, R. (2014). *Coverage Is Not Strongly Correlated with Test Suite Effectiveness*. ICSE 2014, 435–445. [https://doi.org/10.1145/2568225.2568271](https://doi.org/10.1145/2568225.2568271)
30.  DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). *Hints on Test Data Selection: Help for the Practicing Programmer*. IEEE Computer, 11(4), 34–41. [https://doi.org/10.1109/C-M.1978.218136](https://doi.org/10.1109/C-M.1978.218136)
31.  Petrović, G., & Ivanković, M. (2018). *State of Mutation Testing at Google*. ICSE-SEIP 2018, 163–171. [https://doi.org/10.1145/3183519.3183521](https://doi.org/10.1145/3183519.3183521)
32.  Claessen, K., & Hughes, J. (2000). *QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs*. ICFP 2000, 268–279. [https://doi.org/10.1145/351240.351266](https://doi.org/10.1145/351240.351266)
33.  Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). *A Survey on Metamorphic Testing*. IEEE Transactions on Software Engineering, 42(9), 805–824. [https://doi.org/10.1109/TSE.2016.2532875](https://doi.org/10.1109/TSE.2016.2532875)
34.  McKeeman, W. M. (1998). *Differential Testing for Software*. Digital Technical Journal, 10(1), 100–107.
35.  Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). *Measuring GitHub Copilot's Impact on Productivity*. Communications of the ACM, 67(3), 54–63. [https://doi.org/10.1145/3633453](https://doi.org/10.1145/3633453)
36.  Von Arx, S., Chan, L., & Barnes, E. (2025). *Recent Frontier Models Are Reward Hacking*. METR. [https://metr.org/blog/2025-06-05-recent-reward-hacking/](https://metr.org/blog/2025-06-05-recent-reward-hacking/)
37.  MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). *Natural Emergent Misalignment from Reward Hacking in Production RL*. arXiv. [https://arxiv.org/abs/2511.18397](https://arxiv.org/abs/2511.18397)
38.  Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). *Defining and Characterizing Reward Hacking*. NeurIPS 2022. [https://arxiv.org/abs/2209.13085](https://arxiv.org/abs/2209.13085)
39.  Gao, L., Schulman, J., & Hilton, J. (2022). *Scaling Laws for Reward Model Overoptimization*. arXiv; ICML 2023. [https://arxiv.org/abs/2210.10760](https://arxiv.org/abs/2210.10760)
40.  Thompson, K. (1984). *Reflections on Trusting Trust*. Communications of the ACM, 27(8), 761–763. [https://doi.org/10.1145/358198.358210](https://doi.org/10.1145/358198.358210)
41.  Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). *Firecracker: Lightweight Virtualization for Serverless Applications*. NSDI 2020, 419–434. [https://www.usenix.org/conference/nsdi20/presentation/agache](https://www.usenix.org/conference/nsdi20/presentation/agache)
42.  Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). *The True Cost of Containing: A gVisor Case Study*. HotCloud 2019. [https://www.usenix.org/conference/hotcloud19/presentation/young](https://www.usenix.org/conference/hotcloud19/presentation/young)
43.  Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). *Teaching Large Language Models to Self-Debug*. arXiv; ICLR 2024. [https://arxiv.org/abs/2304.05128](https://arxiv.org/abs/2304.05128)
44.  Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). *Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena*. NeurIPS 2023 Datasets and Benchmarks. [https://arxiv.org/abs/2306.05685](https://arxiv.org/abs/2306.05685)
45.  Wang, P., Li, L., Chen, L., et al. (2023). *Large Language Models are not Fair Evaluators*. arXiv; ACL 2024. [https://arxiv.org/abs/2305.17926](https://arxiv.org/abs/2305.17926)
46.  Stechly, K., Marquez, M., & Kambhampati, S. (2023). *GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems*. arXiv. [https://arxiv.org/abs/2310.12397](https://arxiv.org/abs/2310.12397)
47.  Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). *Can Large Language Models Really Improve by Self-critiquing Their Own Plans?* arXiv. [https://arxiv.org/abs/2310.08118](https://arxiv.org/abs/2310.08118)
48.  Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). *LLMs cannot find reasoning errors, but can correct them given the error location*. Findings of ACL 2024. [https://arxiv.org/abs/2311.08516](https://arxiv.org/abs/2311.08516)
49.  Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). *When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs*. Transactions of the ACL, 12. [https://arxiv.org/abs/2406.01297](https://arxiv.org/abs/2406.01297)
50.  Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). *Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement*. ACL 2024. [https://arxiv.org/abs/2402.11436](https://arxiv.org/abs/2402.11436)
51.  Meli, M., McNiece, M. R., & Reaves, B. (2019). *How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories*. NDSS 2019. [https://doi.org/10.14722/ndss.2019.23418](https://doi.org/10.14722/ndss.2019.23418)
52.  Anthropic (2026). *Plugin manifest reference*. Claude Code documentation. [https://code.claude.com/docs/en/plugins-reference](https://code.claude.com/docs/en/plugins-reference)
53.  Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). *Realizing quality improvement through test driven development: results and experiences of four industrial teams*. Empirical Software Engineering, 13(3), 289–302. [https://doi.org/10.1007/s10664-008-9062-z](https://doi.org/10.1007/s10664-008-9062-z)
54.  OpenAI (2024). *Introducing SWE-bench Verified*. [https://openai.com/index/introducing-swe-bench-verified/](https://openai.com/index/introducing-swe-bench-verified/)
55.  Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). *Large Language Monkeys: Scaling Inference Compute with Repeated Sampling*. arXiv. [https://arxiv.org/abs/2407.21787](https://arxiv.org/abs/2407.21787)
56.  Touvron, H., et al. (2023). *Llama 2: Open Foundation and Fine-Tuned Chat Models*. arXiv. [https://arxiv.org/abs/2307.09288](https://arxiv.org/abs/2307.09288)
57.  Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). *Are Mutants a Valid Substitute for Real Faults in Software Testing?* FSE 2014, 654–665. [https://doi.org/10.1145/2635868.2635929](https://doi.org/10.1145/2635868.2635929)
58.  Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). *SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents*. arXiv; NeurIPS 2025. [https://arxiv.org/abs/2505.20411](https://arxiv.org/abs/2505.20411)
59.  Zhao, A., Wu, Y., Yue, Y., et al. (2025). *Absolute Zero: Reinforced Self-play Reasoning with Zero Data*. arXiv. [https://arxiv.org/abs/2505.03335](https://arxiv.org/abs/2505.03335)
60.  Hinton, G., Vinyals, O., & Dean, J. (2015). *Distilling the Knowledge in a Neural Network*. NeurIPS 2014 Deep Learning Workshop. [https://arxiv.org/abs/1503.02531](https://arxiv.org/abs/1503.02531)
61.  Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). *Direct Preference Optimization: Your Language Model is Secretly a Reward Model*. NeurIPS 2023. [https://arxiv.org/abs/2305.18290](https://arxiv.org/abs/2305.18290)
62.  Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). *Hindsight Experience Replay*. NeurIPS 2017. [https://arxiv.org/abs/1707.01495](https://arxiv.org/abs/1707.01495)
63.  Anthony, T., Tian, Z., & Barber, D. (2017). *Thinking Fast and Slow with Deep Learning and Tree Search*. NeurIPS 2017. [https://arxiv.org/abs/1705.08439](https://arxiv.org/abs/1705.08439)
64.  Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). *Mastering the game of Go without human knowledge*. Nature, 550, 354–359. [https://doi.org/10.1038/nature24270](https://doi.org/10.1038/nature24270)
65.  Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). *Reinforced Self-Training (ReST) for Language Modeling*. arXiv. [https://arxiv.org/abs/2308.08998](https://arxiv.org/abs/2308.08998)
66.  Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). *Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models*. arXiv; TMLR 2024. [https://arxiv.org/abs/2312.06585](https://arxiv.org/abs/2312.06585)
67.  Li, Y., Choi, D., Chung, J., et al. (2022). *Competition-level code generation with AlphaCode*. Science, 378(6624), 1092–1097. [https://doi.org/10.1126/science.abq1158](https://doi.org/10.1126/science.abq1158)
68.  Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). *A Long Way to Go: Investigating Length Correlations in RLHF*. arXiv; COLM 2024. [https://arxiv.org/abs/2310.03716](https://arxiv.org/abs/2310.03716)
69.  Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). *Tülu 3: Pushing Frontiers in Open Language Model Post-Training*. arXiv. [https://arxiv.org/abs/2411.15124](https://arxiv.org/abs/2411.15124)
70.  Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe)
71.  Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501)
72.  Yu, Q., Zhang, Z., Zhu, R., et al. (2025). *DAPO: An Open-Source LLM Reinforcement Learning System at Scale*. arXiv. [https://arxiv.org/abs/2503.14476](https://arxiv.org/abs/2503.14476)
73.  Qwen Team (2025). *Qwen3-Coder: Agentic Coding in the World*. [https://qwenlm.github.io/blog/qwen3-coder/](https://qwenlm.github.io/blog/qwen3-coder/)
74.  Yue, Y., Chen, Z., Lu, R., et al. (2025). *Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?* arXiv; NeurIPS 2025. [https://arxiv.org/abs/2504.13837](https://arxiv.org/abs/2504.13837)
75.  Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). *LoRA: Low-Rank Adaptation of Large Language Models*. arXiv; ICLR 2022. [https://arxiv.org/abs/2106.09685](https://arxiv.org/abs/2106.09685)
76.  Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). *QLoRA: Efficient Finetuning of Quantized LLMs*. NeurIPS 2023. [https://arxiv.org/abs/2305.14314](https://arxiv.org/abs/2305.14314)
77.  Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). *LoRA Learns Less and Forgets Less*. Transactions on Machine Learning Research. [https://arxiv.org/abs/2405.09673](https://arxiv.org/abs/2405.09673)
78.  Schulman, J., & Thinking Machines Lab (2025). *LoRA Without Regret*. [https://thinkingmachines.ai/blog/lora/](https://thinkingmachines.ai/blog/lora/)
79.  McCloskey, M., & Cohen, N. J. (1989). *Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem*. Psychology of Learning and Motivation, 24, 109–165. [https://doi.org/10.1016/S0079-7421(08)60536-8](https://doi.org/10.1016/S0079-7421\(08\)60536-8)
80.  Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). *An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning*. arXiv. [https://arxiv.org/abs/2308.08747](https://arxiv.org/abs/2308.08747)
81.  Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). *Simple and Scalable Strategies to Continually Pre-train Large Language Models*. arXiv; TMLR. [https://arxiv.org/abs/2403.08763](https://arxiv.org/abs/2403.08763)
82.  Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). *Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time*. ICML 2022. [https://arxiv.org/abs/2203.05482](https://arxiv.org/abs/2203.05482)
83.  Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). *TIES-Merging: Resolving Interference When Merging Models*. NeurIPS 2023. [https://arxiv.org/abs/2306.01708](https://arxiv.org/abs/2306.01708)
84.  Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). *AI models collapse when trained on recursively generated data*. Nature, 631, 755–759. [https://doi.org/10.1038/s41586-024-07566-y](https://doi.org/10.1038/s41586-024-07566-y)
85.  Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). *Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data*. arXiv. [https://arxiv.org/abs/2404.01413](https://arxiv.org/abs/2404.01413)
86.  Shao, R., Li, S. S., Xin, R., et al. (2025). *Spurious Rewards: Rethinking Training Signals in RLVR*. arXiv. [https://arxiv.org/abs/2506.10947](https://arxiv.org/abs/2506.10947)
87.  Zhou, C., Liu, P., Xu, P., et al. (2023). *LIMA: Less Is More for Alignment*. NeurIPS 2023. [https://arxiv.org/abs/2305.11206](https://arxiv.org/abs/2305.11206)
88.  Muennighoff, N., Yang, Z., Shi, W., et al. (2025). *s1: Simple test-time scaling*. arXiv. [https://arxiv.org/abs/2501.19393](https://arxiv.org/abs/2501.19393)
89.  Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). *LIMO: Less is More for Reasoning*. arXiv (v1, February 2025); COLM 2025. [https://arxiv.org/abs/2502.03387](https://arxiv.org/abs/2502.03387)
90.  Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). *FireAct: Toward Language Agent Fine-tuning*. arXiv. [https://arxiv.org/abs/2310.05915](https://arxiv.org/abs/2310.05915)
91.  Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). *AgentTuning: Enabling Generalized Agent Abilities for LLMs*. arXiv. [https://arxiv.org/abs/2310.12823](https://arxiv.org/abs/2310.12823)
92.  Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). *When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method*. ICLR 2024. [https://arxiv.org/abs/2402.17193](https://arxiv.org/abs/2402.17193)
93.  Chen, L., Li, S., Yan, J., et al. (2023). *AlpaGasus: Training a Better Alpaca with Fewer Data*. arXiv; ICLR 2024. [https://arxiv.org/abs/2307.08701](https://arxiv.org/abs/2307.08701)
94.  Humble, J., & Farley, D. (2010). *Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation*. Addison-Wesley.
95.  Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). *Investigating Data Contamination in Modern Benchmarks for Large Language Models*. NAACL 2024. [https://arxiv.org/abs/2311.09783](https://arxiv.org/abs/2311.09783)
96.  Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). *SWE-Bench+: Enhanced Coding Benchmark for LLMs*. arXiv. [https://arxiv.org/abs/2410.06992](https://arxiv.org/abs/2410.06992)
97.  Zan, D., Huang, Z., Liu, W., et al. (2025). *Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving*. arXiv. [https://arxiv.org/abs/2504.02605](https://arxiv.org/abs/2504.02605)
98.  Kwon, W., Li, Z., Zhuang, S., et al. (2023). *Efficient Memory Management for Large Language Model Serving with PagedAttention*. SOSP 2023. [https://arxiv.org/abs/2309.06180](https://arxiv.org/abs/2309.06180)
99.  Sheng, Y., Cao, S., Li, D., et al. (2023). *S-LoRA: Serving Thousands of Concurrent LoRA Adapters*. arXiv; MLSys 2024. [https://arxiv.org/abs/2311.03285](https://arxiv.org/abs/2311.03285)
100.  Carlini, N., Tramèr, F., Wallace, E., et al. (2021). *Extracting Training Data from Large Language Models*. USENIX Security 2021. [https://arxiv.org/abs/2012.07805](https://arxiv.org/abs/2012.07805)
101.  Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). *Quantifying Memorization Across Neural Language Models*. ICLR 2023. [https://arxiv.org/abs/2202.07646](https://arxiv.org/abs/2202.07646)
102.  Sculley, D., Holt, G., Golovin, D., et al. (2015). *Hidden Technical Debt in Machine Learning Systems*. NeurIPS 2015.
103.  Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). *Training Verifiers to Solve Math Word Problems*. arXiv. [https://arxiv.org/abs/2110.14168](https://arxiv.org/abs/2110.14168)
104.  Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). *CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning*. NeurIPS 2022. [https://arxiv.org/abs/2207.01780](https://arxiv.org/abs/2207.01780)
105.  Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). *RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2410.02089](https://arxiv.org/abs/2410.02089)
106.  Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). *R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents*. arXiv. [https://arxiv.org/abs/2504.07164](https://arxiv.org/abs/2504.07164)
107.  Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). *SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering?* arXiv. [https://arxiv.org/abs/2502.12115](https://arxiv.org/abs/2502.12115)
108.  Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). *Getafix: Learning to Fix Bugs Automatically*. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. [https://doi.org/10.1145/3360585](https://doi.org/10.1145/3360585)
109.  Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). *SapFix: Automated End-to-End Repair at Scale*. ICSE-SEIP 2019, 269–278. [https://doi.org/10.1109/ICSE-SEIP.2019.00039](https://doi.org/10.1109/ICSE-SEIP.2019.00039)
110.  Maniatis, P., & Tarlow, D. (2023). *Large sequence models for software development activities*. Google Research Blog. [https://research.google/blog/large-sequence-models-for-software-development-activities/](https://research.google/blog/large-sequence-models-for-software-development-activities/)
111.  Tabachnyk, M., & Nikolov, S. (2022). *ML-Enhanced Code Completion Improves Developer Productivity*. Google Research Blog. [https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/](https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/)
112.  Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). *Building LLMs for Code Repair*. Replit. [https://replit.com/blog/code-repair](https://replit.com/blog/code-repair)
113.  Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). *End to End Learning for Self-Driving Cars*. arXiv. [https://arxiv.org/abs/1604.07316](https://arxiv.org/abs/1604.07316)
114.  Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). *A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning*. AISTATS 2011. [https://arxiv.org/abs/1011.0686](https://arxiv.org/abs/1011.0686)
115.  Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). *On the Naturalness of Software*. ICSE 2012, 837–847. [https://doi.org/10.1109/ICSE.2012.6227135](https://doi.org/10.1109/ICSE.2012.6227135)
116.  Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). *A Survey of Machine Learning for Big Code and Naturalness*. ACM Computing Surveys, 51(4), Article 81. [https://arxiv.org/abs/1709.06182](https://arxiv.org/abs/1709.06182)
117.  Kimi Team (2025). *Kimi K2: Open Agentic Intelligence*. arXiv. [https://arxiv.org/abs/2507.20534](https://arxiv.org/abs/2507.20534)
118.  Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). *SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering*. NeurIPS 2024. [https://arxiv.org/abs/2405.15793](https://arxiv.org/abs/2405.15793)
119.  Wang, X., Li, B., Song, Y., et al. (2024). *OpenHands: An Open Platform for AI Software Developers as Generalist Agents*. arXiv; ICLR 2025. [https://arxiv.org/abs/2407.16741](https://arxiv.org/abs/2407.16741)
