Coding agents, Self-improving models, Autonomous agents
Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces
Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.
Reading paths
Shorter routes through the book. Each one stands on its own.
The case in half an hour
15 minBuild the oracle
28 minRun the pipeline
22 minApprove the budget
18 minContents
7 parts in reading order. Each row gives the chapter's claim and how long it takes.
I
The argument
Traces are the asset. Rent buys none.
11 min in 2 chapters
II
The oracle
A verdict the model cannot touch or argue with.
28 min in 4 chapters
- 3The oracle problemA verdict the model does not control: deterministic, independent, and hidden, with the fail-to-pass flip stated as a rule.6 min
- 4Organizational oraclesThe checks a team already has, ordered by how directly each one can label a trace.7 min
- 5The air gapThe agent is an adversary with a shell. The oracle runs where the shell cannot reach.8 min
- 6The one-bit verdictWhy PASS or FAIL is all the oracle says, what that costs inside a session, and the research on models grading themselves.7 min
III
Collecting traces
The harness already has hooks. Use them.
20 min in 3 chapters
- 7Harness hooksThe hooks a harness already exposes, turned into a trace collector and a Stop gate without building an agent.6 min
- 8The reference pluginA walkthrough of oracle-flip: five hooks, one constant reason, and an oracle in a container with no network.7 min
- 9The flip rateWrite the test first, dispatch small, sample more than once, manufacture tasks, and keep the failures.7 min
IV
Training
Hundreds of flips move a model. Thousands keep paying.
17 min in 3 chapters
- 10From traces to training dataSelect, deduplicate, format, mask, hold out, version.5 min
- 11Training methodsSupervised fine-tuning on flips, then preference pairs, then reinforcement learning. Adapters, forgetting, collapse, and a shuffled-label control.8 min
- 12Data volumeWhat 491, 5,016, and 8,209 verified trajectories bought in the literature, and a worksheet for your own rate.4 min
V
Delivery
Weights as a release artifact with gates and a rollback.
17 min in 4 chapters
- 13The delivery pipelineWeights as a release artifact: ingest, validate, train, evaluate, gate, canary in shadow, promote, roll back.5 min
- 14EvaluationHeld-out flips, contamination, weak tests, hacking scans, and the drift measures that catch what the flip rate misses.4 min
- 15Serving and the economicsRouting between your model and the rented one, and when owning the weights becomes cheaper than the rent.4 min
- 16Governance and failure modesSecrets, memorization, oracle tampering, pipeline debt, a starving flip signal, and people.4 min
VI
Precedents and closing
Who tried this before and what to do this quarter.
7 min in 2 chapters
VII
Back matter
Schemas, specifications, and sources.
3 min in 3 chapters
- App. ATrace event schemaThe event types and fields in a session's trace file.1 min
- App. BOracle container specificationWhat the oracle container mounts, fixes, applies, runs, and reports.1 min
- App. CData volume worksheetFour numbers from your team and the arithmetic that turns them into a date.1 min
- SourcesSourcesThe works cited, in order of first citation.reference
Introduction
Your organization runs coding agents every day. Each session produces a record: the task, every file the agent read, every command it ran, every edit it made, and whether the work held up. Most teams throw that record away the moment the session ends. This book argues that the record is the most valuable asset the session produces, and that a team which keeps it can, within a year, fine-tune an open-weight model that does its own work as well as the rented frontier model does today.
The book is long because the argument has many parts and each part has evidence behind it. It is written for the engineer who will build the pipeline and for the person who has to approve it. Chapters 1 and 2 make the case. Chapters 3 to 6 are about the oracle, the program that decides whether a trace counts. Chapters 7 to 9 are about collecting traces from the agent harness you already use, with a working Claude Code plugin you can install today. Chapters 10 to 12 are about turning traces into training data and how much you need. Chapters 13 to 16 are about the delivery pipeline that ships new weights on a schedule. Chapter 17 lists the earlier systems that tried the same idea, so you can see what held up.
Three words carry most of the weight, so here they are up front.
- A trace is the full record of one agent session: prompts, tool calls, tool results, edits, and the final state of the repository.
- An oracle is a program that reads the result of a session and returns one verdict, PASS or FAIL, without the agent's help. In this book the oracle is a hidden test that the agent never sees, run in a container the agent cannot reach.
- A flip is the event the whole pipeline is built around. The oracle said FAIL before the agent started. The oracle says PASS after the agent finished. The trace between those two verdicts is a worked example of the task being done right, graded by something the model did not control.
The argument in one page#
Rented tokens are a cost that never ends and a capability you never own. The price per token has fallen fast, and it will keep falling, but the thing you are paying for is a model that knows nothing about your systems on the day you start and nothing more on the day you stop. Every improvement happens in the vendor's weights, under the vendor's terms, on the vendor's schedule.
Open-weight models are now close enough to the frontier that the gap no longer decides the outcome. In late 2024, the best open model trailed the best closed model by about a year.1 In the year to early 2025, the gap between the best open and closed models on a public human-preference leaderboard narrowed from about 8 percent to under 2 percent.2 Models released under permissive licenses by DeepSeek, Alibaba, Meta, Mistral, and in August 2025 by OpenAI itself, run on hardware you can rent by the hour or buy outright.3 A general open model is a floor. What raises the floor on your tasks is training on your tasks.
Traces are that training data, but only when something outside the model grades them. The research on models grading their own work is consistent and unkind. Left to judge themselves, models prefer their own output, cannot reliably find their own errors, and often get worse when they try to self-correct without an outside signal.45 The methods that do produce lasting gains all share one feature: a verifier the model does not control, such as an answer key, a compiler, or a test.67 For software, that verifier already exists. It is the test suite, the type checker, the build, and the hidden test you write before the agent starts.
The oracle has to be deterministic, air-gapped, and silent. Deterministic, because a flaky verdict is noise in the training set, and noise at this stage is expensive to remove later. Air-gapped, because frontier models have been caught editing tests, stubbing functions, and exiting early to make a check pass, and a reward signal the agent can reach is a signal it will eventually reach.89 Silent, meaning it says only PASS or FAIL, because a verdict that explains itself leaks the hidden test into the agent's context, and a test the agent can see is a test it can overfit. The theory for this is older than language models: a holdout that answers with low information stays valid under many adaptive queries, and one that answers with high information does not.1011
You do not have to build the agent. Every serious coding harness now exposes hooks: programs that run before and after each tool call, and when the agent tries to stop. A hook that appends each event to a file is a trace collector. A hook that runs the oracle on Stop and refuses to let the agent finish on FAIL is the flip detector. The plugin in Chapter 8 is about 400 lines.
You need fewer flips than you think to start, and more than you think to finish. Published results put the first large gains at a few hundred verified trajectories on a 32-billion-parameter open model, and the best open results at several thousand.121314 A team of fifty engineers running agents daily reaches the first number in a week and the second in a quarter. The chapters on volume give the arithmetic.
The pipeline is ordinary continuous delivery with weights as the artifact. Ingest, redact, validate, version, train, evaluate against held-out flips the model never saw, gate, package, canary in shadow mode against the live oracle, promote, roll back. None of these stages is new. What is new is that the artifact learns.
Footnotes#
-
Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). How far behind are open models? Epoch AI. https://epoch.ai/blog/open-models-report ↩
-
Stanford Institute for Human-Centered Artificial Intelligence (2025). AI Index Report 2025, Chapter 2: Technical Performance. https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance ↩
-
OpenAI (2025). gpt-oss-120b and gpt-oss-20b Model Card. arXiv. https://arxiv.org/abs/2508.10925 ↩
-
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv; ICLR 2024. https://arxiv.org/abs/2310.01798 ↩
-
Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv; NeurIPS 2024. https://arxiv.org/abs/2404.13076 ↩
-
Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. arXiv. https://arxiv.org/abs/2203.14465 ↩
-
DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948 ↩
-
Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv. https://arxiv.org/abs/2503.11926 ↩
-
Denison, C., MacDiarmid, M., Barez, F., et al. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv. https://arxiv.org/abs/2406.10162 ↩
-
Blum, A., & Hardt, M. (2015). The Ladder: A Reliable Leaderboard for Machine Learning Competitions. ICML 2015, PMLR 37, 1006–1014. https://arxiv.org/abs/1502.04585 ↩
-
Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636–638. https://doi.org/10.1126/science.aaa9375 ↩
-
Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩
-
Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩
-
Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290 ↩
Cite this
Anderson, M. (2026, October 9). Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces. macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models
BibTeX
@online{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
year = {2026},
date = {2026-10-09},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models}
}Updates by email
New research reaches subscribers first.