Part VI: Precedents and closing
17Precedents
Expert iteration to SWE-Gym and Getafix to DIDACT: what each established and what it left to do.
Chapter 17 of 22
Part VI: Precedents and closing
In this chapter
- Verifier-filtered self-training
- Tests as the verifier for code
- Benchmarks built from real repositories
- Learning from an organization's own development process
- Learning from traces in other fields
- Keeping a holdout honest
- Agents trained on agent trajectories
- What is new
The idea in this book is not new. It is the combination of several ideas that are each at least a decade old, applied to a kind of data that did not exist until coding agents did. This chapter lists the precedents, what each one established, and what it left for this pipeline to add.
Verifier-filtered self-training#
Expert iteration, from Anthony, Tian, and Barber in 2017, is the general form: a slow, strong solver produces solutions, a fast policy is trained to imitate them, and the loop repeats with the improved policy as the new starting point.1 AlphaGo Zero, the same year, ran the loop with the outcome of the game as the only verifier and reached superhuman play from random weights.2 The verifier was perfect, cheap, and deterministic, which is the ideal this book's oracle approximates.
STaR brought the loop to language models in 2022, with an answer key as the verifier.3 ReST and ReST-EM scaled it.45 The reinforcement-learning-with-verifiable-rewards line, from Tülu 3 through DeepSeek-R1 and DAPO, is the same loop with a gradient step instead of a fine-tune on the filtered set.678 Cobbe and colleagues' 2021 GSM8K work made the case that a trained verifier scales better than fine-tuning alone, which is the argument for the verifier in Chapter 9.9
What these established: a verifier the model does not control produces lasting gains, and the model's own judgment does not. What they left: the verifier was an answer key or a game. Software has no answer key. It has tests.
Tests as the verifier for code#
AlphaCode, in 2022, sampled up to a million programs per competition problem and filtered them through the problem's example tests before clustering and submitting.10 CodeRL, the same year, used unit-test results as the reward for training a code model with an actor-critic method.11 Meta's RLEF, in 2024, trained a code model with execution feedback as the reward and showed large gains in sample efficiency.12 SWE-RL, DeepSWE, and the Nebius work brought it to repository-scale tasks.131415
What these established: a test is a usable reward, and a model trained against tests gets better at passing tests. What they left: every one of them used public problems with public tests. The tests were visible, or at least public, and the problems were nobody's in particular.
Benchmarks built from real repositories#
SWE-bench, in 2023, defined the fail-to-pass and pass-to-pass construction and built 2,294 tasks from real GitHub issues and the pull requests that closed them.16 SWE-bench Verified had human annotators remove the tasks whose descriptions were unclear or whose tests were unfair, leaving 500.17 SWE-Gym, SWE-smith, R2E-Gym, and SWE-rebench turned the construction into pipelines that produce thousands of executable tasks.18192021 Multi-SWE-bench extended it to seven languages, and SWE-Lancer graded real freelance tasks with end-to-end tests.2223 SWE-bench+ showed how often the construction leaks the answer or accepts a wrong one.24
What these established: the flip is a reliable unit of value for software work, and tasks with hidden tests can be produced at scale from version history. What they left: the tasks are public, so every model has seen them, and the tests are released, so every agent can be shown them. A team's own history is the unlimited supply of tasks that are not public.
Learning from an organization's own development process#
This is the closest precedent and the least cited.
Facebook's Getafix, in 2019, learned fix patterns from the history of human fixes to static-analysis warnings in Facebook's own codebase and proposed fixes for new warnings, which engineers accepted at a high rate.25 SapFix, the same year, generated candidate patches for crashes found by an automated testing system and used that system's tests as the oracle, end to end, in production.26 Both learned from, and were graded by, the company's own code and tests.
Google's DIDACT, described in 2023, trained models on the process of software development inside Google rather than on finished code: the edit histories, the build-error fixes, the code-review comments and their resolutions.27 Google had earlier reported that a completion model trained on its internal code reduced coding iteration time by 6 percent and was accepted for about 3 percent of new code.28 The data was the developers' own activity, and the organization kept it.
GitHub's Copilot research found that acceptance rate was the best available predictor of developers' perceived productivity, which made acceptance a usable, if weak, oracle at scale.29 Replit, in 2024, trained a 7-billion-parameter code-repair model on data built from its own platform: language-server diagnostics, with the state of the file reconstructed by replaying the edit history.30
What these established: an organization's own development activity is training data, and models trained on it perform well on that organization's work. What they left: each was built by a company with a research team and a bespoke agent or tool. The point of Chapters 7 and 8 is that the harness hooks make this available to a team with neither.
Learning from traces in other fields#
End-to-end driving models from 2016 onward trained on recorded human driving, with the steering angle as the label.31 The traces were cheap to collect from vehicles already on the road, and the fleet's data became the moat. The shadow-mode canary in Chapter 13 is borrowed from this field.
Imitation learning has a known failure: a policy trained on an expert's trajectories drifts into states the expert never visited, and its errors compound. DAgger, from 2011, fixes it by running the learner, having the expert label the states the learner reached, and training on those.32 The analogue here is the fallback in Chapter 15: when the team's model fails a task, the rented model solves it from the same starting point, and the trace is a label on a state the team's model reached. Hindsight experience replay, from 2017, relabels failed episodes as successes for the goals they did reach, which Chapter 9 proposed for sessions that did not flip.33
Research on machine learning for code, surveyed by Allamanis and colleagues in 2018, rests on the observation from Hindle and colleagues in 2012 that software is natural: repetitive and predictable enough that statistical models of it work.3435 A team's codebase is more repetitive and more predictable than the public corpus, which is why a model trained on it does well there.
Keeping a holdout honest#
The one-bit verdict in Chapter 6 rests on the Ladder and the reusable holdout, both from 2015.3637 Both were written about machine-learning competitions and scientific data analysis. Neither mentions an agent. The problem they solved, an adaptive optimizer hill-climbing on a holdout it is only supposed to be measured by, is the problem an agent iterating against an oracle has, and the solution transfers without modification.
Agents trained on agent trajectories#
AgentTuning and FireAct, in 2023, fine-tuned open models on a few hundred to a couple of thousand agent trajectories generated by a stronger model and showed that the result generalized.3839 Kimi K2's technical report described a large-scale pipeline for synthesizing agentic data with tool use and verifying it before training.40 The open-weight agent frameworks, SWE-agent and OpenHands, standardized the tool interface that these trajectories are recorded in.4142
What these established: agent behavior transfers through trajectories, and a few hundred are enough to see it. What they left: the trajectories came from public tasks solved by public models. The trajectories this book is about come from your tasks, solved on your code, graded by your tests.
What is new#
Four things, and only four.
Hooks in the harness make trace collection free. The organization does not build the agent, and the agent does not have to be modified.
The hidden test as an air-gapped, one-bit oracle makes the grading trustworthy under optimization pressure, which the public-benchmark work did not have to worry about and the internal-tool work handled with bespoke infrastructure.
Open-weight models are close enough to the frontier, and licensed permissively enough, that the fine-tuned result is competitive on the team's distribution. In 2019 the models were not there. In 2023 the licenses often were not.
And the delivery pipeline treats weights as a release artifact with gates, canaries, and rollback, which is ordinary engineering applied to a thing that used to be a research project.
Everything else in this book was established by someone else, and the footnotes say who.
Footnotes#
-
Anthony, T., Tian, Z., & Barber, D. (2017). Thinking Fast and Slow with Deep Learning and Tree Search. NeurIPS 2017. https://arxiv.org/abs/1705.08439 ↩
-
Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550, 354–359. https://doi.org/10.1038/nature24270 ↩
-
Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. arXiv. https://arxiv.org/abs/2203.14465 ↩
-
Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). Reinforced Self-Training (ReST) for Language Modeling. arXiv. https://arxiv.org/abs/2308.08998 ↩
-
Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models. arXiv; TMLR 2024. https://arxiv.org/abs/2312.06585 ↩
-
Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv. https://arxiv.org/abs/2411.15124 ↩
-
DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948 ↩
-
Yu, Q., Zhang, Z., Zhu, R., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv. https://arxiv.org/abs/2503.14476 ↩
-
Cobbe, K., Kosaraju, V., Bavarian, M., et al. (2021). Training Verifiers to Solve Math Word Problems. arXiv. https://arxiv.org/abs/2110.14168 ↩
-
Li, Y., Choi, D., Chung, J., et al. (2022). Competition-level code generation with AlphaCode. Science, 378(6624), 1092–1097. https://doi.org/10.1126/science.abq1158 ↩
-
Le, H., Wang, Y., Gotmare, A. D., Savarese, S., & Hoi, S. C. H. (2022). CodeRL: Mastering Code Generation through Pretrained Models and Deep Reinforcement Learning. NeurIPS 2022. https://arxiv.org/abs/2207.01780 ↩
-
Gehring, J., Zheng, K., Copet, J., Mella, V., Cohen, T., & Synnaeve, G. (2024). RLEF: Grounding Code LLMs in Execution Feedback with Reinforcement Learning. arXiv. https://arxiv.org/abs/2410.02089 ↩
-
Wei, Y., Duchenne, O., Copet, J., et al. (2025). SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv; NeurIPS 2025. https://arxiv.org/abs/2502.18449 ↩
-
Agentica & Together AI (2025). DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL. https://www.together.ai/blog/deepswe ↩
-
Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning. arXiv. https://arxiv.org/abs/2508.03501 ↩
-
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv; ICLR 2024. https://arxiv.org/abs/2310.06770 ↩
-
OpenAI (2024). Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/ ↩
-
Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩
-
Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩
-
Jain, N., Singh, J., Shetty, M., Zheng, L., Sen, K., & Stoica, I. (2025). R2E-Gym: Procedural Environments and Hybrid Verifiers for Scaling Open-Weights SWE Agents. arXiv. https://arxiv.org/abs/2504.07164 ↩
-
Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. arXiv; NeurIPS 2025. https://arxiv.org/abs/2505.20411 ↩
-
Zan, D., Huang, Z., Liu, W., et al. (2025). Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving. arXiv. https://arxiv.org/abs/2504.02605 ↩
-
Miserendino, S., Wang, M., Patwardhan, T., & Heidecke, J. (2025). SWE-Lancer: Can Frontier LLMs Earn $1 Million from Real-World Freelance Software Engineering? arXiv. https://arxiv.org/abs/2502.12115 ↩
-
Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv. https://arxiv.org/abs/2410.06992 ↩
-
Bader, J., Scott, A., Pradel, M., & Chandra, S. (2019). Getafix: Learning to Fix Bugs Automatically. Proceedings of the ACM on Programming Languages, 3(OOPSLA), Article 159. https://doi.org/10.1145/3360585 ↩
-
Marginean, A., Bader, J., Chandra, S., Harman, M., Jia, Y., Mao, K., Mols, A., & Scott, A. (2019). SapFix: Automated End-to-End Repair at Scale. ICSE-SEIP 2019, 269–278. https://doi.org/10.1109/ICSE-SEIP.2019.00039 ↩
-
Maniatis, P., & Tarlow, D. (2023). Large sequence models for software development activities. Google Research Blog. https://research.google/blog/large-sequence-models-for-software-development-activities/ ↩
-
Tabachnyk, M., & Nikolov, S. (2022). ML-Enhanced Code Completion Improves Developer Productivity. Google Research Blog. https://research.google/blog/ml-enhanced-code-completion-improves-developer-productivity/ ↩
-
Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3), 54–63. https://doi.org/10.1145/3633453 ↩
-
Singhal, M., Carelli, R., Segato, G., Kumar, V., & Catasta, M. (2024). Building LLMs for Code Repair. Replit. https://replit.com/blog/code-repair ↩
-
Bojarski, M., Del Testa, D., Dworakowski, D., et al. (2016). End to End Learning for Self-Driving Cars. arXiv. https://arxiv.org/abs/1604.07316 ↩
-
Ross, S., Gordon, G. J., & Bagnell, J. A. (2011). A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning. AISTATS 2011. https://arxiv.org/abs/1011.0686 ↩
-
Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). Hindsight Experience Replay. NeurIPS 2017. https://arxiv.org/abs/1707.01495 ↩
-
Hindle, A., Barr, E. T., Su, Z., Gabel, M., & Devanbu, P. (2012). On the Naturalness of Software. ICSE 2012, 837–847. https://doi.org/10.1109/ICSE.2012.6227135 ↩
-
Allamanis, M., Barr, E. T., Devanbu, P., & Sutton, C. (2018). A Survey of Machine Learning for Big Code and Naturalness. ACM Computing Surveys, 51(4), Article 81. https://arxiv.org/abs/1709.06182 ↩
-
Blum, A., & Hardt, M. (2015). The Ladder: A Reliable Leaderboard for Machine Learning Competitions. ICML 2015, PMLR 37, 1006–1014. https://arxiv.org/abs/1502.04585 ↩
-
Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636–638. https://doi.org/10.1126/science.aaa9375 ↩
-
Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv. https://arxiv.org/abs/2310.12823 ↩
-
Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv. https://arxiv.org/abs/2310.05915 ↩
-
Kimi Team (2025). Kimi K2: Open Agentic Intelligence. arXiv. https://arxiv.org/abs/2507.20534 ↩
-
Yang, J., Jimenez, C. E., Wettig, A., Lieret, K., Yao, S., Narasimhan, K., & Press, O. (2024). SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering. NeurIPS 2024. https://arxiv.org/abs/2405.15793 ↩
-
Wang, X., Li, B., Song, Y., et al. (2024). OpenHands: An Open Platform for AI Software Developers as Generalist Agents. arXiv; ICLR 2025. https://arxiv.org/abs/2407.16741 ↩
Cite this
Anderson, M. (2026). Precedents. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 17). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/precedents
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {Precedents},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {17},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/precedents}
}Updates by email
New research reaches subscribers first.