Mac Anderson

Part IV: Training

11Training methods

Supervised fine-tuning on flips, then preference pairs, then reinforcement learning. Adapters, forgetting, collapse, and a shuffled-label control.

Mac Anderson8 min read1,767 words
View markdown

With a training set in hand, the question is what to do with it. This chapter covers the three families of methods that have produced results on agent trajectories, in the order a team should adopt them, and the choices inside each: adapter or full fine-tune, how to avoid forgetting, and how to know the signal is real.

Supervised fine-tuning on flips#

The first method is the simplest. Take the flipped trajectories, formatted and masked as Chapter 10 describes, and continue training the base model on them with the standard next-token loss. This is supervised fine-tuning, and when the examples were selected by a verifier, it has a second name: rejection sampling fine-tuning.

The method has a lineage. Expert iteration, from 2017, alternates between a slow expert that solves problems and a fast policy that imitates the solutions, and uses the expert's successes as the training set.1 AlphaGo Zero's self-play is the same loop with the game's outcome as the verifier.2 STaR, from 2022, applied it to language models: sample reasoning, keep the samples whose answers match the key, fine-tune, repeat.3 ReST and ReST-EM scaled it with a reward model and with answer keys.45 For code, AlphaCode filtered thousands of samples per problem through the problem's tests before choosing which to submit.6 Llama 2's post-training used rejection sampling as a step before reinforcement learning.7 The agent-trajectory results from Chapter 1 are this method applied to software: SWE-Gym, SWE-smith, and Skywork-SWE are all supervised fine-tuning on verifier-selected trajectories.8910

It works, it is cheap, and it is the right place to start. Its limit is that it can only teach the model to do what some session already did. A task no session ever flipped contributes nothing to the supervised set.

Preference optimization on pairs#

The second method uses the failures. Direct preference optimization takes pairs of trajectories on the same task, one preferred and one not, and trains the model to assign higher likelihood to the preferred one relative to a reference model.11 It needs no reward model and no sampling during training, which makes it almost as cheap as supervised fine-tuning.

The preference set from Chapter 10 is the input: flipped versus not flipped on the same task. The pairs teach the model something the supervised set cannot, which is what a wrong trajectory looks like on this codebase. A team that has run two attempts per task has pairs for every task where exactly one flipped.

Two cautions. Preference optimization is known to drift toward longer outputs unless the pairs are controlled for length, and agent trajectories vary in length a lot.12 Match pair lengths roughly or use a length-regularized variant. And the pairs should be real contrasts. A pair where the rejected trajectory failed because the oracle timed out is not a lesson about the code.

Reinforcement learning with the oracle as the reward#

The third method puts the oracle in the training loop. The model attempts a task, the oracle grades the attempt, and the model is updated to make PASS more likely. This is reinforcement learning with verifiable rewards, named as such in the Tülu 3 report and used at scale in DeepSeek-R1.1314 For software agents, SWE-RL used a patch-similarity reward on mined pull requests, DeepSWE used a pass-or-fail reward on executable environments, and the Nebius team took a 72-billion-parameter model from 11.4 percent to 39.0 percent on SWE-bench Verified with rejection sampling followed by a variant of the DAPO algorithm.15161718

The appeal is that reinforcement learning can learn from tasks no session has flipped yet, because it searches. The costs are real. Each training step needs many fresh attempts, each attempt needs an oracle run, and the oracle runs in a container. Qwen's team described a system running 20,000 environments in parallel for its coding model's reinforcement learning stage.19 A team does not need that scale to see gains; DeepSWE used about 4,500 tasks.16 It does need an oracle that can grade hundreds of attempts an hour, which means the oracle has to be a service, not a hook.

There is also a result to read before committing. Yue and colleagues found that reinforcement learning with verifiable rewards sharpens a model toward answers its base model could already produce: the trained model wins when it gets one attempt, and the base model wins when both get many attempts, because the base model's attempts are more varied.20 For a team, that argues for a sequence: supervised fine-tuning on flips first, to widen what the model can do on your code; reinforcement learning second, to make it do it reliably on the first try.

Adopt the methods in that order. Supervised fine-tuning on flips as soon as there are a few hundred. Preference optimization as soon as there are pairs. Reinforcement learning when the oracle is a service and the flip rate has plateaued.

Training methods in the order a team should adopt them

  1. Supervised fine-tuning on flipped trajectoriesRejection sampling fine-tuning. Cheapest. Learns what some session already did.
  2. Preference optimization on flipped versus failed pairsUses the failures. Learns what wrong looks like here.
  3. Reinforcement learning with the oracle as rewardSearches. Can learn tasks no session flipped. Needs the oracle as a service.

Cost and what it can learn

Each rung needs the one below it. The data each rung needs comes from the same trace store.

Adapter or full fine-tune#

Low-rank adaptation freezes the base model's weights and trains small matrices added to them.21 The adapter is a few percent of the model's size, trains on less hardware, and can be swapped at inference. QLoRA goes further by quantizing the frozen base to 4 bits, which brought fine-tuning of 65-billion-parameter models onto a single large card.22

The question is whether the adapter learns as much. Biderman and colleagues compared the two on code and math and found that full fine-tuning learned more on the target domain and forgot more of what the base model knew; low-rank adaptation learned less and forgot less, and acted as a regularizer.23 Schulman and colleagues at Thinking Machines then reported that the gap closes when the adapter is applied to all layers rather than only attention, and that for small-to-medium post-training sets, the size a team's trace store will be for its first year, the adapter matches the full fine-tune. Their result for reinforcement learning is sharper: a rank-1 adapter matched full fine-tuning, which they explain by noting that a policy-gradient step absorbs about one bit of information per episode, so the capacity needed is tiny.24

That last point connects to this book's oracle. A one-bit verdict is one bit per episode. A method that learns one bit per episode needs many episodes and almost no adapter capacity. That is the regime the pipeline is in, and the adapter is the right tool for it. Start with adapters on all layers. Move to full fine-tuning only if an evaluation shows the adapter is the bottleneck, which for the first several thousand flips it is unlikely to be.

Forgetting#

A model fine-tuned on traces from one codebase can get worse at everything else. The phenomenon is catastrophic forgetting, known since 1989 and measured in language models at every scale.2526 For a coding model, it looks like a fine-tune that resolves the team's tasks and can no longer write a shell script.

The standard defenses apply.

  • Replay. Mix a fraction of general instruction data into every training run. Ibrahim and colleagues showed that replay combined with re-warming the learning rate lets a model take on new data while matching a full retrain on the old.27
  • Low learning rate and few epochs. One to three passes over the flips at a learning rate an order of magnitude below pretraining.
  • Adapters. The frozen base cannot forget. The adapter can be removed. This is the regularization Biderman and colleagues measured.23
  • Merge, don't stack. When several adapters have been trained on different slices, merge them with a method that keeps the base model's weights and averages or resolves the deltas, such as model soups or TIES, rather than fine-tuning one on top of another.2829
  • Evaluate on a general benchmark. Chapter 14 puts a public coding benchmark in the gate for this reason: a model that gained on the team's tasks and lost on the public set has forgotten, and the gate should see it.

Collapse#

Training a model on its own output can make it worse over generations. Shumailov and colleagues showed that models trained recursively on their own generations lose the tails of the distribution and converge on a narrow set of outputs, a process they called model collapse.30 A pipeline that trains a model on traces produced by the previous version of itself is, on its face, exactly that loop.

Two things make it different. First, the oracle. Collapse happens when generated data replaces real data without selection. The traces in this pipeline are selected by a verifier the model does not control, so the distribution being trained on is the distribution of correct solutions, not of the model's output. Gerstgrasser and colleagues showed that collapse is also avoided when generated data accumulates alongside the original data rather than replacing it.31 Second, the provenance field from Chapter 10. A training run that knows which traces came from a stronger external model and which from its own predecessor can weight them, cap the self-generated share, and watch the ratio over time.

The warning sign is a model whose flips get shorter, more uniform, and more alike across tasks. The frozen evaluation set from Chapter 10 is the instrument.

Is the signal real#

One last check belongs in every training run, and it is cheap. Shao and colleagues found that training a particular family of math models with random rewards, rewards that had nothing to do with correctness, improved its benchmark scores by more than 20 points.32 The gain came from the training procedure nudging the model toward behaviors it already had, not from the reward. The same procedure did nothing for other model families. The result does not say verifiable rewards are fake. It says a gain after training is not, on its own, evidence that the reward carried information.

The control is to shuffle the labels. Train the same recipe on the same traces with the flip labels randomly permuted, so that half the "flipped" trajectories are failures. If the shuffled run gains as much as the real run on the held-out evaluation, the gain is not coming from the oracle, and something else in the recipe is doing the work. A team should run this control once when it sets up the pipeline, and again whenever the recipe changes.

Footnotes#

  1. Anthony, T., Tian, Z., & Barber, D. (2017). Thinking Fast and Slow with Deep Learning and Tree Search. NeurIPS 2017. https://arxiv.org/abs/1705.08439 ↩

  2. Silver, D., Schrittwieser, J., Simonyan, K., et al. (2017). Mastering the game of Go without human knowledge. Nature, 550, 354–359. https://doi.org/10.1038/nature24270 ↩

  3. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. arXiv. https://arxiv.org/abs/2203.14465 ↩

  4. Gulcehre, C., Le Paine, T., Srinivasan, S., et al. (2023). Reinforced Self-Training (ReST) for Language Modeling. arXiv. https://arxiv.org/abs/2308.08998 ↩

  5. Singh, A., Co-Reyes, J. D., Agarwal, R., et al. (2023). Beyond Human Data: Scaling Self-Training for Problem-Solving with Language Models. arXiv; TMLR 2024. https://arxiv.org/abs/2312.06585 ↩

  6. Li, Y., Choi, D., Chung, J., et al. (2022). Competition-level code generation with AlphaCode. Science, 378(6624), 1092–1097. https://doi.org/10.1126/science.abq1158 ↩

  7. Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv. https://arxiv.org/abs/2307.09288 ↩

  8. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩

  9. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩

  10. Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290 ↩

  11. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. https://arxiv.org/abs/2305.18290 ↩

  12. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). A Long Way to Go: Investigating Length Correlations in RLHF. arXiv; COLM 2024. https://arxiv.org/abs/2310.03716 ↩

  13. Lambert, N., Morrison, J., Pyatkin, V., et al. (2024). Tülu 3: Pushing Frontiers in Open Language Model Post-Training. arXiv. https://arxiv.org/abs/2411.15124 ↩

  14. DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948 ↩

  15. Wei, Y., Duchenne, O., Copet, J., et al. (2025). SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv; NeurIPS 2025. https://arxiv.org/abs/2502.18449 ↩

  16. Agentica & Together AI (2025). DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL. https://www.together.ai/blog/deepswe ↩ ↩2

  17. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning. arXiv. https://arxiv.org/abs/2508.03501 ↩

  18. Yu, Q., Zhang, Z., Zhu, R., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv. https://arxiv.org/abs/2503.14476 ↩

  19. Qwen Team (2025). Qwen3-Coder: Agentic Coding in the World. https://qwenlm.github.io/blog/qwen3-coder/ ↩

  20. Yue, Y., Chen, Z., Lu, R., et al. (2025). Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model? arXiv; NeurIPS 2025. https://arxiv.org/abs/2504.13837 ↩

  21. Hu, E. J., Shen, Y., Wallis, P., Allen-Zhu, Z., Li, Y., Wang, S., Wang, L., & Chen, W. (2021). LoRA: Low-Rank Adaptation of Large Language Models. arXiv; ICLR 2022. https://arxiv.org/abs/2106.09685 ↩

  22. Dettmers, T., Pagnoni, A., Holtzman, A., & Zettlemoyer, L. (2023). QLoRA: Efficient Finetuning of Quantized LLMs. NeurIPS 2023. https://arxiv.org/abs/2305.14314 ↩

  23. Biderman, D., Portes, J., Gonzalez Ortiz, J. J., et al. (2024). LoRA Learns Less and Forgets Less. Transactions on Machine Learning Research. https://arxiv.org/abs/2405.09673 ↩ ↩2

  24. Schulman, J., & Thinking Machines Lab (2025). LoRA Without Regret. https://thinkingmachines.ai/blog/lora/ ↩

  25. McCloskey, M., & Cohen, N. J. (1989). Catastrophic Interference in Connectionist Networks: The Sequential Learning Problem. Psychology of Learning and Motivation, 24, 109–165. https://doi.org/10.1016/S0079-7421(08)60536-8 ↩

  26. Luo, Y., Yang, Z., Meng, F., Li, Y., Zhou, J., & Zhang, Y. (2023). An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning. arXiv. https://arxiv.org/abs/2308.08747 ↩

  27. Ibrahim, A., Thérien, B., Gupta, K., et al. (2024). Simple and Scalable Strategies to Continually Pre-train Large Language Models. arXiv; TMLR. https://arxiv.org/abs/2403.08763 ↩

  28. Wortsman, M., Ilharco, G., Gadre, S. Y., et al. (2022). Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time. ICML 2022. https://arxiv.org/abs/2203.05482 ↩

  29. Yadav, P., Tam, D., Choshen, L., Raffel, C., & Bansal, M. (2023). TIES-Merging: Resolving Interference When Merging Models. NeurIPS 2023. https://arxiv.org/abs/2306.01708 ↩

  30. Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755–759. https://doi.org/10.1038/s41586-024-07566-y ↩

  31. Gerstgrasser, M., Schaeffer, R., Dey, A., et al. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv. https://arxiv.org/abs/2404.01413 ↩

  32. Shao, R., Li, S. S., Xin, R., et al. (2025). Spurious Rewards: Rethinking Training Signals in RLVR. arXiv. https://arxiv.org/abs/2506.10947 ↩

Cite this

Anderson, M. (2026). Training methods. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 11). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/training-methods

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Training methods},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {11},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/training-methods}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.