Mac Anderson

Part IV: Training

12Data volume

What 491, 5,016, and 8,209 verified trajectories bought in the literature, and a worksheet for your own rate.

Mac Anderson4 min read920 words
View markdown

How many flips does it take? This chapter collects the published numbers, states the pattern they show, and gives a worksheet a team can fill in with its own rates.

What the literature reports#

The results below are for different models, methods, and benchmarks, so the table is a set of reference points, not a curve. Read it for order of magnitude.

DataMethodModelResultSource
1,000 curated prompt-response pairsSupervised fine-tuningLLaMA 65BPreferred to or tied with GPT-4 responses in 43 percent of human comparisonsLIMA1
1,000 reasoning questions with tracesSupervised fine-tuningQwen2.5-32B-InstructCompetitive with o1-preview on competition math; 57 percent on AIME24 with budget forcings12
817 curated math problems with solutionsSupervised fine-tuningQwen2.5-32B-Instruct57.1 percent on AIME24 and 94.8 percent on MATH500 in the first releaseLIMO3
500 agent trajectoriesSupervised fine-tuningLlama 2 7B77 percent relative gain on a question-answering agent taskFireAct4
1,866 agent trajectories across six tasksSupervised fine-tuning with general data mixed inLlama 2 7B to 70BAgent abilities generalize to held-out tasksAgentTuning5
491 verified software trajectoriesSupervised fine-tuningQwen2.5-Coder-32B7.0 to 20.6 percent on SWE-bench Verified; 32.0 with a trained verifier and 16 samplesSWE-Gym6
5,016 verified software trajectoriesSupervised fine-tuningQwen2.5-Coder-32B40.2 percent on SWE-bench VerifiedSWE-smith7
8,209 verified software trajectoriesSupervised fine-tuningQwen2.5-Coder-32B6.4 to 38.0 percent; 47.0 with best-of-8 and a critic; log-linear in data with no plateauSkywork-SWE8
About 4,500 executable tasksReinforcement learning onlyQwen3-32B23 to 42.2 percent; 59.0 with test-time scalingDeepSWE9
Rejection sampling, then RL on executable tasksBothQwen2.5-72B-Instruct11.4 to 20.5 to 39.0 percentNebius10

Three patterns run through the table.

Hundreds move a model. Every result in the first half of the table used about a thousand examples or fewer and produced a large change. LIMA's authors argued that almost all of a model's knowledge comes from pretraining and that alignment needs only a small set of examples to teach format and style.1 The agent results say something stronger for this domain: 491 verified trajectories nearly tripled a 32-billion-parameter model's resolve rate.6 The first useful model is closer than most teams expect.

Thousands keep paying. Skywork-SWE's scaling curve is the most direct measurement: 2,000 trajectories gave 31.8 percent, 6,000 gave 36.1, and 8,209 gave 38.0, with the curve still rising.8 Zhang and colleagues found the same shape across fine-tuning tasks and model sizes and fit it as a power law in the amount of fine-tuning data.11 The returns diminish per example and do not stop.

Quality beats quantity at every scale. AlpaGasus trained on 9,000 examples filtered from a 52,000-example set and beat the model trained on all 52,000.12 LIMA, s1, and LIMO are all arguments that a small curated set beats a large uncurated one. For this pipeline, the oracle is the curation. A flip is, by construction, an example that was verified. The volume question is how many verified examples, not how many sessions.

Verified trajectories behind each open-model result on SWE-bench Verified

ResultTrajectories
SWE-Gym, 20.6% after fine-tuning491
Skywork-SWE at 31.8%2000
SWE-smith, 40.2%5016
Skywork-SWE at 36.1%6000
Skywork-SWE, 38.0%8209
Sources: Pan et al. (2024), Yang et al. (2025), Zeng et al. (2025). The same 32-billion-parameter base model in every row. The Skywork rows are points on one scaling curve; the other two are separate recipes.

The worksheet#

Four numbers set the time to a given training set size.

  1. Sessions per day. Engineers times sessions each. A team of 50 running 10 each is 500.
  2. Oracle share. The fraction of sessions dispatched on tasks that have a hidden test. A team starting out might reach 20 percent. A team that writes the test first for every bug and most features can reach 60.
  3. Flip rate. The fraction of oracle sessions that end on PASS. With a silent oracle and a rented frontier model on tasks of reasonable size, 30 to 50 percent is a working assumption; a team should measure it in its first week.
  4. Attempts per task. Independent sessions dispatched per oracle task. Each attempt costs tokens and raises the chance that at least one flips.

Flips per day is roughly: sessions × oracle share × flip rate, adjusted upward for attempts. With the numbers above and one attempt, 500 × 0.2 × 0.33 is about 33 flips a day. With three attempts and a per-attempt flip rate of 33 percent, the chance a task flips at least once is about 70 percent, so flips per task-day rises to about 70 of 100 oracle tasks, at three times the token cost.

Working days to reach a training set size, one attempt per task

20 percent oracle share60 percent oracle share
500 flips, first fine-tune15days5days
2,000 flips60days20days
5,000 flips150days50days
8,000 flips240days80days
Arithmetic for 500 sessions a day at a 33 percent flip rate. Sampling three attempts per task roughly halves every number at three times the token cost. These are planning figures, not measurements.

The reading of the chart is that a mid-sized team reaches the SWE-Gym regime in weeks and the Skywork regime within a year, and that the oracle share is the lever that matters most. Every bug fixed without a hidden test is a session that could have been a flip and was not.

Limits of volume#

More flips do not fix a bad oracle. A thousand traces graded by a flaky test are a thousand noisy labels, and the noise does not average out; it teaches. More flips do not fix a narrow task distribution. Five thousand flips on one service teach that service. And more flips do not fix a leaked hidden test; they make the leak worse, because every trace fitted to the leaked test is one more example of fitting.

The order of operations is: get the oracle right, get the task supply broad, then grow the volume. Chapter 13 is the pipeline that does the third once the first two are in place.

Footnotes#

  1. Zhou, C., Liu, P., Xu, P., et al. (2023). LIMA: Less Is More for Alignment. NeurIPS 2023. https://arxiv.org/abs/2305.11206 ↩ ↩2

  2. Muennighoff, N., Yang, Z., Shi, W., et al. (2025). s1: Simple test-time scaling. arXiv. https://arxiv.org/abs/2501.19393 ↩

  3. Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). LIMO: Less is More for Reasoning. arXiv (v1, February 2025); COLM 2025. https://arxiv.org/abs/2502.03387 ↩

  4. Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv. https://arxiv.org/abs/2310.05915 ↩

  5. Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv. https://arxiv.org/abs/2310.12823 ↩

  6. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩ ↩2

  7. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩

  8. Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290 ↩ ↩2

  9. Agentica & Together AI (2025). DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL. https://www.together.ai/blog/deepswe ↩

  10. Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning. arXiv. https://arxiv.org/abs/2508.03501 ↩

  11. Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method. ICLR 2024. https://arxiv.org/abs/2402.17193 ↩

  12. Chen, L., Li, S., Yan, J., et al. (2023). AlpaGasus: Training a Better Alpaca with Fewer Data. arXiv; ICLR 2024. https://arxiv.org/abs/2307.08701 ↩

Cite this

Anderson, M. (2026). Data volume. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 12). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Data volume},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {12},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.