Part IV: Training
12Data volume
What 491, 5,016, and 8,209 verified trajectories bought in the literature, and a worksheet for your own rate.
Chapter 12 of 22
How many flips does it take? This chapter collects the published numbers, states the pattern they show, and gives a worksheet a team can fill in with its own rates.
What the literature reports#
The results below are for different models, methods, and benchmarks, so the table is a set of reference points, not a curve. Read it for order of magnitude.
| Data | Method | Model | Result | Source |
|---|---|---|---|---|
| 1,000 curated prompt-response pairs | Supervised fine-tuning | LLaMA 65B | Preferred to or tied with GPT-4 responses in 43 percent of human comparisons | LIMA1 |
| 1,000 reasoning questions with traces | Supervised fine-tuning | Qwen2.5-32B-Instruct | Competitive with o1-preview on competition math; 57 percent on AIME24 with budget forcing | s12 |
| 817 curated math problems with solutions | Supervised fine-tuning | Qwen2.5-32B-Instruct | 57.1 percent on AIME24 and 94.8 percent on MATH500 in the first release | LIMO3 |
| 500 agent trajectories | Supervised fine-tuning | Llama 2 7B | 77 percent relative gain on a question-answering agent task | FireAct4 |
| 1,866 agent trajectories across six tasks | Supervised fine-tuning with general data mixed in | Llama 2 7B to 70B | Agent abilities generalize to held-out tasks | AgentTuning5 |
| 491 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 7.0 to 20.6 percent on SWE-bench Verified; 32.0 with a trained verifier and 16 samples | SWE-Gym6 |
| 5,016 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 40.2 percent on SWE-bench Verified | SWE-smith7 |
| 8,209 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 6.4 to 38.0 percent; 47.0 with best-of-8 and a critic; log-linear in data with no plateau | Skywork-SWE8 |
| About 4,500 executable tasks | Reinforcement learning only | Qwen3-32B | 23 to 42.2 percent; 59.0 with test-time scaling | DeepSWE9 |
| Rejection sampling, then RL on executable tasks | Both | Qwen2.5-72B-Instruct | 11.4 to 20.5 to 39.0 percent | Nebius10 |
Three patterns run through the table.
Hundreds move a model. Every result in the first half of the table used about a thousand examples or fewer and produced a large change. LIMA's authors argued that almost all of a model's knowledge comes from pretraining and that alignment needs only a small set of examples to teach format and style.1 The agent results say something stronger for this domain: 491 verified trajectories nearly tripled a 32-billion-parameter model's resolve rate.6 The first useful model is closer than most teams expect.
Thousands keep paying. Skywork-SWE's scaling curve is the most direct measurement: 2,000 trajectories gave 31.8 percent, 6,000 gave 36.1, and 8,209 gave 38.0, with the curve still rising.8 Zhang and colleagues found the same shape across fine-tuning tasks and model sizes and fit it as a power law in the amount of fine-tuning data.11 The returns diminish per example and do not stop.
Quality beats quantity at every scale. AlpaGasus trained on 9,000 examples filtered from a 52,000-example set and beat the model trained on all 52,000.12 LIMA, s1, and LIMO are all arguments that a small curated set beats a large uncurated one. For this pipeline, the oracle is the curation. A flip is, by construction, an example that was verified. The volume question is how many verified examples, not how many sessions.
The worksheet#
Four numbers set the time to a given training set size.
- Sessions per day. Engineers times sessions each. A team of 50 running 10 each is 500.
- Oracle share. The fraction of sessions dispatched on tasks that have a hidden test. A team starting out might reach 20 percent. A team that writes the test first for every bug and most features can reach 60.
- Flip rate. The fraction of oracle sessions that end on PASS. With a silent oracle and a rented frontier model on tasks of reasonable size, 30 to 50 percent is a working assumption; a team should measure it in its first week.
- Attempts per task. Independent sessions dispatched per oracle task. Each attempt costs tokens and raises the chance that at least one flips.
Flips per day is roughly: sessions × oracle share × flip rate, adjusted upward for attempts. With the numbers above and one attempt, 500 × 0.2 × 0.33 is about 33 flips a day. With three attempts and a per-attempt flip rate of 33 percent, the chance a task flips at least once is about 70 percent, so flips per task-day rises to about 70 of 100 oracle tasks, at three times the token cost.
Working days to reach a training set size, one attempt per task
| 20 percent oracle share | 60 percent oracle share | |
|---|---|---|
| 500 flips, first fine-tune | 15days | 5days |
| 2,000 flips | 60days | 20days |
| 5,000 flips | 150days | 50days |
| 8,000 flips | 240days | 80days |
The reading of the chart is that a mid-sized team reaches the SWE-Gym regime in weeks and the Skywork regime within a year, and that the oracle share is the lever that matters most. Every bug fixed without a hidden test is a session that could have been a flip and was not.
Limits of volume#
More flips do not fix a bad oracle. A thousand traces graded by a flaky test are a thousand noisy labels, and the noise does not average out; it teaches. More flips do not fix a narrow task distribution. Five thousand flips on one service teach that service. And more flips do not fix a leaked hidden test; they make the leak worse, because every trace fitted to the leaked test is one more example of fitting.
The order of operations is: get the oracle right, get the task supply broad, then grow the volume. Chapter 13 is the pipeline that does the third once the first two are in place.
Footnotes#
-
Zhou, C., Liu, P., Xu, P., et al. (2023). LIMA: Less Is More for Alignment. NeurIPS 2023. https://arxiv.org/abs/2305.11206 ↩ ↩2
-
Muennighoff, N., Yang, Z., Shi, W., et al. (2025). s1: Simple test-time scaling. arXiv. https://arxiv.org/abs/2501.19393 ↩
-
Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). LIMO: Less is More for Reasoning. arXiv (v1, February 2025); COLM 2025. https://arxiv.org/abs/2502.03387 ↩
-
Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). FireAct: Toward Language Agent Fine-tuning. arXiv. https://arxiv.org/abs/2310.05915 ↩
-
Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). AgentTuning: Enabling Generalized Agent Abilities for LLMs. arXiv. https://arxiv.org/abs/2310.12823 ↩
-
Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩ ↩2
-
Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩
-
Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290 ↩ ↩2
-
Agentica & Together AI (2025). DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL. https://www.together.ai/blog/deepswe ↩
-
Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning. arXiv. https://arxiv.org/abs/2508.03501 ↩
-
Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method. ICLR 2024. https://arxiv.org/abs/2402.17193 ↩
-
Chen, L., Li, S., Yan, J., et al. (2023). AlpaGasus: Training a Better Alpaca with Fewer Data. arXiv; ICLR 2024. https://arxiv.org/abs/2307.08701 ↩
Cite this
Anderson, M. (2026). Data volume. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 12). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {Data volume},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {12},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume}
}Updates by email
New research reaches subscribers first.