# Chapter 12: Data volume

> What 491, 5,016, and 8,209 verified trajectories bought in the literature, and a worksheet for your own rate.

From *How to Own Intelligence* by Mac Anderson. Canonical page: https://macanderson.com/research/how-to-own-intelligence/data-volume

How many flips does it take? This chapter collects the published numbers, states the pattern they show, and gives a worksheet a team can fill in with its own rates.

## What the literature reports

The results below are for different models, methods, and benchmarks, so the table is a set of reference points, not a curve. Read it for order of magnitude.

| Data | Method | Model | Result | Source |
| --- | --- | --- | --- | --- |
| 1,000 curated prompt-response pairs | Supervised fine-tuning | LLaMA 65B | Preferred to or tied with GPT-4 responses in 43 percent of human comparisons | LIMA[1](#user-content-fn-lima) |
| 1,000 reasoning questions with traces | Supervised fine-tuning | Qwen2.5-32B-Instruct | Competitive with o1-preview on competition math; 57 percent on AIME24 with budget forcing | s1[2](#user-content-fn-s1) |
| 817 curated math problems with solutions | Supervised fine-tuning | Qwen2.5-32B-Instruct | 57.1 percent on AIME24 and 94.8 percent on MATH500 in the first release | LIMO[3](#user-content-fn-limo) |
| 500 agent trajectories | Supervised fine-tuning | Llama 2 7B | 77 percent relative gain on a question-answering agent task | FireAct[4](#user-content-fn-fireact) |
| 1,866 agent trajectories across six tasks | Supervised fine-tuning with general data mixed in | Llama 2 7B to 70B | Agent abilities generalize to held-out tasks | AgentTuning[5](#user-content-fn-agenttuning) |
| 491 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 7.0 to 20.6 percent on SWE-bench Verified; 32.0 with a trained verifier and 16 samples | SWE-Gym[6](#user-content-fn-swe-gym) |
| 5,016 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 40.2 percent on SWE-bench Verified | SWE-smith[7](#user-content-fn-swe-smith) |
| 8,209 verified software trajectories | Supervised fine-tuning | Qwen2.5-Coder-32B | 6.4 to 38.0 percent; 47.0 with best-of-8 and a critic; log-linear in data with no plateau | Skywork-SWE[8](#user-content-fn-skywork-swe) |
| About 4,500 executable tasks | Reinforcement learning only | Qwen3-32B | 23 to 42.2 percent; 59.0 with test-time scaling | DeepSWE[9](#user-content-fn-deepswe) |
| Rejection sampling, then RL on executable tasks | Both | Qwen2.5-72B-Instruct | 11.4 to 20.5 to 39.0 percent | Nebius[10](#user-content-fn-nebius-rl) |

Three patterns run through the table.

**Hundreds move a model.** Every result in the first half of the table used about a thousand examples or fewer and produced a large change. LIMA's authors argued that almost all of a model's knowledge comes from pretraining and that alignment needs only a small set of examples to teach format and style.[1](#user-content-fn-lima) The agent results say something stronger for this domain: 491 verified trajectories nearly tripled a 32-billion-parameter model's resolve rate.[6](#user-content-fn-swe-gym) The first useful model is closer than most teams expect.

**Thousands keep paying.** Skywork-SWE's scaling curve is the most direct measurement: 2,000 trajectories gave 31.8 percent, 6,000 gave 36.1, and 8,209 gave 38.0, with the curve still rising.[8](#user-content-fn-skywork-swe) Zhang and colleagues found the same shape across fine-tuning tasks and model sizes and fit it as a power law in the amount of fine-tuning data.[11](#user-content-fn-zhang-scaling) The returns diminish per example and do not stop.

**Quality beats quantity at every scale.** AlpaGasus trained on 9,000 examples filtered from a 52,000-example set and beat the model trained on all 52,000.[12](#user-content-fn-alpagasus) LIMA, s1, and LIMO are all arguments that a small curated set beats a large uncurated one. For this pipeline, the oracle is the curation. A flip is, by construction, an example that was verified. The volume question is how many verified examples, not how many sessions.

> **Verified trajectories behind each open-model result on SWE-bench Verified**
>
> | Result | Trajectories |
> | --- | --- |
> | SWE-Gym, 20.6% after fine-tuning | 491 |
> | Skywork-SWE at 31.8% | 2000 |
> | SWE-smith, 40.2% | 5016 |
> | Skywork-SWE at 36.1% | 6000 |
> | Skywork-SWE, 38.0% | 8209 |
>
> *Sources: Pan et al. (2024), Yang et al. (2025), Zeng et al. (2025). The same 32-billion-parameter base model in every row. The Skywork rows are points on one scaling curve; the other two are separate recipes.*

## The worksheet

Four numbers set the time to a given training set size.

1.  **Sessions per day.** Engineers times sessions each. A team of 50 running 10 each is 500.
2.  **Oracle share.** The fraction of sessions dispatched on tasks that have a hidden test. A team starting out might reach 20 percent. A team that writes the test first for every bug and most features can reach 60.
3.  **Flip rate.** The fraction of oracle sessions that end on PASS. With a silent oracle and a rented frontier model on tasks of reasonable size, 30 to 50 percent is a working assumption; a team should measure it in its first week.
4.  **Attempts per task.** Independent sessions dispatched per oracle task. Each attempt costs tokens and raises the chance that at least one flips.

Flips per day is roughly: sessions × oracle share × flip rate, adjusted upward for attempts. With the numbers above and one attempt, 500 × 0.2 × 0.33 is about 33 flips a day. With three attempts and a per-attempt flip rate of 33 percent, the chance a task flips at least once is about 70 percent, so flips per task-day rises to about 70 of 100 oracle tasks, at three times the token cost.

> **Working days to reach a training set size, one attempt per task**
>
> |  | 20 percent oracle share | 60 percent oracle share |
> | --- | --- | --- |
> | 500 flips, first fine-tune | 15days | 5days |
> | 2,000 flips | 60days | 20days |
> | 5,000 flips | 150days | 50days |
> | 8,000 flips | 240days | 80days |
>
> *Arithmetic for 500 sessions a day at a 33 percent flip rate. Sampling three attempts per task roughly halves every number at three times the token cost. These are planning figures, not measurements.*

The reading of the chart is that a mid-sized team reaches the SWE-Gym regime in weeks and the Skywork regime within a year, and that the oracle share is the lever that matters most. Every bug fixed without a hidden test is a session that could have been a flip and was not.

## Limits of volume

More flips do not fix a bad oracle. A thousand traces graded by a flaky test are a thousand noisy labels, and the noise does not average out; it teaches. More flips do not fix a narrow task distribution. Five thousand flips on one service teach that service. And more flips do not fix a leaked hidden test; they make the leak worse, because every trace fitted to the leaked test is one more example of fitting.

The order of operations is: get the oracle right, get the task supply broad, then grow the volume. Chapter 13 is the pipeline that does the third once the first two are in place.

## Footnotes

1.  Zhou, C., Liu, P., Xu, P., et al. (2023). *LIMA: Less Is More for Alignment*. NeurIPS 2023. [https://arxiv.org/abs/2305.11206](https://arxiv.org/abs/2305.11206) [↩](#user-content-fnref-lima) [↩2](#user-content-fnref-lima-2)
    
2.  Muennighoff, N., Yang, Z., Shi, W., et al. (2025). *s1: Simple test-time scaling*. arXiv. [https://arxiv.org/abs/2501.19393](https://arxiv.org/abs/2501.19393) [↩](#user-content-fnref-s1)
    
3.  Ye, Y., Huang, Z., Xiao, Y., Chern, E., Xia, S., & Liu, P. (2025). *LIMO: Less is More for Reasoning*. arXiv (v1, February 2025); COLM 2025. [https://arxiv.org/abs/2502.03387](https://arxiv.org/abs/2502.03387) [↩](#user-content-fnref-limo)
    
4.  Chen, B., Shu, C., Shareghi, E., Collier, N., Narasimhan, K., & Yao, S. (2023). *FireAct: Toward Language Agent Fine-tuning*. arXiv. [https://arxiv.org/abs/2310.05915](https://arxiv.org/abs/2310.05915) [↩](#user-content-fnref-fireact)
    
5.  Zeng, A., Liu, M., Lu, R., Wang, B., Liu, X., Dong, Y., & Tang, J. (2023). *AgentTuning: Enabling Generalized Agent Abilities for LLMs*. arXiv. [https://arxiv.org/abs/2310.12823](https://arxiv.org/abs/2310.12823) [↩](#user-content-fnref-agenttuning)
    
6.  Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). *Training Software Engineering Agents and Verifiers with SWE-Gym*. arXiv; ICML 2025. [https://arxiv.org/abs/2412.21139](https://arxiv.org/abs/2412.21139) [↩](#user-content-fnref-swe-gym) [↩2](#user-content-fnref-swe-gym-2)
    
7.  Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). *SWE-smith: Scaling Data for Software Engineering Agents*. arXiv. [https://arxiv.org/abs/2504.21798](https://arxiv.org/abs/2504.21798) [↩](#user-content-fnref-swe-smith)
    
8.  Zeng, L., Li, Y., Xiao, Y., et al. (2025). *Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs*. arXiv. [https://arxiv.org/abs/2506.19290](https://arxiv.org/abs/2506.19290) [↩](#user-content-fnref-skywork-swe) [↩2](#user-content-fnref-skywork-swe-2)
    
9.  Agentica & Together AI (2025). *DeepSWE: Training a Fully Open-sourced, State-of-the-Art Coding Agent by Scaling RL*. [https://www.together.ai/blog/deepswe](https://www.together.ai/blog/deepswe) [↩](#user-content-fnref-deepswe)
    
10.  Golubev, A., Trofimova, M., Polezhaev, S., et al. (2025). *Training Long-Context, Multi-Turn Software Engineering Agents with Reinforcement Learning*. arXiv. [https://arxiv.org/abs/2508.03501](https://arxiv.org/abs/2508.03501) [↩](#user-content-fnref-nebius-rl)
     
11.  Zhang, B., Liu, Z., Cherry, C., & Firat, O. (2024). *When Scaling Meets LLM Finetuning: The Effect of Data, Model and Finetuning Method*. ICLR 2024. [https://arxiv.org/abs/2402.17193](https://arxiv.org/abs/2402.17193) [↩](#user-content-fnref-zhang-scaling)
     
12.  Chen, L., Li, S., Yan, J., et al. (2023). *AlpaGasus: Training a Better Alpaca with Fewer Data*. arXiv; ICLR 2024. [https://arxiv.org/abs/2307.08701](https://arxiv.org/abs/2307.08701) [↩](#user-content-fnref-alpagasus)
