Mac Anderson

Part III: Collecting traces

9The flip rate

Write the test first, dispatch small, sample more than once, manufacture tasks, and keep the failures.

Mac Anderson7 min read1,612 words
View markdown

The pipeline's throughput is the number of flips per day. This chapter is about raising it: choosing tasks that can flip, dispatching work in a shape that flips, sampling more than once, and manufacturing tasks when the natural supply runs low. It ends with what to do with the sessions that do not flip, which is most of them.

Start with the test#

A task can flip only if a hidden test fails on the base commit. The highest-leverage change a team can make is to write the test before dispatching the task. This is test-driven development with the roles split: a person or a separate agent writes the red test, and the solving agent is dispatched with the task description and never sees it.

The practice has a long history under the name red-green-refactor, and the evidence that it improves defect rates predates language models.1 What is new is the reason to do it. In a team with an oracle, every red test is a potential verified trajectory. In a team without one, a red test is just a test.

Some task shapes come with a red test for free.

  • A bug report with a reproduction. The reproduction is the hidden test. Clean it up, assert the correct behavior, confirm it fails on the base commit, and dispatch the bug.
  • A failing test in continuous integration. The test already exists and already fails. Hide it from the agent by giving the agent a checkout where the test is removed, and keep the test as the oracle.
  • A feature with an acceptance criterion. If the criterion can be stated as an assertion, it can be a hidden test.
  • A refactor. The existing suite is the oracle, with one added differential test: outputs on a set of inputs must match the previous build.

Task shapes that do not come with a red test, such as "improve the error messages" or "clean up this module," are not oracle tasks as stated. They can be made into oracle tasks by adding a check, such as a snapshot test of the messages or a lint rule the cleanup must satisfy, or they can be done without the oracle and their traces kept as unlabeled data.

Dispatch small#

A flip is binary. A task with three requirements and one hidden test that covers all three flips only when all three are done. Split it into three tasks with one hidden test each. The agent gets more feedback across the same work, the traces are shorter and cleaner, and a session that completes two of three is two flips instead of none.

Small also means a short session. Long sessions drift, run out of context, and end on a FAIL that is as much about the session's length as about the task. SWE-bench Verified's human annotators excluded tasks whose descriptions were underspecified or whose tests were unfair, and the benchmark's resolve rates roughly doubled for the same models once the unfair tasks were removed.2 The lesson for a team is that a clear, bounded task description raises the flip rate without changing the agent.

Sample more than once#

Repeated sampling is the single largest lever on the flip rate per task, and its cost is tokens.

Brown and colleagues measured how coverage, the share of problems solved by at least one of k attempts, grows with k. On SWE-bench Lite, an open model that solved 15.9 percent of problems with one attempt solved 56 percent with 250 attempts.3 The oracle is what makes this usable: with a deterministic verifier, the team keeps the attempt that passes and discards the rest. Without one, the team has 250 patches and no way to choose.

This is rejection sampling, and it is how most of the verified-trajectory datasets in the literature were built. SWE-Gym sampled many trajectories per task and kept the 491 that resolved their tasks.4 SWE-smith kept 5,016 out of thousands more.5 Llama 2's post-training used rejection sampling against a reward model as a core step.6 For a team, the practical version is: dispatch each oracle task to several independent sessions with fresh context, and let the gate find the flip.

The sessions must be independent. If one session's output leaks into another's context, the attempts are correlated and coverage grows slower. Fresh context per attempt and a silent oracle give independence. A verbose oracle that told each attempt why the last one failed would make the attempts a single long session in disguise.

Coverage grows with attempts when a verifier picks the winner

Attempts per taskDeepSeek-Coder-V2-Instruct on SWE-bench Lite
115.9%
25056%
Source: Brown et al. (2024). Only the two reported endpoints are plotted; the paper shows a roughly log-linear curve between them. Every added attempt costs tokens and, with a hidden oracle, each attempt is an independent draw.

A verifier trained on your own traces makes sampling cheaper. Pan and colleagues trained a verifier on the same SWE-Gym trajectories and used it to pick the best of 16 attempts, which raised their model from 20.6 to 32.0 percent.4 The verifier does not replace the oracle; the oracle still grades the chosen attempt. The verifier reduces how many attempts need to reach the oracle.

Manufacture tasks#

When the natural supply of red tests is smaller than the agent capacity, tasks can be made.

Break working code. Take a module with good tests. Introduce a bug: delete a branch, flip a comparison, drop a null check. The existing tests that now fail are the hidden tests. The task description is the symptom, written from the test's point of view without naming the test. This is how SWE-smith built 50,000 task instances from 128 repositories, and the authors found that the synthetic bugs trained a model that transferred to real issues.5 Mutation tools generate these bugs automatically, and the mutation literature has catalogued which mutants resemble real faults.7

Mine history. Every past commit that changed source and tests together is a candidate task: the tests it added are the hidden tests, the parent commit is the base, and the commit message or linked issue is the task description. This is how SWE-bench was built and how SWE-rebench automated the construction into a pipeline that produced more than 21,000 tasks.89 A team's own history is a supply of tasks on its own code that no public dataset contains.

Reverse a fix. For a merged bug fix with a regression test, revert the source change and keep the test. The agent is dispatched to fix the bug again. The trace is a worked example on real code with a real test.

Let a model propose tasks under an executor. Zhao and colleagues' Absolute Zero had a model propose coding tasks and solve them, with a code executor checking both that the task was well-formed and that the answer was right.10 The executor is the oracle. This is the most speculative item on the list and the one that produces the least realistic tasks; it is here because it shows that the supply of oracle tasks is not bounded by the supply of issues.

Use the strong model on the hard tail#

Some tasks will not flip with the model you are training. Dispatch them to a rented frontier model and keep the trace. A flipped trace from a stronger model is a distillation example, and distillation from a stronger model into a weaker one is the oldest trick in the post-training book.11 SWE-smith's 5,016 trajectories came from Claude 3.7 Sonnet, and the model trained on them was a 32-billion-parameter Qwen.5 A team already paying for the frontier model is already producing these traces. The oracle is what sorts them.

Keep the failures#

Most sessions will not flip. Those traces are not waste.

A session that failed on the same task where another session flipped is half of a preference pair. Direct preference optimization trains a model to prefer the chosen trajectory over the rejected one, and it needs exactly this data.12 A team that keeps only flips has a supervised set. A team that keeps everything has a supervised set and a preference set.

A session that failed is also a record of what went wrong, for people. If a task fails ten sessions in a row, the oracle log will say whether the hidden test is wrong, the task description is unclear, or the task is beyond the model. Each of those has a different fix, and the trace is how you find out which.

Hindsight relabeling, from robotics, is the formal version of this idea: an episode that failed its goal succeeded at whatever it did reach, and can be relabeled as a success for that.13 A session that did not flip the hidden test but did make the existing suite pass after breaking it, or did fix a different bug on the way, has a relabeled success in it. The pipeline in Chapter 10 does not do this automatically, and it is the kind of thing a team adds once the basic loop runs.

The arithmetic#

A team of 50 engineers runs an agent on perhaps 10 tasks a day each. Call it 500 sessions a day. If one in five sessions is on an oracle task, that is 100 oracle sessions a day. If one in three of those flips, that is about 33 flips a day. In a working month, around 700. Chapter 12 says what 700 buys. Doubling the oracle-task share doubles it; sampling three attempts per task roughly doubles it again. The levers are the share of work that has a red test, the size of each task, and the number of attempts. None of them is the model.

Footnotes#

  1. Nagappan, N., Maximilien, E. M., Bhat, T., & Williams, L. (2008). Realizing quality improvement through test driven development: results and experiences of four industrial teams. Empirical Software Engineering, 13(3), 289–302. https://doi.org/10.1007/s10664-008-9062-z ↩

  2. OpenAI (2024). Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/ ↩

  3. Brown, B., Juravsky, J., Ehrlich, R., Clark, R., Le, Q. V., Ré, C., & Mirhoseini, A. (2024). Large Language Monkeys: Scaling Inference Compute with Repeated Sampling. arXiv. https://arxiv.org/abs/2407.21787 ↩

  4. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩ ↩2

  5. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩ ↩2 ↩3

  6. Touvron, H., et al. (2023). Llama 2: Open Foundation and Fine-Tuned Chat Models. arXiv. https://arxiv.org/abs/2307.09288 ↩

  7. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). Are Mutants a Valid Substitute for Real Faults in Software Testing? FSE 2014, 654–665. https://doi.org/10.1145/2635868.2635929 ↩

  8. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv; ICLR 2024. https://arxiv.org/abs/2310.06770 ↩

  9. Badertdinov, I., Golubev, A., Nekrashevich, M., et al. (2025). SWE-rebench: An Automated Pipeline for Task Collection and Decontaminated Evaluation of Software Engineering Agents. arXiv; NeurIPS 2025. https://arxiv.org/abs/2505.20411 ↩

  10. Zhao, A., Wu, Y., Yue, Y., et al. (2025). Absolute Zero: Reinforced Self-play Reasoning with Zero Data. arXiv. https://arxiv.org/abs/2505.03335 ↩

  11. Hinton, G., Vinyals, O., & Dean, J. (2015). Distilling the Knowledge in a Neural Network. NeurIPS 2014 Deep Learning Workshop. https://arxiv.org/abs/1503.02531 ↩

  12. Rafailov, R., Sharma, A., Mitchell, E., Ermon, S., Manning, C. D., & Finn, C. (2023). Direct Preference Optimization: Your Language Model is Secretly a Reward Model. NeurIPS 2023. https://arxiv.org/abs/2305.18290 ↩

  13. Andrychowicz, M., Wolski, F., Ray, A., et al. (2017). Hindsight Experience Replay. NeurIPS 2017. https://arxiv.org/abs/1707.01495 ↩

Cite this

Anderson, M. (2026). The flip rate. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 9). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-flip-rate

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {The flip rate},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {9},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-flip-rate}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.