Part IV: Training
10From traces to training data
Select, deduplicate, format, mask, hold out, version.
Chapter 10 of 22
A trace store is not a training set. Between the two sit a series of transformations that decide what the model learns and what it does not. This chapter walks through them in the order the pipeline applies them.
Select#
The first decision is which sessions to include. The rule for the supervised set is simple: a session is included when its outcome is flipped and its verdict record verifies. Everything else is excluded from the supervised set, and the reason is recorded.
A session that flipped is not automatically a good example. Three further checks are cheap and worth running.
- The diff touched source. A session whose only changes were to paths the oracle dropped did not flip because of anything the agent wrote to the code under test. If the oracle still reported PASS, the baseline was wrong. Exclude the session and quarantine the task.
- The session did not exceed the length budget. A session with 400 tool calls that flipped is a worked example of flailing until something stuck. Set a budget per task class and exclude sessions over it, or truncate to the final successful stretch if the trace shows a clear restart.
- The agent did not attempt to touch excluded paths repeatedly. The oracle's dropped-hunk log shows an agent that kept editing test configuration. That session may have flipped on the merits, but its trace teaches the behavior you least want. Exclude it, and look at whether the task description invited it.
For the preference set, pair each flipped session on a task with each session on the same task that did not flip. Where there are many of each, sample pairs rather than taking the full cross product, so one task does not dominate.
Deduplicate#
Two sessions on the same task that flipped with nearly identical trajectories add little to each other. Hash the sequence of tool names and the final diff; where two sessions share both, keep one. Where they share the diff but not the path to it, keep both, because the paths are the data.
Across tasks, deduplicate on the task itself. A team that manufactured 200 tasks by mutating one module will have 200 near-identical trajectories, and a model trained on them will learn that module. Cap the number of flips per source file or per task family.
Format#
A trace is a sequence of events. A training example is a conversation in the chat format the base model was trained on, with tool calls in the model's native tool-call syntax. The conversion is mechanical but has choices in it.
The system prompt should be the one the agent ran with, including the rule files that were loaded, because that is the context the model will see at inference. If the team's rule files change often, consider training on a canonical version and keeping the diff as metadata.
The user turn is the task description.
Each assistant turn is either a tool call, with the tool name and arguments, or a message. Where the harness exposes the model's reasoning, include it as the base model's format expects; where it does not, the turn is the call alone.
Each tool result is a tool turn, with the redacted output.
The final assistant turn is the agent's completion message.
Convert to the exact chat template of the base model you will fine-tune. A mismatch between the training template and the inference template is the most common silent failure in fine-tuning, and it produces a model that looks trained and behaves untrained.
Mask#
The loss, the quantity the training run minimizes, should be computed only on the tokens the model is supposed to produce. In a trace, that is the assistant's turns: its reasoning, its tool calls, and its messages. The system prompt, the user's task, and every tool result are context, not targets.
Masking the tool results matters more than it might seem. Tool outputs are long: a file read returns hundreds of lines, a test run returns pages. Unmasked, they are most of the tokens in the example, and the model spends its capacity learning to predict file contents and test logs instead of learning to act. A model trained without the mask will also learn to hallucinate tool results, because it was trained to produce them. The reference export tool emits a mask field per turn for this reason.
Handle length#
Agent trajectories are long. A session that flipped after 60 tool calls, each returning a few kilobytes, is a hundred thousand tokens. Base models with long context windows handle this; training on sequences that long is expensive and some frameworks do not support it well.
Three strategies, in order of preference:
- Train on the full trajectory when the framework and hardware allow it. This preserves the behavior you want: the model learns to carry a plan across many steps.
- Truncate tool outputs, not turns. A file read can be cut to the lines around the ones the agent later edited, with a marker. A test log can be cut to the failures. The trajectory keeps its shape and loses bulk.
- Window the trajectory into overlapping segments, each with the system prompt and task prepended and a summary of the dropped prefix. This loses long-range structure and should be the fallback.
Do not drop the dead ends. A trajectory in which the agent tried something, saw it fail, and backed out is a trajectory that teaches recovery. SWE-Gym's authors kept full trajectories including unproductive steps and reported that it worked; cleaning them to the shortest path is a reasonable experiment but not the default.1
Hold out#
Before anything is trained, split by task, not by session. All sessions on a task go to the same side of the split. Set aside a fraction of tasks, with their hidden tests, as the evaluation set, and never train on any session from them. A model evaluated on tasks it saw during training will look better than it is, and the leak is undetectable afterward.
Keep two evaluation sets. A frozen set fixed at the start of the program, so that every model version is scored against the same tasks and the trend is comparable. A rolling set of recent tasks, so that the evaluation tracks the work the team is doing now. Chapter 14 covers how to read them.
Version#
Every training set is a manifest: the list of session ids included, the hash of each trace, the hash of each verdict, the filter rules applied, the template used, and the split. Store the manifest beside the weights it produced. When a model regresses, the manifest is how you find the traces that taught it.
A training set that is reproducible from its manifest is the data equivalent of a reproducible build. The same traces and the same rules give the same examples, and a change in either is visible as a change in the hash.
Record provenance#
Each example carries the identity of the oracle that graded it, as Chapter 4 said, and the identity of the model that produced the trace. The second matters more than it looks. A training set that is half traces from a rented frontier model and half from the team's own fine-tuned model is a mixture of distillation and self-improvement, and the two have different failure modes. Chapter 11 covers the self-improvement risk. The provenance field is what lets you tell them apart later.
Footnotes#
-
Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩
Cite this
Anderson, M. (2026). From traces to training data. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 10). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/from-traces-to-training-data
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {From traces to training data},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {10},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/from-traces-to-training-data}
}Updates by email
New research reaches subscribers first.