Part I: The argument
2The trace
The record of one session, why the tool calls carry the signal, and the flip as the unit of value.
Chapter 2 of 22
A trace is the record of one agent session. The word is borrowed from systems tracing, where a trace is the record of one request as it passes through many services. The analogy is close. An agent session passes through many tool calls, and the trace records all of them in order with their inputs and outputs.
Anatomy of a trace#
A coding-agent session has a regular structure. The agent receives a task. It takes a turn: it thinks, then it either calls a tool or writes a message. A tool call runs and returns a result. The agent takes another turn. This continues until the agent decides it is done, or a person stops it, or a limit is reached.
The trace records each of these events:
- The task as the agent received it, including the system prompt, the rule files loaded into context, and the user's instruction.
- Each tool call: the tool name, the arguments the agent passed, the result the tool returned, and how long it took.
- Each message the agent wrote to the person, including the final message where it says the work is done.
- The repository state at the start and end: the base commit, the final diff, and the list of files touched.
- The oracle verdicts: the verdict before the agent started, and the verdict at each point the agent tried to stop.
- Metadata: the model, the harness and its version, the session id, timestamps, token counts, and cost.
Of these, the tool calls carry the learning signal. They show the model how an expert in this codebase finds the relevant file, which command verifies a change, what a failing test looks like here, and how a fix is shaped. Training on them teaches the model to take those actions on this codebase. The final diff alone would not teach that. A diff is an answer; a trace is a worked solution.
The records in a trace and how they relate
- Sessionhas oneSystem prompt, rule files, user instructionBase commit of the repositoryTask
- Sessionhas manyOne model responseEnds in a tool call or a messageTurn
- Turnhas zero or oneName, input, output, durationOutput redacted before storageTool call
- Sessionhas manyPASS or FAIL and nothing elseOne at the start, one per stop attemptOracle verdict
- Sessionhas oneflipped, no flip, or abortedFinal diff and files touchedOutcome
The episode and the flip#
In reinforcement learning, an episode is one complete attempt at a task from start to a terminal state. A trace is an episode. What makes an episode useful for training is a reward: a number that says how well it went. For most of what organizations do, there is no clean reward. For software, there is one, and it has a name in the benchmark literature.
SWE-bench, the benchmark most coding-agent papers report against, grades a candidate patch with two sets of tests.1 The FAIL_TO_PASS tests fail on the repository before the fix and pass after the reference fix. The PASS_TO_PASS tests pass both before and after; they are there to catch a patch that fixes the issue by breaking something else. A patch resolves the task when every FAIL_TO_PASS test passes and every PASS_TO_PASS test still passes.
That is the flip. Before the agent starts, the oracle runs the hidden tests and records FAIL. After the agent says it is done, the oracle applies the agent's diff to a clean copy, runs the same hidden tests, and records PASS or FAIL. A session whose verdict goes from FAIL to PASS, with no regression in the tests that already passed, has flipped. The trace of that session is a verified trajectory. A session that ends on FAIL is also kept, because a failed attempt beside a successful one on the same task is a preference pair, and preference pairs are training data too. Chapter 10 covers that.
The flip is the unit of value for three reasons.
It is binary, so it cannot be argued with. There is no rubric, no score from a judge model, no partial credit for a patch that "looks right." Chapter 6 is about why that matters.
It is cheap, so it can be run on every session. A test suite that runs in minutes grades a session for cents. A human review of the same session would cost more than the session.
It is yours. The hidden test encodes what your organization means by correct for this task. A public benchmark encodes what a benchmark author meant.
Other records#
A trace is not a transcript. Harnesses keep transcripts for the person to scroll back through, and those files are useful, but they mix the agent's messages with display formatting, and they may not include the final message when the session ends.2 A trace is written by hooks as events happen, in a schema you control, with redaction applied before anything touches disk.
A trace is not a log. Logs are for debugging a system. Traces are for training a model. A log can be lossy and unstructured. A trace that drops the tool output for one call has a hole in the worked example at exactly the point the model needs to learn from.
A trace is not the diff. The diff is the answer. A model trained only on diffs learns to produce patches that look like your patches. It does not learn to find the file, run the test, read the failure, and try again. Pan and colleagues trained on full agent trajectories, each a sequence of tool calls and observations, and moved their model from 7.0 to 20.6 percent with 491 of them.3 The actions are the data.
The schema#
Appendix A gives a full schema. The shape is a JSON Lines file per session, one event per line, with a small set of event types.
session_start session_id, task, base_commit, model, harness, started_at
tool_call turn, tool_name, tool_input, tool_output_redacted, duration_ms
message turn, role, text
oracle_verdict attempt, verdict (PASS|FAIL), tests_hash, ran_at
session_end outcome (flipped|no_flip|aborted), final_diff_hash, ended_at
Two design choices in the schema deserve a sentence each. The verdict record carries a hash of the hidden test list, not the list. That lets you prove later which oracle graded a trace without putting the test names anywhere the agent's process can read. And the tool output is stored after redaction, never before. A trace store is a secret store if you let it be one, and Chapter 16 explains why that is the failure that ends programs like this.
Why the harness should write it#
You could write a coding agent and have it write its own traces. Several groups have, and their agents are good. But the harness your team already uses runs thousands of sessions a week today, and it exposes hooks. A hook is a program the harness runs at a fixed point in the session: before a tool call, after a tool call, when the agent tries to stop, when the session starts and ends. Each hook receives the event as JSON on standard input. A hook that appends that JSON to a file is a trace collector, and it took the author of this book an afternoon to write. Chapter 7 covers hooks in detail. The reason to mention them here is that the most common objection to collecting traces, "we would have to build our own agent," is false.
Footnotes#
-
Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv; ICLR 2024. https://arxiv.org/abs/2310.06770 ↩
-
Anthropic (2026). Hooks reference. Claude Code documentation. https://code.claude.com/docs/en/hooks ↩
-
Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩
Cite this
Anderson, M. (2026). The trace. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 2). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-trace
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {The trace},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {2},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-trace}
}Updates by email
New research reaches subscribers first.