Mac Anderson

Part III: Collecting traces

7Harness hooks

The hooks a harness already exposes, turned into a trace collector and a Stop gate without building an agent.

Mac Anderson6 min read1,442 words
View markdown

You do not need to build a coding agent to collect traces from one. The harness your team already runs exposes hooks, and hooks see everything a trace needs. This chapter explains what a hook is, what each one sees, and how to turn a set of hooks into a trace collector and an oracle gate. The examples use Claude Code because its hook interface is documented in detail and the reference plugin targets it.1 Other harnesses have equivalents, and the last section covers them.

What a hook is#

A hook is a program the harness runs at a fixed point in a session. The harness passes the event to the program as JSON on standard input, waits for it to exit, and reads its exit code and anything it printed. A hook can observe the event, add context for the model, or, for some events, change what happens next.

The events that matter for tracing are the ones at the boundaries of a turn and a session.

EventWhen it runsWhat it carriesWhat it can decide
SessionStartA session begins or resumesSession id, working directory, model, how the session startedNothing. It can add context.
UserPromptSubmitThe person sends a promptThe prompt textIt can block the prompt.
PreToolUseBefore a tool runsTool name and inputIt can allow, deny, or ask.
PostToolUseAfter a tool runsTool name, input, output, durationIt can add context beside the result.
StopThe agent is about to end its turnThe agent's final message, whether a Stop hook already continued the turnIt can refuse to let the agent stop.
SessionEndThe session endsThe reason the session endedNothing. It runs cleanup.

Every event also carries the session id, the path to the harness's own transcript, the working directory, and the permission mode. The session id is the join key for everything the trace collector writes.

The trace collector#

A trace collector is three hooks.

On SessionStart, write a session_start event: the session id, the working directory, the base commit of the repository, the model, and the time. Do not run anything slow here. Hook budgets at session start are short, and a container run does not fit.

On UserPromptSubmit, write a message event with the role user and the prompt text. This is the task as the agent received it, which the training set needs as the first turn.

On PostToolUse, write a tool_call event: the tool name, the input, the output, and the duration. Redact the input and output before writing. Cap the size of each and store a hash of the full value beside the truncated one, so a trace that was cut can be identified later.

That is the whole collector. The agent's own messages between tool calls are not delivered by PostToolUse, so the collector adds a fourth hook.

On Stop, write a message event with the role assistant and the agent's final message, which the Stop event carries as last_assistant_message. Then run the oracle gate, below.

The collector does not read the harness's transcript file. The transcript is for the person, its format is the harness's to change, and the harness documentation notes that the file is not guaranteed to include the final message at the moment Stop fires.1 The hook events are the record.

The oracle gate#

The Stop hook is where the flip is detected and where an agent is kept from calling the work done.

When the agent decides it has finished, the harness fires Stop. The gate hook does the following.

  1. Reads the event. If the agent is already continuing because of a previous Stop hook, the event says so in a field named stop_hook_active. The gate uses this, with its own attempt counter, to decide whether to keep going.
  2. Builds the agent's diff against the base commit recorded at session start, including new files.
  3. Sends the diff and the task id to the oracle. The oracle returns one bit.
  4. Writes an oracle_verdict event with the attempt number, the verdict, and the hash of the test list the oracle reported.
  5. On PASS, exits normally. The agent stops. The session has flipped.
  6. On FAIL, returns a decision that blocks the stop, with a reason the harness shows to the agent. The reason is a fixed string. It says the oracle reported FAIL and the agent should keep working. It does not say why.

The blocking mechanism is specific. In Claude Code, a Stop hook refuses the stop by printing a JSON object with decision set to block and a reason, or by exiting with code 2 and the reason on standard error.1 The agent sees the reason as the explanation for why it must continue. With a one-bit oracle, the reason is the same every time.

What the harness limits#

Two limits in the harness shape how the gate behaves, and a team should know both before relying on it.

First, the harness caps consecutive continuations. In Claude Code, after Stop hooks have continued the turn eight times in a row, the harness overrides the next block and ends the turn. The count resets whenever the agent calls a tool, so an agent that keeps working does not hit it, and an agent that keeps saying "done" without doing anything does.1 The cap is configurable. The gate should treat a session that ends this way as no_flip, and the trace should record that the cap, not the oracle, ended it.

Second, hooks have timeouts. A Stop hook has minutes, not seconds, which is enough for an oracle that runs a test suite in a container. It is not enough for a suite that takes an hour. For slow oracles, the gate should submit the diff to the oracle host, return a provisional block with the fixed reason, and let the next Stop pick up the verdict. The reference plugin runs the oracle inline because most suites that are fit for an oracle run in minutes.

Redaction#

The tool output is where secrets live. A cat .env, a failing test that prints a connection string, a curl response with a token: all of it passes through PostToolUse. The trace collector redacts before it writes.

Redaction at this stage is pattern-based and should be treated as a floor, not a guarantee. Patterns for the common shapes, such as AKIA prefixed AWS keys, GitHub tokens, private key blocks, bearer headers, and KEY=value lines where the key name contains SECRET, TOKEN, or PASSWORD, catch most of what appears in practice. They do not catch a password that looks like a word. Meli, McNiece, and Reaves found secrets leaking into public repositories at a rate of thousands of new unique secrets a day, across more than a hundred thousand repositories, which is a measure of how often they appear in code and output that people thought was fine.2 The collector redacts, and the pipeline runs a second pass with a dedicated scanner before anything reaches training, and the model is kept private to the organization whose traces it learned from. Chapter 16 covers the memorization research that makes the last rule a hard one.

Storage#

Traces are written to a directory the plugin owns, not to the repository and not to the plugin's install directory. In Claude Code, that is the plugin's data directory, which the harness exposes as CLAUDE_PLUGIN_DATA and which survives plugin updates.3 A sync job moves finished traces to an object store under a path keyed by organization, repository, and date. The store is append-only. Nothing in the pipeline edits a trace after it is written; corrections are new records that reference the old.

Each verdict record is signed by the oracle, as Chapter 5 describes. A trace whose verdict does not verify is excluded from training and flagged. This is the check that stops a compromised or misconfigured agent host from writing its own PASS.

The same idea in other harnesses#

The hook design is not specific to one product. OpenAI's Codex CLI, Cursor's agent, OpenCode, and others expose lifecycle events with the same shape: before and after tool calls, and at the end of a turn. The names differ. The stdin JSON differs. The decision mechanism for refusing a stop differs, and in some harnesses does not exist, in which case the gate has to run as a wrapper around the harness instead of inside it.

What stays constant is the architecture. One hook writes events. One hook, at the end of a turn, asks the oracle and either lets the agent stop or sends it back. The trace schema is the harness-neutral part, and a team that runs more than one harness should normalize events into the same schema at write time, so the training set does not depend on which tool produced it.

Footnotes#

  1. Anthropic (2026). Hooks reference. Claude Code documentation. https://code.claude.com/docs/en/hooks ↩ ↩2 ↩3 ↩4

  2. Meli, M., McNiece, M. R., & Reaves, B. (2019). How Bad Can It Git? Characterizing Secret Leakage in Public GitHub Repositories. NDSS 2019. https://doi.org/10.14722/ndss.2019.23418 ↩

  3. Anthropic (2026). Plugin manifest reference. Claude Code documentation. https://code.claude.com/docs/en/plugins-reference ↩

Cite this

Anderson, M. (2026). Harness hooks. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 7). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/harness-hooks

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Harness hooks},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {7},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/harness-hooks}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.