Part III: Collecting traces
8The reference plugin
A walkthrough of oracle-flip: five hooks, one constant reason, and an oracle in a container with no network.
Chapter 8 of 22
This chapter walks through oracle-flip, a Claude Code plugin that does what Chapters 2 through 7 describe. It is public at github.com/macanderson/oracle-flip under the MIT license. It is a reference, meant to be read and adapted, and it is small enough to read in one sitting.
Install#
/plugin marketplace add macanderson/oracle-flip
/plugin install oracle-flip@macanderson
To try it without installing, clone the repository and start Claude Code with the plugin loaded from disk:
git clone https://github.com/macanderson/oracle-flip
claude --plugin-dir ./oracle-flip
Layout#
oracle-flip/
.claude-plugin/
plugin.json name, version, description
marketplace.json lets /plugin marketplace add find it
hooks/
hooks.json the five hook registrations
scripts/
common.py paths, event writer, redaction, config
session_start.py SessionStart: write session_start
trace.py UserPromptSubmit and PostToolUse: write message and tool_call
gate.py Stop: diff, oracle, verdict, block or allow
session_end.py SessionEnd: write session_end with the outcome
oracle/
run.sh reference oracle runner: container, allowlist, hidden tests, one bit
grade.sh runs inside the container
filter_diff.py drops hunks outside the allowlist
Dockerfile pinned base image; the container runs with no network
example/ a sample task with a hidden test
tools/
export_sft.py traces plus verdicts to SFT and preference-pair JSONL
schema/
trace-event.schema.json
tests/
test_redaction.py, test_filter_diff.py, test_gate.py
The scripts are Python with no dependencies outside the standard library, so the plugin runs anywhere python3 runs. The oracle runner is shell plus one Python script, because it is meant to be replaced by whatever the team's continuous-integration system already does.
The hook registrations#
The plugin's hooks/hooks.json wraps the event map in a hooks key, which is the shape plugins use.1 Each hook is in exec form: a command and an args array, with the plugin root substituted by the harness, so paths with spaces need no quoting.
{
"description": "oracle-flip: trace every tool call and gate Stop on a hidden oracle",
"hooks": {
"SessionStart": [
{ "hooks": [ { "type": "command", "command": "python3",
"args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_start.py"], "timeout": 20 } ] }
],
"UserPromptSubmit": [
{ "hooks": [ { "type": "command", "command": "python3",
"args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 10 } ] }
],
"PostToolUse": [
{ "matcher": "*",
"hooks": [ { "type": "command", "command": "python3",
"args": ["${CLAUDE_PLUGIN_ROOT}/scripts/trace.py"], "timeout": 20 } ] }
],
"Stop": [
{ "hooks": [ { "type": "command", "command": "python3",
"args": ["${CLAUDE_PLUGIN_ROOT}/scripts/gate.py"], "timeout": 600 } ] }
],
"SessionEnd": [
{ "hooks": [ { "type": "command", "command": "python3",
"args": ["${CLAUDE_PLUGIN_ROOT}/scripts/session_end.py"], "timeout": 20 } ] }
]
}
}
Configuration#
Configuration is split by who controls the file. Anything the agent can edit holds nothing the gate trusts.
The project the agent works in holds one small file, .oracle-flip/config.json. It names the task and sets the attempt budget.
{ "task_id": "billing-proration-rounding", "base_commit": null, "max_attempts": 6 }
base_commit left null means the commit at session start. max_attempts is the number of FAIL verdicts after which the gate stops blocking and records no_flip. It exists so a session that will never flip does not run the oracle forever, and it is set below the harness's own cap of eight consecutive continuations so the plugin, not the harness, decides when to give up. The SessionStart hook reads this file once and copies it into the session's state. The gate reads the state, never the project, so an agent that edits the file mid-session changes nothing. An agent that edits it before the session starts can at most point the session at a task whose baseline passes, which the gate records as no_task, never as a flip.
The oracle command, the oracle store, and the container image live in the plugin's user settings, which Claude Code stores outside every repository and passes to hook processes as environment variables. The oracle command receives the mode, the task id, the base commit, the path of the diff file, and a hash of the repository's remote URL, and it must exit 0 for PASS and 1 for FAIL. Any other exit code is an oracle error, which the gate records and treats as FAIL for the purpose of blocking, with a different fixed reason.
The allowlist and denylist of paths the diff may touch live in the oracle store beside the hidden tests, one file each, and the oracle applies them before it touches a file. The gate's host does the filtering nowhere, because the gate's host is the agent's host. The hidden tests are in the same store, keyed by task id, and nowhere in the repository.
The gate#
The gate is the file to read if you read one. Its structure follows Chapter 7.
def decide(event, state, *, oracle=run_oracle, diff=write_diff):
write_event(assistant_message(event)) # last_assistant_message, redacted
if state.task_id is None:
return allow() # no oracle task; the plugin only traces
if state.outcome is not None:
return allow() # already decided; never loop on a settled outcome
if state.baseline is None:
baseline = oracle(state.oracle_command, mode="baseline", ...)
state.baseline = baseline.verdict
record_verdict(state, "baseline", baseline)
if baseline.verdict == "PASS":
state.outcome = "no_task" # the hidden test already passes
return allow()
if baseline.verdict == "ERROR":
state.outcome = "oracle_error"
return allow()
diff_path = diff(event["cwd"], state.base_commit, ...)
verdict = oracle(state.oracle_command, mode="grade", diff_path=diff_path, ...)
state.attempts += 1
record_verdict(state, "grade", verdict)
if verdict.verdict == "PASS":
state.outcome = "flipped"
return allow()
if state.attempts >= state.max_attempts:
state.outcome = "no_flip"
return allow()
return block(BLOCK_REASON_FAIL if verdict.verdict == "FAIL" else BLOCK_REASON_ERROR)
Four details are worth pointing at.
The gate reads everything from the session state the SessionStart hook wrote: the task id, the base commit, the attempt budget, and the oracle command. It reads nothing from the project.
The baseline is computed on the first Stop, not at session start, because the oracle is slow and the baseline is deterministic. It is stored, so later Stops in the same session reuse it. A team with many sessions per task should compute it once per task and share it; the plugin keeps it per session for simplicity.
The block reason is one of two constants. Neither includes the oracle's output. The oracle's output is written by the oracle, on the oracle's side, and the gate never sees it. The gate's own tests check that the reasons contain no test name and no assertion.
stop_hook_active is honored by the attempt counter rather than by an early return, because the harness sets it on every Stop after the first block, and returning early on it would mean the gate only ever runs once. The attempt counter plus max_attempts is what prevents the loop, and decide takes the oracle and the diff builder as parameters so the loop logic is tested without a container.
The oracle runner#
oracle/run.sh is the reference oracle. It expects a store with one directory of hidden tests per task and a mirror of the repository, and it runs the grade inside a container with networking disabled. The core of it:
task_dir="$store/$ORACLE_TASK_ID"
log_dir="$store/.log/$ORACLE_TASK_ID/$(date -u +%Y%m%dT%H%M%SZ)-$ORACLE_MODE"
args=(run --rm --network none
-e TZ=UTC -e LC_ALL=C.UTF-8 -e PYTHONHASHSEED=0 -e SOURCE_DATE_EPOCH=1700000000
-e ORACLE_MODE="$ORACLE_MODE" -e ORACLE_BASE_COMMIT="$ORACLE_BASE_COMMIT"
-v "$task_dir:/task:ro" -v "$mirror:/mirror:ro" -v "$log_dir:/log")
[ "$ORACLE_MODE" = "grade" ] && args+=(-v "$ORACLE_DIFF:/in/diff.patch:ro")
docker "${args[@]}" "$image" bash /oracle/grade.sh > "$log_dir/stdout.txt" 2> "$log_dir/stderr.txt"
status=$?
grep -E '^tests_hash=' "$log_dir/stdout.txt" || true
exit "$status"
grade.sh runs inside the container. It clones the repository from the read-only mirror, checks out the base commit, filters the diff through the task's allowlist and denylist with filter_diff.py, applies what is left, copies the hidden tests from /task/hidden into place, runs the existing suite and then the hidden tests, prints tests_hash= followed by a hash of the hidden test list, and exits 0 or 1. In baseline mode it runs without a diff and exits 1 when the suite passes and the hidden tests fail, which is the only baseline that makes a task. Its full output goes to the log directory on the oracle's side. Only the exit code and the tests_hash line return to the gate.
The Dockerfile pins its base image by digest and installs the repository's dependencies from a lockfile at build time, so the run-time container needs no network and gets none.
Limits of the plugin#
It runs the oracle on the same machine as the agent, in a container. Chapter 5 explained why that protects the oracle's run and not the oracle's secrets. On a laptop, the agent's shell can read ORACLE_STORE. The README says so. The step to a separate host is to replace oracle/run.sh with a script that submits the diff to a job runner the agent cannot log into and waits for the bit.
It does not sign verdicts. Signing needs a key on the oracle's side and a verifier in the pipeline, and the plugin has neither because it has no pipeline side. The verdict event has a field for the signature and the export tool has a flag that requires it.
It redacts with patterns. That is a floor. Run a real secret scanner on the trace store before training.
It has not been tested with a real container run on the author's machine at the time of writing, for a reason unrelated to the plugin, and the README says that too. The hook scripts are tested by piping sample events into them, and claude plugin validate passes.
The export tool#
tools/export_sft.py reads a directory of traces and their verdicts and writes two files.
The first is the supervised set: one JSON object per flipped session, in a chat format with tool calls, with a mask that marks which turns to train on. User prompts and tool outputs are masked out; the model is trained to produce the assistant's turns and tool calls, not to reproduce the environment's replies. Chapter 10 explains why.
The second is the preference set: for every task with at least one flipped session and at least one that did not flip, a pair with the flipped trace as the preferred one. Chapter 11 explains what to do with it.
Both files exclude any session whose verdict record is missing, and, with --require-signature, any whose verdict does not verify.
Footnotes#
-
Anthropic (2026). Plugin manifest reference. Claude Code documentation. https://code.claude.com/docs/en/plugins-reference ↩
Cite this
Anderson, M. (2026). The reference plugin. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 8). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-reference-plugin
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {The reference plugin},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {8},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-reference-plugin}
}Updates by email
New research reaches subscribers first.