Mac Anderson

Part V: Delivery

14Evaluation

Held-out flips, contamination, weak tests, hacking scans, and the drift measures that catch what the flip rate misses.

Mac Anderson4 min read999 words
View markdown

A pipeline that trains weights needs a way to know whether the new weights are better. This chapter is about building that measurement so that it stays honest as the pipeline optimizes against it.

Held-out flips are the metric#

The primary metric is the flip rate on tasks the model never trained on, graded by the same oracle that grades production sessions. It is the metric that matches the goal. A public benchmark measures how well the model solves public tasks. Held-out flips measure how well it solves yours.

Two sets, as Chapter 10 said. The frozen set is drawn once, at the start, from tasks across the team's repositories, and never changes. Every model version is scored against it, so the series is comparable. The rolling set is the most recent few hundred oracle tasks, refreshed each cycle, so the score tracks the current work. A model that gains on the frozen set and not the rolling set has learned the past. A model that gains on the rolling set and not the frozen set has learned something narrow about recent work. The gate wants both.

Contamination#

A held-out task that the model saw during training is not held out. The leak is easy to create and impossible to detect after the fact.

Deng and colleagues showed that language models can reproduce the missing parts of benchmark items they were trained on, which is how contamination shows up in public evaluations.1 The SWE-bench+ study found that about a third of one agent's passing patches on SWE-bench had the solution available in the issue text or its comments, and another third passed because the tests were too weak to catch a wrong fix; filtering those out dropped the measured resolve rate from 12.47 percent to 3.97 percent.2 A team's internal evaluation is exposed to the same two failures: the hidden test can leak, and the hidden test can be weak.

Four rules keep the held-out set clean.

  1. Split by task, with every session on a task on the same side.
  2. Never train on a trace from a held-out task, including traces from the rented frontier model. The provenance field makes this checkable.
  3. Never show the held-out hidden tests to any agent, including the one generating training traces. The one-bit oracle enforces this for the agent; the pipeline has to enforce it for the people.
  4. Retire a held-out task when its hidden test is released into the repository, which Chapter 6 recommended once a task has been flipped in production. A test the agents can see is no longer held out.

Weak tests#

A hidden test that passes for a wrong fix inflates the flip rate and teaches the wrong fix. Mutation testing measures this directly: generate mutants of the code the task touches and check that the hidden test kills them.3 A held-out task whose hidden test kills few mutants near the change is a weak oracle, and its flips mean less. The evaluation report should carry the mutation score of each task's hidden test beside the flip rate, so a gain concentrated on weak tasks is visible.

The public benchmark#

A fixed subset of a public coding benchmark sits in the gate for one purpose: catching forgetting. The team's model should not get worse at general coding while it gets better at the team's code. SWE-bench Verified, or a multilingual benchmark if the team's code is not Python, serves.45 The score itself is secondary. The change in the score between versions is the signal.

Do not optimize for it. A pipeline that gates on a public benchmark and also trains on data derived from public repositories will, over time, find the benchmark's tasks in its training data. The public set is a thermometer, not a target.

Watching for hacking#

The oracle is air-gapped, so the agent cannot change the verdict. The agent can still produce a trace that flipped for a reason the team would not endorse, and a model trained on such traces learns the reason.

Baker and colleagues found that a monitor reading the agent's chain of thought caught most reward hacking, and that a weaker model was an adequate monitor.6 In this pipeline, the equivalent is a scan of each flipped trace for patterns that should not be there: edits to paths on the denylist, even though the oracle dropped them; commands that search for the hidden tests; tool calls that read the oracle's configuration; test runs that were made to pass by changing the test. The scan flags, a person reads, and flagged traces are excluded from training until cleared. The scan should run on traces from the rented model too, because distillation copies behavior.

Track the dropped-hunk rate from the oracle log. A rising share of diffs with hunks on the denylist means the agents are learning to touch the tests, and the next model will learn it faster.

Drift in what the model produces#

Three cheap measurements catch most regressions that the flip rate misses.

  • Length. Mean tokens per session and per tool call. Preference optimization and reward-driven training both tend toward verbosity unless controlled.7 A model whose sessions get longer without flipping more often is spending the team's tokens.
  • Tool mix. The distribution of tool calls per session. A model that stops running tests, or starts reading every file in the repository, has changed in a way the flip rate will show late.
  • Similarity. Pairwise similarity between the model's trajectories on different tasks. Rising similarity is the early sign of collapse from Chapter 11.

The report#

Each evaluation produces one report with: flip rate on the frozen set and the rolling set, each with its uncertainty from repeated runs; the public-benchmark score and its change; the mutation score distribution of the held-out tests; the counts of flagged traces by reason; the three drift measurements; and the manifest and recipe hashes. The gate reads the report. So do people. A report that a person cannot read in five minutes is too long.

Footnotes#

  1. Deng, C., Zhao, Y., Tang, X., Gerstein, M., & Cohan, A. (2024). Investigating Data Contamination in Modern Benchmarks for Large Language Models. NAACL 2024. https://arxiv.org/abs/2311.09783 ↩

  2. Aleithan, R., Xue, H., Mohajer, M. M., Nnorom, E., Uddin, G., & Wang, S. (2024). SWE-Bench+: Enhanced Coding Benchmark for LLMs. arXiv. https://arxiv.org/abs/2410.06992 ↩

  3. Just, R., Jalali, D., Inozemtseva, L., Ernst, M. D., Holmes, R., & Fraser, G. (2014). Are Mutants a Valid Substitute for Real Faults in Software Testing? FSE 2014, 654–665. https://doi.org/10.1145/2635868.2635929 ↩

  4. OpenAI (2024). Introducing SWE-bench Verified. https://openai.com/index/introducing-swe-bench-verified/ ↩

  5. Zan, D., Huang, Z., Liu, W., et al. (2025). Multi-SWE-bench: A Multilingual Benchmark for Issue Resolving. arXiv. https://arxiv.org/abs/2504.02605 ↩

  6. Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv. https://arxiv.org/abs/2503.11926 ↩

  7. Singhal, P., Goyal, T., Xu, J., & Durrett, G. (2023). A Long Way to Go: Investigating Length Correlations in RLHF. arXiv; COLM 2024. https://arxiv.org/abs/2310.03716 ↩

Cite this

Anderson, M. (2026). Evaluation. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 14). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/evaluation

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Evaluation},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {14},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/evaluation}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.