Mac Anderson

Part V: Delivery

16Governance and failure modes

Secrets, memorization, oracle tampering, pipeline debt, a starving flip signal, and people.

Mac Anderson4 min read960 words
View markdown

A pipeline that collects everything an agent did and trains a model on it has new ways to go wrong. This chapter lists them, with what the research says and what the pipeline does about each. The failures are ordered by how much damage they do, not by how likely they are.

Secrets in the trace store#

The trace store holds the output of every command the agents ran. If a secret passed through a terminal, it is in a trace unless redaction caught it. A trace store is the most complete record of a team's operational secrets that has ever existed in one place, and it should be protected like one.

The defenses are layered. Redaction at write time, with patterns. A dedicated scanner at ingest, with quarantine. Access control on the store, with the training job as the only automated reader. Encryption at rest. And a short retention period for raw tool output, after which only the selected and scanned training examples remain. The last one is a tradeoff: it forecloses re-processing old traces with better redaction. A team should decide its retention consciously.

Memorization#

A model trained on traces can reproduce them. Carlini and colleagues extracted training examples from a deployed language model by prompting it, including names, contact details, and code.1 In a follow-up, they measured how memorization scales: it grows with model size, with how many times an example was duplicated in training, and with how much of the example's prefix the prompt supplies.2 A fine-tune on a few thousand trajectories, each seen for several epochs, is in the regime where memorization is expected.

Two consequences. First, a secret that survived redaction into training can be extracted from the model. The secret scanning in the pipeline is the control, and it has to be good. Second, the model is a copy of the team's code in a form that can be queried. It is a private asset and should be served privately. A team that fine-tunes on its traces and then exposes the model to people outside the team has published its code in a lossy format. Chapter 1 said the traces were the asset nobody else has. That is only true while the model trained on them stays inside.

Tampering with the oracle#

Chapter 5 built the air gap. The governance question is who can change what is inside it. The hidden tests, the allowlist, the container image, and the signing key are the pipeline's root of trust. Changes to them should go through the same review as changes to production code, and the oracle log should record every verdict with the hash of the test list that produced it, so a change in the tests is visible as a change in the hash. Rotate the hidden tests for a task once it has flipped in production, as Chapter 6 said, and retire the task from the held-out set.

Training on the wrong lesson#

The oracle grades the result. It does not grade the path. A session that flipped by reading the hidden test's name from a stray log line, by copying a fix from a sibling repository that happened to be checked out, or by special-casing the inputs it guessed the test used, is a flip with a bad path. The scan in Chapter 14 catches some of these. The rest are caught by reading flipped traces, which a person should do for a sample every cycle. A pipeline nobody reads is a pipeline that trains on whatever got through.

The pipeline itself#

Sculley and colleagues catalogued the ways machine-learning systems accumulate debt that ordinary software does not: data dependencies that nobody tracks, feedback loops where the model's output changes its own training data, configuration that grows without review, and pipelines glued together from pieces nobody owns.3 This pipeline has every one of those. Its training data comes from agents running its own previous model, which is a feedback loop by design. Its configuration is a recipe file, an allowlist, a template, and a set of hidden tests, each of which changes the model when it changes. Its stages are a trace collector, an oracle, a trainer, an evaluator, and a router, built from different tools.

The defenses are the ordinary ones, applied without exception. Every input to a stage is versioned. Every stage records what it did. Every change to configuration is reviewed. Every model can be traced to its manifest. And the frozen evaluation set, which never changes, is the one fixed point against which drift in everything else is measured.

Loss of the flip signal#

If the oracle share falls, because the team stops writing tests first, the flip rate falls with it and the pipeline starves. If the task supply narrows, because one repository produces most of the oracle tasks, the model narrows with it. If the fallback rate stops falling, the team's model has plateaued and the recipe needs to change. Each of these is visible in the pipeline's own records, and the weekly report should carry them: oracle share, flip rate, tasks per repository, fallback rate.

People#

The last failure mode is the one the research does not cover. A pipeline that grades every agent session on a hidden test is also, if someone chooses to read it that way, a pipeline that grades the engineers who dispatched the sessions. It should not be used that way. The flip rate is a property of the task, the oracle, and the model. The moment it becomes a measure of a person, people will stop dispatching tasks that might not flip, the oracle share will fall, and the pipeline will starve. Say this in writing when the pipeline is introduced, and keep the per-person numbers out of the report.

Footnotes#

  1. Carlini, N., Tramèr, F., Wallace, E., et al. (2021). Extracting Training Data from Large Language Models. USENIX Security 2021. https://arxiv.org/abs/2012.07805 ↩

  2. Carlini, N., Ippolito, D., Jagielski, M., Lee, K., Tramèr, F., & Zhang, C. (2023). Quantifying Memorization Across Neural Language Models. ICLR 2023. https://arxiv.org/abs/2202.07646 ↩

  3. Sculley, D., Holt, G., Golovin, D., et al. (2015). Hidden Technical Debt in Machine Learning Systems. NeurIPS 2015. ↩

Cite this

Anderson, M. (2026). Governance and failure modes. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 16). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/governance-and-failure-modes

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Governance and failure modes},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {16},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/governance-and-failure-modes}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.