Part V: Delivery
13The delivery pipeline
Weights as a release artifact: ingest, validate, train, evaluate, gate, canary in shadow, promote, roll back.
Chapter 13 of 22
Continuous delivery is the practice of keeping software in a state where any change can be released at any time, by automating the path from a commit to production and gating each step on checks.1 This chapter applies it to weights. The artifact is a model. The commit is a new batch of verified traces. The release is a new model serving the team's agents. Everything between them is a stage with an input, an output, and a gate.
The stages#
From trace store to serving model
- IngestSync traces and signed verdicts from agent hosts and oracle hosts
- Redact and scanSecond-pass secret scan; quarantine hits
- ValidateSchema check; verdict signatures; baseline sanity
- Select and versionFilters from Chapter 10; write the manifest
- TrainAdapter on all layers; replay mix; fixed seed
- EvaluateHeld-out flips, frozen and rolling; public benchmark subset
- GateBeat the serving model on held-out flips; no regression on the public set
- PackageMerge or ship the adapter; quantize; record the manifest hash
- CanaryShadow mode on live tasks; the oracle scores both models
- Promote or roll backRoute traffic; keep the previous weights warm
Repeat every batch
Ingest. A scheduled job pulls finished trace files from each agent host's plugin data directory and verdict records from the oracle host into the trace store. Files are content-addressed: the path includes the hash of the contents, so a re-upload is a no-op and a tampered file lands at a different path. The job is safe to run twice.
Redact and scan. The hook redacted with patterns at write time. This stage runs a dedicated secret scanner over every trace and quarantines any with a hit. Quarantined traces are not deleted; a person reviews them, because a false positive on a test fixture is common and a true positive is a credential to rotate.
Validate. Every event parses against the schema. Every verdict record's signature verifies against the oracle's public key. Every session has a session_start, a session_end, and at least one verdict, or it is marked incomplete and excluded. A session whose baseline was PASS is excluded and its task is flagged.
Select and version. The filters from Chapter 10 run. The output is a manifest: the list of sessions in the supervised set, the pairs in the preference set, the held-out task list, the template version, and the hash of each. The manifest is the thing that is versioned. The training data is derived from it.
Train. The training job takes a manifest and a recipe and produces weights. The recipe is a file: base model and its hash, adapter rank and target modules, learning rate, epochs, replay mix and its source, seed. The job records the recipe hash and the manifest hash in the weights' metadata. A training run with the same manifest, recipe, and seed should produce the same weights, within the limits of the hardware's determinism, and the pipeline should check that it does once.
Evaluate. The new weights run against the held-out tasks. For each task, the model attempts it in a fresh sandbox with the same harness the team uses, and the oracle grades the result. The metric is the flip rate on the frozen set and on the rolling set. The weights also run against a fixed subset of a public coding benchmark, for the forgetting check. Chapter 14 is about reading these.
Gate. The new weights are promoted only if the frozen-set flip rate is at least the serving model's, the rolling-set flip rate is higher, and the public-set score has not fallen by more than a set tolerance. A run that fails the gate is kept, with its evaluation, so the trend is visible even when nothing ships.
Package. The adapter is merged into the base weights or shipped as a separate adapter, depending on how the serving layer loads models. The weights are quantized if the serving hardware needs it, and the quantized weights are re-evaluated on the frozen set, because quantization can cost more on a fine-tuned model than on its base. The package carries the manifest hash and the recipe hash.
Canary. Before the new model serves anyone, it runs in shadow. For a sample of live oracle tasks, both the serving model and the candidate attempt the task in parallel sandboxes. The oracle grades both. The person sees only the serving model's result. After enough tasks, the candidate's live flip rate against the serving model's is the number that decides. This is the same design as shadow mode in driving systems, where a new model runs alongside the one in control and its decisions are compared without being acted on.
Promote or roll back. Promotion is a routing change. The previous weights stay loaded for a period, and rollback is the same routing change reversed. A rollback is not a failure of the pipeline; it is the pipeline working.
Cadence#
How often to run depends on the flip rate. A team producing 30 flips a day has 200 new examples a week, which is enough to retrain weekly and expect to see movement. A team producing 5 a day should batch monthly. A training run on an adapter for a 32-billion-parameter model over a few thousand trajectories takes hours on a single node with eight large accelerators; the evaluation, which runs an agent on each held-out task, often takes longer than the training. Budget for both.
The cadence should be a schedule, not a trigger. A pipeline that trains whenever enough new traces arrive produces models at irregular intervals that are hard to compare. A pipeline that trains every Sunday night and evaluates on Monday produces a weekly series.
What is deterministic and what is not#
The oracle is deterministic by construction. Training is deterministic up to the hardware. Evaluation is not: the model samples, and the harness is a live process. Fix the sampling temperature and seed for evaluation, run each held-out task more than once, and report the mean. Two models within noise of each other are tied, and the gate should say so instead of promoting on a coin flip.
The record#
Every stage writes a record to the same store the traces live in: what it took in, what it produced, the hashes, the time, and the outcome. A model in production can be traced back to the manifest that trained it, to the sessions in the manifest, to the oracle that graded each session, and to the hidden test behind each verdict. That chain is what makes a regression debuggable and what makes the pipeline auditable to anyone who asks where the model came from.
This is where the author's own work connects, and the connection is disclosed. Oxagen, the company the author founded, records agent runs with their tool calls, their outcomes, and their costs as a product. A run record is a trace in this book's sense. The pipeline in this chapter does not depend on it; the plugin writes its own traces to its own files. But the reason the record is kept beside the verdict, rather than inside the model's context, is the same reason Oxagen keeps the outcome outside the agent that produced it.
Footnotes#
-
Humble, J., & Farley, D. (2010). Continuous Delivery: Reliable Software Releases through Build, Test, and Deployment Automation. Addison-Wesley. ↩
Cite this
Anderson, M. (2026). The delivery pipeline. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 13). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-delivery-pipeline
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {The delivery pipeline},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {13},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-delivery-pipeline}
}Updates by email
New research reaches subscribers first.