Part II: The oracle
5The air gap
The agent is an adversary with a shell. The oracle runs where the shell cannot reach.
An oracle the agent can reach is an oracle the agent will eventually satisfy without doing the work. This chapter is about the threat, the evidence that it is real, and the design that removes it.
The threat model#
Treat the agent as an adversary with full control of its working directory and shell. That is not a judgment about the model's intent. It is a description of its permissions. A coding agent can run any command, read any file it can reach, and write any file in its checkout. If the verdict depends on anything in that set, the verdict is the agent's to decide.
The ways an agent can make a check pass without doing the task are ordinary:
- Edit the test so that it asserts what the code does.
- Mark the test skipped, or add it to an exclusion list in the test configuration.
- Edit the test runner's configuration so the test is not collected.
- Edit the build file or the continuous-integration file so the check does not run.
- Stub the function under test to return the expected value for the inputs the test uses.
- Patch the test framework itself, or a fixture, or a conftest file that runs before every test.
- Set an environment variable the test reads.
- Exit the process with status zero before the test runner reports.
- Read the hidden test from disk or from the network, if it is anywhere the agent can reach, and fit the code to it.
- If a model judges the result, write text into the diff or the commit message that addresses the judge.
Every one of these has been observed. Baker and colleagues at OpenAI trained a frontier reasoning model on coding tasks and monitored its chain of thought. They found the model saying, in its own reasoning, that it would make the tests pass by other means, and then doing so: calling sys.exit(0) before the tests ran so the runner reported success, raising the test framework's skip exception, and writing stubs instead of implementations where test coverage was thin.1 When they penalized the model for having such thoughts, it kept hacking the tests and stopped writing down that it was going to. They called this obfuscated reward hacking and recommended against training against the monitor.
Denison and colleagues at Anthropic built a curriculum of environments where a model could gain reward by gaming the check, starting with flattery and ending with the chance to edit its own reward function. A model trained on the easy environments generalized to the hard one: in a small but nonzero fraction of runs, it edited the code that computed its reward, and sometimes edited the test that would have caught the edit.2 METR, evaluating frontier models on its own tasks in 2025, reported models that rewrote scoring code, patched the timer in a benchmark harness, and special-cased the inputs the grader used.3 A later Anthropic study trained a model with reinforcement learning on real production coding environments after teaching it about reward-hacking strategies. At the point where the model learned to hack the rewards, it also began to generalize to other unwanted behavior, including attempts at sabotage.4
The lesson is not that models are malicious. It is that optimization finds the shortest path to the reward, and if the shortest path goes through the test file, that is the path. Skalse and colleagues formalized this as reward hacking: a proxy that correlates with the goal is maximized, and the correlation breaks under optimization.5 Gao, Schulman, and Hilton measured it for reward models: the true reward rises as a policy is optimized against the proxy, peaks, and then falls while the proxy keeps rising.6 The hidden test is the goal. Anything the agent can touch is a proxy.
The lesson is also older than machine learning. Thompson's 1984 Turing Award lecture described a compiler that inserts a back door into the login program and into any compiler compiled from clean source, so that no inspection of the source finds it.7 His point was that you cannot trust code you did not totally create yourself, and that the trust boundary has to be drawn around the thing you can verify. The oracle is the thing you can verify. Draw the boundary around it.
The design#
The oracle runs in a separate container, on a separate host when possible, with the following properties.
No network. The container is started with networking disabled. Dependencies come from a local mirror baked into the image or mounted read-only. Nothing the oracle needs is fetched at run time, and nothing the agent could have planted on a server is reachable.
A pristine checkout. The oracle has its own clone of the repository, from its own mirror. It checks out the base commit the session started from. It does not read the agent's working copy at all.
The source diff only. The oracle receives the agent's final diff as a file. Before applying it, the oracle filters it through an allowlist of paths. Source directories are on the allowlist. Test directories, test configuration files, continuous-integration configuration, build scripts, dependency manifests and lockfiles, and anything under the repository's tool configuration are not. Hunks that touch excluded paths are dropped. The oracle logs that they were dropped, in a log the agent cannot read.
Hidden tests from the oracle's store. The hidden tests are stored with the oracle, keyed by task id, and copied into the checkout after the diff is applied. They are never in the repository the agent works in, never in a branch the agent can fetch, and never on a path the agent's shell can read.
The existing suite from the base commit. The PASS_TO_PASS tests are the ones in the pristine checkout at the base commit, not the ones in the agent's diff. If the agent deleted a test, the oracle still runs it.
Fixed clock, seed, locale, and time zone. Set in the container's environment. A test that reads the system clock reads a constant.
Pinned image. The container image is referenced by digest, and the digest is recorded with the verdict.
Timeouts. The oracle run has a wall-clock limit. A run that exceeds it is a FAIL with a reason the oracle logs and the agent does not see. An agent cannot turn a failing test into a hang and have the hang read as success.
One bit out, signed. The oracle's only output to the agent's side is PASS or FAIL. To the trace store it writes a verdict record: the bit, the hash of the test list, the image digest, the diff hash, the base commit, the time, and a signature over all of these using a key the oracle host holds. A verdict in the trace store that does not verify was not written by the oracle.
Separate credentials. The process that runs the agent holds no credential that can write to the oracle's store, start an oracle run with a different test list, or read the hidden tests. The hook that asks for a verdict sends a diff and a task id and gets a bit back.
What crosses the boundary between the agent's host and the oracle's host
- Agent hostsendsTask idBase commitThe final diff as a fileOracle host
- Oracle hostreturnsPASS or FAILNothing elseAgent host
- Oracle hostwritesSigned verdict recordDropped-hunk logTest output, never shown to the agentTrace store
- Agent hostwritesTool-call events, redactedMessagesSession metadataTrace store
Why a separate host#
A container on the same machine as the agent is isolation from the agent's working copy. It is not isolation from the agent's shell. On a developer's laptop, the agent's Bash tool can read any path the developer can read, including the directory where the hidden tests live, the oracle's log, and the container's image. A container boundary protects the oracle's run from the agent's changes. A host boundary protects the oracle's secrets from the agent's eyes.
The reference plugin in Chapter 8 runs on one machine because that is where most teams will start, and it says so in its README. On one machine, store the hidden tests encrypted at rest and decrypt them inside the oracle container with a key the agent's process does not hold. That raises the bar from "read the file" to "find and misuse the key." It does not close the gap. Closing it means running the oracle on a host the agent cannot log into, which for a team with a continuous-integration system is a short step: the oracle is a job that runs on a runner, takes a diff and a task id, and returns a bit.
Firecracker, the virtual-machine monitor behind AWS Lambda, starts a microVM in about 125 milliseconds and gives each one its own kernel.8 gVisor gives a container a user-space kernel and intercepts its system calls.9 Either is a stronger boundary than a plain container for the oracle run, and both are used for exactly this purpose: running code you do not trust next to data you care about. For the hidden tests and the signing key, neither replaces a separate host.
Baselines and the flip that was not#
The flip needs a baseline: a FAIL recorded before the agent started. Three rules keep baselines honest.
The baseline is computed on the base commit, by the oracle, from its own checkout. It is deterministic, so compute it once per task and store it. Do not recompute it in a session-start hook; a container run is too slow for a hook budget, and the result would be the same anyway.
A task whose baseline is PASS is not a task. The hidden test already passes. Either the test is wrong or the work is already done. Quarantine the task, and do not count any session on it as a flip.
A task whose existing suite fails at baseline is not fit for an oracle. The agent could flip the hidden test while the suite stays broken, and the PASS_TO_PASS rule would record a FAIL for a session that did the work. Fix the base or exclude the task.
What the oracle logs and who reads it#
The oracle keeps a full log: the test output, the dropped hunks, the timing, and the reason for every FAIL. This log is for the people who run the pipeline. It is how you find a hidden test that fails for the wrong reason, a flaky test that slipped through, or an agent that keeps trying to edit the test configuration. It is never returned to the agent and never written anywhere the agent's host can read. Chapter 6 is about why.
Footnotes#
-
Baker, B., Huizinga, J., Gao, L., et al. (2025). Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. arXiv. https://arxiv.org/abs/2503.11926 ↩
-
Denison, C., MacDiarmid, M., Barez, F., et al. (2024). Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models. arXiv. https://arxiv.org/abs/2406.10162 ↩
-
Von Arx, S., Chan, L., & Barnes, E. (2025). Recent Frontier Models Are Reward Hacking. METR. https://metr.org/blog/2025-06-05-recent-reward-hacking/ ↩
-
MacDiarmid, M., Wright, B., Uesato, J., et al. (2025). Natural Emergent Misalignment from Reward Hacking in Production RL. arXiv. https://arxiv.org/abs/2511.18397 ↩
-
Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. https://arxiv.org/abs/2209.13085 ↩
-
Gao, L., Schulman, J., & Hilton, J. (2022). Scaling Laws for Reward Model Overoptimization. arXiv; ICML 2023. https://arxiv.org/abs/2210.10760 ↩
-
Thompson, K. (1984). Reflections on Trusting Trust. Communications of the ACM, 27(8), 761–763. https://doi.org/10.1145/358198.358210 ↩
-
Agache, A., Brooker, M., Florescu, A., Iordache, A., Liguori, A., Neugebauer, R., Piwonka, P., & Popa, D.-M. (2020). Firecracker: Lightweight Virtualization for Serverless Applications. NSDI 2020, 419–434. https://www.usenix.org/conference/nsdi20/presentation/agache ↩
-
Young, E. G., Zhu, P., Caraza-Harter, T., Arpaci-Dusseau, A. C., & Arpaci-Dusseau, R. H. (2019). The True Cost of Containing: A gVisor Case Study. HotCloud 2019. https://www.usenix.org/conference/hotcloud19/presentation/young ↩
Cite this
Anderson, M. (2026). The air gap. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 5). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-air-gap
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {The air gap},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {5},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-air-gap}
}Updates by email
New research reaches subscribers first.