Mac Anderson

Part II: The oracle

3The oracle problem

A verdict the model does not control: deterministic, independent, and hidden, with the fail-to-pass flip stated as a rule.

Mac Anderson6 min read1,477 words
View markdown

Software testing has a name for the thing that decides whether a program's output is correct: the test oracle. Barr, Harman, McMinn, Shahbaz, and Yoo surveyed the field in 2015 and called the difficulty of building one "the oracle problem."1 Generating inputs to a program is easy. Knowing what the program should have done with them is hard. Every automated check, from a unit test to a type checker, is a partial answer to that problem.

This book uses the word in the same sense, narrowed. An oracle here is a program that takes the state of a repository after an agent has worked on it and returns PASS or FAIL. It runs without the agent. It does not take the agent's word for anything. The quality of everything downstream, every training set and every evaluation, is bounded by the quality of the oracle, so this chapter is about what makes a good one.

Deterministic#

A deterministic oracle returns the same verdict every time it runs on the same input. That sounds like a low bar. In practice, most test suites do not clear it.

Luo, Hariri, Eloussi, and Marinov studied flaky tests, tests that pass and fail on the same code, across 51 open-source projects and classified the causes in 201 fixing commits. The leading causes were waiting on asynchronous work, concurrency, and dependence on test order.2 At Google, Micco reported that about 1.5 percent of all test runs gave a flaky result, that almost 16 percent of tests showed some flakiness, and that about 84 percent of transitions from pass to fail involved a flaky test.3 A flaky test in a benchmark is an annoyance. A flaky test in an oracle that labels training data is corruption: a session is recorded as a flip when the agent did nothing, or as a failure when the agent succeeded.

The fixes are the fixes the testing literature has recommended for a decade, applied strictly because the stakes are higher.

  • Run the oracle in a fresh container from a pinned image, so the environment is identical every time. Reproducible builds give the same guarantee for the artifact under test: the same source produces the same binary, bit for bit, so a verdict is about the code and not the build machine.4
  • Pin every dependency by hash, with a local mirror, and give the container no network. A test that downloads anything is not deterministic.
  • Fix the clock, the random seed, the locale, and the time zone inside the container. A test that depends on the wall clock is a test that depends on when the oracle ran.
  • Run the hidden tests twice on the base commit before the agent starts. If the two verdicts disagree, the task is not fit for an oracle. Quarantine it.
  • Record the test list as a hash with the verdict, so a later change to the tests cannot be confused with a change in the code.

The cost of determinism is that some tests cannot be oracles. Integration tests against a live service, tests that depend on timing, and tests that depend on data that changes are excluded. That is a loss, and Chapter 4 is partly about what you can use instead.

Independent#

An oracle must be independent of the agent in two senses.

It must not run in the agent's environment. The agent's working copy is under the agent's control. The agent can edit any file there, including the tests, the test configuration, the build file, and the environment variables the test runner reads. An oracle that runs pytest in the agent's directory is asking the agent whether it passed. Chapter 5 describes the air gap that fixes this.

It must not be a model grading itself. A judge model that reads the diff and decides whether the task is done is not an oracle. It is a second opinion from a system with the same blind spots as the first. Chapter 6 reviews the evidence.

Hidden#

The hidden test is the version of the oracle this book recommends. The team writes a test that fails on the current code and passes when the task is done correctly. The agent never sees it. The agent sees the task description and the existing test suite, and can write and run its own tests freely. When the agent says it is done, the oracle applies the agent's changes to a clean copy, runs the hidden test and the existing suite, and records the verdict.

This is the SWE-bench construction.5 The FAIL_TO_PASS tests are hidden from the agent during the attempt. The agent reads the issue, not the test. The construction is also how the Defects4J database of real Java bugs has been used for a decade: each bug comes with at least one test that exposes it and passes after the developer's fix.6

Hidden tests matter because visible tests are a weaker oracle than they look. Qi, Long, Achour, and Rinard examined patches that three automated repair systems had reported as fixing bugs, meaning the patches made the visible tests pass. Most of the patches were wrong.7 They passed the tests by deleting the functionality the tests happened not to cover. Smith, Barr, Le Goues, and Brun showed the same thing with a controlled experiment and gave it a name: patch overfitting.8 A patch that passes the tests the agent can see has been fitted to those tests, and nothing more has been shown. A patch that passes a test the agent could not see has been shown something.

The flip, formally#

With the pieces named, the flip can be stated as a rule.

Let H be the hidden tests for the task and S be the existing suite, both frozen as a list with a hash. Let base be the commit the agent started from and diff be the agent's final change.

  1. On base, run H and S in a fresh container. Require that H fails and S passes. If H passes, there is no task; if S fails, the base is broken and the task is not fit for an oracle. Record the verdict as the baseline.
  2. On base with diff applied to source paths only, run H and S in a fresh container. Record PASS if every test in H passes and every test in S passes. Record FAIL otherwise.
  3. A session has flipped when the baseline is FAIL and the final verdict is PASS.

Step 2 says "source paths only." The diff the agent produced may include changes to test files, test configuration, continuous-integration files, and build scripts. The oracle does not apply those. It applies the agent's changes to the code under test, then runs its own copies of the tests against them. Chapter 5 explains the allowlist that does this.

One oracle run

  1. Fresh containerPinned image, no network, fixed clock and seed
  2. Clean checkoutThe base commit, from the oracle's own mirror
  3. Apply source diffOnly paths on the allowlist
  4. Inject hidden testsFrom the oracle's store, never from the workspace
  5. Run H and SHidden tests and the existing suite
  6. One bit outPASS or FAIL, plus a hash of the test list
Everything the agent could have touched is replaced before the tests run. The agent's only contribution to the oracle run is the source diff.

Tests are a weak oracle too#

Hidden tests are the best oracle most teams can build quickly. They are not a perfect one. Inozemtseva and Holmes showed that code coverage, the share of lines a test suite runs, is not strongly correlated with how many faults the suite detects once suite size is accounted for.9 A hidden test that runs a line is not a hidden test that checks it. The same study of patch overfitting that argues for hidden tests also argues for better ones.

Mutation testing is the standard way to measure how good a test is. A mutation tool makes small changes to the code, such as flipping a comparison or deleting a statement, and checks whether the tests fail. A test that does not fail on a mutant is not checking the mutated behavior. The idea is from 1978 and has been used at Google at scale since at least 2018, where mutants are shown to developers during code review.1011 A hidden test that kills the mutants near the code the task touches is a stronger oracle than one that merely runs. Chapter 9 uses mutation the other way around, to manufacture tasks.

The practical rule is: a hidden test must fail on the base commit for the right reason. Before you accept a task into the oracle pool, read the failure. If the test fails because of an import error or a fixture that is missing, the agent can flip it by fixing the fixture. If it fails because the feature is missing, the agent has to build the feature.

Footnotes#

  1. Barr, E. T., Harman, M., McMinn, P., Shahbaz, M., & Yoo, S. (2015). The Oracle Problem in Software Testing: A Survey. IEEE Transactions on Software Engineering, 41(5), 507–525. https://doi.org/10.1109/TSE.2014.2372785 ↩

  2. Luo, Q., Hariri, F., Eloussi, L., & Marinov, D. (2014). An Empirical Analysis of Flaky Tests. FSE 2014, 643–653. https://doi.org/10.1145/2635868.2635920 ↩

  3. Micco, J. (2016). Flaky Tests at Google and How We Mitigate Them. Google Testing Blog. https://testing.googleblog.com/2016/05/flaky-tests-at-google-and-how-we.html ↩

  4. Lamb, C., & Zacchiroli, S. (2022). Reproducible Builds: Increasing the Integrity of Software Supply Chains. IEEE Software, 39(2), 62–70. https://arxiv.org/abs/2104.06020 ↩

  5. Jimenez, C. E., Yang, J., Wettig, A., Yao, S., Pei, K., Press, O., & Narasimhan, K. (2023). SWE-bench: Can Language Models Resolve Real-World GitHub Issues? arXiv; ICLR 2024. https://arxiv.org/abs/2310.06770 ↩

  6. Just, R., Jalali, D., & Ernst, M. D. (2014). Defects4J: A Database of Existing Faults to Enable Controlled Testing Studies for Java Programs. ISSTA 2014, 437–440. https://doi.org/10.1145/2610384.2628055 ↩

  7. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems. ISSTA 2015, 24–36. https://doi.org/10.1145/2771783.2771791 ↩

  8. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair. ESEC/FSE 2015, 532–543. https://doi.org/10.1145/2786805.2786825 ↩

  9. Inozemtseva, L., & Holmes, R. (2014). Coverage Is Not Strongly Correlated with Test Suite Effectiveness. ICSE 2014, 435–445. https://doi.org/10.1145/2568225.2568271 ↩

  10. DeMillo, R. A., Lipton, R. J., & Sayward, F. G. (1978). Hints on Test Data Selection: Help for the Practicing Programmer. IEEE Computer, 11(4), 34–41. https://doi.org/10.1109/C-M.1978.218136 ↩

  11. Petrović, G., & Ivanković, M. (2018). State of Mutation Testing at Google. ICSE-SEIP 2018, 163–171. https://doi.org/10.1145/3183519.3183521 ↩

Cite this

Anderson, M. (2026). The oracle problem. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 3). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-oracle-problem

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {The oracle problem},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {3},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-oracle-problem}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.