Part II: The oracle
4Organizational oracles
The checks a team already has, ordered by how directly each one can label a trace.
Chapter 4 of 22
The hidden test is one oracle. An organization has many. This chapter is a catalogue, ordered from the oracles that give the cleanest signal to the ones that give the noisiest, with a note on how each can be used. The ordering matters because the cleaner the oracle, the more directly its verdict can be used as a training label. The noisier the oracle, the more it belongs in a preference pair or a filter rather than as a label.
Hard oracles#
A hard oracle is deterministic, runs in minutes, and returns a verdict a program can read. These can label a trace directly.
Hidden unit and integration tests. The reference oracle, covered in Chapter 3. Written by a person or by a separate model before the task is dispatched. The strongest version is a test that was written to expose a real bug or specify a real feature, because it encodes a real requirement.
The existing test suite as PASS_TO_PASS. Every task gets the whole existing suite as a regression check for free. A flip requires that nothing already passing breaks. On a repository with a large suite, this alone rules out most of the bad patches that a visible test would accept.
Type checkers and compilers. A change that does not compile, or that fails tsc --noEmit, mypy --strict, or cargo check, has failed. These are deterministic, fast, and hard to game without touching configuration, which the allowlist excludes. On their own they are a weak oracle, since code that compiles can still be wrong, but as a component of the oracle they remove a class of failures cheaply.
Linters with pinned rule sets. A lint failure is a weak signal of incorrectness and a strong signal of style drift. Use it as a gate, not a label: a trace that flips the hidden test but introduces lint errors is a flip with a defect, and you can decide per repository whether that counts.
Property-based tests. Instead of a fixed input and expected output, a property test states a rule, such as "decoding what you encoded gives the original," and a library generates many inputs to check it.1 With a fixed seed, a property test is deterministic. It is a stronger oracle than an example test because it checks many cases, and a harder one for an agent to fit to because the agent cannot see the cases.
Metamorphic tests. When the correct output is unknown but a relationship between outputs is known, a metamorphic test checks the relationship. If a search for "a" returns results, a search for "a OR b" should return at least as many. Metamorphic testing was proposed for programs without an oracle and is well suited to data pipelines, search, and numerical code.2
Differential tests against the previous build. Run the same inputs through the version before the change and the version after. For a refactor, the outputs must match. For a bug fix, they must differ only on the inputs that exposed the bug. McKeeman's differential testing of compilers is the origin of the method.3 The previous binary is an oracle you already have.
Mutation score. Chapter 3 covered mutation as a test-quality measure. It can also be an oracle for a task of the form "add tests for this module." The hidden check is whether the mutants the task names are killed after the change. This is one of the few ways to make test-writing itself a flippable task.
Reproducible build hash. For a task that must not change behavior, such as a dependency bump with no code change, the oracle can be that the build output is byte-identical to a reference, or differs only where expected.4
Schema and migration round trips. For a database change: apply the migration to a copy of the schema, run the down migration, and diff the result against the original. For a data pipeline: run the transform on a frozen fixture and compare row counts and checksums against a stored expectation.
Infrastructure plan idempotence. For infrastructure-as-code: after applying the change, a second plan must report nothing to do. This is a deterministic check of the form most infrastructure tools provide directly.
Formal checks. Where a module has a specification in a checkable form, a solver or a proof assistant is the strongest oracle there is. Few organizations have these. The ones that do should use them.
Soft and delayed oracles#
A soft oracle is a signal that correlates with correctness but is not deterministic, is not available for minutes or days, or depends on a person. These cannot label a trace as a flip. They can rank traces, build preference pairs, and filter.
Pull request merged. A merged change passed review and continuous integration. This is the oracle most code-model training has used, in the form of mined commits and pull requests. Meta's SWE-RL built its training corpus from about eleven million pull requests and used similarity to the merged patch as the reward.5 It is a real signal and a slow one, and reviewers miss things.
Reverted within N days. A change that was merged and then reverted was a failure the review did not catch. The pair "merged, then reverted" against "merged, kept" is a clean preference pair with a delay of days to weeks.
Incident linked to the change. Rarer and stronger than a revert. A change that caused an incident is a hard negative example.
Review comments. A change that received requests for changes before merge is weaker than one approved on the first pass. Review text also says what was wrong, which is useful for a data scientist and useless as a training label.
Acceptance of a suggestion. For completion models, whether the developer accepted the suggestion was found to be the best available predictor of perceived productivity in GitHub Copilot telemetry.6 For agents, the equivalent is whether the person kept the agent's change or discarded the session. It is a weak oracle and an abundant one.
Time to next edit of the same lines. If a person edits the lines an agent wrote within an hour of the session, the agent's work was probably incomplete. This is a proxy with many false positives and is best used to flag traces for a closer look.
A judge model. A second model that reads the diff and scores it. Chapter 6 argues that this is the weakest oracle in the catalogue and should not be used as a label. It can be used as a filter for obvious garbage, and even then its errors should be measured against a hard oracle on a sample.
Oracles by how directly their verdict can label a trace
- Judge modelA model scores the change. Not a label. At most a coarse filter, and even then measured against a hard oracle.
- Acceptance, next edit, review commentsHuman behavior signals. Rank and flag, do not label.
- Merged, reverted, incidentDelayed by days. Build preference pairs.
- Type check, lint, build hashFast and deterministic. Necessary, not sufficient. Use as gates inside the oracle.
- Existing suite as PASS_TO_PASSFree regression check on every task.
- Hidden tests, property and metamorphic testsDeterministic, fast, independent, unseen. Labels a flip.
Signal quality and speed
Oracles outside the code#
The method is not limited to software, but the hard oracles mostly are. For the sake of completeness, here is what the same construction looks like elsewhere, with the caveat that each of these is a soft oracle and should be treated as one.
- Data work. A transform's output on a frozen input has a checksum. Row counts and null rates have expectations. Tools that run assertions on data in a pipeline exist and can be hidden from the agent in the same way tests are.
- Documentation. A link checker, a prose checker with a pinned rule set, and a build that fails on a broken reference are deterministic. Whether the document is clear is not.
- Support. A ticket resolved and not reopened within two weeks is a delayed soft oracle on the agent's answer.
- Operations. A runbook step that leaves a system in a state a monitoring check accepts is close to a hard oracle, if the check is deterministic and the system is a copy.
A team should start with the hard oracles in its code, because that is where the construction is cleanest and the data is largest. Once the pipeline runs, the soft oracles can be added as preference signals.
Which oracle wrote the label#
Every trace should carry the identity of the oracle that graded it: the kind of oracle, the hash of its test list or rule set, the container image digest, and the time. A training set built from traces graded by different oracles at different times is a training set whose labels mean different things. When a later evaluation shows a regression, the first question is which oracle labeled the traces the model learned it from. Without the record, the question cannot be answered.
Footnotes#
-
Claessen, K., & Hughes, J. (2000). QuickCheck: A Lightweight Tool for Random Testing of Haskell Programs. ICFP 2000, 268–279. https://doi.org/10.1145/351240.351266 ↩
-
Segura, S., Fraser, G., Sanchez, A. B., & Ruiz-Cortés, A. (2016). A Survey on Metamorphic Testing. IEEE Transactions on Software Engineering, 42(9), 805–824. https://doi.org/10.1109/TSE.2016.2532875 ↩
-
McKeeman, W. M. (1998). Differential Testing for Software. Digital Technical Journal, 10(1), 100–107. ↩
-
Lamb, C., & Zacchiroli, S. (2022). Reproducible Builds: Increasing the Integrity of Software Supply Chains. IEEE Software, 39(2), 62–70. https://arxiv.org/abs/2104.06020 ↩
-
Wei, Y., Duchenne, O., Copet, J., et al. (2025). SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv; NeurIPS 2025. https://arxiv.org/abs/2502.18449 ↩
-
Ziegler, A., Kalliamvakou, E., Li, X. A., Rice, A., Rifkin, D., Simister, S., Sittampalam, G., & Aftandilian, E. (2024). Measuring GitHub Copilot's Impact on Productivity. Communications of the ACM, 67(3), 54–63. https://doi.org/10.1145/3633453 ↩
Cite this
Anderson, M. (2026). Organizational oracles. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 4). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/organizational-oracles
BibTeX
@incollection{anderson2026continuousdeliveryof,
author = {Anderson, Mac},
title = {Organizational oracles},
booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
chapter = {4},
year = {2026},
url = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/organizational-oracles}
}Updates by email
New research reaches subscribers first.