Mac Anderson

Part II: The oracle

6The one-bit verdict

Why PASS or FAIL is all the oracle says, what that costs inside a session, and the research on models grading themselves.

Mac Anderson7 min read1,709 words
View markdown

The oracle says PASS or FAIL. It does not say which test failed, what the assertion was, or what the output looked like. This is the design choice readers push back on most, so this chapter takes it slowly: first the cost, then the three reasons, then the research on self-grading that the third reason depends on.

The cost, stated plainly#

Richer feedback helps an agent fix a bug in the current session. Chen, Lin, Schärli, and Zhou compared three kinds of feedback for a model debugging its own code: simple feedback, meaning only whether the code passed; unit-test feedback, meaning the execution results; and a self-written explanation of the code. Unit-test feedback produced the largest gains.1 A verdict with no explanation is the "simple feedback" condition in that study, and it is the weakest of the three for raising the pass rate inside one session.

So the one-bit verdict lowers the flip rate per session. A team that adopts it will see more sessions end on FAIL than a team that shows the agent the failing test. That is a real cost. Three things make it worth paying.

Reason one: the hidden test stays valid#

A test the agent can see is a test the agent can fit to. Chapter 3 covered the patch-overfitting results: patches that pass visible tests by deleting what the tests do not check.23 An oracle that returns the name of the failing test and the assertion that failed has shown the agent the test. The agent will fix the assertion. Whether it fixed the feature is now unknown, which is where you started.

The theory for this is from statistics, not software. Blum and Hardt studied machine-learning competitions where participants submit many models and see a score on a holdout set each time. With enough submissions, a participant can overfit the holdout without ever seeing its labels, by hill-climbing on the score. Their fix, the Ladder, releases a new score only when it improves on the previous best by more than a threshold, and otherwise repeats the old score. That turns a high-information answer into a low-information one and makes the leaderboard reliable under adaptive submissions.4 Dwork, Feldman, Hardt, Pitassi, Reingold, and Roth proved the general result: a holdout answered with limited information, through a mechanism they called Thresholdout, stays statistically valid under far more adaptive queries than one answered exactly.5

An agent iterating against an oracle is a participant iterating against a holdout. A verdict of PASS or FAIL is about as low-information as an answer gets. An execution trace with the failing assertion is about as high-information as one gets. The Ladder result says which one keeps the hidden test meaningful.

Reason two: the diagnosis is the data#

The point of collecting traces is to train a model on them. The behavior you want the model to learn is not "read the failing assertion and change the code until it passes." It is "form a hypothesis about what is wrong, write a test that checks it, run the test, read the result, and fix the cause." That is what expert engineers do, and the trace of an expert doing it is what you want in the training set.

If the oracle tells the agent why it failed, the trace after that point is the agent copying the oracle's diagnosis. The model trained on it learns to wait for a diagnosis. If the oracle says only FAIL, the trace after that point is the agent's own diagnosis: it has to write its own tests, run them, read its own failures, and reason about what the hidden requirement might be. That is the behavior worth learning, and it only appears in the trace when the oracle withholds the answer.

The agent is not working blind. It has the task description, the whole repository, the existing test suite, and a shell. It can write any test it wants and run it as often as it wants, with full output. The one-bit rule applies to the hidden oracle only. In-session feedback from the agent's own tests is as rich as the agent cares to make it. The oracle's silence forces the agent to generate that feedback for itself, which is the skill.

Reason three: the alternative is a model grading a model#

The usual proposal for richer feedback that does not leak the test is a judge: a second model that reads the diff, the test output, and the task, and writes an explanation for the agent. The judge sees the hidden test; the agent sees the judge's prose. This is where the research on self-grading matters, because the judge is a model with the same training and the same blind spots as the agent.

Zheng and colleagues, in the paper that introduced MT-Bench and the LLM-as-a-judge method, documented three biases in model judges: position bias, where the judge prefers whichever answer appears first; verbosity bias, where it prefers the longer answer; and self-enhancement bias, where it prefers answers it wrote.6 Wang and colleagues showed that swapping the order of two answers could reverse a judge's verdict.7 Panickssery, Bowman, and Feng showed that a model's preference for its own output is not an accident of style. Models can recognize their own text above chance, and the ones that recognize it best prefer it most; fine-tuning a model to recognize its own output more accurately made its self-preference stronger.8 A judge from the same family as the agent is a judge that likes the agent.

The research on self-correction is worse. Huang and colleagues reviewed the claim that models can correct their own reasoning and found that the gains reported in earlier work depended on an oracle: the model was told whether its answer was right before being asked to reconsider. Without that signal, asking the model to check its work made the answers worse more often than better.9 That paper is sometimes read as an argument against external feedback. It is the opposite. It shows that a binary external signal is the thing that made self-correction work in the papers that reported it working, and that the model's own judgment was not contributing.

Stechly, Marquez, and Kambhampati tested GPT-4 as a critic of its own graph-coloring solutions and found it could not reliably tell a correct coloring from an incorrect one; iterating on its own critique did not help, and a simple external checker did.10 Valmeekam, Marquez, and Kambhampati found the same for planning: a model critiquing its own plans lowered the success rate, and an external verifier raised it.11 Tyen and colleagues separated two skills and found that models are poor at locating the error in a chain of reasoning but can often fix it once the location is given.12 Kamoi and colleagues surveyed the self-correction literature and concluded that no reliable self-correction has been shown without external feedback, and that many reported successes used unrealistic setups where the model was given information it would not have in practice.13 Xu and colleagues found that self-refinement loops amplify the model's own biases over iterations.14

None of these results says a judge model is useless. They say a judge model's errors are correlated with the agent's errors, so using one to explain failures to the agent does not add an independent signal. It adds a confident one. The hidden test is independent. Its one bit is worth more than the judge's paragraph.

Feedback channels from the oracle to the agent

  1. PASS or FAILThe Ladder regime. The hidden test stays valid. The agent's own diagnosis is in the trace.
  2. Which test failedNames the requirement. The agent now knows what to fit to.
  3. The assertion and the valuesThe agent can special-case the inputs. Patch overfitting territory.
  4. A judge model's explanationAdds a correlated, biased reading of the test to the leak. The worst of both.
  5. The test sourceThe hidden test is no longer hidden. The oracle measures nothing.

Information leaked about the hidden test

Each rung down leaks more of the hidden test and leaves less of the agent's own reasoning in the trace. The book recommends the top rung and names the cost.

Recovering some of the cost#

There are ways to give back part of the in-session success rate without giving up the properties above.

Disclose to the trace, never to the agent. The oracle writes the full test output to its own log, which the data team reads. When a task fails many sessions in a row, a person reads the log and either fixes a hidden test that fails for the wrong reason or rewrites the task description so the requirement is clearer. The agent gets a better task, not a leaked test.

A visible test the agent must write. Make the task description say that the change must come with a test that fails before and passes after. The agent writes its own red-green pair, which is rich in-session feedback and good trace content. The hidden test stays hidden and checks that the agent's test checked the right thing.

Staged tasks. If a task's hidden test covers three requirements, split it into three tasks with one hidden test each. The agent gets three bits instead of one across the same work, and each bit still says nothing about its test.

More attempts, fresh context. The flip rate per session is lower; the flip rate per task does not have to be. Chapter 9 covers sampling several sessions per task and keeping the one that flips. With the oracle silent, the attempts are independent, which is what repeated sampling needs.

Disclose after the fact. Once a task has been flipped by some session, its hidden test can be released into the repository as an ordinary test. It has done its job as an oracle. Future sessions on nearby code get it as a visible regression test, and new hidden tests are written for new tasks.

The rule#

The hidden oracle says PASS or FAIL. Every other channel from the oracle to the agent is closed. Everything the oracle knows goes to the trace store, for people, under access control. If a team decides to open a wider channel, it should do so knowing which rung of the ladder it has moved to and what it has traded for the higher flip rate.

Footnotes#

  1. Chen, X., Lin, M., Schärli, N., & Zhou, D. (2023). Teaching Large Language Models to Self-Debug. arXiv; ICLR 2024. https://arxiv.org/abs/2304.05128 ↩

  2. Qi, Z., Long, F., Achour, S., & Rinard, M. (2015). An Analysis of Patch Plausibility and Correctness for Generate-and-Validate Patch Generation Systems. ISSTA 2015, 24–36. https://doi.org/10.1145/2771783.2771791 ↩

  3. Smith, E. K., Barr, E. T., Le Goues, C., & Brun, Y. (2015). Is the Cure Worse Than the Disease? Overfitting in Automated Program Repair. ESEC/FSE 2015, 532–543. https://doi.org/10.1145/2786805.2786825 ↩

  4. Blum, A., & Hardt, M. (2015). The Ladder: A Reliable Leaderboard for Machine Learning Competitions. ICML 2015, PMLR 37, 1006–1014. https://arxiv.org/abs/1502.04585 ↩

  5. Dwork, C., Feldman, V., Hardt, M., Pitassi, T., Reingold, O., & Roth, A. (2015). The reusable holdout: Preserving validity in adaptive data analysis. Science, 349(6248), 636–638. https://doi.org/10.1126/science.aaa9375 ↩

  6. Zheng, L., Chiang, W.-L., Sheng, Y., et al. (2023). Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. NeurIPS 2023 Datasets and Benchmarks. https://arxiv.org/abs/2306.05685 ↩

  7. Wang, P., Li, L., Chen, L., et al. (2023). Large Language Models are not Fair Evaluators. arXiv; ACL 2024. https://arxiv.org/abs/2305.17926 ↩

  8. Panickssery, A., Bowman, S. R., & Feng, S. (2024). LLM Evaluators Recognize and Favor Their Own Generations. arXiv; NeurIPS 2024. https://arxiv.org/abs/2404.13076 ↩

  9. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv; ICLR 2024. https://arxiv.org/abs/2310.01798 ↩

  10. Stechly, K., Marquez, M., & Kambhampati, S. (2023). GPT-4 Doesn't Know It's Wrong: An Analysis of Iterative Prompting for Reasoning Problems. arXiv. https://arxiv.org/abs/2310.12397 ↩

  11. Valmeekam, K., Marquez, M., & Kambhampati, S. (2023). Can Large Language Models Really Improve by Self-critiquing Their Own Plans? arXiv. https://arxiv.org/abs/2310.08118 ↩

  12. Tyen, G., Mansoor, H., Cărbune, V., Chen, P., & Mak, T. (2024). LLMs cannot find reasoning errors, but can correct them given the error location. Findings of ACL 2024. https://arxiv.org/abs/2311.08516 ↩

  13. Kamoi, R., Zhang, Y., Zhang, N., Han, J., & Zhang, R. (2024). When Can LLMs Actually Correct Their Own Mistakes? A Critical Survey of Self-Correction of LLMs. Transactions of the ACL, 12. https://arxiv.org/abs/2406.01297 ↩

  14. Xu, W., Zhu, G., Zhao, X., Pan, L., Li, L., & Wang, W. Y. (2024). Pride and Prejudice: LLM Amplifies Self-Bias in Self-Refinement. ACL 2024. https://arxiv.org/abs/2402.11436 ↩

Cite this

Anderson, M. (2026). The one-bit verdict. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 6). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-one-bit-verdict

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {The one-bit verdict},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {6},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-one-bit-verdict}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.