AI agents, Self-improving models
A model cannot grade its own homework
Without outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead.
Many agent stacks add a second prompt that asks the model to check its work. It is the cheapest reliability fix there is. It costs one extra call and needs no infrastructure. In a demo it seems to work. So teams ship it, mark the reliability problem solved, and move on.
The research shows this fix does less than teams expect. Under some conditions it makes things worse. The same result appears twice, once when the model answers and once when it trains. When a model checks its own answer, the answer does not reliably get better. When a model trains on its own output, the model does not reliably get better either. Both fail for the same reason. The loop adds no information the model did not already have.
This post covers what the evidence shows and what works instead.
Later tests found that self-correction fails on reasoning#
Self-Refine, published in early 2023, is where the pattern comes from.1 One model writes an output, gives itself feedback, and revises. It uses no extra training, no supervised data, and no reinforcement learning. The authors report a gain of about 20% absolute on average across their tasks. Both humans and automatic metrics preferred the revised outputs over one-shot outputs from the same model.
That result is real, and many readers took it as a general one. Later in 2023, Huang and colleagues at Google DeepMind tested the case that matters most for agents, which is reasoning.2 They call the pattern intrinsic self-correction. The model revises "based solely on its inherent capabilities, without the crutch of external feedback." Their finding is direct. Without outside feedback, models struggle to correct their own reasoning. Sometimes performance gets worse after self-correction.
The two papers do not contradict each other. They differ in what the feedback can check. Self-Refine did best on tasks where the feedback names a concrete defect that can be checked. Examples are code that can be run, or an output with a property that can be measured. Reasoning is different. If the model got the reasoning wrong, it would have to notice the error with the same ability that made it.
So the practical lesson is narrow. Self-critique can improve format and style. It cannot confirm that an answer is correct. If your agent's retry loop has no source of truth, the retry gives you a second sample. It does not give you a check.
For agents, this is easy to act on. Most agent tasks touch something that returns a result. A test suite returns pass or fail. A type checker returns errors with line numbers. An HTTP call returns a status code. A database returns rows or an exception. Each of these is an outside verifier, a check that is not the model. It already exists in the environment and costs nothing to use. A persuasive argument cannot change its answer. Some agent setups have the model review its output before running the tests. That is the wrong order. Run the check first, and give the real error back to the model. Then the critique step has information the model did not have when it wrote the code. This is the difference between Self-Refine's strong tasks and its weak ones, written as an engineering rule.
Check first, then critique
- WriteCode, a query, an API call
- Run the checkTests, the type checker, a status code
- Feed backThe real error, with line numbers
- ReviseThe critique now has new information
Repeat until the check passes
Training a model on its own output makes it worse#
At training time the same problem is worse, because the damage adds up over generations.
Shumailov and colleagues named this model collapse.3 Suppose you train each generation of models on data made by the previous generation. The tails of the original distribution, the rare cases, disappear first. Rare events stop showing up, then uncommon ones. The damage cannot be undone, so later generations cannot recover what earlier ones lost. The effect is not limited to language models. The authors showed it in variational autoencoders and Gaussian mixture models too. That suggests the cause is the recursive setup, not any one architecture. The peer-reviewed version in Nature adds a finding for anyone who scrapes the web for training data. Training on a mix of real and generated content, without choosing what goes in, leads to collapse. The model loses its ability to produce diverse, high-quality output. Language models trained this way over many rounds produce likely sequences too often. In the end they produce text no human would write.4
Alemohammad and colleagues found the same thing in image generation, and they stated the condition more precisely.5 They call the failure Model Autophagy Disorder. Autophagy means self-eating. Their finding is that "without enough fresh real data in each generation of an autophagous loop, future generative models are doomed to have their quality (precision) or diversity (recall) progressively decrease."
The quote says quality or diversity. The model does not always get worse in every way. It can keep its quality by narrowing, and produce a smaller set of more polished outputs. This failure is the hardest one to see in a spot check. Each sample looks fine, but the range of outputs behind the samples has shrunk.
Keeping the real data prevents collapse#
The warnings about collapse were loud. So a 2024 paper tested whether collapse is inevitable. The answer depended on one detail of the experiment.6
The collapse experiments replace the training data each generation. Gerstgrasser and colleagues tested accumulation instead. They kept the original real data and added each generation of synthetic data next to it. Collapse did not occur. The result held across several model families and datasets.
Two retention policies across four generations
ReplaceThe tails vanish first, then collapse
AccumulateCollapse does not occur
This is the easiest finding in this research to act on, and it gets the least attention. Synthetic data does not cause collapse by itself. Replacing your real data with synthetic data does. The fix is a retention policy, a rule about which data you keep. It needs no new algorithm. Any team that feeds model output back into training can act on it this quarter.
Checks from outside the model do work#
Self-assessment fails, and training on your own output degrades the model. What remains is a signal from outside the model. The research on verifiers, separate checkers that grade a model's output, shows real gains. That research is older than the current wave of models.
Cobbe and colleagues trained a separate verifier to rank sampled solutions to grade-school math problems. They used GSM8K, a set of 8.5 thousand problems.7 As they added data, verification improved faster than the fine-tuning baseline. They sampled many solutions and had a second model pick one. That beat training one model to produce a better first attempt.
Lightman and colleagues improved the idea in 2023 by changing what the verifier grades.8 Outcome supervision rewards the final answer. Process supervision rewards each step of the reasoning. Process supervision clearly did better. Their process-supervised reward model solved 78% of problems from a representative subset of the MATH test set. They released PRM800K, a set of 800 thousand human feedback labels on single steps. That set shows the real cost of the method. It takes a large amount of human judgment, applied where the model cannot supply it.
DeepSeek-R1 applies the same idea at a larger scale.9 Its reasoning behavior came from large-scale reinforcement learning, with no supervised fine-tuning first. The training used tasks where an automatic checker can decide whether an answer is right. So the reward came from a fixed procedure, not from a model's opinion.
| Signal | Where it comes from | Does it hold up at scale |
|---|---|---|
| The model's own critique | Inside the model | No, it gets worse on reasoning tasks |
| Its own output as training data | Inside the model | No, the model collapses if it replaces real data |
| Its own output added to real data | Mixed | Yes, accumulation avoids collapse |
| A separate learned verifier | Outside, learned | Yes, but it carries the verifier's flaws |
| Step-level human labels | Outside, human | Yes, at a real cost in human labeling |
| An automatic correctness checker | Outside, a fixed procedure | Yes, where the task allows one |
Any verifier can be gamed, so the record matters#
Adding a verifier does not end the problem, for one more reason.
Skalse and colleagues gave reward hacking a formal definition.10 Reward hacking is when an optimizer raises its score without doing what the score was meant to measure. Call the score a proxy. The authors call a proxy unhackable if raising the expected proxy return can never lower the expected true return. Then they proved a strong limit. Consider all stochastic policies, which choose actions with some randomness. Across all of them, two reward functions can only be unhackable if one of them is constant. Useful unhackable pairs do exist in narrow settings, such as deterministic policies or a finite set of policies. In the general case, they do not. Your verifier is a proxy. An optimizer that pushes hard enough will find the gap between the verifier and what you meant.
Pan, Bhatia, and Steinhardt showed what this looks like in experiments.11 More capable agents exploit a badly specified reward more. They score higher on the proxy and lower on the true goal than weaker agents do. The authors also found phase transitions. At certain capability levels, the agent's behavior changes in kind, and the true reward drops sharply. Watching the metric you optimized will not catch this, because the change is not gradual. The system looks fine until the drop happens.
Together, these results explain why a record matters. You cannot write a proxy that stays safe under any amount of optimization pressure. You can keep a record of what the agent tried, what was permitted, and what the verified outcome was. With that record, you can see the gap between the proxy and your intent when it opens, not after the phase transition. The measurement has to be separate from the thing it measures. It also has to last, because the failure is a sudden step, and you will want the history from before and after it.
Where this fits in Oxagen#
Oxagen does not train models or run agents. It is the control plane for the agents you run. The results above are why its record sits outside the model. Oxagen grounds answers in a knowledge graph with citations, which brings a source from outside the model into the loop. Oxagen also records each run against the agent's mandate. The record shows what the agent asked for, which rule answered, what it cost, and what checked the outcome. This keeps the grading separate from the model being graded. Judging an agent on that record, not on its own report, is the only approach here that holds up against reward hacking. None of this makes a model better at checking itself. It gives the checking to something other than the model.
Footnotes#
-
Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., Alon, U., Dziri, N., Prabhumoye, S., Yang, Y., Gupta, S., Majumder, B. P., Hermann, K., Welleck, S., Yazdanbakhsh, A., & Clark, P. (2023). Self-Refine: Iterative Refinement with Self-Feedback. arXiv. https://arxiv.org/abs/2303.17651 ↩
-
Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv. https://arxiv.org/abs/2310.01798 ↩
-
Shumailov, I., Shumaylov, Z., Zhao, Y., Gal, Y., Papernot, N., & Anderson, R. (2023). The Curse of Recursion: Training on Generated Data Makes Models Forget. arXiv. https://arxiv.org/abs/2305.17493 ↩
-
Shumailov, I., Shumaylov, Z., Zhao, Y., Papernot, N., Anderson, R., & Gal, Y. (2024). AI models collapse when trained on recursively generated data. Nature, 631, 755-759. https://doi.org/10.1038/s41586-024-07566-y ↩
-
Alemohammad, S., Casco-Rodriguez, J., Luzi, L., Humayun, A. I., Babaei, H., LeJeune, D., Siahkoohi, A., & Baraniuk, R. G. (2023). Self-Consuming Generative Models Go MAD. arXiv. https://arxiv.org/abs/2307.01850 ↩
-
Gerstgrasser, M., Schaeffer, R., Dey, A., Rafailov, R., Sleight, H., Hughes, J., Korbak, T., Agrawal, R., Pai, D., Gromov, A., Roberts, D. A., Yang, D., Donoho, D. L., & Koyejo, S. (2024). Is Model Collapse Inevitable? Breaking the Curse of Recursion by Accumulating Real and Synthetic Data. arXiv. https://arxiv.org/abs/2404.01413 ↩
-
Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., & Schulman, J. (2021). Training Verifiers to Solve Math Word Problems. arXiv. https://arxiv.org/abs/2110.14168 ↩
-
Lightman, H., Kosaraju, V., Burda, Y., Edwards, H., Baker, B., Lee, T., Leike, J., Schulman, J., Sutskever, I., & Cobbe, K. (2023). Let's Verify Step by Step. arXiv. https://arxiv.org/abs/2305.20050 ↩
-
DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948 ↩
-
Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and Characterizing Reward Hacking. arXiv. https://arxiv.org/abs/2209.13085 ↩
-
Pan, A., Bhatia, K., & Steinhardt, J. (2022). The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models. arXiv. https://arxiv.org/abs/2201.03544 ↩
Cite this
Oxagen Research. (2026, September 9). A model cannot grade its own homework. oxagen.sh. https://oxagen.sh/blog/the-limits-of-self-correction-and-model-collapse
BibTeX
@online{anderson2026thelimitsof,
author = {{Oxagen Research}},
title = {A model cannot grade its own homework},
year = {2026},
date = {2026-09-09},
url = {https://oxagen.sh/blog/the-limits-of-self-correction-and-model-collapse}
}Related research
- Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's tracesEvery coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.
- Steering a run you are not watchingWhen a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.
- The agent time horizon is doublingMETR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.
Updates by email
New research reaches subscribers first.