Mac Anderson

Self-improving models

From STaR to DeepSeek-R1: what self-improvement means

A vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one.

Oxagen Research8 min readFirst published on oxagen.sh
View markdown

On a call, someone tells your team that the model improves itself. Nobody in the room can say whether that means a training loop, a prompt trick, or just a slide. So nobody questions the claim. Six weeks later, you are debugging a failure nobody predicted, because nobody knew what was running.

The research behind that claim is real. It goes back about four years, and it is more specific than the sales pitch. Each published method that works is a loop with three steps. First, the model produces candidate outputs. Second, something decides which candidates are good. Third, the model trains on the ones that pass. The second step matters most. What does the deciding sets what the loop can learn, how far it gets, and where it stops improving.

Below, each method is covered in turn, with the signal that drives it and the limit that stops it.

Each self-improvement method follows this loop

  1. GenerateThe model produces candidate outputs
  2. DecideAn answer key, a reward model, or a judge
  3. TrainFine-tune on the outputs that passed

Repeat run it again

The methods below differ mainly in what does step 2. That choice sets what the loop can learn and where it stops improving.

STaR keeps the reasoning that reaches the right answer#

STaR, published in 2022, is the simplest version of the idea.1 A model gets a few worked examples and a question. It writes out its reasoning step by step, called a rationale. If the final answer is right, the rationale is kept. If the answer is wrong, the model is shown the correct answer and asked to write a rationale that reaches it. The authors call this step rationalization. The model is then fine-tuned, meaning trained a little more, on every rationale that was kept. Then the whole process runs again.

loop:
  rationales = model.generate(questions)
  keep = [r for r in rationales if r.answer == answer_key[r.question]]
  retry = rationalize(model, wrong_ones, answer_key)
  model = finetune(model, keep + retry)

The signal is the answer key. The model's confidence and its own judgment play no part. The answer key is an outside label that the model cannot argue with. So the loop ends somewhere useful instead of drifting.

The answer key is also the limit. If the model never gets a problem right, it has no correct rationale to train on. So STaR can learn only from problems it already solves some of the time. Rationalization helps by giving the model the answer first. But then some rationales reach the right answer through bad reasoning.

Two later methods follow the same pattern. The first is rejection sampling fine-tuning (RFT), from a 2023 scaling study by Yuan and colleagues. It samples many reasoning paths from a model trained on labeled examples. It keeps the paths that are correct and distinct, and adds them to the training data.2 The paper reports larger gains when the kept paths are more varied. It also reports larger gains for weaker starting models. So the method helps weaker models catch up more than it pushes the strongest models further.

The second is ReST, from DeepMind the same year. It splits the loop into two steps.3 The Grow step samples outputs from the current model. The Improve step scores those samples with a reward model, a separate model trained to score outputs. It keeps the high scorers and trains on them with offline reinforcement learning, which learns from a fixed set of saved samples. Because the samples are saved, several Improve passes can reuse them. So the authors describe ReST as more efficient than typical online reinforcement learning from human feedback, which needs new samples each round. They show it on machine translation.

ReST replaces the answer key with a reward model, and that change drove the next five years of work. A reward model costs less and covers more tasks. But it is a model, so it can be wrong in ways an answer key cannot.

Self-Instruct has the model write its own training data#

Self-Instruct, also from 2022, moves one step earlier.4 STaR generates reasoning for questions you already have. Self-Instruct generates the questions too. A model starts from a small set of tasks written by people. It writes new instructions, then writes inputs and outputs for them. A filter drops the invalid ones and the near-duplicates. What is left becomes training data that teaches a model to follow instructions.

Most of the open instruction datasets that came later were built this way. So it helps to be precise about what the method checks. The filters are rules of thumb. They check that an instruction is well-formed and not too close to one already in the set. No step checks that the generated answer is correct.

That works for what the method is for, which is teaching a model the form of following an instruction. It does not check correctness. Treating it as if it does is a common and costly mistake.

Some methods have a model supply the reward#

Constitutional AI, from Anthropic in 2022, replaces the human labeler for one purpose.5 The model critiques and revises its own outputs against a written list of principles. Those AI-made comparisons then train a preference model, and the preference model drives reinforcement learning. The paper makes a narrow claim. It trains a harmless assistant through self-improvement "without any human labels identifying harmful outputs."

People still took part. They wrote the constitution. The signal still comes from outside the model. It arrives once, as text at the start, instead of as labels all the way through.

RLAIF, from Google in 2023, tested how far this approach reaches.6 It used an off-the-shelf model to label which of two answers was better. On summarization and dialogue, it reports results comparable to reinforcement learning from human feedback. It also beat the model trained only on labeled examples, even when the labeling model was the same size as the model being trained. That last result is the surprising one. A model no larger than the one being trained still gives a usable training signal.

Self-Rewarding Language Models, from Meta in 2024, closes the loop fully.7 One model writes candidate responses and then judges them, using a prompt that asks it to act as a judge (the LLM-as-a-Judge method). It then trains on its own preferences. Three rounds of this on Llama 2 70B produced a model that beat Claude 2, Gemini Pro, and GPT-4 0613 on the AlpacaEval 2.0 leaderboard. Both skills improved together. The model got better at following instructions and better at judging.

Read how it was scored before you read the result as open-ended. AlpacaEval 2.0 is a preference benchmark, and a model scores it. So the loop improved what a model judge rewards, as measured by a model judge. That is a real result about the style of following instructions. It does not show that a model can train itself into being right about facts. And the paper ran three rounds, not thirty.

Quiet-STaR adds reasoning at every token#

Quiet-STaR, from 2024, makes reasoning happen at every token instead of in a separate prompt.8 A token is a word or part of a word. At each position, the model writes short hidden rationales in parallel. The training signal is whether a rationale helped the model predict the text that comes next. Rationales that help are reinforced. The others are dropped.

The reported gains are zero-shot, meaning the model saw no examples of the task, and they came without task-specific fine-tuning. GSM8K rose from 5.9% to 10.9%, and CommonsenseQA rose from 36.3% to 47.2%.

Quiet-STaR, zero-shot

BeforeAfter Quiet-STaR
GSM8K5.9%10.9%
CommonsenseQA36.3%47.2%
Source: Zelikman et al. (2024). No task-specific fine-tuning. The only training signal is whether a thought helped predict the text that follows.

Here the checker is the training text itself. That has a clear benefit, because ordinary text is plentiful and needs no labels. It is also the limit. Predicting the next token better stands in for reasoning better, but the two are different. A thought that makes the next sentence easier to predict may be a sound inference. It may also be a good guess about the writer's habits.

DeepSeek-R1 uses reinforcement learning on answers a program can check#

DeepSeek-R1, released in January 2025, is the current reference point.9 R1-Zero was trained with large-scale reinforcement learning, without a fine-tuning step on labeled examples first. Reasoning behaviors came from the reward alone. The Nature version of the paper describes self-reflection, verification, and changing strategies as they appeared, without reasoning examples labeled by people.10

The preprint is direct about the cost. It says R1-Zero has "poor readability, and language mixing." The released R1 fixes this. It adds training in several stages, and it adds a set of starting examples (cold-start data) before the reinforcement learning phase. So the pure loop worked, but people found its output hard to read. The authors added a stage with examples chosen by people to make it usable.

The key detail is what the reward measured. Reinforcement learning at that scale works on tasks where an automatic checker can score an answer. Examples are mathematics, competitive programming, and other fields where a program can decide correctness without a person. This is the same kind of signal STaR used, applied with far more compute. The common thread is an outside check, not the model's view of its own work. Cobbe and colleagues argued this in 2021 using GSM8K, a set of 8.5 thousand grade-school math problems. They trained a verifier, a model that ranks sampled solutions. As they added data, the verifier improved results more than fine-tuning did.11

MethodSignal sourceVerifierKnown limit
STaR (2022)Rationales for labeled questionsThe answer keyOnly learns problems it already sometimes solves
Self-Instruct (2022)Instruction data the model writesRule-of-thumb filters onlyNothing checks whether the answer is correct
Constitutional AI (2022)Self-critique against written principlesPrinciples written by peopleAims at harmlessness, not correctness
RFT (2023)Sampled reasoning paths, kept if correctThe answer keyGains shrink as the starting model gets stronger
ReST (2023)Samples from the current modelA trained reward modelCarries every flaw in the reward model
RLAIF (2023)A language model labeling pairs of answersThe labeling modelGives a preference, not a fact
Self-Rewarding LM (2024)The model judging its own outputThe same modelJudge and student miss the same things
Quiet-STaR (2024)Whether a thought helps predict the next textThe training textBetter prediction stands in for better reasoning
DeepSeek-R1-Zero (2025)Reward on tasks a program can checkAn automatic checkerPoor readability, language mixing

Read the Verifier column from top to bottom. Each method with a lasting gain in skill has something in that column the model does not control. Where the column points back to the model itself, the gain shows up only in things a model measures. Take that difference into your next call with a vendor. It is also where these methods fail. Models struggle to correct their own reasoning without outside feedback, and they sometimes get worse when they try.12

Where Oxagen fits#

Oxagen does not train models and does not run agents. It is the agent control plane for the agents you run. The Verifier column shows why that matters for your business. An agent's mandate binds its identity, its knowledge scope, its permitted actions, and its commercial terms into one object. The run's record keeps the checked outcome beside them. So the outcome signal is kept outside the model that produced it. The meter prices that same record. To know whether a self-improving system is improving, you need an outcome log the system cannot write for itself.

Footnotes#

  1. Zelikman, E., Wu, Y., Mu, J., & Goodman, N. D. (2022). STaR: Bootstrapping Reasoning With Reasoning. arXiv. https://arxiv.org/abs/2203.14465 ↩

  2. Yuan, Z., Yuan, H., Li, C., Dong, G., Lu, K., Tan, C., Zhou, C., & Zhou, J. (2023). Scaling Relationship on Learning Mathematical Reasoning with Large Language Models. arXiv. https://arxiv.org/abs/2308.01825 ↩

  3. Gulcehre, C., Le Paine, T., Srinivasan, S., Konyushkova, K., Weerts, L., Sharma, A., Siddhant, A., Ahern, A., Wang, M., Gu, C., Macherey, W., Doucet, A., Firat, O., & de Freitas, N. (2023). Reinforced Self-Training (ReST) for Language Modeling. arXiv. https://arxiv.org/abs/2308.08998 ↩

  4. Wang, Y., Kordi, Y., Mishra, S., Liu, A., Smith, N. A., Khashabi, D., & Hajishirzi, H. (2022). Self-Instruct: Aligning Language Models with Self-Generated Instructions. arXiv. https://arxiv.org/abs/2212.10560 ↩

  5. Bai, Y., et al. (2022). Constitutional AI: Harmlessness from AI Feedback. arXiv. https://arxiv.org/abs/2212.08073 ↩

  6. Lee, H., et al. (2023). RLAIF vs. RLHF: Scaling Reinforcement Learning from Human Feedback with AI Feedback. arXiv. https://arxiv.org/abs/2309.00267 ↩

  7. Yuan, W., Pang, R. Y., Cho, K., Li, X., Sukhbaatar, S., Xu, J., & Weston, J. (2024). Self-Rewarding Language Models. arXiv. https://arxiv.org/abs/2401.10020 ↩

  8. Zelikman, E., Harik, G., Shao, Y., Jayasiri, V., Haber, N., & Goodman, N. D. (2024). Quiet-STaR: Language Models Can Teach Themselves to Think Before Speaking. arXiv. https://arxiv.org/abs/2403.09629 ↩

  9. DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948 ↩

  10. DeepSeek-AI (2025). DeepSeek-R1 incentivizes reasoning in LLMs through reinforcement learning. Nature, 645, 633-638. https://doi.org/10.1038/s41586-025-09422-z ↩

  11. Cobbe, K., Kosaraju, V., Bavarian, M., Chen, M., Jun, H., Kaiser, L., Plappert, M., Tworek, J., Hilton, J., Nakano, R., Hesse, C., & Schulman, J. (2021). Training Verifiers to Solve Math Word Problems. arXiv. https://arxiv.org/abs/2110.14168 ↩

  12. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. arXiv. https://arxiv.org/abs/2310.01798 ↩

Cite this

Oxagen Research. (2026, September 9). From STaR to DeepSeek-R1: what self-improvement means. oxagen.sh. https://oxagen.sh/blog/self-improving-models-from-star-to-self-rewarding

BibTeX
@online{anderson2026selfimprovingmodels,
  author  = {{Oxagen Research}},
  title   = {From STaR to DeepSeek-R1: what self-improvement means},
  year    = {2026},
  date    = {2026-09-09},
  url     = {https://oxagen.sh/blog/self-improving-models-from-star-to-self-rewarding}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.