Mac Anderson

AI agents, Self-evolving agents

Self-evolving agents: what the evidence shows

What changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve.

Oxagen Research8 min readFirst published on oxagen.sh
View markdown

Your agent got a task wrong last Tuesday, and it will get the same task wrong next Tuesday. It does not know it has seen the task before. Each run starts from the same prompt, the same tools, and the same empty memory. So the hundredth run costs the same as the first and is no better. That is the first problem, and it is the costly one.

The second problem is the opposite worry. A system that changes itself can change in ways nobody asked for and nobody noticed. Ask an engineer whether they want their agent to rewrite its own tool code, and they will hesitate for a long time.

Both reactions make sense. There is now enough research on self-evolving agents to say which worry fits which case. A 2025 survey sorts the field by four questions: what changes, when it changes, how it changes, and where.1 This post answers the first and third questions with numbers. The measured gains differ by more than ten times, depending on what changes and what checks the change.

Four things an agent can change about itself#

Almost every system in this research changes one of four things. They are listed here from least to most that can go wrong.

The prompt or the context. The agent writes something down and reads it back on its next attempt. Reflexion is the clearest example. After a failed attempt, the agent writes a short note on what went wrong. It stores the note in a buffer, and the note goes into the next attempt's input.2 The model's weights do not change. Self-Refine does this within one turn. The same model critiques its own output and then rewrites it.3

The memory. Generative Agents stores a full plain-language record of what happened. It then combines those records into higher-level reflections and looks them up to plan.4 The agent does not learn a new skill. It builds a short summary of its own history.

The skills or tools. Voyager keeps a growing library of code it can run. When it works out how to craft an item, it stores the steps as a function. Later it uses that function to build harder items.5 Agent Workflow Memory does the same thing with routines instead of functions. It finds reusable workflows in its past attempts. It can do this offline from training examples, or online from new tasks as they arrive.6

The code. At the far end, the agent changes its own program. STOP starts with a simple program that improves code, called an improver. It then runs the improver on itself, which produces a better improver.7 In Automated Design of Agentic Systems (ADAS), one agent writes new agent programs and adds them to a growing archive.8 The Darwin Gödel Machine edits its own code over and over and keeps an archive of versions.9

As you go down that list, more can go wrong, and the measured gain also rises. The two rise together for a reason, and the larger gain has a cost.

Four things an agent can change about itself

  1. Prompt or contextReflexion, Self-Refine
  2. MemoryGenerative Agents
  3. Skills or toolsVoyager, Agent Workflow Memory
  4. CodeSTOP, ADAS, Darwin Gödel Machine

What can go wrong and the measured gain both rise

The rungs show order only and are not measured. Each rung names the systems in this post that change that layer.

The check on each change decides how the agent improves#

Each system above has a loop that makes a change and then checks it. If you change the check, the whole system behaves differently.

The strong systems all check against something outside the model. Reflexion's coding results come from running unit tests. Voyager adds a skill to its library only after the skill passes its own check and runs without error in the game. Errors from running go back into the next attempt. AlphaEvolve needs an automated evaluator that scores each candidate program. So it works on problems with a goal a machine can check, and not on problems without one.10 The Darwin Gödel Machine tests each self-edit on coding benchmarks. In all four cases, the loop relies on a checker that the model cannot argue with.

Each loop keeps only what passes a check outside the model

  1. ProposeA reflection, a skill, a workflow, or a code edit
  2. VerifyTests, running the code, or an evaluator outside the model
  3. KeepOnly what passed enters memory, the library, or the archive
  4. Next attemptStarts from what was kept

Repeat the agent improves only as far as step 2 checks correctly

Reflexion, Voyager, AlphaEvolve, and the Darwin Gödel Machine all run step 2 outside the model. Without that outside check, self-correction stalls or gets worse.

The weak version is a model grading itself with no outside signal, and the evidence on it is poor. Huang and colleagues tested this kind of self-correction on reasoning tasks. The model revised its answers using only its own judgment. Models struggled to correct themselves without outside feedback. Sometimes they did worse after revising.11 Self-Refine reports about 20 points of absolute gain on average across seven tasks. But its feedback is written for each task, and some of its tasks have clear scoring rules.3 The difference between the two results comes mostly from what does the checking.

This leads to a practical rule. An agent improves only as far as its checker measures the right thing. If your checker is a test suite, the agent gets better at passing tests. If your checker is a benchmark score, the agent gets better at the benchmark. You have to say what "better" means before the agent can get better.

The measured gains#

This table covers the field. Each number comes from the cited paper or its official write-up.

System (year)What changesWhat checks the changeMeasured gain
Reflexion (2023)2Reflection notes stored in a bufferUnit tests, reward from the environment91% pass@1 on HumanEval, against 80% for the GPT-4 baseline
Self-Refine (2023)3The output, rewrittenThe same model's own feedbackAbout 20 points absolute on average across 7 tasks
Voyager (2023)5A library of skills stored as codeIts own check plus running the code in Minecraft3.3x more unique items, 2.3x longer distances, tech-tree milestones up to 15.3x faster
Generative Agents (2023)4Memory and combined reflectionsPeople rating how believable the behavior isRemoving reflection makes behavior less believable
STOP (2023)7The program around the modelA supplied scoring functionThe improved improver beats the starting improver on later tasks
Agent Workflow Memory (2024)6Reusable workflows found in past attemptsTask success over 1,000+ tasks in 200+ domains24.6% relative success gain on Mind2Web, 51.1% on WebArena
ADAS / Meta Agent Search (2024)8The agent program, added to an archiveBenchmark score in the target domainDiscovered agents beat the best hand-designed ones, and still do when moved to other domains and models
Darwin Gödel Machine (2025)9Its own codeSWE-bench and Polyglot scoresSWE-bench 20.0% to 50.0%, Polyglot 14.2% to 30.7%
AlphaEvolve (2025)10Candidate program codeAutomated evaluatorsA scheduling heuristic in production for over a year recovering 0.7% of Google's worldwide compute, up to 32.5% faster FlashAttention kernel, and 4x4 complex matrix multiplication in 48 scalar multiplications

Two results stand out. First, the cheap changes can have large effects. Agent Workflow Memory adds no new model and no new tools. It spots a reusable routine in a past attempt, and that raises WebArena success by about half.6 Second, AlphaEvolve's results are the only ones here that run in production. The others are benchmark numbers. Its 4x4 complex matrix multiplication uses 48 scalar multiplications. That is the first improvement in that setting over Strassen's algorithm in 56 years.10

What the evidence does not show#

Keep three limits in mind when you decide.

The checker is usually a benchmark, and benchmarks have gaps. Kapoor and colleagues audited agent benchmarks and how they are used. They found a narrow focus on accuracy with no attention to cost. They found benchmarks with weak holdout sets, the test cases kept apart from tuning, or none at all. And they found results were often hard to reproduce.12 If an agent evolves against a benchmark with a weak holdout set, it overfits to that benchmark. So each result in the table above is only as good as the benchmark that scored it.

Papers rarely report cost next to accuracy. The same audit found that leading agents are more complex and costly than they need to be, because no one was trying to lower cost.12 Self-evolution loops cost more than any other agent design. They run the task many times, and the archive-based ones run many versions of the agent. If you cannot see the cost of each accepted improvement, you cannot tell whether the loop is helping or only spending tokens.

None of this is recursive self-improvement, and the papers say so. The STOP authors state that the language model itself never changes, so this is not full recursive self-improvement.7 Schmidhuber's original design required a proof that a self-edit helps before making it. Such a proof cannot be found in practice. So the Darwin Gödel Machine drops that requirement. It tests edits instead and runs with sandboxing and human oversight.9 Today's agents improve the program around a fixed model, and a fixed checker judges them. That is useful. It is less than what the phrase "self-evolving" makes people picture.

What to do first#

If your agent repeats the same failure, do not start by letting it rewrite its own code. Start by giving it a place to write down what happened. Add a checker that can tell it whether the next attempt was better. That covers the top two rows of the table. It costs almost nothing, and a large share of the reported gain comes from it.

Next, try finding workflows. Look at your successful attempts and check whether a routine is hidden in them. Agent Workflow Memory does this. It has the best gain for its risk in all of this research. A workflow found this way is a readable file, so you can inspect it before you trust it.6

Let the agent change its own code last. The largest benchmark jumps come from this layer. It is also the only layer where the agent can change what it is permitted to do. So it is a governance problem before it is an engineering problem.

What Oxagen records#

Oxagen is the agent control plane for the agents you run. It does not run them. The part of this that falls to Oxagen is the record. The record shows which version of an agent produced a result and what knowledge it was scoped to. It shows what the agent asked for, which rule answered, and what checked the outcome. With that record, you can show that an agent got better, instead of only saying so. The meter matters too. A self-evolution loop is the design most likely to spend real money in the background. The cost of each accepted improvement tells you whether to keep the loop running.

Footnotes#

  1. Gao, H., Geng, J., Hua, W., Hu, M., Juan, X., Liu, H., et al. (2025). A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence. arXiv. https://arxiv.org/abs/2507.21046 ↩

  2. Shinn, N., Cassano, F., Berman, E., Gopinath, A., Narasimhan, K., & Yao, S. (2023). Reflexion: Language Agents with Verbal Reinforcement Learning. NeurIPS 2023. https://arxiv.org/abs/2303.11366 ↩ ↩2

  3. Madaan, A., Tandon, N., Gupta, P., Hallinan, S., Gao, L., Wiegreffe, S., et al. (2023). Self-Refine: Iterative Refinement with Self-Feedback. NeurIPS 2023. https://arxiv.org/abs/2303.17651 ↩ ↩2 ↩3

  4. Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. https://arxiv.org/abs/2304.03442 ↩ ↩2

  5. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. arXiv. https://arxiv.org/abs/2305.16291 ↩ ↩2

  6. Wang, Z. Z., Mao, J., Fried, D., & Neubig, G. (2024). Agent Workflow Memory. arXiv. https://arxiv.org/abs/2409.07429 ↩ ↩2 ↩3 ↩4

  7. Zelikman, E., Lorch, E., Mackey, L., & Kalai, A. T. (2023). Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. COLM 2024. https://arxiv.org/abs/2310.02304 ↩ ↩2 ↩3

  8. Hu, S., Lu, C., & Clune, J. (2024). Automated Design of Agentic Systems. arXiv. https://arxiv.org/abs/2408.08435 ↩ ↩2

  9. Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv. https://arxiv.org/abs/2505.22954 ↩ ↩2 ↩3

  10. Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., Wagner, A. Z., et al. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv, and Google DeepMind blog (14 May 2025). https://arxiv.org/abs/2506.13131 ↩ ↩2 ↩3

  11. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2024). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798 ↩

  12. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). AI Agents That Matter. arXiv. https://arxiv.org/abs/2407.01502 ↩ ↩2

Cite this

Oxagen Research. (2026, September 9). Self-evolving agents: what the evidence shows. oxagen.sh. https://oxagen.sh/blog/self-evolving-agents-what-the-evidence-shows

BibTeX
@online{anderson2026selfevolvingagents,
  author  = {{Oxagen Research}},
  title   = {Self-evolving agents: what the evidence shows},
  year    = {2026},
  date    = {2026-09-09},
  url     = {https://oxagen.sh/blog/self-evolving-agents-what-the-evidence-shows}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.