Mac Anderson

Code graphs, Self-evolving agents

Governing an agent that rewrites itself

An agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts.

Oxagen Research9 min readFirst published on oxagen.sh
View markdown

Your compliance team has one simple question about the agent: what is it allowed to do? For a normal service, the answer is a document. The service changes only when somebody merges a pull request, so the document stays true. An agent that changes its own behavior is different. Each pass of its loop can change it, so the document is true only until the next pass.

Systems like this exist today. The Darwin Gödel Machine edits its own code over and over and checks each edit on coding benchmarks. It raised its SWE-bench score from 20.0% to 50.0% and its Polyglot score from 14.2% to 30.7%.1 So the open question is no longer whether an agent can improve itself. It is what you agree to when you let one run inside your company.

Darwin Gödel Machine, before and after editing itself

Initial agentAfter self-edits
SWE-bench20%50%
Polyglot14.2%30.7%
Source: Zhang, Hu, Lu, Lange, and Clune (2025). The system checks each edit on the same benchmarks that score it. This post argues that a benchmark cannot do that job alone.

Safety researchers have worked on this question for a decade, and their answers are concrete. Most of those answers say you cannot skip the step you were hoping to skip.

Self-changing systems go wrong by chasing a flawed score#

In practice, self-changing systems go wrong in a plain way. They improve exactly what you measured, and what you measured was not what you meant.

Krakovna and colleagues at DeepMind collected examples of this and called it specification gaming. In one, the goal was to stack a red block on a blue one. The agent was rewarded for how high the underside of the red block was. So it flipped the red block over. In another, a boat-racing agent was rewarded for hitting green blocks along the course. It drove in circles hitting the same blocks over and over and never finished the race.2 Neither agent was broken. Both did exactly what the score asked.

Skalse and colleagues then made this precise. A proxy is a measure you optimize in place of the real goal. They call a proxy unhackable if raising the expected proxy score can never lower the expected true score. That is the property you want from any metric your agent optimizes. Their result is bad news. Across all possible random policies (the rules an agent uses to pick actions), two reward functions can be unhackable only if one of them is constant.3 So you cannot write a proxy metric that is safe to push without limit. Unhackable pairs where neither function is constant exist only when you limit the set of policies. In other words, safety comes from limiting what the agent can do. A better score does not provide it.

That is the main design lesson, and it is older than today's agents. A decade ago, "Concrete Problems in AI Safety" sorted the ways systems fail into five groups: avoiding side effects, avoiding reward hacking, scalable supervision, safe exploration, and distributional shift.4 Distributional shift means the conditions the system meets drift away from the ones it was tested on. A self-evolving agent touches all five groups at once. Self-modification also makes the last one loop back on itself. The agent that shifts the conditions is also the agent being tested under them.

A sandbox is needed, but it is not enough on its own#

Every serious self-modification paper runs its system in a sandbox, an isolated space it should not be able to leave. The Darwin Gödel Machine authors say plainly that all experiments used safety measures, including sandboxing and human oversight.1

At least one paper measured how well its sandbox held. STOP points a language-model improver at its own code. Its authors measured how often the generated code bypassed the sandbox.5 That is the right experiment. They ran it because the answer might not be zero. A sandbox is a boundary, and you should measure a boundary instead of assuming it holds.

There is a deeper reason not to rely on the sandbox alone. The original Gödel machine design had a real answer. It could rewrite any part of its own code only after it found a proof that the rewrite helps. That proof started from axioms describing the utility function, the hardware, and the starting code.6 That is a real guarantee. It is also why no one ships it. Such a proof cannot be found in practice. So the Darwin Gödel Machine replaces it with tests on benchmarks.1

Trading a proof for a benchmark score is reasonable. But it moves your whole safety case onto the benchmark, and careful audits show benchmarks cannot carry that load. Kapoor and colleagues found agent benchmarks with weak holdout sets or none at all. They found agents that take shortcuts and overfit. They found that results are often hard to reproduce.7 Suppose the only check in an evolution loop is a benchmark score. Then the loop will find that benchmark's weak spots, because finding what raises the score is what the loop does best.

Limiting what an agent can reach differs from shaping what it aims for#

These are two separate controls. Teams often use one and assume they have both, so it helps to name each.

Capability control limits what the agent can reach: the sandbox, the list of allowed tools, the scope of its credentials, and the data it can query. You can enforce it and test it. Skalse's result says this is the control that carries the guarantee, since safety comes from limiting the set of policies.3

Incentive control shapes what the agent aims for: the reward, the score, and the test a self-edit must pass. It feels like the stronger control. Skalse's result shows it cannot close every gap on its own.

Oversight sits on top of both, and the evidence here is better than many expect. Bowman and colleagues tested a deliberately simple setup for scalable oversight. Humans chatted with an unreliable model assistant to answer questions. The pair did much better than the model alone and the humans alone on MMLU and time-limited QuALITY.8 This does not mean humans should review every self-edit. It means a person working with a real interface does better than a person reading a summary. So design the review step with care instead of treating it as a formality.

A governance model for a system that changes itself#

These limits together lead to one answer with four parts.

Typed contracts instead of documents. A system that changes between readings cannot be held to a written policy. So the unit of governance has to be an object a machine can check. Bind six things together in one signed record: identity, knowledge scope, permitted action, commercial terms, verified outcome, and audit record.

{
  "identity": "agent:invoice-triage@7f2c1a",
  "knowledge_scope": { "ontology": "finance/ap", "as_of": "2026-09-09T00:00:00Z" },
  "permitted_actions": ["read_invoice", "propose_gl_code"],
  "commercial_terms": { "meter": "governed_action", "budget_usd_month": 400 },
  "verified_outcome": { "check": "gl_code_matches_approved_ledger" },
  "audit_record": "required"
}

The version hash in the identity is the part everything depends on. When an agent rewrites itself, it gets a new identity. A new identity means a new contract. The old contract does not carry over to it.

A knowledge scope defined by an ontology. It is easy to list the permitted actions. It is hard to list what the agent may know about. Most access-control models give up here and hand over a whole database. An ontology solves this. It is a model of the business with typed entities and typed relationships between them. With it, the scope can be a statement about the graph instead of a list of rows. The query below, written in Cypher (the Neo4j query language), returns the entities in the agent's domain that are valid on a given date.

MATCH (a:Agent {id: $agentId})-[:SCOPED_TO]->(d:Domain)
MATCH (d)<-[:IN_DOMAIN]-(n:Entity)
WHERE n.valid_from <= $asOf AND coalesce(n.valid_to, $asOf) >= $asOf
RETURN n

A compliance team can read that query as a boundary. It also holds up when the agent changes itself. The agent can change its own code, but that does not change what the domain contains. Limiting what the agent can see is capability control, written as structure.

Self-edits go through the same checks as any other change. Teams often skip this part. A self-edit changes a production system. So it takes the same path as every other change: a proposal, a diff, a policy check, a test run, a decision, and a record. A model writing the change does not earn it a faster path. Much self-evolution research already produces output a person can review, so this costs less than it sounds. Agent Workflow Memory builds named, readable workflows from past attempts instead of hidden weight updates.9 ADAS stores the agents it discovers as code in an archive.10 You can diff both. Review what you can read, and do not promote what you cannot read.

A self-edit takes the path every change takes

  1. ProposalThe diff, and the new identity it produces
  2. Policy checkAgainst the contract in force
  3. Test runEvidence as well as a score
  4. DecisionA person who has the evidence
  5. RecordAppend-only, each entry linked by hash to the one before
The agent's actions pass through the same checks. If this path is weaker than that one, the easiest way for the agent to take a forbidden action is to edit itself.

A record of every accepted change that cannot be edited. New entries are only added to the end. Each entry holds a hash of the one before, so a changed entry shows. Each entry holds the contract version, the proposal, the evidence for accepting it, and the approver. This is not for show. Reproducibility is the known weak point of this whole field.7 Say someone asks in the future why the agent did something in March. You can answer only if you wrote it down in March. The 2025 survey of self-evolving agents describes the open problems the same way. It names evaluation and safety as the limits on the field, not capability.11

What this does not solve#

None of this makes a self-evolving agent safe in the strong sense. There is no proof, and the original Gödel machine shows why you should not expect one.6 What it does is make the system readable. At any moment, you can state with evidence what the agent is, what it may know, what it may do, what it cost, and what changed since last week.

The evolution loop also does not justify this extra work everywhere. AlphaEvolve's results are real and in production. One is a scheduling heuristic that recovered 0.7% of Google's worldwide compute. But they come from problems with automated evaluators that score each candidate exactly.12 If you have no definition of "better" that a machine can check, the loop has nothing to judge its changes. Without a judge, the loop changes without improving.

How the four parts map to Oxagen#

Oxagen is the agent control plane for the agents you run. It does not run them. The four parts above map to the mandate an agent works under. The typed contract is the mandate itself. It holds identity, knowledge scope, permitted action, commercial terms, outcome, and audit record in one object, checked on the calls routed through Oxagen. The knowledge scope bounded by the ontology is the equipment the agent is given. That is why the graph tracks time, instead of being a vector index. The change record that cannot be edited is the record. The budget in the contract is the budget clause. For an evolution loop the meter matters a great deal, because the cost of each accepted improvement tells you whether to keep the loop running. Oxagen does not decide whether a proposed self-edit is a good idea. A person with the evidence still decides that.

Footnotes#

  1. Zhang, J., Hu, S., Lu, C., Lange, R., & Clune, J. (2025). Darwin Godel Machine: Open-Ended Evolution of Self-Improving Agents. arXiv. https://arxiv.org/abs/2505.22954 ↩ ↩2 ↩3

  2. Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020). Specification gaming: the flip side of AI ingenuity. Google DeepMind. https://deepmind.google/discover/blog/specification-gaming-the-flip-side-of-ai-ingenuity/ ↩

  3. Skalse, J., Howe, N. H. R., Krasheninnikov, D., & Krueger, D. (2022). Defining and Characterizing Reward Hacking. NeurIPS 2022. https://arxiv.org/abs/2209.13085 ↩ ↩2

  4. Amodei, D., Olah, C., Steinhardt, J., Christiano, P., Schulman, J., & Mané, D. (2016). Concrete Problems in AI Safety. arXiv. https://arxiv.org/abs/1606.06565 ↩

  5. Zelikman, E., Lorch, E., Mackey, L., & Kalai, A. T. (2023). Self-Taught Optimizer (STOP): Recursively Self-Improving Code Generation. COLM 2024. https://arxiv.org/abs/2310.02304 ↩

  6. Schmidhuber, J. (2003). Goedel Machines: Self-Referential Universal Problem Solvers Making Provably Optimal Self-Improvements. arXiv. https://arxiv.org/abs/cs/0309048 ↩ ↩2

  7. Kapoor, S., Stroebl, B., Siegel, Z. S., Nadgir, N., & Narayanan, A. (2024). AI Agents That Matter. arXiv. https://arxiv.org/abs/2407.01502 ↩ ↩2

  8. Bowman, S. R., Hyun, J., Perez, E., Chen, E., Pettit, C., Heiner, S., et al. (2022). Measuring Progress on Scalable Oversight for Large Language Models. arXiv. https://arxiv.org/abs/2211.03540 ↩

  9. Wang, Z. Z., Mao, J., Fried, D., & Neubig, G. (2024). Agent Workflow Memory. arXiv. https://arxiv.org/abs/2409.07429 ↩

  10. Hu, S., Lu, C., & Clune, J. (2024). Automated Design of Agentic Systems. arXiv. https://arxiv.org/abs/2408.08435 ↩

  11. Gao, H., Geng, J., Hua, W., Hu, M., Juan, X., Liu, H., et al. (2025). A Survey of Self-Evolving Agents: What, When, How, and Where to Evolve on the Path to Artificial Super Intelligence. arXiv. https://arxiv.org/abs/2507.21046 ↩

  12. Novikov, A., Vũ, N., Eisenberger, M., Dupont, E., Huang, P.-S., Wagner, A. Z., et al. (2025). AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv, and Google DeepMind blog (14 May 2025). https://arxiv.org/abs/2506.13131 ↩

Cite this

Oxagen Research. (2026, September 9). Governing an agent that rewrites itself. oxagen.sh. https://oxagen.sh/blog/governing-an-agent-that-rewrites-itself

BibTeX
@online{anderson2026governinganagent,
  author  = {{Oxagen Research}},
  title   = {Governing an agent that rewrites itself},
  year    = {2026},
  date    = {2026-09-09},
  url     = {https://oxagen.sh/blog/governing-an-agent-that-rewrites-itself}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.