Mac Anderson

AI agents, Autonomous agents

What an agent should be told before it starts

A prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things.

Oxagen Research8 min readFirst published on oxagen.sh
View markdown

Most teams that run agents end up with one instruction file. It starts as four lines about the test command. Then it gains a section on which directory things go in. Someone adds a paragraph at 2am after an incident. Within a quarter the file is 600 lines long, and nobody has read all of it since March.

The file does real work, but its shape is wrong for that work. Its author is whoever edited it last. It has no scope and no dates. It cannot say that rule three applies to one repository and rule nine applies to everyone. It also keeps no record of whether the agent read line 412 during its four hours in your codebase. So when something goes wrong, you cannot say what the agent was told. You can only show what the file says now.

Why a longer prompt does not help#

The obvious fix is to write more into the file. Two research findings rule that out.

Liu and colleagues measured how models use long inputs.1 Accuracy was highest when the relevant passage sat near the start or the end of the input. It dropped when the passage sat in the middle. This U-shaped pattern held even for models built for long inputs. In your file, a rule's position depends on when someone added it. So how well a rule works depends on when it was written, not on how much it matters.

The second finding rules out the other fix, which is to repeat the rule during the run. Laban and colleagues gave models a full instruction all at once. They also gave the same instruction in pieces across a conversation. Across six generation tasks, the split version scored 39 percent lower on average.2 Most of the drop came from less consistent answers, not from lost skill. The models made an assumption early and then relied on it too much. So telling an agent at hour six what it needed at hour zero is the worst of the options.

This means the standing rules must be in place before the run starts. They must also be put in a chosen position, not added to the end. A plain text file cannot do that.

What the memory research found#

Research on agent memory agreed on one idea some time ago. An agent's context should be built by picking the right pieces for each step, not pasted in whole.

Sumers and colleagues proposed CoALA, a framework based on cognitive architectures, which are older models of how a mind is organised.3 CoALA describes a language agent as separate memory parts plus a set of actions. Some actions work on the agent's memory, and some work on the outside world. The specific parts matter less than the main idea. An agent knows several kinds of things, and each kind lasts a different length of time. If you treat them all as one long transcript, you lose the differences that decide what the agent should keep in view.

Packer and colleagues built MemGPT on the same idea.4 They borrowed tiered memory from operating systems. The system keeps the right subset of memory in the limited context window and moves items in and out as needed. Park and colleagues stored a full record of an agent's experiences in plain language.5 Their system wrote higher-level reflections from that record. Then it retrieved the relevant pieces when the agent planned, instead of keeping everything in view. In each case, a component with rules decides what the agent sees on each turn. A person with a text editor does not.

Wang and colleagues showed what stored knowledge is worth when it builds up over time. Their agent, Voyager, explores by following a curriculum of tasks.6 It saves the skills it learns as code in a growing library, and it retrieves them later.

Voyager with a saved skill library, compared with earlier agents

MeasureImprovement
Unique items collected3.3x
Distance travelled2.3x
Tech tree milestones unlockedup to 15.3x faster
Source: Wang et al. (2023), Voyager in Minecraft compared with earlier state-of-the-art agents. The authors also report that the learned skill library transfers to new worlds, where competing approaches struggled.

The transfer result matters most here. The library worked in worlds it was not built in. Knowledge saved as a named item that can be found again still helps when the task changes. Knowledge kept in one run's prompt does not.

What a rule needs#

Go through the file line by line. For each line, ask what it would need to count as a rule.

It needs an owner, so that someone answers for it. It needs a scope, because a rule about the staging database is not a rule for every repository in the organisation. It needs a start date and a way to end, because much of the file describes systems that no longer exist. It needs a source, because the incident that caused a rule is often the most useful thing to know about it. Last, it needs a way to take effect that is more than one person editing a line. If an agent could rewrite a rule, the rule would not bind the agent.

In Oxagen, that object is a steering record. A record holds one thing that an agent or a person states, learns, proposes, or decides. Once written, a record never changes. Oxagen computes a hash, a short fingerprint of the content, from a standard form of the record. So two copies are either the same record or visibly different records. Each record has a scope, from one user up to the whole organisation. Each record also points back to the frame, the recorded event in a run, where it was learned. So the rule and the incident behind it are one link apart.

Three records over the same repository, read as of today

FactHoldsStatus (Today)
Migrations run against the shared planeJan to JunNo longer holds
Each organisation resolves its own storeJun onwardHolds
Release notes wait for a named approverMar onwardHolds
Illustrative. A record has a start date and can be retracted. So you can ask what the agent was told in April separately from what it would be told now.

The first two rows contradict each other, and each was correct at its own time. A text file cannot hold both. It holds only the newest line. When someone edits that line, the record of what the agent worked under in April is lost.

An agent proposes a record and a person merges it#

Records work as governance, and not only as storage, because of where they live and how they take effect.

A workspace keeps its records as files in its linked repository. A record takes effect only through a pull request. An agent that learns something during a run can propose a record, but it cannot put one into effect. The proposal opens a pull request. Checks run against it. A person reviews it under the workspace's review mode. The merge puts the record into effect. So Git decides which records are active. "Which rules were active on the 14th" has the same answer as "what was in the tree on the 14th." Your existing tools already answer that question.

How a proposed record becomes one the agent runs under

  1. ProposeAn agent or a person writes a record, with its scope and its source.
  2. CheckOxagen checks the format, the link to the source, and the hash. It scans for secrets and for conflicts with active records.
  3. ReviewA person approves under the workspace's review mode.
  4. MergeThe commit puts the record into effect.
  5. PromoteOxagen writes a promotion event to a ledger. Each entry carries the hash of the entry before it.
  6. DeliverOxagen adds the record to the context of the next run.
The steering PR path. Without the review step, a record would be only a note an agent wrote to itself.

The review step answers a problem this pillar keeps returning to. It is the self-conditioning problem, where an agent builds on its own earlier mistakes. If an agent could write its own standing rules, one wrong conclusion at hour three would become a rule for every later run. The merge requirement does not make the agent's proposals better. It lets a person review them while they are still proposals.

How records reach a run#

Oxagen delivers records in a way that deals with the position problem directly.

Records that must hold go at the front of the stable part of the system context. The research above found that models read the start of a long input best. Oxagen picks lower-priority records for each run by relevance instead of sending all of them. A budget that lets everything in does not limit anything. Some records arrive between turns. Oxagen adds each of these as a separate message right after the cached system block, and does not edit it into that block. This keeps the cached start of the context the same across runs. A record can also be delivered as a context frame that points back to the record and to the commit that made it active. So an answer can cite the rule it came from.

Two limits follow. First, delivering a record does not make the model obey it. Nothing here makes a model comply. The check that the work was right still has to come from outside the model. Huang and colleagues found that when models corrected themselves without outside help, their answers sometimes got worse.7 Second, a record is only as good as its scope. Someone may promote a rule to the whole organisation because that was easier than scoping it. That rule will be wrong somewhere.

Where this meets Oxagen#

Oxagen is workforce management for autonomous agents. Steering records belong to the equipment clause of an agent's mandate. That clause covers the business context the agent may read and the steering it runs under. The engineers accountable for the agent set it. The scope on a record and the scope in the mandate use the same mechanism. So an agent gets what its work needs, not everything the organisation knows.

For an operator, this means "what was this agent told" has an answer. When a run goes wrong, the answer is a set of records. Each one has an owner, a scope, a hash, and the commit that made it active. Without records, the answer is a file edited over time by whoever was on call. When an agent learns something worth keeping, it proposes a record and does not decide alone. When a rule turns out to be wrong, you retract it with a pull request, so the retraction has a date too.

The next post covers the other half: what to do when a run has started and is going the wrong way.

Footnotes#

  1. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). Lost in the Middle: How Language Models Use Long Contexts. TACL. https://arxiv.org/abs/2307.03172 ↩

  2. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). LLMs Get Lost In Multi-Turn Conversation. arXiv. https://arxiv.org/abs/2505.06120 ↩

  3. Sumers, T. R., Yao, S., Narasimhan, K., & Griffiths, T. L. (2023). Cognitive Architectures for Language Agents. TMLR. https://arxiv.org/abs/2309.02427 ↩

  4. Packer, C., Wooders, S., Lin, K., Fang, V., Patil, S. G., Stoica, I., & Gonzalez, J. E. (2023). MemGPT: Towards LLMs as Operating Systems. arXiv. https://arxiv.org/abs/2310.08560 ↩

  5. Park, J. S., O'Brien, J. C., Cai, C. J., Morris, M. R., Liang, P., & Bernstein, M. S. (2023). Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. https://arxiv.org/abs/2304.03442 ↩

  6. Wang, G., Xie, Y., Jiang, Y., Mandlekar, A., Xiao, C., Zhu, Y., Fan, L., & Anandkumar, A. (2023). Voyager: An Open-Ended Embodied Agent with Large Language Models. TMLR. https://arxiv.org/abs/2305.16291 ↩

  7. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798 ↩

Cite this

Oxagen Research. (2026, September 16). What an agent should be told before it starts. oxagen.sh. https://oxagen.sh/blog/what-an-agent-should-be-told-before-it-starts

BibTeX
@online{anderson2026whatanagent,
  author  = {{Oxagen Research}},
  title   = {What an agent should be told before it starts},
  year    = {2026},
  date    = {2026-09-16},
  url     = {https://oxagen.sh/blog/what-an-agent-should-be-told-before-it-starts}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.