Mac Anderson

AI agents, Autonomous agents

Steering a run you are not watching

When a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.

Oxagen Research7 min readFirst published on oxagen.sh
View markdown

An agent has been running for nine hours. You open the transcript over breakfast. Around hour three, the agent decided the caching layer was the problem. Since then it has rewritten the wrong component, with tests and documentation.

You have two buttons. One stops the run. That throws away nine hours, including the part of the work that was fine. The other lets you type a message. Most people pick this one, and the research shows it works badly.

Corrections typed as chat often do not work#

A correction typed as chat only helps if the agent can take a correction that way. Laban and colleagues tested this. They gave models an instruction over several turns instead of all at once. Scores dropped by 39 percent on average across six generation tasks.1 The authors split the drop into a small loss in skill and a large rise in unreliability. They found the cause. Models commit to an assumption early and then rely on it too much.

So your message at hour nine lands in a context that already holds nine hours of confident work in the wrong direction. Sinha and colleagues explain why that matters. When a model's context holds its own earlier errors, the model becomes more likely to make new mistakes. They call this self-conditioning.2 Your correction is one paragraph. The rest of the context was written by the model, and it points the other way.

Backlund and Petersson found the strongest form of this in long-running agents. In Vending-Bench, runs go past 20 million tokens. Agents fall into what the authors call tangential meltdown loops, and they rarely recover. The authors found no clear link between these failures and the context window filling up.3 So the runs that most need a correction are the runs least able to act on one sent as a suggestion.

People still need to correct long runs. The correction has to arrive through some channel other than one more chat turn.

Interruption has to be built into the loop#

The formal study of this question is older than language agents. Hadfield-Menell and colleagues modeled the off-switch as a game. An agent with a fixed goal cannot reach that goal while it is switched off. So a rational agent with a fixed goal has a reason to disable its off-switch.4 The authors showed what it takes for the agent to prefer keeping the switch. The agent has to be unsure how good the outcome is. It also has to treat the human's actions as evidence about that.

This result says nothing about what today's models want. It says where the ability to interrupt should live. Suppose the agent has to choose to check for messages. Then interruption is a behavior, and long runs tend to lose behaviors. Suppose instead the interrupt is part of the loop the agent runs inside. Then it works in any state the agent is in. So the control belongs in the harness, the program that runs the agent's loop, and not in the prompt.

Takerngsaksiri and colleagues built an applied version at Atlassian. Their system, HULA, puts software engineers in the loop of an LLM agent. The engineers refine and guide the coding plan and the source code, and they review each step. The authors report that HULA is deployed internally on Jira.5 The useful finding is where the human comes in. The human enters at set points in the agent's process. The human does not have to win an argument with the agent.

A steer is a message with a delivery mode#

In Oxagen, a correction you send to a running agent is called a steer. A steer is free text sent to a run. It enters the run at a model request. Nothing else in a run can receive text, and the surface is kept small on purpose. Oxagen never carries out a steer as a command. The steer is content that reaches the model. Each harness decides for itself whether to treat that content as an instruction.

A steer differs from a chat box in one way. The sender picks how soon the steer lands, and each choice has a stated cost.

Three ways a steer reaches a running agent

  1. Next stepThe default. The steer goes out with the next model request. Nothing already running is disturbed, and nothing extra is billed.
  2. Turn boundaryThe steer waits for the current turn to end. The agent finishes its current thought, then reads the steer.
  3. InterruptStops the response that is streaming. The partial tokens are billed. A pending tool call is dropped only if it can be undone.

Cost of delivery

The delivery modes an operator can choose. If a pending tool call cannot be undone, an interrupt waits for the next step instead. It does not drop a call whose effect has already happened.

Think about an interrupt that arrives while the agent is halfway through a payment, a deploy, or a delete. Stopping there does not undo the action. It leaves the action half done. So when the pending call cannot be undone, the interrupt waits for the next step boundary. This keeps the control from causing the incident it was meant to stop.

Each steer has a recorded status#

A chat box also gives you no status. A steer has one.

The states a steer moves through

  1. QueuedAccepted, waiting for its delivery point.
  2. SentSent to the run.
  3. ReceivedThe run has it.
  4. AcknowledgedIt reached a model request.
  5. AppliedThe model read it in a turn the record names.
The path a delivered steer takes. A steer can also be cancelled, expire, or fail. Oxagen records each of those as a state too.

The list of states is fixed, and Oxagen records every state. People usually guess whether a message got through. With a steer, they can look it up. The record names the model request each steer landed on. You can run a query to see whether anyone corrected a run before it went wrong. A steer that expired before delivery shows as expired. It does not look the same as a steer the model read and ignored.

Oxagen records each steer with the operator who sent it. The steer is stored as a frame, one entry in the run's record, next to the model requests and tool calls. This keeps what the operator said on the same timeline as what the agent did. That is the lasting part of what watching a run used to give you.

Operators and other agents send messages with different authority#

Once more than one agent is running, the channel for corrections can carry instructions from anyone who can reach it. Oxagen tells the two kinds of message apart by who sent them. It does not judge the content.

Two kinds of message that arrive through the same channel, with different authority

  • An operatorsteersEnters with operator authorityRecorded as a frame, with the sender's nameA running agent
  • Another agentmessagesEnters quoted and marked as untrusted (tainted)Evidence the model may weigh, never an instructionA running agent
One channel, two levels of authority. A message from another agent is content to consider. The record says where it came from.

With one agent running, this split costs nothing. With several, it becomes the main design choice. If one agent can instruct another, then someone can use the first agent to instruct the second. The text that does it will come from a web page, a ticket, or a file someone else wrote. So Oxagen marks where the message came from and does not promote it to an instruction. This does not filter the content. It records what kind of message it is. That decision keeps working even when the content is convincing.

What steering replaces#

The first post in this pillar described a loop: assign, watch, correct, review. Watching was not the useful part. It delivered two useful things. You could catch a wrong turn early. You could also say something the run would act on.

When a run lasts a week, both have to come from something other than a person's attention. Checks that run during the work, not after it, catch wrong turns early. A channel with a delivery mode, a status, and an owner carries messages the run will act on. Neither is as good as a senior engineer sitting beside the agent for a week. Both keep working while the agent runs for a week and the engineer does other work. Teams make this trade whether or not they write it down.

Where this fits in Oxagen#

Oxagen is workforce management for autonomous agents. Each agent works under a mandate. Steering sits in the mandate's equipment clause, beside the tools and skills the agent may use and the business context it may read. The rules that decide which requests stop and wait for a person sit in the budget and rules clause. Both are set before a run starts and applied while it runs.

For actions routed through Oxagen, this gives an operator three things at hour nine. First, a rule may already have stopped the request that started the rewrite and routed it to a named person. Second, a steer can reach the run at the next model request, at the next turn, or right away. Each choice has a stated cost, and an interrupt waits if the pending call cannot be undone. Third, the record holds the request, the rule that answered it, the steer, its sender, and the turn it landed on. So the next time this happens, you start from a query instead of nine hours of reading.

Footnotes#

  1. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). LLMs Get Lost In Multi-Turn Conversation. arXiv. https://arxiv.org/abs/2505.06120 ↩

  2. Sinha, A., Arun, A., Goel, S., Staab, S., & Geiping, J. (2025). The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs. arXiv. https://arxiv.org/abs/2509.09677 ↩

  3. Backlund, A., & Petersson, L. (2025). Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents. arXiv. https://arxiv.org/abs/2502.15840 ↩

  4. Hadfield-Menell, D., Dragan, A., Abbeel, P., & Russell, S. (2016). The Off-Switch Game. IJCAI 2017. https://arxiv.org/abs/1611.08219 ↩

  5. Takerngsaksiri, W., Pasuksmit, J., Thongtanunam, P., Tantithamthavorn, C., Zhang, R., Jiang, F., Li, J., Cook, E., Chen, K., & Wu, M. (2024). Human-In-the-Loop Software Development Agents. arXiv. https://arxiv.org/abs/2411.12924 ↩

  6. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798 ↩

Cite this

Oxagen Research. (2026, September 16). Steering a run you are not watching. oxagen.sh. https://oxagen.sh/blog/steering-a-run-you-are-not-watching

BibTeX
@online{anderson2026steeringarun,
  author  = {{Oxagen Research}},
  title   = {Steering a run you are not watching},
  year    = {2026},
  date    = {2026-09-16},
  url     = {https://oxagen.sh/blog/steering-a-run-you-are-not-watching}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.