Mac Anderson

AI agents, Autonomous agents

The agent time horizon is doubling

METR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.

Oxagen Research7 min readFirst published on oxagen.sh
View markdown

You give an agent a ticket and go get coffee. Ten minutes is about the right length for a run like that. The agent has time to do something useful, and you can read the whole transcript when you come back. Because runs are this short, most teams still treat an agent as a faster autocomplete. They watch it because watching costs little.

Runs are getting longer, and the measurements are public. When a run lasts a day, no person reads the whole transcript anymore. When a run lasts a week, the habits built around watching have already stopped working, whether or not anyone replaced them.

What METR measures#

Kwa and colleagues proposed a metric that turns a benchmark score into something an operator can use. It is called the 50 percent task-completion time horizon. Take the tasks a model finishes with a 50 percent success rate. The horizon is how long a human with the right domain skills takes on those tasks.1 The authors timed people on RE-Bench, HCAST, and 66 shorter tasks they wrote for the study. Then they scored models on the same set.

They found two things. When they wrote the paper, frontier models had a 50 percent horizon of about 50 minutes. The frontier horizon had also doubled about every seven months since 2019.

METR revised the dataset in January 2026.2 The task suite grew to 228 tasks. The number of tasks that take eight hours or more went from 14 to 31. The revised trend puts the doubling time at 196.5 days across the whole period. For models released since 2023, it is 130.8 days. The highest measured 50 percent horizon in that release was 320 minutes. Its confidence interval runs from 170 to 729 minutes.

METR states the limits of these numbers plainly. Human baseline times were measured for only 5 of the 31 long tasks. The rest are estimates. The confidence intervals are wide. The public tracker notes that measurements above 16 hours are unreliable with the current task suite.3 That upper limit matters most. No one has a reliable measurement of an agent that works for a week, because no one has a task suite that can grade one.

What the doubling times add up to#

A doubling time is a small number, and it is easy to misjudge in planning. Three doublings sounds small, but it is a factor of eight.

Doublings from today's highest measured 50 percent horizon

  1. About 5 hoursThe highest measured 50 percent horizon in January 2026 (320 minutes).
  2. A working dayLess than one doubling away. About 7 months at 196 days, about 4 at 131.
  3. A working weekAbout three doublings away. About 19 months at 196 days, about 13 at 131.
  4. A month of workAbout five doublings away. About 32 months at 196 days, about 22 at 131.

Length of task

Arithmetic on the doubling times METR reports (Time Horizon 1.1, January 2026), starting from its highest measured point. For illustration only. A trend line is not a forecast, and METR states that its measurements above 16 hours are unreliable.

Read the table as arithmetic, not as a roadmap. A trend that held for seven years can stop in the eighth. The confidence intervals on the recent points are wide enough to change every row. The problem of grading long tasks will also arrive before agents can do them. The table helps with one thing. It shows how much of your current practice depends on runs staying short. If most of it does, your plans look less far ahead than the capability trend does.

Production needs more than a 50 percent success rate#

If you ship the result, half is an odd place to draw the line. It is the right place for a research metric. That part of the curve is the steepest, so it shows a change in the model most clearly. It is the wrong place for a deployment decision. There, you are deciding whether to let a run finish while you sleep.

Yao and colleagues made the same point with a different metric. The usual metric is pass@k, the chance that at least one of k attempts succeeds. They define pass^k instead, the chance that all k independent attempts at a task succeed.4 In their study, the best function-calling agents succeeded on under 50 percent of tasks. In the retail domain, pass^8 fell below 25 percent. More attempts raise pass@k and lower pass^k. So a system tuned on pass@k can get less consistent while its headline score improves.

A five-hour horizon at 50 percent success means the agent gets a five-hour task right about half the time. Run one such task a day for a working week. The chance that all five succeed is about 3 percent. This is not a reason to avoid long runs. It shows that the horizon number measures what an agent can reach. It does not measure what you can ship, and those are separate questions.

Long runs fail differently from short runs#

You might expect long runs to fail because the context window fills up. The measurements do not support that.

Backlund and Petersson built Vending-Bench to test whether an agent stays coherent over time, rather than what it can do. In the benchmark, an agent runs a vending machine business. It manages inventory, places orders, sets prices, and pays daily costs.5 Each task is simple, but the run is long. Single runs go past 20 million tokens. The agents misread delivery schedules and forget orders they placed. They also fall into what the authors call tangential meltdown loops, and they rarely recover. The authors found no clear link between failures and the point where the context window fills. Runs of the same model on the same task also varied a lot.

Sinha and colleagues describe a cause that fits. They studied long tasks and found a self-conditioning effect. When a model's context already holds its own errors from earlier turns, the model becomes more likely to make new mistakes.6 The agent reads its own bad work from earlier and treats it as settled. The same paper also has good news. Small gains in single-step accuracy add up to large gains in how long a task a model can complete. That is why the horizon keeps growing.

Together, these findings show how a long run goes wrong. It does not get slowly worse toward the end. It works correctly, then takes one wrong turn. It writes that turn into its own record. Then it spends hours staying consistent with it. By the time a person looks, the wrong turn is forty steps back, and everything after it fits with it.

Watching does not scale to long runs#

Most teams supervise an agent the way they supervise a new engineer. That is the process they already had.

The supervision loop most teams are running

  1. AssignA ticket, a prompt, a branch.
  2. WatchRead the transcript as it streams.
  3. CorrectInterrupt, re-prompt, restart.
  4. ReviewRead the diff, approve or reject.

Repeat Repeat per task, while someone is awake

Illustrative. The loop works because a person is there for the middle two steps. When a run takes a week, nobody is.

The loop works because a person is there for the middle two steps. Everything the team relies on comes from that. The person catches the wrong turn early. The correction is cheap. The reviewer has watched enough of the run to know what the diff means. None of this holds when a run is longer than the reviewer can sit through. The loop does not fail in an obvious way. It keeps going with the watching step skipped. The reviewer then reads a week of work with no idea which forty steps mattered.

Instead of watching, you decide in advance what you used to decide in the moment. You decide what the agent may reach, what it may spend, what it is told before it starts, and which requests stop and wait for a person. These are the same decisions, moved from the transcript to the mandate. The next posts in this pillar cover them one at a time.

Where this fits in Oxagen#

Oxagen is workforce management for autonomous agents. It does not run them. It holds each agent's mandate. The mandate covers the identity the agent acts as, the systems and data it may request, the budget and rules it works under, the tools and skills it is equipped with, and the record of what it did.

Three parts of the mandate matter most for longer runs. First, Oxagen answers each request at the moment of use. A run that asks for a system at hour forty gets the same rule it would have got at minute one. A rule can allow the request, deny it, or route it to a named person while the run waits. Second, Oxagen prices each governed action. It records the cost against the person, the agent, the run, the turn, and the step. So a week-long run has a cost you can read step by step, not one line on a monthly bill. Third, the record keeps each run as a series of frames, one per recorded event, next to its mandate. Finding the wrong turn becomes a query instead of a reread.

This applies to actions routed through Oxagen. Oxagen does not govern a call that does not pass through it. No record makes an agent correct. The record lets a person inspect a long run afterward. The watching step used to give you that.

Footnotes#

  1. Kwa, T., West, B., Becker, J., Deng, A., Garcia, K., Hasin, M., Jawhar, S., Kinniment, M., Rush, N., Von Arx, S., Bloom, R., Broadley, T., Du, H., Goodrich, B., Jurkovic, N., Miles, L. H., Nix, S., Lin, T., Painter, C., Parikh, N., Rein, D., Sato, L. J. K., Wijk, H., Ziegler, D. M., Barnes, E., & Chan, L. (2025). Measuring AI Ability to Complete Long Tasks. arXiv. https://arxiv.org/abs/2503.14499 ↩

  2. METR (2026). Time Horizon 1.1. https://metr.org/blog/2026-1-29-time-horizon-1-1/ ↩

  3. METR. Task-Completion Time Horizons of Frontier AI Models. https://metr.org/time-horizons/ ↩

  4. Yao, S., Shinn, N., Razavi, P., & Narasimhan, K. (2024). tau-bench: A Benchmark for Tool-Agent-User Interaction in Real-World Domains. arXiv. https://arxiv.org/abs/2406.12045 ↩

  5. Backlund, A., & Petersson, L. (2025). Vending-Bench: A Benchmark for Long-Term Coherence of Autonomous Agents. arXiv. https://arxiv.org/abs/2502.15840 ↩

  6. Sinha, A., Arun, A., Goel, S., Staab, S., & Geiping, J. (2025). The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs. arXiv. https://arxiv.org/abs/2509.09677 ↩

Cite this

Oxagen Research. (2026, September 16). The agent time horizon is doubling. oxagen.sh. https://oxagen.sh/blog/the-agent-time-horizon-is-doubling

BibTeX
@online{anderson2026theagenttime,
  author  = {{Oxagen Research}},
  title   = {The agent time horizon is doubling},
  year    = {2026},
  date    = {2026-09-16},
  url     = {https://oxagen.sh/blog/the-agent-time-horizon-is-doubling}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.