# Agents are waiting on a process built for people

> An agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour.

Published 2026-09-30 by Oxagen Research. Canonical page: https://oxagen.sh/blog/agents-are-waiting-on-a-process-built-for-people

An agent can write a change to your code in a few minutes. In most teams, that change then waits. It waits for a person to open it, read it, and say yes. At night, it waits until morning. On a Friday evening, it waits until Monday. The agent could start the next change, but the next change waits in the same line.

The agent is not the slow part. The process around it is. Most software processes were made for people, and they still expect a person at every step.

## The wait

A week has 168 hours. A 40-hour work week covers 40 of them. A process that needs a person at every step can only move in those 40 hours. For the other 128 hours, the work sits still.

The wait inside the work week is long too, and it has been measured for years. A 2018 study of code review at Google found a median wait of under an hour for first feedback on a small change and about 5 hours on the largest ones. The median for a whole review, all sizes together, was under 4 hours. The same paper lists the earlier figures it beat: a median time to approval of 17.5 hours at AMD, 15.7 hours for Chrome OS, and between 14.7 and 19.8 hours on three Microsoft projects.[1](#user-content-fn-1) Meta measured how long changes wait for a reviewer. In early 2021, the typical change waited a few hours. The slowest quarter of changes waited as much as a day.[2](#user-content-fn-2) Meta then built a tool that reminds reviewers about old changes. It cut the average wait by 7 percent. It cut the number of changes that waited more than three days by 12 percent.[2](#user-content-fn-2) Microsoft built the same kind of tool. In a trial across 147 repositories, its reminders cut the time to resolve 8,500 pull requests by 60 percent, and it went on to send 210,000 reminders a year across 8,000 repositories.[3](#user-content-fn-3) Those are real gains. They are gains on a wait counted in hours and days, for work an agent does in minutes.

The wait is not only for the first look. At Google, reviewers leave millions of comments a year, and an author spends about 60 minutes of active work on a change between sending it for review and submitting it.[4](#user-content-fn-4) Reading also has limits. A SmartBear study of a Cisco team found that a person reads code best in chunks of 200 to 400 lines. Past 400 lines, people find fewer problems. The study also advises reading for no more than 60 minutes at a time.[5](#user-content-fn-5) One agent task can produce more than 400 lines. Read well, that is about an hour of one person's time, for each task.

A change an agent wrote waits longer than one a person wrote. A 2026 benchmark over 8.1 million pull requests from 4,800 teams found that a pull request opened by an agent waited 5.3 times longer for a reviewer to pick it up than one written without AI. Once picked up, AI-written pull requests were reviewed twice as fast, and fewer of them were accepted: 32.7 percent against 84.4 percent.[6](#user-content-fn-6) A study of 33,596 agent-authored pull requests on GitHub found that 61 percent received no recorded review at all. Of the review comments the rest did get, 72 percent came from other agents.[7](#user-content-fn-7)

## Faster writing

The writing itself did get faster, in most measurements. In a 2023 trial, 95 programmers given an AI code completion tool finished a task 55.8 percent faster than those without one.[8](#user-content-fn-8) Three field experiments at Microsoft, Accenture, and a Fortune 100 company, covering 4,867 developers, found 26 percent more completed tasks.[9](#user-content-fn-9) A trial with 96 Google engineers found a gain of about 21 percent, with a wide confidence interval.[10](#user-content-fn-10) One trial found the opposite. Sixteen experienced open-source developers took 19 percent longer with AI on repositories they knew well.[11](#user-content-fn-11) An earlier post, [The problem is not slop, it is your process](/blog/the-problem-is-not-slop-it-is-your-process), looks at that result.

Agents are also taking on bigger jobs. METR measures the length of task an AI model can finish, counted in the time a skilled person would need. That length has doubled about every seven months since 2019.[12](#user-content-fn-12) Longer tasks make bigger changes, and bigger changes take longer to read.

## Bigger changes

DORA's 2024 report shows what happens when teams add AI and keep the old process. When a team's use of AI went up by a quarter, delivery throughput fell by an estimated 1.5 percent. Delivery stability fell by an estimated 7.2 percent.[13](#user-content-fn-13) The report names large batches as one cause. DORA's 2025 report, from a survey of nearly 5,000 technology professionals, found 90 percent of respondents using AI at work. This time, more AI went with higher throughput. It still went with lower stability. The report names what is missing: automated tests, mature version control practice, and fast feedback, without which more change brings more instability.[14](#user-content-fn-14)

A 2025 telemetry study shows where the time goes. Across more than 10,000 developers in 1,255 teams, developers using AI completed 21 percent more tasks and merged 98 percent more pull requests. The average pull request grew by 154 percent. Review time per pull request rose by 91 percent, and bugs per developer rose by 9 percent. At the company level, the study found no significant link between AI adoption and delivery metrics.[15](#user-content-fn-15) Each developer got faster. The process absorbed the gain.

## Process changes at eight companies

Some companies changed the process instead of waiting on it. Each change below moves a step that needed a person onto a machine, and keeps a person for the rule, the exception, or the final click.

### The fix beside the comment

Google trained a model on its own review history to turn a reviewer's comment into a suggested edit. The author sees the edit next to the comment and applies it with one click. In production, authors resolve 7.5 percent of all reviewer comments this way.[4](#user-content-fn-4) By mid 2024, AI assistance addressed more than 8 percent of review comments, and code completion wrote 50 percent of the code characters at Google, with engineers accepting 37 percent of its suggestions.[16](#user-content-fn-16) The reviewer still writes the comment. The author no longer writes the fix.

### Tests filtered before review

Meta built TestGen-LLM to extend existing unit tests. A generated test reaches an engineer only after it builds, passes reliably, and raises coverage. On Instagram's Reels and Stories, 75 percent of generated tests built, 57 percent passed reliably, and 25 percent raised coverage. Engineers accepted 73 percent of the tests that passed the filter for production.[17](#user-content-fn-17) The filter did the first review. The person did the last.

### Migrations as a loop

Airbnb moved about 3,500 React test files from one test framework to another. The estimate by hand was 1.5 years. The migration took six weeks. A pipeline sent each file through a set of steps. When a step failed, it sent the errors and the current file back to the model and tried again, up to a retry limit. Seventy-five percent of the files finished in four hours. Raising the retry limit, to between 50 and 100 attempts for the hardest files, took the total to 97 percent. People finished the last 3 percent.[18](#user-content-fn-18)

Amazon used its own tool to move production Java applications from Java 8 and 11 to Java 17. An upgrade that typically took 50 developer-days took a few hours. In under six months, more than half of Amazon's production Java systems moved. Developers shipped 79 percent of the generated changes without editing them. Amazon puts the saving at 4,500 developer-years and 260 million dollars a year.[19](#user-content-fn-19) Google ran 39 code migrations over twelve months with three developers. Of the 595 changes submitted, 74 percent were written by the model, and the developers put the time saved at half.[20](#user-content-fn-20) In each case the loop runs without a person, and a person reads the result.

### A machine reviewer on every change

Uber's uReview reads more than 90 percent of the roughly 65,000 changes its engineers post each week and comments on them before a person does. Engineers mark 75 percent of its comments as useful and address more than 65 percent of them.[21](#user-content-fn-21) A comment is not a decision. The person still decides.

Not every rollout saves time. One company added an automated reviewer across ten projects and 238 practitioners. Of 4,335 pull requests, 1,568 got an automated review. Engineers resolved 73.8 percent of the automated comments, and the average time to close a pull request rose from 5 hours 52 minutes to 8 hours 20 minutes.[22](#user-content-fn-22) A reviewer that only adds comments adds work. The gain comes when the review changes what a person has to read.

### Approval by rule

Google's large-scale change tooling has worked this way for years. A change that touches thousands of files is split by ownership into pieces that can land on their own. Each piece goes through its own test, mail, and submit pipeline. Reviewers for the whole change use pattern-based tools to approve the pieces that match what they expect, and read only the anomalies, such as a merge conflict. The pipeline lands more than 700 changes touching more than 15,000 files a day.[23](#user-content-fn-23)

Two companies now approve pull requests by rule. At Intercom, more than 93 percent of pull requests in the two main codebases are agent-driven. More than 19 percent are approved with no person in the loop, and time to approval at the 75th percentile improved by 6 to 16 times. Any engineer can ask for a human review on any change.[24](#user-content-fn-24) At Rewind, a review bot approves about 40 percent of pull requests. Approval is not merge. The bot posts an approval the way a person would and does not press the merge button, and a pull request an agent wrote cannot merge until a person clicks. The policy is one YAML file, which code owners must approve and which changes through the same review as any other code. A verdict engine of ordinary tested code, not a model, turns the findings into the decision.[25](#user-content-fn-25)

GitHub's coding agent shows the default without such a rule. The agent works in a GitHub Actions environment and pushes commits to a draft pull request. It opened more than a million pull requests between May and September 2025.[26](#user-content-fn-26) Each one waits for a person's approval before its automated checks even run.[27](#user-content-fn-27)

### The shared shape

In every case, a machine does the first pass, and a person does something smaller than reading every line: writes the rule, handles the exception, or clicks the final approval. The person is still in charge. The person is no longer in the line.

## A process made for agents

A process made for agents keeps people in charge and takes them out of the line. It has four parts.

-   **Rules.** People write the rules before the work starts. The rules say what an agent may do, what it may spend, and when it must ask a person.
-   **Checks.** Automatic checks build and test every change. Review agents read it and rate each problem they find. A serious problem blocks the change.
-   **Decisions.** A rule can send a request to a person: more budget, a new tool, or a change to production data. The run waits for that answer, and only for that answer.
-   **The record.** Every change leaves a record of what it did, which checks passed, and what it cost. A person reads the record each day instead of reading every line.

Each change then moves through ten stages. Agents run all ten. The figure follows one change, #17, through them.

> **Ten stages of the agent SDLC**
>
> 1.  **Issue**: The issue says what done looks like
> 2.  **Triage**: A triage agent sets priority and size
> 3.  **Claim**: One change, one writer
> 4.  **Build**: The change ships with its tests
> 5.  **Check**: Automatic checks build and test it
> 6.  **Review**: Review agents rate each finding
> 7.  **Merge**: Green checks and no serious finding
> 8.  **Release**: Database changes go first
> 9.  **Watch**: A broken main opens an issue
> 10.  **Learn**: Minutes planned against minutes spent
>
> Repeat Each run ends with a note that the next issue can use.
>
> *Schematic. One change, #17, moves through the ten stages and the loop back to the next issue. Each stage is described in full at sdlc.oxagen.sh.*

## Oxagen's own process

Oxagen builds its own product this way. Four agent tools do the work: Claude Code, Codex, Cursor, and Stella. They follow one set of written rules, kept in the code next to the product. These are some of the rules:

-   Every change starts as an issue, and the issue says what done looks like.
-   A triage agent sets each issue's priority and size. Whoever files the issue does not.
-   One change has one writer. An agent posts a claim before it writes, and the claim lasts 90 minutes.
-   Automatic checks are the only place the code is built and tested.
-   An agent looks at each open change every 60 seconds. It fixes failed checks, answers review comments, and clears conflicts.
-   When the main branch breaks, a check opens a top-priority issue. The fix goes straight to the main branch. The check closes the issue when the branch is green again.

From September 1 to 29, 2026, 1,064 changes merged into Oxagen's main branch.[28](#user-content-fn-28) Of those, 685 merged outside weekday work hours. That is 64 percent. 283 merged on a Saturday or a Sunday, and 289 merged between 10 pm and 6 am Pacific time.

> **Changes merged into Oxagen, September 1 to 29, 2026**
>
> | When the change merged | Changes |
> | --- | --- |
> | Weekdays, 9 am to 6 pm | 379 |
> | Weekdays, other hours | 402 |
> | Saturday and Sunday | 283 |
>
> *Source: GitHub search over Oxagen's repository, pull requests merged September 1 to 29, 2026, counted by UTC date. Hours are Pacific time.*

These numbers show when changes merged. They do not show how good each change was. The checks and the review findings answer that for each change, and the record keeps the answer.

Oxagen is workforce management for autonomous agents. In this process, Oxagen is where each agent gets its own identity and works under a mandate: its access, its budget and rules, and its tools and skills. For actions routed through Oxagen, a rule answers each request, and the record shows what each agent did and what it cost.

## A copy for your team

The full process is at [sdlc.oxagen.sh](https://sdlc.oxagen.sh). It has the ten stages, the rules, the checks, and a four-week plan to start. Type your company name below to get a copy with your name in it, ready to edit.

Company name

FormatWordGoogle DocsPDF

Google Docs opens the Word file. The page that opens shows you how.

## Limits

-   The merge counts come from one repository over 29 days. They describe when Oxagen's changes merged. They do not predict what another team will see.
-   The company figures come from each company's own engineering posts and papers. Each company measured its own code with its own tools, and none of the figures has been reproduced elsewhere.
-   The Google, Meta, and SmartBear review numbers describe people reading code that people wrote. They were measured before agents wrote much of the code.
-   The trial results measure how fast one person finishes a task. They do not measure delivery.
-   The pull request benchmark and the telemetry study come from vendors of engineering analytics, over the teams that use their products.
-   DORA reports links in survey data. A link is not proof of a cause.

## Footnotes

1.  Sadowski, C., Söderberg, E., Church, L., Sipko, M., & Bacchelli, A. (2018). *Modern Code Review: A Case Study at Google*. ICSE SEIP 2018. [https://research.google/pubs/modern-code-review-a-case-study-at-google/](https://research.google/pubs/modern-code-review-a-case-study-at-google/) [↩](#user-content-fnref-1)
    
2.  Riggs, P. (2022). *Improving code review time at Meta*. Engineering at Meta. [https://engineering.fb.com/2022/11/16/culture/meta-code-review-time-improving/](https://engineering.fb.com/2022/11/16/culture/meta-code-review-time-improving/) [↩](#user-content-fnref-2) [↩2](#user-content-fnref-2-2)
    
3.  Maddila, C., Upadhyaya, S. S., Bansal, C., Nagappan, N., Gousios, G., & van Deursen, A. (2022). *Nudge: Accelerating Overdue Pull Requests Towards Completion*. ACM TOSEM. [https://arxiv.org/abs/2011.12468](https://arxiv.org/abs/2011.12468) [↩](#user-content-fnref-3)
    
4.  Frömmgen, A., Austin, J., Choy, P., et al. (2024). *Resolving Code Review Comments with Machine Learning*. ICSE SEIP 2024. [https://research.google/pubs/resolving-code-review-comments-with-machine-learning/](https://research.google/pubs/resolving-code-review-comments-with-machine-learning/) [↩](#user-content-fnref-4) [↩2](#user-content-fnref-4-2)
    
5.  SmartBear. *Best practices for code review*, reporting a SmartBear study of a Cisco Systems programming team. [https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/](https://smartbear.com/learn/code-review/best-practices-for-peer-code-review/) [↩](#user-content-fnref-5)
    
6.  LinearB (2026). *2026 Software Engineering Benchmarks Report*, a study of 8.1 million pull requests from 4,800 engineering teams. [https://linearb.io/resources/engineering-benchmarks](https://linearb.io/resources/engineering-benchmarks) [↩](#user-content-fnref-6)
    
7.  Duma, K., Wróblewski, P., Bobińska, J., Winiarska, J., & Przymus, P. (2026). *These Aren't the Reviews You're Looking For: How Humans Review AI-Generated Pull Requests*. arXiv. [https://arxiv.org/abs/2605.02273](https://arxiv.org/abs/2605.02273) [↩](#user-content-fnref-7)
    
8.  Peng, S., Kalliamvakou, E., Cihon, P., & Demirer, M. (2023). *The Impact of AI on Developer Productivity*. arXiv. [https://arxiv.org/abs/2302.06590](https://arxiv.org/abs/2302.06590) [↩](#user-content-fnref-8)
    
9.  Cui, Z. K., Demirer, M., Jaffe, S., Musolff, L., Peng, S., & Salz, T. (2025). *The Effects of Generative AI on High-Skilled Work: Evidence from Three Field Experiments with Software Developers*. [https://economics.mit.edu/sites/default/files/inline-files/draft\_copilot\_experiments.pdf](https://economics.mit.edu/sites/default/files/inline-files/draft_copilot_experiments.pdf) [↩](#user-content-fnref-9)
    
10.  Paradis, E., Grey, K., Hoang, Q., et al. (2024). *How much does AI impact development speed? An enterprise-based randomized controlled trial*. arXiv. [https://arxiv.org/abs/2410.12944](https://arxiv.org/abs/2410.12944) [↩](#user-content-fnref-10)
     
11.  Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). *Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity*. arXiv. [https://arxiv.org/abs/2507.09089](https://arxiv.org/abs/2507.09089) [↩](#user-content-fnref-11)
     
12.  Kwa, T., West, B., Becker, J., et al. (2025). *Measuring AI Ability to Complete Long Software Tasks*. METR. [https://arxiv.org/abs/2503.14499](https://arxiv.org/abs/2503.14499) [↩](#user-content-fnref-12)
     
13.  DORA (2024). *Announcing the 2024 DORA report*. Google Cloud. [https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report](https://cloud.google.com/blog/products/devops-sre/announcing-the-2024-dora-report) [↩](#user-content-fnref-13)
     
14.  DORA (2025). *Announcing the 2025 DORA report*. Google Cloud. [https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report](https://cloud.google.com/blog/products/ai-machine-learning/announcing-the-2025-dora-report) [↩](#user-content-fnref-14)
     
15.  Faros AI (2025). *The AI Productivity Paradox Report 2025*, telemetry from more than 10,000 developers across 1,255 teams. [https://www.faros.ai/blog/ai-software-engineering](https://www.faros.ai/blog/ai-software-engineering) [↩](#user-content-fnref-15)
     
16.  Chandra, S., & Tabachnyk, M. (2024). *AI in software engineering at Google: Progress and the path ahead*. Google Research. [https://research.google/blog/ai-in-software-engineering-at-google-progress-and-the-path-ahead/](https://research.google/blog/ai-in-software-engineering-at-google-progress-and-the-path-ahead/) [↩](#user-content-fnref-16)
     
17.  Alshahwan, N., Chheda, J., Finegenova, A., et al. (2024). *Automated Unit Test Improvement using Large Language Models at Meta*. arXiv. [https://arxiv.org/abs/2402.09171](https://arxiv.org/abs/2402.09171) [↩](#user-content-fnref-17)
     
18.  Covey-Brandt, C. (2025). *Accelerating Large-Scale Test Migration with LLMs*. The Airbnb Tech Blog. [https://medium.com/airbnb-engineering/accelerating-large-scale-test-migration-with-llms-9565c208023b](https://medium.com/airbnb-engineering/accelerating-large-scale-test-migration-with-llms-9565c208023b) [↩](#user-content-fnref-18)
     
19.  Jassy, A. (2024). Post on the Amazon Q Developer Java upgrades, August 2024. [https://www.linkedin.com/posts/andy-jassy-8b1615\_one-of-the-most-tedious-but-critical-tasks-activity-7232374162185461760-AdSz/](https://www.linkedin.com/posts/andy-jassy-8b1615_one-of-the-most-tedious-but-critical-tasks-activity-7232374162185461760-AdSz/) See also Arisoy Cholkar, A. (2024). *Amazon Q Developer just reached a 260 million dollar milestone*. AWS DevOps Blog. [https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone](https://aws.amazon.com/blogs/devops/amazon-q-developer-just-reached-a-260-million-dollar-milestone) [↩](#user-content-fnref-19)
     
20.  Ziftci, C., Nikolov, S., Sjövall, A., Kim, B., Codecasa, D., & Kim, M. (2025). *Migrating Code At Scale With LLMs At Google*. arXiv. [https://arxiv.org/abs/2504.09691](https://arxiv.org/abs/2504.09691) [↩](#user-content-fnref-20)
     
21.  Mahajan, A., Roy Choudhary, S., Bond, M., Wang, Z., & Utture, A. (2025). *uReview*. Uber Engineering. [https://www.uber.com/blog/ureview/](https://www.uber.com/blog/ureview/) [↩](#user-content-fnref-21)
     
22.  Cihan, U., Haratian, V., İnce, A., et al. (2025). *Automated Code Review In Practice*. ICSE SEIP 2025. [https://arxiv.org/abs/2412.18531](https://arxiv.org/abs/2412.18531) [↩](#user-content-fnref-22)
     
23.  Wright, H. (2020). *Large-Scale Changes*, chapter 22 of *Software Engineering at Google*. O'Reilly. [https://abseil.io/resources/swe-book/html/ch22.html](https://abseil.io/resources/swe-book/html/ch22.html) [↩](#user-content-fnref-23)
     
24.  Mykhailov, K., & Young, N. (2026). *AI is approving our pull requests: Here's how we made it safe*. Intercom. [https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/](https://www.intercom.com/blog/ai-is-approving-our-pull-requests-heres-how-we-made-it-safe/) [↩](#user-content-fnref-24)
     
25.  North, D. (2026). *How we let AI approve pull requests (safely)*. Rewind. [https://rewind.com/blog/ai-approve-pull-requests-safely/](https://rewind.com/blog/ai-approve-pull-requests-safely/) [↩](#user-content-fnref-25)
     
26.  GitHub (2025). *Octoverse: A new developer joins GitHub every second as AI leads TypeScript to #1*. The GitHub Blog. [https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/](https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-github-every-second-as-ai-leads-typescript-to-1/) [↩](#user-content-fnref-26)
     
27.  Dohmke, T. (2025). *Meet the new coding agent*. The GitHub Blog. [https://github.blog/news-insights/product-news/github-copilot-meet-the-new-coding-agent/](https://github.blog/news-insights/product-news/github-copilot-meet-the-new-coding-agent/) [↩](#user-content-fnref-27)
     
28.  Oxagen (2026). GitHub search `repo:macanderson/oxagen is:pr is:merged merged:2026-09-01..2026-09-29`, run on September 30, 2026. It returned 1,064 pull requests. [↩](#user-content-fnref-28)
