Mac Anderson

AI agents, Autonomous agents

The problem is not slop, it is your process

The worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work.

Oxagen Research7 min readFirst published on oxagen.sh
View markdown

People who complain about AI slop are complaining about output. They mean low-quality pull requests, documentation nobody asked for, and long, plausible text with a mistake in the middle. The complaint is fair. But the published measurements do not find the cost in the output. So output quality is the wrong thing to organise a team around.

The measurements find the cost in the process around the output. Teams take a workflow built for people and give it to an agent. They keep the parts that only worked because a person was doing the work. The agent then fails in exactly the places where the workflow assumed a person.

In one study, every forecast was wrong in the same direction#

Becker and colleagues ran a randomised controlled trial with 16 experienced open-source developers across 246 tasks. Each developer worked on a repository where they had about five years of prior experience.1 Half the tasks allowed AI tools, and half did not.

Developers using AI took 19 percent longer to complete tasks.

The forecasts stand out. Before the study, the developers expected AI to cut completion time by 24 percent. Economists asked to predict the result said 39 percent. Machine-learning experts said 38 percent. After the tasks, the developers had been slower. They still estimated that AI had made them 20 percent faster.

Forecast change in task completion time, against the measured change

ForecastMeasured
Economists-39%19%
Machine-learning experts-38%19%
The developers, beforehand-24%19%
The developers, afterwards-20%19%
Source: Becker et al. (2025), 16 experienced open-source developers over 246 randomised tasks. Negative is faster. Every forecast was on the wrong side of zero, including the one made after the work.

None of the people in that chart describe slop. The developers accepted the model's output. They did not report being buried in bad output. The time went to other things. It went to prompting, reading, waiting, and small corrections that each seem minor but add up. The developers' own reports also stayed positive after the work. Teams usually decide whether a practice works by asking the people who do it. Here, asking would have given the opposite of the measurement.

Bigger batches made delivery worse#

DORA's 2024 report found the same pattern across whole organisations. In their survey population, a 25 percent increase in AI adoption was linked to an estimated 1.5 percent drop in delivery throughput. It was also linked to an estimated 7.2 percent drop in delivery stability. This was the second year in a row that AI adoption was linked to worse delivery performance.2 In the same survey, about three quarters of respondents reported productivity gains.

DORA points to batch size as the cause. AI makes code cheaper to write, so changesets get bigger. Bigger changesets carry more risk through a review and release process tuned for smaller ones. The generated code does not have to be bad for this to happen. The team only has to keep a process that was safe because batches were human-sized. Then AI removes the limit that kept batches that size.

Put simply, the old workflow had a hidden limit, which was how much one person could do. The agent removed that limit, and nobody replaced it.

Habits that work between people fail with agents#

Most of a working software process is unwritten. Most of the unwritten part is about how people handle vague requests. You ask a colleague for something vague. They ask two questions, and you agree on what you meant. You send them a long document and say the answer is in there, and they find it. Halfway through, they notice they have been building the wrong thing, and they back out. None of that is in the ticket, but the process depends on all of it.

Research has measured each of these habits with models.

Four habits a team hands to an agent, and what each one assumes

  • Clarify as we goassumes39 percent average drop, multi-turn against single-turnLaban et al. (2025)Recovery from a wrong turn
  • It is in the specassumesAccuracy is highest at the start and end, lowest in the middleLiu et al. (2023)Even attention across a long document
  • Double-check your workassumesPerformance can degrade after unaided self-correctionHuang et al. (2023)Useful self-review
  • You will notice the mistakeassumesErrors already in context raise the rate of later errorsSinha et al. (2025)Error does not become evidence
Each row pairs a habit that works between people with a published finding that the habit fails when handed to an agent.

The first row does the most damage. Laban and colleagues gave models a fully specified instruction all at once. They also gave the same instruction spread across a conversation. The conversation led to an average drop of 39 percent across six generation tasks.3 The authors split the drop into a small loss in skill and a large rise in unreliability. They describe the cause. Models make assumptions in early turns and try a final answer too soon. Then they rely on that answer too much. So the next message does not fix an early wrong turn. The model builds on it.

The other three findings add to the problem. Liu and colleagues showed that a model is most accurate on a long input when the relevant passage is near the start or the end. Accuracy drops for passages in the middle.4 So a key constraint on page nine of the spec may not reach the model. Huang and colleagues found that models struggle to correct their own reasoning without outside feedback. Self-correction without help sometimes made the answer worse.5 So asking the agent to check itself does not work as a check. Sinha and colleagues describe self-conditioning. When a model's own earlier errors sit in its context, it becomes more likely to make new errors.6 So the longer the run, the more the model treats an early mistake as fact.

Slop comes from a missing check#

Output quality still matters. It depends on the process that produces it.

A team gets slop when it has no check outside the model and no limit on batch size. Take a process whose only real check was a person reading everything. Remove the person as the limit, and the result is more output than anyone can evaluate. Most teams respond by reviewing harder. That means more review, by the same people, on more material. That is the part that already did not scale.

The measurements support a different response, and a duller one. Move the check off the model and off the reviewer's patience, and onto something that runs. Write the limit down, instead of leaving it to someone's attention. State the standing rules once, in a place the agent reads at the start of every run. Do not state them in the middle of a conversation the agent will mishandle. When the run is going wrong, change the run instead of arguing with it.

This pillar works through that response as four practices.

  1. State the standing rules once, in a lasting place. Do not put them in a prompt that is rewritten each time, or deep in a document the model reads unevenly. The rules an agent works under should be records with an owner, a scope, and a date. Then you can look up what the agent was told instead of piecing it together.
  2. Correct the run itself. A correction sent as another chat turn meets the same multi-turn failure it is trying to fix. So change what the run is operating under instead.
  3. Check with something outside the model. A test suite, a schema, a database state, or a named person can do this. The check has to return a verdict the agent cannot write itself.
  4. Record each step. When a week-long run goes wrong, you need to know which step went wrong. Either you query a record, or a person rereads the transcript.

Where this fits in Oxagen#

Oxagen is workforce management for autonomous agents. It holds the mandate each agent works under. The mandate covers the identity the agent acts as, the systems and data it may request, the budget and rules it works under, the tools and skills it is equipped with, and the record of what it did.

The first two practices above map to two of those clauses. The business context an agent may read and the steering it runs under belong to the equipment clause. The engineers accountable for the agent set it. Approval thresholds and decision rules belong to the budget and rules clause. The people who own the spend and the systems set it. Both are written before a run starts. Oxagen applies them to governed requests as the run makes them. A rule that answers at the moment of use works even if the agent did not read it carefully. It also works when no person is awake to restate it.

The next two posts cover the first two practices. One covers what an agent should be told before it starts. The other covers what to do when a run you are not watching starts going the wrong way.

Footnotes#

  1. Becker, J., Rush, N., Barnes, E., & Rein, D. (2025). Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv. https://arxiv.org/abs/2507.09089 ↩

  2. DORA (2024). Accelerate State of DevOps Report 2024. Google Cloud. https://dora.dev/research/2024/dora-report/ ↩

  3. Laban, P., Hayashi, H., Zhou, Y., & Neville, J. (2025). LLMs Get Lost In Multi-Turn Conversation. arXiv. https://arxiv.org/abs/2505.06120 ↩

  4. Liu, N. F., Lin, K., Hewitt, J., Paranjape, A., Bevilacqua, M., Petroni, F., & Liang, P. (2023). Lost in the Middle: How Language Models Use Long Contexts. TACL. https://arxiv.org/abs/2307.03172 ↩

  5. Huang, J., Chen, X., Mishra, S., Zheng, H. S., Yu, A. W., Song, X., & Zhou, D. (2023). Large Language Models Cannot Self-Correct Reasoning Yet. ICLR 2024. https://arxiv.org/abs/2310.01798 ↩

  6. Sinha, A., Arun, A., Goel, S., Staab, S., & Geiping, J. (2025). The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs. arXiv. https://arxiv.org/abs/2509.09677 ↩

Cite this

Oxagen Research. (2026, September 16). The problem is not slop, it is your process. oxagen.sh. https://oxagen.sh/blog/the-problem-is-not-slop-it-is-your-process

BibTeX
@online{anderson2026theproblemis,
  author  = {{Oxagen Research}},
  title   = {The problem is not slop, it is your process},
  year    = {2026},
  date    = {2026-09-16},
  url     = {https://oxagen.sh/blog/the-problem-is-not-slop-it-is-your-process}
}

Related research

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.