Mac Anderson

Part 9: Orchestration and memory

Why coordinator agents don't scale

Most multi-agent systems pay a coordination tax and call it architecture.

Mac Anderson5 min read1,063 words3 sources cited
View markdown

You split the work. A manager agent decomposes the ticket, three specialists execute, and the manager stitches the results back together. Then the token bill climbs, the wall clock stretches, and when the patch comes back wrong you read four transcripts to find which hand-off dropped the constraint. The org chart is not an architecture.

The tax, itemized#

  • Token duplication. Each sub-agent needs enough context to act, so you re-send and re-bill the shared background once per agent, then re-serialize every result back through the coordinator's window.
  • Latency amplification. Coordinator → worker → coordinator turns serialize model calls. A hand-off costs seconds, not microseconds.
  • Lossy hand-offs. Each hand-off compresses intent into natural language. Errors compound across hops.
  • Conflicts found at merge time. Parallel agents cannot see each other's decisions. You pay for both branches, then pay a model call to adjudicate the collision.
  • Fan-out with no budget. A misbehaving branch multiplies cost recursively, and an instruction written in prose does not stop it.
Evidence · Three independent sources, one conclusion

Anthropic runs multi-agent systems in production, so it is not a hostile witness. Its report: multi-agent systems consume ~15× the tokens of chat, are economical only where task value justifies it, and are explicitly "not a good fit" for domains with tight dependencies between agents. Their example of a poor fit is coding.3

Cognition (Devin) published "Don't Build Multi-Agents." In their production experience, parallel sub-agents with fragmented context make conflicting assumptions, and reliability comes from a single agent with continuous, fully-shared context. Context engineering, not agent proliferation.4

Berkeley (Cemri et al., NeurIPS 2025) annotated 1,600+ traces from 7 popular multi-agent frameworks. Performance gains over single agents were often minimal. The team catalogued 14 recurring failure modes in 3 classes: specification failures, inter-agent misalignment, and verification failures. Most stem from system design, not model capability, and better models will not fix them.5

Anthropic engineering, 2025 · Cognition, 2025 · Cemri et al., MAST (the Berkeley failure taxonomy), NeurIPS 2025

Deterministic orchestration, the boring alternative#

Look at what the coordinator agent does all day: routing, sequencing, retrying, aggregating. None of those four needs judgment. All four have mature, deterministic machinery.

◌ Coordinator agent

  • Control flow decided per run by inference, so it differs every time
  • State lives in a context window, and a crash loses it
  • Retries, timeouts, and budgets improvised in prose
  • Debugging means reading several transcripts and reconstructing the order

● State machine / workflow engine

  • Control flow is code: explicit states, typed transitions, reviewable in a PR
  • State is durable and replayable (Temporal-class engines, LangGraph-style graphs)
  • Retries, timeouts, budgets, and circuit breakers are first-class primitives
  • Debugging means inspecting a state machine's history

The pattern that works is deterministic skeleton, probabilistic muscles. A workflow engine owns the graph, localize → retrieve → patch → test → review, and calls the model only inside the nodes where judgment is required. Event-driven orchestration handles fan-out where sub-tasks are independent, the one regime where even Anthropic's data says parallelism pays.3 Where sub-agents do earn their keep, the framing to keep is Cognition's: the reliable ones are mostly read-only context gatherers, closer to expensive tool calls than to collaborating colleagues.4

Written down, the skeleton fits on a page. Every line below is code you can read in a diff, and the budget is a number the engine enforces rather than a sentence the model may ignore.

workflow fix_ticket:
  budget tokens=120_000  wall_clock=8m  attempts=3

  localize : deterministic   # parser and code graph, no model call
  retrieve : deterministic   # graph traversal, sliced to budget
  patch    : model           # judgment lives here
  test     : deterministic   # test runner, structured failure
  review   : model           # judgment lives here

  on test.fail   → retry patch with the sliced failure, up to attempts
  on budget.out  → halt, persist state, page a person

That leaves one question worth asking per node: does this step need a model at all?

Use a workflow node whenUse a sub-agent call when
The step has a decidable rule (route by file type, retry on exit code 1)The step needs judgment over ambiguous text
The step's output is typed and checkableThe output is prose a person or a test will grade
Later steps depend on this one's exact resultThe sub-task is independent and read-only
You want the same answer twiceYou can afford variance and will verify the result
Do this week
  1. Draw the graph your coordinator improvises. Take ten recent multi-agent runs, write down the node sequence each one actually followed, and count how many distinct sequences you get. More than two or three for the same task class tells you the routing is inference, not design.
  2. Move routing into code. Replace the manager's decomposition prompt with a function that maps task type to a fixed node list. Done looks like a run whose node sequence you can predict before you start it.
  3. Port one task class to a workflow engine. Pick a durable engine (Temporal-class) or a graph library (LangGraph-style) and encode localize, retrieve, patch, test, and review as typed states. Done looks like a replayable run history you can open after a crash.
  4. Put the budget in the engine. Set a token ceiling, a wall-clock ceiling, and an attempt ceiling per run, and make the engine halt on breach. Done looks like a halted run with persisted state, not a surprise invoice.
  5. Demote the survivors. Any sub-agent still standing should be read-only. If it writes files or decides control flow, it is a node in disguise.
Measure it
  • coordination_token_share: tokens spent on hand-off messages and re-sent background, divided by total tokens for the run. Tag each model call with its node, then sum the routing and aggregation nodes. Down is good.
  • handoff_count: model calls whose input is mostly another model's output. Count them per completed task. Down is good, and a run with zero is a run with no telephone game in it.
  • path_entropy: the number of distinct node sequences observed across runs of the same task class. Log the sequence on every run. Down is good, and 1 means the control flow is code.
  • budget_halt_rate: runs the engine stopped on a ceiling, divided by all runs. You want this small and nonzero. Zero usually means no ceiling is set.
Takeaway

Coordination is a solved, deterministic problem: state machines, workflow engines, queues. Spend inference on judgment inside the nodes, and keep it off the edges between them.

Cite this

Anderson, M. (2026). Why coordinator agents don't scale. In Engineering Deterministic AI Coding Agents (2nd ed., Part 9). Oxagen Inc. https://macanderson.com/manual/why-coordinator-agents-do-not-scale

BibTeX
@incollection{anderson2026whycoordinatoragents,
  author    = {Anderson, Mac},
  title     = {Why coordinator agents don't scale},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  chapter   = {9},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/why-coordinator-agents-do-not-scale}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.