# AI agents are not expensive. Bad architecture is.

> Most of an agent's token bill comes from the system around the model, not from the model.

Part 1: The economics. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/agents-are-not-expensive-bad-architecture-is

Your agent works, and the invoice is bigger than the team that uses it. The first suspects are the obvious ones: the per-token price, the reasoning overhead, the size of the context window. Then you instrument the agent, log every request, and look at where the tokens went. The model did what it was asked to do. The system asked it to do far too much.

### Measure where tokens actually go

Run an agent framework against a real repository and log every request. In the runs this book draws on, the dominant cost is not *reasoning about the problem*. It is *rebuilding context the system already had*: re-reading files it read two turns ago, carrying 400 lines of stack trace when 6 frames mattered, pushing the whole conversation history into every tool call, and walking the filesystem with `ls` and `cat` because nothing indexed the repository up front.

The published numbers point the same way. Anthropic's engineering team, describing their production multi-agent research system, reported that agents consume roughly **4× the tokens of a chat interaction**, and multi-agent systems roughly **15×**, and that on their internal evaluations **token usage alone explained about 80% of performance variance**.[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") The largest lever they found was not model choice or prompt wording. It was how many tokens the architecture decided to spend.

> **Evidence · Simplicity wins the leaderboard**
>
> The clearest data point in this book is **Agentless** (Xia et al., UIUC). Instead of an autonomous agent choosing its own actions with tools, Agentless runs a fixed three-phase pipeline: hierarchical localization, patch generation, validation. On SWE-bench Lite it became the top open-source approach of its time at **27.3% solved for $0.34 per issue** (later 32.0% at $0.70), while contemporary agent-based systems spent roughly 10× more per issue for comparable or worse results.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") The model did not choose the next action. The pipeline chose, and the model filled in the parts that needed judgment.
>
> Xia et al., "Agentless: Demystifying LLM-based Software Engineering Agents," FSE 2025

### Context construction against reasoning cost

Split your agent's token bill into two buckets. **Reasoning tokens** are the model thinking about the actual problem: the diff, the root cause, the design decision. **Context-construction tokens** are everything spent getting the model to the point where reasoning can begin: file dumps, search results, log output, retries after the model lost the thread. In an unindexed architecture the second bucket dominates, often by an order of magnitude. Most of what sits in that second bucket is computable without a model. Parsing, indexing, slicing, and graph traversal are deterministic operations. They run on CPU, return the same answer on the same input, and you can test them like any other code.

You cannot split the bill you do not record, so record it at the call site. The caller knows why it is making the call, and the caller is the only part of the system that does.

```
# One row per model call, written when the call returns.
log_call(
    task_id=task.id,
    purpose="context",                  # or "reasoning", set by the caller
    source="file_read:billing/processors.py",
    input_tokens=usage.input,
    cached_input_tokens=usage.cache_read,
    output_tokens=usage.output,
)
# context_token_share = input where purpose == "context" / all input
# Group by source to rank what is filling the window.
```

> **Evidence · Cost must be a first-class metric**
>
> Princeton's "AI Agents That Matter" (Kapoor, Stroebl, Siegel, Nadgir, Narayanan) showed that agent research had been ignoring cost, and that when you evaluate on a joint cost and accuracy Pareto frontier, simple baselines such as retrying a model call matched or beat elaborate agent architectures on HumanEval at a fraction of the price. Their conclusion: accuracy alone cannot identify progress, and state of the art (SOTA) agents are frequently more complex and more costly than the task requires.[2](https://macanderson.com/manual/sources#r2 "Kapoor, Stroebl, Siegel, Nadgir & Narayanan (Princeton). \"AI Agents That Matter.\" TMLR 2025. arXiv:2407.01502")
>
> Kapoor et al., "AI Agents That Matter," TMLR 2025

### Thinking harder against needing less thinking

The default answer to agent failure is to think harder: a bigger model, a longer reasoning chain, more retries, more sub-agents. Each of those multiplies token spend and adds another probabilistic step, which is another place for variance to enter. The engineering answer runs the other way. Shrink the problem before the model sees it. Every ambiguity you resolve deterministically is reasoning the model no longer has to do, tokens you no longer pay for, and variance you no longer ship. Determinism is not a philosophical preference. It is a cost model.

Attributing that spend across the teams who run the agents is part 17.

> **Do this week**
>
> 1.  **Log usage on every model call.** Write one row per call with `task_id`, `purpose`, `source`, input tokens, cached input tokens, and output tokens. The artifact is a table you can group by, not a log line you can read.
> 2.  **Tag each call reasoning or context.** Set the tag at the call site, where the caller knows why the call is happening. Done looks like every call carrying a tag and none defaulting to unknown.
> 3.  **Rank your context sources.** Group last week's rows by `source` and sum input tokens. The top three sources are your work queue for parts 2 to 4.
> 4.  **Turn on prompt caching for the stable prefix.** Put the system prompt, tool definitions, and repository map ahead of anything that changes per turn, then watch the cache-read share rise.
> 5.  **Put a token budget in code.** Give each task a ceiling on input tokens and fail the run when it trips, so a runaway loop ends in a test failure instead of an invoice.

> **Measure it**
>
> -   `context_token_share`: context-construction input tokens divided by all input tokens, per task. Down is good. Above 0.8 means the model is paying to rediscover what your code already knows.
> -   `tokens_per_completed_task`: input plus output tokens divided by tasks that passed their check. Down is good, and it is the only token metric that survives a change in task mix.
> -   `cache_read_share`: cached input tokens divided by all input tokens. Up is good. A falling share usually means something volatile moved into the prefix.
> -   `cost_per_completed_task`: dollars divided by tasks that passed. Down is good, and it is the number to put next to accuracy when you compare two architectures.

> **Takeaway**
>
> Reduce uncertainty before you invoke intelligence. Every token of context your system constructs deterministically is a token the model does not spend rediscovering it probabilistically.
