# Measuring agent intelligence

> Benchmarking on coding challenges misses what matters in production.

Part 13: Measurement. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/measuring-an-agent-in-production

Someone asks how good your agent is, and the number you have is a benchmark score. It tells you one thing about one distribution of tasks under lab conditions. It does not tell you what a task costs, how long it takes, how much collateral change it makes, or whether it will do the same thing twice. Benchmarks measure capability ceilings. Production needs efficiency, precision, and repeatability, and those take different instruments.

> **Evidence · Accuracy-only evaluation misleads**
>
> The Princeton team showed this formally. Without cost as a co-equal axis, leaderboards reward scientifically meaningless spend, including retry loops and complexity that buy accuracy at any price, and the community drew wrong conclusions about why agents improved. Their prescription, jointly optimizing on a **cost-accuracy Pareto frontier**, revealed simple designs matching complex ones at a fraction of the cost.[2](https://macanderson.com/manual/sources#r2 "Kapoor, Stroebl, Siegel, Nadgir & Narayanan (Princeton). \"AI Agents That Matter.\" TMLR 2025. arXiv:2407.01502") Agentless made the same point empirically by reporting *% resolved, average cost, and average tokens together*, and separately audited SWE-bench Lite itself, finding problem-quality issues (such as issues whose text leaks the solution) that inflate naive readings of benchmark scores.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") Anthropic's variance analysis completes the picture. If token usage explains ~80% of performance variance,[3](https://macanderson.com/manual/sources#r3 "Anthropic Engineering. \"How we built our multi-agent research system.\" June 2025. anthropic.com/engineering/multi-agent-research-system") then a score reported without token counts is mostly measuring budget, not intelligence.
>
> Kapoor et al., TMLR 2025 · Xia et al., FSE 2025 · Anthropic, 2025

### The production scorecard

##### files\_touched / lines\_changed

Surgical or scattered. Two agents both solve the ticket. One edits 2 files, the other rewrites 14. Blast radius is review cost, merge risk, and regression surface.

##### tokens\_consumed per task

The direct cost of thinking. Tracked per phase (retrieval, generation, repair), it tells you which part of the architecture is wasteful.

##### tool\_calls and graph\_queries

The exploration ratio. A high tool-call count with a low graph-query count means the agent is foraging probabilistically for what an index should hand it.

##### time\_to\_green

Wall clock from task start to passing CI. The metric people feel, and the one a latency-amplifying architecture (Part 9) quietly destroys.

##### test\_pass\_rate and first\_attempt\_rate

Not just eventually green but green in how many attempts. Each repair loop multiplies cost and proxies for context quality.

##### cost\_per\_completed\_task

The number finance cares about, inclusive of retries, abandoned runs, and human rework. Report it like cost of goods sold (COGS), because it is.

##### retrieval\_precision

Of the context supplied, how much did the final patch depend on? Intersect provided spans with edited and read spans. The best single health metric for your deterministic layer.

##### context\_utilization

The inverse view: tokens supplied against tokens that mattered. Utilization under 10% means you are paying attention-degradation costs (Part 7) for nothing.

##### variance / repeatability

Run the same task 10 times. Determinism-heavy systems cluster tightly on cost and outcome. Agent-heavy systems scatter. Variance is risk, and risk is cost.

These do not replace benchmarks. They are the instrument panel a benchmark cannot be. A capability score tells you which model to buy. The scorecard tells you whether your system is improving: whether this quarter's retrieval work moved precision from 8% to 40%, whether graph queries are displacing exploratory tool calls, whether cost per task is falling while pass rates hold. Optimizing these numbers is engineering. Optimizing a single leaderboard score tells you about the model you bought, not the system you built.

### The scorecard on one page

Copy this into your runbook. Each row is a metric, the computation behind it, and the failure it usually points at.

| Metric | How to compute it | What a bad reading usually means |
| --- | --- | --- |
| `files_touched` | Count distinct files in the final diff, per completed task | Retrieval is too broad, so the agent edits what it read rather than what the ticket named |
| `tokens_per_task` | Sum prompt plus completion tokens per run, tagged by phase | One phase dominates. Repair usually means bad context, retrieval usually means no budget |
| `graph_query_ratio` | Graph queries divided by graph queries plus exploratory file reads | The index is missing edges, so the agent forages instead of looking up |
| `time_to_green` | Wall clock from task start to the first passing CI run | Hand-offs between model calls dominate, which is the Part 9 tax |
| `first_attempt_rate` | Tasks green on attempt 1 divided by all completed tasks | The context bundle lacks the contract, so the model guesses the behavior |
| `retrieval_precision` | Supplied spans that the final patch read or edited, divided by all supplied spans | The retriever is shipping neighborhoods rather than slices |
| `cost_per_completed_task` | Total spend including retries and abandoned runs, divided by tasks merged | Abandoned runs are uncounted, and the true figure is higher than the one you quote |
| `outcome_variance` | Run one task 10 times, take the spread of cost and of pass or fail | Control flow is decided by inference, so the same input takes different paths |

Cost per task is the number that travels furthest outside engineering, and it raises a question this part does not answer: which agent, which run, and which person does a given dollar belong to. Part 17 takes that up.

> **Do this week**
>
> 1.  **Emit one span per run.** Wrap each agent run in an OpenTelemetry span carrying task id, model, prompt tokens, completion tokens, tool calls, and outcome. Done looks like a trace you can group by task class.
> 2.  **Log the bundle and the diff together.** Record the spans you supplied and the spans the patch read or edited, so `retrieval_precision` becomes an intersection rather than an estimate. Done looks like one precision number for last week.
> 3.  **Count abandoned runs in the cost.** Add halted, timed-out, and human-rewritten runs to the denominator of `cost_per_completed_task`. Done looks like a figure higher than the one you were quoting, and true.
> 4.  **Run a repeatability probe.** Pick three representative tasks, run each 10 times on the same commit, and record the spread of cost and outcome. Done looks like a variance number you can put next to your pass rate.
> 5.  **Publish the eight-row table weekly.** Post it where the team reads it, with last week's values beside this week's. Done looks like an argument about one row instead of an argument about the agent.

> **Measure it**
>
> -   `retrieval_precision`: supplied context spans the final patch depended on, divided by all supplied spans. Up is good, and it moves when your index improves rather than when the model changes.
> -   `cost_per_completed_task`: all spend, including retries and abandoned runs, divided by tasks merged. Down is good. Compute it weekly, because a monthly figure hides a bad week.
> -   `first_attempt_rate`: tasks green on the first attempt, divided by all completed tasks. Up is good, and each repair loop you remove pays twice, in tokens and in wall clock.
> -   `outcome_variance`: the spread of cost and pass or fail across 10 runs of one task on one commit. Down is good, and a wide spread tells you inference is deciding something code should decide.

> **Takeaway**
>
> A benchmark score measures the model's ceiling. Tokens, tool calls, blast radius, time to green, and retrieval precision measure your architecture. One of those is under your control, so instrument it.
