Mac Anderson

Part 17: Operating the workforce

Spend you can attribute

A total is not an answer. The answer is which agent spent what, and on whose behalf.

Mac Anderson5 min read1,136 words4 sources cited
View markdown

The model invoice arrives and the number is larger than last month. You can see the total by API key and by day. The question from finance is different: which agent spent it, for which person, on which task, under which budget? Can you attribute last month's agent spend to the agent, the run, and the person who started the run? If the honest answer is an estimate built from a spreadsheet, the data to answer it was never recorded.

Why the invoice cannot answer#

A provider bills the account that made the call. It has no field for your agent, your run, or your colleague, because you never sent them. Attribution cannot be reconstructed afterward from a total. It is recorded at the call, by the layer the call passes through, or it does not exist.

Part 13 made cost per completed task a first-class metric, and the Princeton work behind it showed why accuracy reported without cost misleads.2 That chapter measured the architecture. This one measures the organization: the same rows, with four more keys on each.

The meter is a row#

Write one row per priced step. A step is one model call or one tool call. Each row carries the keys that every later question groups by.

ColumnWhy it is there
personWho started the task. The "on whose behalf" column.
agentThe identity from part 14.
run, turn, stepWhere in the work the cost fell. A run is one agent on one task. A turn is one prompt through to the point the agent stops. A step is one call.
model, input_tokens, cached_tokens, output_tokensThe quantities the price applies to. Cached reads price differently, so they need their own column.16
price_versionWhich price list was applied. Prices change, and a rate negotiated in March should not reprice February.
cost_usd, cost_basisThe amount, and whether it was measured, reported, or estimated.
rule_idThe rule that answered the request behind this step, from part 16.

The cost_basis column carries more weight than it looks. Measured means your layer saw the tokens and applied the price. Reported means the harness or the provider told you a figure. Estimated means you inferred it. A finance lead can work with all three as long as each one is labelled. A blended number with no label is the one that fails review.

Budgets belong beside the agent#

A budget in a cloud console alerts someone after the money is spent, and it alerts on the account, not the agent. A budget in the mandate is checked before the next step and names the agent it stops. Three settings are enough to start:

  • A monthly limit per agent, in currency, set by the person who owns the number.
  • A per-run ceiling, so one run that loops cannot spend the month. Part 9 called this a circuit breaker, and this is where it gets its number.
  • An at_limit action: hold the agent and notify the operator, route the next step to a person, or continue and flag.

The finance lead and the operator then read the same rows. One reads them grouped by cost center, the other grouped by run.

Evidence · Allocation comes before optimization

The FinOps Foundation's framework treats allocation as a core capability: assign cost to the teams and workloads that incur it, using consistent metadata, so that owners are accountable for their own usage.24 The order matters. A saving is hard to verify when nobody can say whose spend fell. Anthropic's production report gives agent teams a reason to care about the same order: on their internal evaluations, token usage alone explained about 80% of performance variance.3 Spend is a description of what the architecture did. Attributed spend says which part did it.

FinOps Foundation, FinOps Framework · Anthropic engineering, 2025

A worked example#

One agent, one week, read three ways from the same rows:

Grouped byWhat it showsWho acts
agentagt_refund_triage spent $212 of a $400 limitThe operator: on pace, no change
runTwo runs cost $61 together. The median run costs $1.90.The engineer: both runs looped on a failing migration, so set a per-run ceiling and fix the schema slice from part 8
personOne person started 70% of the runsFinance: charge the payments cost center, not platform

The figures are illustrative. The point is that three people took three different actions from one table, and none of them needed the invoice.

Counterweight · What this does not claim

Attribution is not savings. Recording who spent what does not lower the bill. It shows where the bill comes from, which is the precondition for any change you can verify afterward. If you later claim a reduction, state the workload, the baseline, the model and price versions, the sample size, the quality measure, and the costs including retries. Without those, the claim you can support is visibility.

Where Oxagen fits

In Oxagen every governed action is priced and attributed to the person, the agent, the run, the turn, and the step. The Spend page and the Run page read the same rows, and measured, reported, and estimated costs are labelled as such. The budget an agent runs under, what it has spent this month, and the rule that stops it sit beside the agent. This covers calls that pass through Oxagen. A harness whose model calls do not pass through the gateway reports its spend, or is not metered, and the page says which.

Do this week
  1. Add four keys to your model-call log. person, agent, run, and step. If you proxy model calls, add them there. If you do not, add them in the wrapper that starts the run.
  2. Separate cached input tokens. Log them in their own column and price them at the cached rate.
  3. Version your price list. A small table keyed by model and effective date. Stamp each row with the version used.
  4. Label the cost basis. Go through each source of cost data and mark it measured, reported, or estimated.
  5. Set one per-run ceiling. Pick the agent with the widest spread between median and maximum run cost. Set the ceiling at ten times the median and route to the operator when a run reaches it.
Measure it
  • attributed_spend_ratio: spend carrying person, agent, and run, over total model spend on the invoice. The gap is what you cannot explain.
  • cost_per_run_p50 and p99, per agent. A wide spread points at loops, and part 9 and part 11 are where to look.
  • unpriced_steps: steps with tokens and no price, usually a new model missing from the price list.
  • budget_holds: times an at_limit action fired, per agent, per month.
Takeaway

Record the person, the agent, the run, and the step on every priced call, label how each cost was obtained, and keep the budget beside the agent it limits. Then the question of who spent what has a query for an answer.

Cite this

Anderson, M. (2026). Spend you can attribute. In Engineering Deterministic AI Coding Agents (2nd ed., Part 17). Oxagen Inc. https://macanderson.com/manual/spend-you-can-attribute

BibTeX
@incollection{anderson2026spendyoucan,
  author    = {Anderson, Mac},
  title     = {Spend you can attribute},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  chapter   = {17},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/spend-you-can-attribute}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.