Part 17: Operating the workforce
Spend you can attribute
A total is not an answer. The answer is which agent spent what, and on whose behalf.
The model invoice arrives and the number is larger than last month. You can see the total by API key and by day. The question from finance is different: which agent spent it, for which person, on which task, under which budget? Can you attribute last month's agent spend to the agent, the run, and the person who started the run? If the honest answer is an estimate built from a spreadsheet, the data to answer it was never recorded.
Why the invoice cannot answer#
A provider bills the account that made the call. It has no field for your agent, your run, or your colleague, because you never sent them. Attribution cannot be reconstructed afterward from a total. It is recorded at the call, by the layer the call passes through, or it does not exist.
Part 13 made cost per completed task a first-class metric, and the Princeton work behind it showed why accuracy reported without cost misleads.2 That chapter measured the architecture. This one measures the organization: the same rows, with four more keys on each.
The meter is a row#
Write one row per priced step. A step is one model call or one tool call. Each row carries the keys that every later question groups by.
| Column | Why it is there |
|---|---|
person | Who started the task. The "on whose behalf" column. |
agent | The identity from part 14. |
run, turn, step | Where in the work the cost fell. A run is one agent on one task. A turn is one prompt through to the point the agent stops. A step is one call. |
model, input_tokens, cached_tokens, output_tokens | The quantities the price applies to. Cached reads price differently, so they need their own column.16 |
price_version | Which price list was applied. Prices change, and a rate negotiated in March should not reprice February. |
cost_usd, cost_basis | The amount, and whether it was measured, reported, or estimated. |
rule_id | The rule that answered the request behind this step, from part 16. |
The cost_basis column carries more weight than it looks. Measured means your layer saw the tokens and applied the price. Reported means the harness or the provider told you a figure. Estimated means you inferred it. A finance lead can work with all three as long as each one is labelled. A blended number with no label is the one that fails review.
Budgets belong beside the agent#
A budget in a cloud console alerts someone after the money is spent, and it alerts on the account, not the agent. A budget in the mandate is checked before the next step and names the agent it stops. Three settings are enough to start:
- A monthly limit per agent, in currency, set by the person who owns the number.
- A per-run ceiling, so one run that loops cannot spend the month. Part 9 called this a circuit breaker, and this is where it gets its number.
- An
at_limitaction: hold the agent and notify the operator, route the next step to a person, or continue and flag.
The finance lead and the operator then read the same rows. One reads them grouped by cost center, the other grouped by run.
The FinOps Foundation's framework treats allocation as a core capability: assign cost to the teams and workloads that incur it, using consistent metadata, so that owners are accountable for their own usage.24 The order matters. A saving is hard to verify when nobody can say whose spend fell. Anthropic's production report gives agent teams a reason to care about the same order: on their internal evaluations, token usage alone explained about 80% of performance variance.3 Spend is a description of what the architecture did. Attributed spend says which part did it.
A worked example#
One agent, one week, read three ways from the same rows:
| Grouped by | What it shows | Who acts |
|---|---|---|
agent | agt_refund_triage spent $212 of a $400 limit | The operator: on pace, no change |
run | Two runs cost $61 together. The median run costs $1.90. | The engineer: both runs looped on a failing migration, so set a per-run ceiling and fix the schema slice from part 8 |
person | One person started 70% of the runs | Finance: charge the payments cost center, not platform |
The figures are illustrative. The point is that three people took three different actions from one table, and none of them needed the invoice.
Attribution is not savings. Recording who spent what does not lower the bill. It shows where the bill comes from, which is the precondition for any change you can verify afterward. If you later claim a reduction, state the workload, the baseline, the model and price versions, the sample size, the quality measure, and the costs including retries. Without those, the claim you can support is visibility.
In Oxagen every governed action is priced and attributed to the person, the agent, the run, the turn, and the step. The Spend page and the Run page read the same rows, and measured, reported, and estimated costs are labelled as such. The budget an agent runs under, what it has spent this month, and the rule that stops it sit beside the agent. This covers calls that pass through Oxagen. A harness whose model calls do not pass through the gateway reports its spend, or is not metered, and the page says which.
- Add four keys to your model-call log.
person,agent,run, andstep. If you proxy model calls, add them there. If you do not, add them in the wrapper that starts the run. - Separate cached input tokens. Log them in their own column and price them at the cached rate.
- Version your price list. A small table keyed by model and effective date. Stamp each row with the version used.
- Label the cost basis. Go through each source of cost data and mark it measured, reported, or estimated.
- Set one per-run ceiling. Pick the agent with the widest spread between median and maximum run cost. Set the ceiling at ten times the median and route to the operator when a run reaches it.
attributed_spend_ratio: spend carrying person, agent, and run, over total model spend on the invoice. The gap is what you cannot explain.cost_per_run_p50andp99, per agent. A wide spread points at loops, and part 9 and part 11 are where to look.unpriced_steps: steps with tokens and no price, usually a new model missing from the price list.budget_holds: times anat_limitaction fired, per agent, per month.
Record the person, the agent, the run, and the step on every priced call, label how each cost was obtained, and keep the budget beside the agent it limits. Then the question of who spent what has a query for an answer.
Cite this
Anderson, M. (2026). Spend you can attribute. In Engineering Deterministic AI Coding Agents (2nd ed., Part 17). Oxagen Inc. https://macanderson.com/manual/spend-you-can-attribute
BibTeX
@incollection{anderson2026spendyoucan,
author = {Anderson, Mac},
title = {Spend you can attribute},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
chapter = {17},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/spend-you-can-attribute}
}Updates by email
Get the next edition of the field manual and new research when it is published.