Field manual
How to use this book
Read the two parts that match your problem this week. Come back for the rest.
You run coding agents, or you are about to, and two questions keep coming up. The first is from your own team: why does this cost so much and behave differently each time? The second is from everyone else: what are these agents allowed to do, what did they spend, and who answers for them? Parts 1 to 13 answer the first question. Parts 14 to 20 answer the second. Part 21 explains why both have the same answer.
How each part is laid out#
- The situation. Each part opens on a problem you have probably seen.
- The mechanism. What to build, in order, with a short sample.
- Evidence. A published result, cited. The sources are listed at the back, and the primary papers are better than any summary of them.
- Counterweight. Where the argument stops applying. Not every part has one.
- Do this week. Three to five steps, cheapest first. Each says what done looks like.
- Measure it. The readings that tell you whether the step worked.
Parts 14 to 20 add a short box, "Where Oxagen fits". Oxagen publishes this book and builds a product for that half of the job. The practice in each part stands without the product, and the box says what the product does and where its boundary is. Parts 1 to 13 mention no product.
Three reading paths#
| If you are | Start with | Then |
|---|---|---|
| An engineer building or tuning an agent | Parts 1, 2, and 3 | Part 13 to set up measurement, then the rest of 4 to 12 in the build order below |
| A platform lead or operator answering for several agents | Parts 14, 15, and 16 | Parts 17 and 20, then part 13 for the engineering scorecard |
| A security or finance lead reviewing agent use | Part 16 (security) or part 17 (finance) | Parts 15 and 19, which cover the object you review and the record you review it from |
A build order#
The parts are numbered by topic, not by the order to build them. If you are starting from an agent that works and costs too much, this order pays back fastest, because each step makes the next one measurable.
| Step | Build | Part | Why now |
|---|---|---|---|
| 1 | Token and cost logging per step, with run, agent, and person | 13, 17 | Nothing later can be verified without it |
| 2 | An inventory of agents, each with an operator | 14 | One afternoon, and every later step needs the names |
| 3 | Trace and log slicing | 2 | The largest reduction in context for the least code |
| 4 | Prompt parsing and a symbol index | 3, 4 | Replaces the exploratory calls at the start of a run |
| 5 | Schema slice and test lookup | 8, 10 | Removes two common causes of wrong patches |
| 6 | A mandate for each agent, and rules for the top ten write actions | 15, 16 | By now other teams are asking |
| 7 | A workflow skeleton and tiered memory | 9, 11 | Worth doing once runs are long enough to loop |
| 8 | Per-agent tool assignment | 18 | Needs the offered-against-called data from step 1 |
| 9 | Domain dossiers and the knowledge graph | 5, 12 | The largest build. It federates the indexes from steps 4 and 5 |
| 10 | The run record, chained, and completion checks for bounded tasks | 19, 20 | Turns the rows you already write into something a reviewer can use |
A self-assessment#
Answer yes or no. Count a yes only if you could show the evidence today. Each no points at a part.
| # | Question | If no, read |
|---|---|---|
| 1 | Can you split one run's tokens into context construction and reasoning? | 1 |
| 2 | Does code, not the model, reduce a stack trace before it reaches the window? | 2 |
| 3 | Are file paths and symbols in a prompt resolved against an index before the first model call? | 3, 4 |
| 4 | Does a new run start from stored knowledge of the repository? | 5 |
| 5 | Can every standing instruction fail a check in CI? | 6 |
| 6 | Does every context assembler take a token budget? | 7 |
| 7 | Does the agent see the live schema, with lineage, before it writes a migration? | 8 |
| 8 | Is control flow between steps written as code? | 9 |
| 9 | Can you look up the tests that cover a function? | 10 |
| 10 | Does old content leave the window by policy? | 11 |
| 11 | Can you answer a three-hop question about your system with one query? | 12 |
| 12 | Do you report cost per completed task, including retries? | 13 |
| 13 | Can you list every agent that can write to a system, with a named operator for each? | 14 |
| 14 | Is each agent's authority, budget, and equipment written in one reviewed place? | 15 |
| 15 | For a given write action, can you name the rule that allowed it and its author? | 16 |
| 16 | Can you attribute last month's spend to the agent, the run, and the person? | 17 |
| 17 | Do you know how many tools each agent is offered, and how many it calls? | 18 |
| 18 | Could someone outside your team reconstruct one agent action in 30 minutes? | 19 |
| 19 | Are completion checks for bounded tasks fixed before the run starts? | 20 |
| 20 | Does each ongoing agent have a review date on a calendar? | 20 |
Score yourself now and write the number down. Part 21 asks you to do it again.
Cite this
Anderson, M. (2026). How to use this book. In Engineering Deterministic AI Coding Agents (2nd ed., How to use this book). Oxagen Inc. https://macanderson.com/manual/how-to-use-this-book
BibTeX
@incollection{anderson2026howtouse,
author = {Anderson, Mac},
title = {How to use this book},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/how-to-use-this-book}
}Updates by email
Get the next edition of the field manual and new research when it is published.