Mac Anderson

Field manual

Field kit

Templates and worksheets to copy. Print this section and fill it in with the people who own each line.

Mac Anderson4 min read934 words
View markdown

Every template here is referenced from a part of the book. They are starting points. Change the fields to match your systems, and keep the owners.

Sheet 1 · Agent inventory (part 14)

One row per agent that can write to a system. A row without an operator is the first thing to fix.

Agent idPurpose, one sentenceOperatorHarnessBounded or ongoingCredentials it can reachShared with
Sheet 2 · Mandate (part 15)

One file per agent. Each clause has one owner who can change it by pull request.

agent:
operator:
boundary: "Governs actions routed through ____. Does not govern ____."

access:                      # owner:
  acts_as:
  may_request:
    - system:
      scope:
      actions: []

budget_and_rules:            # owner:
  monthly_limit_usd:
  per_run_ceiling_usd:
  at_limit:                  # hold_and_notify | route_to_operator | continue_and_flag
  rules: []

equipment:                   # owner:
  tools: []
  skills: []
  context_scope: []
  tool_budget: { max_tools: , max_tokens: }

record:
  covers:
  retention_days:
last_reviewed:               # date, and the three owners who were present
Sheet 3 · Decision rule (part 16)
id:
owner:
when:
  action:
  resource:
effect:                      # allow | deny | route
route_to:                    # a role or a named person, when effect is route
timeout:
on_timeout:                  # deny | allow
reason: ""                   # written as the agent's next step

# At the top of the rules directory:
default_when_no_rule_matches:   # deny | identity_alone
Write actionCalls per weekAnswer you wantRouted toRule id
Sheet 4 · Scoped rule object (part 6)
{
  "rule": "Money is integer cents. No float arithmetic on amounts.",
  "applies_to": "billing/**",
  "verified_by": "lint:no-float-money",
  "since": "a1b2c3",
  "owner": "payments-platform"
}

A rule with no verified_by is a preference. Keep it in prose, or write the check.

Sheet 5 · Domain dossier (part 5)
domain: billing
source_sha:                  # the commit this was generated from
entry_points: []             # routes, jobs, and consumers
core_types: []               # with one line each
invariants: []               # each links to the check that holds it
owned_tables: []
covering_tests: []
may_call: []                 # other domains
must_not_call: []
recent_decisions: []         # ADR ids
budget_tokens: 300
Sheet 6 · Cost row (part 17)
person, agent, run, turn, step,
model, input_tokens, cached_tokens, output_tokens,
price_version, cost_usd, cost_basis,      # measured | reported | estimated
rule_id, started_at
Sheet 7 · Production scorecard (part 13)
MetricThis quarterLast quarterOwner
cost_per_completed_task
tokens_consumed per task, by phase
retrieval_precision
first_attempt_rate
time_to_green
files_touched per task
Variance across 10 repeats of one task
attributed_spend_ratio
agents_with_mandate
time_to_answer_review
Sheet 8 · Completion checks for a bounded task (part 20)
task: ""
checks:
  - run:   ""                # a command that must exit 0
  - diff:  { only_paths: [] }
  - file:  { exists: "" }
  - human: { role: , question: "" }
on_fail: { retries: , then: route_to_operator }
# Hash this file when the run opens. Record the digest as the first row.
Sheet 9 · The operator's week (part 20)
WhenDoRead
DailyAnswer routed requests. Look at anything held at a budget limit.open_requests_age, budget_holds
WeeklyRead spend and denials per agent. Review runs that tripped a threshold.requests_by_answer, cost_per_run_p99
MonthlyRe-read each mandate with its owners. Turn repeated approvals into narrower rules. Retire agents nobody can justify.days_since_mandate_review, tool_use_ratio
QuarterlyRun a tabletop review with someone outside the team. Fill in the scorecard.time_to_answer_review, sheet 7
Sheet 10 · A 30, 60, and 90 day plan
By dayEngineering halfOperating halfYou can now answer
30Per-step token and cost logging. Trace slicing in one agent.Inventory complete. Every agent has an operator and an id.Where do the tokens go, and who answers for each agent?
60Prompt parsing, a symbol index, a schema slice, and a test lookup.One mandate per agent. Rules for the top ten write actions. A stated default.What is each agent allowed to do, and which rule said so?
90A workflow skeleton with budgets. Per-agent tool assignment. The scorecard filled in twice.Attributed spend. A chained run record. Completion checks for one bounded task type. Review dates booked.What did it cost, per agent and per person, and can someone else verify what happened?
One task through the whole system

To see how the parts connect, follow one prompt. "The POST /api/v2/refunds endpoint throws DecimalConversionError in billing/processors.py after #4821 merged."

  1. The run opens. The record stores the agent, the operator, the initiator, and the mandate version (parts 14, 15, and 19). The completion checks are hashed (part 20).
  2. Code parses the prompt into a route, an exception class, a file path, and an issue, and resolves each against an index (part 3).
  3. The graph expands the anchors one hop. The schema index adds the two tables. The test lookup adds three covering tests. The billing dossier adds its invariants. Everything is cut to an 8,000 token budget (parts 4, 5, 7, 8, and 10).
  4. The workflow engine runs the failing test under a tracer, and the slicer returns about 1,200 tokens of frames and locals (parts 2 and 9).
  5. The model writes the patch. This is the first step where judgment is needed, and the first large model call.
  6. The runner executes the checks. On a failure, the sliced result feeds one more attempt, and working memory drops the superseded one (parts 10 and 11).
  7. The agent requests github.push to a release branch. The rule routes it to the release owner, who approves. The push goes through a mediated connection (part 16).
  8. Every step wrote a cost row with the person, the agent, and the run (part 17). The run closes, the checks held, and the chain is signed (parts 19 and 20).
  9. The scorecard updates: tokens by phase, retrieval precision, attempts, time to green, and cost (part 13).

The model made one kind of decision in that sequence: what the patch should be. Code, rules, and people made the rest, and each of those decisions is in the record.

Cite this

Anderson, M. (2026). Field kit. In Engineering Deterministic AI Coding Agents (2nd ed., Field kit). Oxagen Inc. https://macanderson.com/manual/field-kit

BibTeX
@incollection{anderson2026fieldkit,
  author    = {Anderson, Mac},
  title     = {Field kit},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/field-kit}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.