Part 19: Operating the workforce
Keep a record another person can read
A log answers the engineer who wrote it. A record answers the person who was not there.
Something an agent did is being reviewed. It might be an incident, an audit sample, or a customer question. The person reviewing was not on the run and does not know your logging conventions. They ask four things: what did the agent read, what did it change, who or what allowed it, and what did it cost? When an agent requests an action in a system another team owns, which rule answers it, and where is that answer recorded? If those answers live in four tools and a chat transcript, the review becomes an investigation.
Logs and records are different artifacts#
Logs are for debugging. They are verbose, unstructured at the edges, sampled under load, and deleted after a few weeks. All of that is correct for debugging. A record has a different reader and a different job. It is complete for the actions it covers, structured so a query can answer a question, ordered so the sequence is not in doubt, and kept as long as the accountability lasts.
◌ A log
- Written for the author of the code
- Free text with some fields
- Sampled and rotated
- Answers "why did this break?"
● A record
- Written for a reader who was not there
- One typed row per event, with stable keys
- Complete for governed actions, and retained
- Answers "what happened, and under whose authority?"
What a run's record holds#
Parts 14 to 18 each produced rows. The record is those rows in one sequence, keyed by the run.
| Event | Fields | From |
|---|---|---|
| Run opened | agent, operator, initiator, task, mandate version | Parts 14 and 15 |
| Context read | what was retrieved, from which scope, at which source commit | Parts 3 to 5 and 18 |
| Request | tool, action, resource, the answer, the rule id, and the approver when routed | Part 16 |
| Step | model or tool call, tokens, cost, cost basis | Part 17 |
| Change | files, rows, or resources written, with a diff or a reference | The tool that made it |
| Run closed | outcome, totals, and the checks that ran if the task was bounded | Part 20 |
The important property is that the request, the rule that answered it, and the cost of the step sit in the same row family under the same keys. A reviewer follows one action back to its authority without joining systems by timestamp.
Order you can check#
A record that can be edited quietly is an account, not evidence. Two inexpensive mechanisms raise the bar. First, chain the rows: each row stores a hash of the row before it, so a removed or altered row breaks the chain at a known point. Second, hash a canonical form. Two serializers can emit the same JSON object with keys in a different order, and the hashes will differ. RFC 8785 defines a canonical JSON serialization for this purpose, so that the same data yields the same bytes and the same digest everywhere.26
row = {
"run": "run_7f3a", "seq": 41, "kind": "request",
"agent": "agt_refund_triage", "initiator": "dana.okafor",
"action": "github.push", "resource": "acme/billing@release/2026.09",
"answer": "routed", "rule": "release-branch", "approver": "priya.n",
"prev": "sha256:9c1e..." # digest of row 40
}
digest = sha256(canonicalize(row)) # RFC 8785, then hash
# At close, sign the last digest. A reviewer with the export can
# recompute the chain without access to the system that wrote it.
This shows integrity, and only integrity. A valid chain means the rows are the rows that were written, in that order. It does not mean the agent's work was correct, and it does not cover an action that never passed through the recording layer.
Google's site reliability practice treats the written postmortem as the unit of learning from an incident: a record of the impact, the actions taken, the root causes, and the follow-up, kept blameless so that people contribute what they know.27 The practice depends on a timeline that people accept as accurate. With agents, the timeline is the hard part, because the actor cannot be interviewed and the transcript is long. Part 13 instrumented the run to improve the architecture. The same instrumentation, kept and ordered, is what a review reads.
Write the record for its three readers#
- The operator wants the run as a sequence: what it is waiting on now, and what it did before that.
- The security reviewer wants requests grouped by answer and by rule, with every routed request showing who signed.
- The finance reader wants the same rows grouped by person and cost center.
One set of rows serves all three when the keys are consistent. Three separate exports, built at three different times, will disagree with each other, and the review will be about the disagreement.
State what the record covers. It holds governed activity: the calls that went through the layer that writes it. It does not hold what an agent did on a path around that layer, and it records an observe-mode call without having enforced anything on it. It also holds data that may be sensitive. Decide what is stored in full, what is stored as a digest, who can read which rows, and how long each kind is kept, before the first review asks.
For governed activity, Oxagen keeps what the run read, what it changed, what it cost, and which rule answered each request. Every governed request is a row: who started the task, which agent asked, which tool it wanted, what data it would reach, which rule answered, and who signed. It is the same row the meter prices. Oxagen calls each recorded event in a run a frame. Every frame is hash-chained to the one before it, and the seal signs the close of the run, so a reviewer with the export can check its integrity outside the product.
- Run a tabletop review. Pick one agent action from last week. Give a colleague who was not involved 30 minutes to answer the four questions in the opening. Write down where they got stuck.
- Pick the run id and thread it through. One id, created when the run opens, present on every log line, request, and cost row.
- Separate the record from the logs. Give it its own store, schema, and retention, even if the first version is one append-only table.
- Chain the rows. Add
seqandprev, and canonicalize before hashing. Write the 20-line verifier at the same time, and run it nightly. - Write the coverage sentence. "This record covers X. It does not cover Y." Put it at the top of the export.
time_to_answer_review: minutes for someone outside the team to answer the four questions for a sampled action. Measure it quarterly. This is the number the whole part exists to lower.record_coverage: write actions present in the record, over write actions seen by the target systems' own audit logs. The gap is the ungoverned path count from part 15, observed.chain_breaks: verifier failures per night. The expected value is 0, and any other value is worth a person's morning.
Keep one ordered, typed record per run that holds the request, the rule that answered it, the change, and the cost under the same keys. Make its order checkable, and say what it covers.
Cite this
Anderson, M. (2026). Keep a record another person can read. In Engineering Deterministic AI Coding Agents (2nd ed., Part 19). Oxagen Inc. https://macanderson.com/manual/keep-a-record-another-person-can-read
BibTeX
@incollection{anderson2026keeparecord,
author = {Anderson, Mac},
title = {Keep a record another person can read},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
chapter = {19},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/keep-a-record-another-person-can-read}
}Updates by email
Get the next edition of the field manual and new research when it is published.