Field manual, second edition
Engineering Deterministic AI Coding Agents
A language model is a probabilistic engine. Almost everything around it can be decided in code: what it retrieves, what stays in its window, which step runs next, what it may do, what it may spend, and what gets recorded. Parts 1 to 13 build that system. Parts 14 to 20 show how to operate the agents you now run.
By Mac Anderson (ORCID 0009-0005-1646-9676). 21 parts, 27 cited sources, 24,568 words. Free to read and share.
Reading paths
Read the two parts that match your problem this week. Come back for the rest.
An engineer building or tuning an agent
- 1Agents are not expensive. Bad architecture is
- 2Stop making agents read your logs
- 3Parse the prompt before the model sees it
Then: Part 13 to set up measurement, then the rest of 4 to 12.
A platform lead or operator answering for several agents
Then: Parts 17 and 20, then part 13.
A security or finance lead reviewing agent use
Then: Parts 15 and 19.
Contents
Deterministic retrieval
- 2Stop making agents read your entire logsLogs are structured data, and treating them as prompt text costs tokens you do not have to spend.4 min
- 3Parse the user prompt before the model ever sees itYour prompt already carries structured information, so extract it before you pay for inference.4 min
- 4Why grep is the wrong retrieval engineSearching is not understanding, and code already carries the structure you need.4 min
- 5Build semantic memory onceA run should start from what the last one learned instead of repeating it.4 min
Context engineering
- 6A rule written in prose cannot fail CIConstraints belong in a representation a validator can check.4 min
- 7Context compression is worth more than a bigger modelRemoving irrelevant information beats buying a larger model.4 min
- 8Agents guess at schema they were not shownA model adds a duplicate column when the context did not carry the existing one.5 min
Orchestration and memory
- 9Why coordinator agents don't scaleMost multi-agent systems pay a coordination tax and call it architecture.5 min
- 10Tests should be first-class retrieval objectsTests describe behavior in executable form, and a test runner's verdict is computed rather than inferred.4 min
- 11Agents need working memory, not bigger context windowsYou do not reread every book you own before fixing a bug, and your agent should not either.4 min
- 12Knowledge graphs beat prompt engineeringKnowledge should be connected, not concatenated.5 min
Operating the workforce
- 14Give every agent its own identityYou cannot set authority for, bill, or review something you cannot name.4 min
- 15Write the mandateFour decisions govern an agent. Most teams have made all four. Few have them in one place.4 min
- 16The agent asks, a rule decidesAuthority is a decision at the moment of use, written down by the team that owns the system.5 min
- 17Spend you can attributeA total is not an answer. The answer is which agent spent what, and on whose behalf.5 min
- 18Equip agents on purposeA tool in the catalog is not a tool in the agent's hands. Assign equipment the way you assign access.5 min
- 19Keep a record another person can readA log answers the engineer who wrote it. A record answers the person who was not there.5 min
- 20Bounded tasks and ongoing workSome work has an endpoint. Other work continues. Manage each in its own way.5 min