# Mac Anderson > Mac Anderson is the founder and CEO of Oxagen, the author of Engineering Deterministic AI Coding Agents, and the creator of Stella, an open-source coding agent. He researches and builds coding agents that ship software on their own. Mac Anderson researches agentic coding systems and builds coding agents. Some of them are self-evolving: they write their own tools and get better at the job. His goal is software that agents deliver on their own, and that is useful, scales, and stays easy to maintain. He lives in Los Angeles. He is the founder and CEO of [Oxagen](https://oxagen.sh), workforce management for autonomous agents. Each agent gets its own identity, a mandate that sets what it may do and spend, the tools it needs, and a record of what it did. He also created [Stella](https://github.com/macanderson/stella), an open-source coding agent written in Rust, and wrote *Engineering Deterministic AI Coding Agents*, a field manual in 21 parts. He wrote software by hand for 16 years. Before Oxagen he co-founded Fonteva, a Salesforce-native software company for associations that Togetherwork acquired in 2021, and founded inTown Technologies, which built software for city governments. In 2026 he handed all of his coding to agents and kept the job of directing them, reading their work, and deciding what merged. Between 26 February and 20 September 2026 his GitHub account recorded 21,861 contributions, against 38 in all of 2025. Every page on this site has a markdown version at the same address with `.md` added, and requests that send `Accept: text/markdown` receive it. An MCP server at https://macanderson.com/mcp offers search and page reads (read-only, no key). ## Field manual - [Engineering Deterministic AI Coding Agents](https://macanderson.com/manual.md): A field manual in 21 parts, by Mac Anderson of Oxagen. Parts 1 to 13 show how to build the deterministic system around a coding agent, with published evidence in each part. Parts 14 to 20 show how to operate the agents you now run: identity, authority, budget, equipment, and a record another person can read. Every part ends with steps for this week and metrics to track. - [How to use this book](https://macanderson.com/manual/how-to-use-this-book.md): Read the two parts that match your problem this week. Come back for the rest. - [Part 1: AI agents are not expensive. Bad architecture is.](https://macanderson.com/manual/agents-are-not-expensive-bad-architecture-is.md): Most of an agent's token bill comes from the system around the model, not from the model. - [Part 2: Stop making agents read your entire logs](https://macanderson.com/manual/stop-making-agents-read-your-logs.md): Logs are structured data, and treating them as prompt text costs tokens you do not have to spend. - [Part 3: Parse the user prompt before the model ever sees it](https://macanderson.com/manual/parse-the-prompt-before-the-model-sees-it.md): Your prompt already carries structured information, so extract it before you pay for inference. - [Part 4: Why grep is the wrong retrieval engine](https://macanderson.com/manual/why-grep-is-the-wrong-retrieval-engine.md): Searching is not understanding, and code already carries the structure you need. - [Part 5: Build semantic memory once](https://macanderson.com/manual/build-semantic-memory-once.md): A run should start from what the last one learned instead of repeating it. - [Part 6: A rule written in prose cannot fail CI](https://macanderson.com/manual/rules-belong-in-data-not-in-prose.md): Constraints belong in a representation a validator can check. - [Part 7: Context compression is worth more than a bigger model](https://macanderson.com/manual/compression-beats-a-bigger-model.md): Removing irrelevant information beats buying a larger model. - [Part 8: Agents guess at schema they were not shown](https://macanderson.com/manual/show-the-agent-the-schema.md): A model adds a duplicate column when the context did not carry the existing one. - [Part 9: Why coordinator agents don't scale](https://macanderson.com/manual/why-coordinator-agents-do-not-scale.md): Most multi-agent systems pay a coordination tax and call it architecture. - [Part 10: Tests should be first-class retrieval objects](https://macanderson.com/manual/tests-as-retrieval-objects.md): Tests describe behavior in executable form, and a test runner's verdict is computed rather than inferred. - [Part 11: Agents need working memory, not bigger context windows](https://macanderson.com/manual/working-memory-not-bigger-windows.md): You do not reread every book you own before fixing a bug, and your agent should not either. - [Part 12: Knowledge graphs beat prompt engineering](https://macanderson.com/manual/knowledge-graphs-over-concatenation.md): Knowledge should be connected, not concatenated. - [Part 13: Measuring agent intelligence](https://macanderson.com/manual/measuring-an-agent-in-production.md): Benchmarking on coding challenges misses what matters in production. - [Part 14: Give every agent its own identity](https://macanderson.com/manual/give-every-agent-its-own-identity.md): You cannot set authority for, bill, or review something you cannot name. - [Part 15: Write the mandate](https://macanderson.com/manual/write-the-mandate.md): Four decisions govern an agent. Most teams have made all four. Few have them in one place. - [Part 16: The agent asks, a rule decides](https://macanderson.com/manual/the-agent-asks-a-rule-decides.md): Authority is a decision at the moment of use, written down by the team that owns the system. - [Part 17: Spend you can attribute](https://macanderson.com/manual/spend-you-can-attribute.md): A total is not an answer. The answer is which agent spent what, and on whose behalf. - [Part 18: Equip agents on purpose](https://macanderson.com/manual/equip-agents-on-purpose.md): A tool in the catalog is not a tool in the agent's hands. Assign equipment the way you assign access. - [Part 19: Keep a record another person can read](https://macanderson.com/manual/keep-a-record-another-person-can-read.md): A log answers the engineer who wrote it. A record answers the person who was not there. - [Part 20: Bounded tasks and ongoing work](https://macanderson.com/manual/bounded-tasks-and-ongoing-work.md): Some work has an endpoint. Other work continues. Manage each in its own way. - [Part 21: Better systems, not better models](https://macanderson.com/manual/better-systems-not-better-models.md): You rent the model. You own the system around it, and the way you operate it. - [Closing thought](https://macanderson.com/manual/closing-thought.md): Intelligence is expensive. Determinism is cheap. Spend the first only where the second cannot do the job. - [Field kit](https://macanderson.com/manual/field-kit.md): Templates and worksheets to copy. Print this section and fill it in with the people who own each line. - [Glossary](https://macanderson.com/manual/glossary.md): The words this book uses in a fixed sense. - [About the author](https://macanderson.com/manual/about-the-author.md): Oxagen is workforce management for autonomous agents: give each agent an identity, set its authority and budget, equip it with tools and skills, and review… - [Sources](https://macanderson.com/manual/sources.md): Every empirical claim in this book traces to one of the sources below. Read the primary literature. It is better than any summary of it, including this one.… ## Research - [Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models.md): Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week. - [Chapter 1: Rented tokens](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/rented-tokens.md): The token bill falls every year. The capability and the data never arrive. What a team owns after a year of renting. - [Chapter 2: The trace](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-trace.md): The record of one session, why the tool calls carry the signal, and the flip as the unit of value. - [Chapter 3: The oracle problem](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-oracle-problem.md): A verdict the model does not control: deterministic, independent, and hidden, with the fail-to-pass flip stated as a rule. - [Chapter 4: Organizational oracles](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/organizational-oracles.md): The checks a team already has, ordered by how directly each one can label a trace. - [Chapter 5: The air gap](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-air-gap.md): The agent is an adversary with a shell. The oracle runs where the shell cannot reach. - [Chapter 6: The one-bit verdict](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-one-bit-verdict.md): Why PASS or FAIL is all the oracle says, what that costs inside a session, and the research on models grading themselves. - [Chapter 7: Harness hooks](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/harness-hooks.md): The hooks a harness already exposes, turned into a trace collector and a Stop gate without building an agent. - [Chapter 8: The reference plugin](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-reference-plugin.md): A walkthrough of oracle-flip: five hooks, one constant reason, and an oracle in a container with no network. - [Chapter 9: The flip rate](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-flip-rate.md): Write the test first, dispatch small, sample more than once, manufacture tasks, and keep the failures. - [Chapter 10: From traces to training data](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/from-traces-to-training-data.md): Select, deduplicate, format, mask, hold out, version. - [Chapter 11: Training methods](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/training-methods.md): Supervised fine-tuning on flips, then preference pairs, then reinforcement learning. Adapters, forgetting, collapse, and a shuffled-label control. - [Chapter 12: Data volume](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume.md): What 491, 5,016, and 8,209 verified trajectories bought in the literature, and a worksheet for your own rate. - [Chapter 13: The delivery pipeline](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/the-delivery-pipeline.md): Weights as a release artifact: ingest, validate, train, evaluate, gate, canary in shadow, promote, roll back. - [Chapter 14: Evaluation](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/evaluation.md): Held-out flips, contamination, weak tests, hacking scans, and the drift measures that catch what the flip rate misses. - [Chapter 15: Serving and the economics](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/serving-and-the-economics.md): Routing between your model and the rented one, and when owning the weights becomes cheaper than the rent. - [Chapter 16: Governance and failure modes](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/governance-and-failure-modes.md): Secrets, memorization, oracle tampering, pipeline debt, a starving flip signal, and people. - [Chapter 17: Precedents](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/precedents.md): Expert iteration to SWE-Gym and Getafix to DIDACT: what each established and what it left to do. - [Closing: What to do this quarter](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/what-to-do-this-quarter.md): Seven steps for this quarter. - [Appendix A: Trace event schema](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/trace-event-schema.md): The event types and fields in a session's trace file. - [Appendix B: Oracle container specification](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/oracle-container-specification.md): What the oracle container mounts, fixes, applies, runs, and reports. - [Appendix C: Data volume worksheet](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/data-volume-worksheet.md): Four numbers from your team and the arithmetic that turns them into a date. - [Sources](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/sources.md): The works cited, in order of first citation. - [Agents are waiting on a process built for people](https://macanderson.com/research/agents-are-waiting-on-a-process-built-for-people.md): An agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour. - [Steering a run you are not watching](https://macanderson.com/research/steering-a-run-you-are-not-watching.md): When a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record. - [The agent time horizon is doubling](https://macanderson.com/research/the-agent-time-horizon-is-doubling.md): METR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision. - [The problem is not slop, it is your process](https://macanderson.com/research/the-problem-is-not-slop-it-is-your-process.md): The worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work. - [What an agent should be told before it starts](https://macanderson.com/research/what-an-agent-should-be-told-before-it-starts.md): A prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things. - [A model cannot grade its own homework](https://macanderson.com/research/the-limits-of-self-correction-and-model-collapse.md): Without outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead. - [Deterministic coding agents: every turn on the record](https://macanderson.com/research/deterministic-coding-agents-every-turn-on-the-record.md): A coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit. - [From STaR to DeepSeek-R1: what self-improvement means](https://macanderson.com/research/self-improving-models-from-star-to-self-rewarding.md): A vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one. - [Governing an agent that rewrites itself](https://macanderson.com/research/governing-an-agent-that-rewrites-itself.md): An agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts. - [Graph-grounded retrieval vs vector search](https://macanderson.com/research/graph-grounded-retrieval-vs-vector-search.md): Vector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference. - [Self-evolving agents: what the evidence shows](https://macanderson.com/research/self-evolving-agents-what-the-evidence-shows.md): What changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve. - [The science of AI agents: from ReAct to tool use](https://macanderson.com/research/the-science-of-ai-agents-from-react-to-tool-use.md): Planning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do. - [What an Ontology Buys an Agent](https://macanderson.com/research/what-an-ontology-buys-an-agent.md): An agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers. - [What SWE-bench Measures, and What It Misses](https://macanderson.com/research/what-swe-bench-measures-and-what-it-misses.md): Coding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you. - [Why agents fail: measuring reliability and cost](https://macanderson.com/research/why-agents-fail-measuring-reliability-and-cost.md): AgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides. ## About Mac - [Site blueprint](https://macanderson.com/blueprint.md): How macanderson.com is built to be read by people and by AI agents, and how to build a site like it. Markdown twins, llms.txt, an MCP server, a guarded assistant, and an installable app. - [Privacy](https://macanderson.com/privacy.md): What this site collects, why, and how to remove it. It sets no tracking cookies and runs no analytics. - [The story](https://macanderson.com/story.md): In 2026 Mac Anderson stopped writing code by hand and let AI agents write all of it. This is what the seven months from 26 February to 20 September looked like, from his own record. - [Work](https://macanderson.com/work.md): What Mac Anderson builds and has built. Oxagen, Stella, Arena, the Context Graph Protocol, his career from Fonteva to Oxagen, and his GitHub activity. ## Optional - [Everything on this site in one file](https://macanderson.com/llms-full.txt) - [Book as PDF](https://macanderson.com/engineering-deterministic-ai-coding-agents.pdf) - [RSS feed](https://macanderson.com/feed.xml)