Story
The story
In 2026 Mac Anderson stopped writing code by hand and let AI agents write all of it. This is what the seven months from 26 February to 20 September looked like, from his own record.

Mac Anderson wrote software by hand for 16 years. In 2026 he made a decision: agents would write all of the code he shipped. His job would be to direct them, read their work, and decide what merged. Some of it went well. A lot of it went wrong. The hardest part turned out to be managing the agents, and he built Oxagen to do that job.
This page follows those seven months from the record he left: his GitHub history, the git history of 13 repositories, and 8,133 prompts he sent to Claude Code. The interactive original, with the charts, is at oxagen.sh/story.
Nine chapters#
Before the log#
26 February to 28 June. GitHub shows the work starting on 26 February. March brought 689 contributions, April 1,846, and May 2,167. On 28 May the first commits land in the Oxagen repository: the foundations and the agent runtime. His prompts from these months were on an earlier laptop. So this chapter uses commits only.
Beat Claude Code#
29 June to 9 July. He sets up a new laptop and works fast on the TypeScript Oxagen command-line agent. He adds caching, effort levels, a code graph, and an SWE-bench test setup. On 3 July he sets the goal: "Stripe for agents." Four days later he drops metered billing and moves to governance, traceability, budgets, and permissions. He also writes "if you see it you own it" into his agent instructions.
i do not want to falsely promote that oxagen is better
Stella is born#
10 July to 19 July. One night of "crazy ideas" leads to the mutation verifier. It checks that each test fails when its fix is removed. The next day the Rust rewrite gets a name, Stella, and a public repository. The witness test becomes the definition of done: a test that fails on the old code and passes on the new code. By 17 July you can install Stella 0.3.0 from Homebrew, and fleets run with no coordinator at all.
determinism is fast, cheap, reproducable, and better than intelligence 10/10 times.
Protocols and proof#
20 July to 31 July. The context protocol is renamed the Context Graph Protocol, and its SDK ships on npm. ArenaBench goes public. Steering records store the agent's standing instructions in git, as TOML files. On 31 July he runs the first head-to-head test against Claude Code on Terminal-Bench 2.1, with its rules published before the run. Both agents use the same model. Stella solves 58 of 89 tasks, and he publishes the run.
verification can not lie
The hard weeks#
1 August to 11 August. Benchmark runs had hidden caps. A verifier model spent four minutes and never wrote its test. On 8 August he found that the code graph had been empty, with no error. Stella's "losses" came from a $1-per-task budget cap. So proving the agent worked took more effort than building it. He made rules from these problems, and they still hold. Prove a defect with a failing test first. Check results with tests, not with a model. Never cap a benchmark. On 11 August he also published a result that made Stella look worse: on Sonnet 5, Stella cost about four times what Claude Code did.
WE HAVE MOVED TO FAST AND I AM WORRIED I WILL HAVE TO THROW AWAY THIS PROJECT
Less is more#
12 August to 19 August. Semantic code search starts working, and Stella's 72 tools shrink to five. Verification moves out of Stella into Vera, a plugin that can block a turn. The core becomes one loop. Everything else becomes a plugin, written in any language. On 19 August he deletes the pipeline crate. Self-driving Stella starts working through its own list of issues.
less is more and way less is best the 5 tools kicked ass
Many agents at once#
20 August to 30 August. He runs ten to fifteen agent conversations at once, each on a long goal that needs little input. So most of his prompts are short: "go," "resume," and "merge." He types fewer prompts, and the output climbs. He moves everything to AWS. He writes down how he steers his agents as standing decisions. He has Stella build a supply-chain demo app on its own. On 30 August a file-by-file review of Oxagen opens 1,059 issues. GitHub counts 1,144 contributions that day, the most of any day.
it is supposed to be hands off autonomous with controls to rollback
Oxagen becomes the control plane#
31 August to 10 September. He removes the generative AI features from Oxagen and makes Stella the engine inside it. Agents get their own identities and permissions, the same way people do. The Terminal-Bench page goes up: Stella solves 58 of 89 tasks, and Claude Code solves 44. He adds wrappers, hooks installed beside an agent, so Oxagen can govern Claude Code and Codex. Oxagen starts pricing by the governed action: one call routed through Oxagen that it checked and recorded.
I think stella has the best agent. I think oxagen has the best control plane.
Starting over from a new spec#
11 September to 20 September. He restarts Oxagen from a new spec. The spec says: wrap any agent, decide what each tool call may do, replay every run, and match spend to the cent. About twelve agent conversations, running side by side, build the demo mockups in one day. Within the week, a macOS installer wraps Claude Code and Codex. Oxagen becomes a gateway for model calls, and every model call on his own laptop goes through it.
they are worried about agents and the damage is done at tool call level.
Learning to manage a fleet#
With one agent, you give it a task, watch it work, and read the diff. With many agents at once, the job changes. It was the hardest thing he learned in seven months, and none of it was about prompting.
He started with one agent on one branch, and he watched it work. That setup cannot grow, because a person can only watch one thing at a time. So he added more agents. In the setup he ended with, six or more agents work in parallel, each in its own git worktree.
More agents caused problems he had not expected. Two agents edited the same branch. A branch merged into main with no conflicts, but it still did not work with code that had merged beside it. He had to repair a broken main three separate times before he understood the cause. A merge with no conflicts says nothing about whether the code works. Agents also finished, reported success, and left their work uncommitted.
What running many agents needed:
- A separate worktree for each agent, so two agents cannot edit the same files.
- A written definition of done. Without one, an agent decides for itself what finished means.
- Checks that run without him. There are 21 of them, so he does not have to review every agent's work by hand.
- A trial merge before the real one, because a merge with no conflicts can still break the code.
What he still could not answer:
- What each agent did, in order, and what it read.
- What that cost, and which piece of work it paid for.
- Who allowed each action, and under which rule.
- Which agent is stuck right now, and on what.
He could build the first list himself. The second list is a management problem, and it got worse with every agent he added. He built Oxagen to answer it.
The code his agents wrote#
Counted from tracked files on 21 September 2026:
- 2,258 test files. Unit, integration, and end to end.
- 272 architecture tests. They fail the build when code breaks a structure rule, such as a layer, import, or boundary rule.
- 21 checks in CI, including one that fails on an em dash.
- 136 decision records, plus 5 schema change records. Each one says what was decided and why.
A human reviewer notices when a file imports across a boundary it should not. Many agents working in parallel do not. Each may break the same rule in a different place on the same afternoon. A test that fails the build checks every change, however many agents make them. So with many agents, the architecture tests matter most.
The counts#
He did not type this code. Agents wrote it. He directed them, read their work, and merged it.
- 21,861 GitHub contributions, 26 February to 20 September 2026. The same account shows 38 in all of 2025.
- 6,341 commits on the default branch of 13 repositories, 28 May to 20 September. He made commits on 114 of those 116 days.
- 4,427 pull requests merged, 28 May to 20 September.
- 8,133 prompts to Claude Code, 29 June to 20 September, across 1,769 conversations on one laptop.
- 3,906,359 source lines added and 1,529,699 removed, counting source files on default branches only.
Read these counts with three facts beside them. The repositories were four months old, with no release schedule, no on-call rotation, and no other engineers. When he opened and merged a pull request himself, no second person reviewed it. The prompt count is the least complete number, because it covers 81 days on one laptop.
How the counts were made#
GitHub contributions come from GitHub's contribution calendar for the account macanderson, private repositories included. Commits and lines come from the default branch of 13 local clones, with each commit counted once by its hash, across the 10 name and email pairs he has used. Prompts come from Claude Code's prompt history on one laptop, loaded into SQLite and read in order. The 58 of 89 result is Stella against Claude Code on Terminal-Bench 2.1, both on the same GLM model, on 31 July 2026.