Mac Anderson

Research

Research on AI agents

Essays on how agents work, fail, and should be operated. Each one starts from published evidence and cites it. Most were first published by Oxagen Research, the research arm of the company Mac founded.

  1. 17-chapter bookBuilding a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's tracesEvery coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.1 h 49 min
  2. Agents are waiting on a process built for peopleAn agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour.11 min
  3. Steering a run you are not watchingWhen a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.7 min
  4. The agent time horizon is doublingMETR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.7 min
  5. The problem is not slop, it is your processThe worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work.7 min
  6. What an agent should be told before it startsA prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things.8 min
  7. A model cannot grade its own homeworkWithout outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead.8 min
  8. Deterministic coding agents: every turn on the recordA coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit.8 min
  9. From STaR to DeepSeek-R1: what self-improvement meansA vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one.8 min
  10. Governing an agent that rewrites itselfAn agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts.9 min
  11. Graph-grounded retrieval vs vector searchVector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference.8 min
  12. Self-evolving agents: what the evidence showsWhat changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve.8 min
  13. The science of AI agents: from ReAct to tool usePlanning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do.8 min
  14. What an Ontology Buys an AgentAn agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers.8 min
  15. What SWE-bench Measures, and What It MissesCoding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you.9 min
  16. Why agents fail: measuring reliability and costAgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides.8 min