# Research by Mac Anderson

> Essays on how AI agents work, fail, and should be operated. Each one cites its sources.

- [Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces](https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models.md) (2026-10-09): Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.
- [Agents are waiting on a process built for people](https://macanderson.com/research/agents-are-waiting-on-a-process-built-for-people.md) (2026-09-30): An agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour.
- [Steering a run you are not watching](https://macanderson.com/research/steering-a-run-you-are-not-watching.md) (2026-09-16): When a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.
- [The agent time horizon is doubling](https://macanderson.com/research/the-agent-time-horizon-is-doubling.md) (2026-09-16): METR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.
- [The problem is not slop, it is your process](https://macanderson.com/research/the-problem-is-not-slop-it-is-your-process.md) (2026-09-16): The worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work.
- [What an agent should be told before it starts](https://macanderson.com/research/what-an-agent-should-be-told-before-it-starts.md) (2026-09-16): A prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things.
- [A model cannot grade its own homework](https://macanderson.com/research/the-limits-of-self-correction-and-model-collapse.md) (2026-09-09): Without outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead.
- [Deterministic coding agents: every turn on the record](https://macanderson.com/research/deterministic-coding-agents-every-turn-on-the-record.md) (2026-09-09): A coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit.
- [From STaR to DeepSeek-R1: what self-improvement means](https://macanderson.com/research/self-improving-models-from-star-to-self-rewarding.md) (2026-09-09): A vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one.
- [Governing an agent that rewrites itself](https://macanderson.com/research/governing-an-agent-that-rewrites-itself.md) (2026-09-09): An agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts.
- [Graph-grounded retrieval vs vector search](https://macanderson.com/research/graph-grounded-retrieval-vs-vector-search.md) (2026-09-09): Vector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference.
- [Self-evolving agents: what the evidence shows](https://macanderson.com/research/self-evolving-agents-what-the-evidence-shows.md) (2026-09-09): What changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve.
- [The science of AI agents: from ReAct to tool use](https://macanderson.com/research/the-science-of-ai-agents-from-react-to-tool-use.md) (2026-09-09): Planning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do.
- [What an Ontology Buys an Agent](https://macanderson.com/research/what-an-ontology-buys-an-agent.md) (2026-09-09): An agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent's answers.
- [What SWE-bench Measures, and What It Misses](https://macanderson.com/research/what-swe-bench-measures-and-what-it-misses.md) (2026-09-09): Coding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you.
- [Why agents fail: measuring reliability and cost](https://macanderson.com/research/why-agents-fail-measuring-reliability-and-cost.md) (2026-09-09): AgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides.
