<?xml version="1.0" encoding="UTF-8"?>
<rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom">
<channel>
<title>Mac Anderson</title>
<link>https://macanderson.com</link>
<atom:link href="https://macanderson.com/feed.xml" rel="self" type="application/rss+xml"/>
<description>Mac Anderson is the founder and CEO of Oxagen, the author of Engineering Deterministic AI Coding Agents, and the creator of Stella, an open-source coding agent. He researches and builds coding agents that ship software on their own.</description>
<language>en-us</language>
<item><title>Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization&apos;s traces</title><link>https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models</link><guid>https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models</guid><pubDate>Fri, 09 Oct 2026 16:00:00 GMT</pubDate><description>Every coding-agent session your team runs produces a trace. Kept and graded by an oracle the agent cannot touch, those traces become the training set for a model you own. This book covers the oracle, the air gap, the one-bit verdict, the harness hooks that collect traces for free, how much data a fine-tune needs, and the pipeline that delivers new weights every week.</description></item>
<item><title>Agents are waiting on a process built for people</title><link>https://macanderson.com/research/agents-are-waiting-on-a-process-built-for-people</link><guid>https://macanderson.com/research/agents-are-waiting-on-a-process-built-for-people</guid><pubDate>Wed, 30 Sep 2026 16:00:00 GMT</pubDate><description>An agent can write a change in minutes. Then the change waits for a person to read it. This post measures that wait, what eight companies changed about it, and how Oxagen ships at every hour.</description></item>
<item><title>Engineering Deterministic AI Coding Agents, second edition</title><link>https://macanderson.com/manual</link><guid>https://macanderson.com/manual</guid><pubDate>Sat, 19 Sep 2026 16:00:00 GMT</pubDate><description>A field manual in 21 parts, by Mac Anderson of Oxagen. Parts 1 to 13 show how to build the deterministic system around a coding agent, with published evidence in each part. Parts 14 to 20 show how to operate the agents you now run: identity, authority, budget, equipment, and a record another person can read. Every part ends with steps for this week and metrics to track.</description></item>
<item><title>Steering a run you are not watching</title><link>https://macanderson.com/research/steering-a-run-you-are-not-watching</link><guid>https://macanderson.com/research/steering-a-run-you-are-not-watching</guid><pubDate>Wed, 16 Sep 2026 16:00:00 GMT</pubDate><description>When a long run goes wrong, most teams can stop it or type at it. Both work badly. A steer is a third option. It is a message with a delivery mode, a status, and a record.</description></item>
<item><title>The agent time horizon is doubling</title><link>https://macanderson.com/research/the-agent-time-horizon-is-doubling</link><guid>https://macanderson.com/research/the-agent-time-horizon-is-doubling</guid><pubDate>Wed, 16 Sep 2026 16:00:00 GMT</pubDate><description>METR measures how long a task an agent can finish on its own. That length has doubled about every seven months since 2019. This post covers what week-long runs mean for supervision.</description></item>
<item><title>The problem is not slop, it is your process</title><link>https://macanderson.com/research/the-problem-is-not-slop-it-is-your-process</link><guid>https://macanderson.com/research/the-problem-is-not-slop-it-is-your-process</guid><pubDate>Wed, 16 Sep 2026 16:00:00 GMT</pubDate><description>The worry about AI slop is about output quality. The measurements point to a different cause. Teams give an agent a workflow built for people and expect it to work.</description></item>
<item><title>What an agent should be told before it starts</title><link>https://macanderson.com/research/what-an-agent-should-be-told-before-it-starts</link><guid>https://macanderson.com/research/what-an-agent-should-be-told-before-it-starts</guid><pubDate>Wed, 16 Sep 2026 16:00:00 GMT</pubDate><description>A prompt file has no owner, no date, no scope, and no record that the agent read it. So it is a poor place for a standing rule. Steering records give each rule those things.</description></item>
<item><title>A model cannot grade its own homework</title><link>https://macanderson.com/research/the-limits-of-self-correction-and-model-collapse</link><guid>https://macanderson.com/research/the-limits-of-self-correction-and-model-collapse</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>Without outside feedback, self-correction fails, and a model trained on its own output gets worse. This post covers what the collapse and verifier research says to do instead.</description></item>
<item><title>Deterministic coding agents: every turn on the record</title><link>https://macanderson.com/research/deterministic-coding-agents-every-turn-on-the-record</link><guid>https://macanderson.com/research/deterministic-coding-agents-every-turn-on-the-record</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>A coding agent changed your code and no one can replay how. What the research says about feedback from running code, random sampling, and turns you can audit.</description></item>
<item><title>From STaR to DeepSeek-R1: what self-improvement means</title><link>https://macanderson.com/research/self-improving-models-from-star-to-self-rewarding</link><guid>https://macanderson.com/research/self-improving-models-from-star-to-self-rewarding</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>A vendor says the model improves itself. This is the research behind that claim, the signal that drives each training loop, and what stops each one.</description></item>
<item><title>Governing an agent that rewrites itself</title><link>https://macanderson.com/research/governing-an-agent-that-rewrites-itself</link><guid>https://macanderson.com/research/governing-an-agent-that-rewrites-itself</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>An agent that edits its own code needs the same review as any other change. What safety research says about reward hacking, sandboxes, oversight, and typed contracts.</description></item>
<item><title>Graph-grounded retrieval vs vector search</title><link>https://macanderson.com/research/graph-grounded-retrieval-vs-vector-search</link><guid>https://macanderson.com/research/graph-grounded-retrieval-vs-vector-search</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>Vector search finds the passage that looks like your question. Graph-grounded retrieval finds the fact that answers it. What the research says about the difference.</description></item>
<item><title>Self-evolving agents: what the evidence shows</title><link>https://macanderson.com/research/self-evolving-agents-what-the-evidence-shows</link><guid>https://macanderson.com/research/self-evolving-agents-what-the-evidence-shows</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>What changes when an agent improves itself, what checks the change, and the measured gain, across eight systems from Voyager to AlphaEvolve.</description></item>
<item><title>The science of AI agents: from ReAct to tool use</title><link>https://macanderson.com/research/the-science-of-ai-agents-from-react-to-tool-use</link><guid>https://macanderson.com/research/the-science-of-ai-agents-from-react-to-tool-use</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>Planning, tool use, memory, and reflection each come from a paper that measured something. This post traces those papers and what agents still cannot do.</description></item>
<item><title>What an Ontology Buys an Agent</title><link>https://macanderson.com/research/what-an-ontology-buys-an-agent</link><guid>https://macanderson.com/research/what-an-ontology-buys-an-agent</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>An agent can answer with confidence from the wrong context. This post covers what classes, relations, constraints, and dated facts add to an agent&apos;s answers.</description></item>
<item><title>What SWE-bench Measures, and What It Misses</title><link>https://macanderson.com/research/what-swe-bench-measures-and-what-it-misses</link><guid>https://macanderson.com/research/what-swe-bench-measures-and-what-it-misses</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>Coding agents are ranked by their SWE-bench resolve rate. This post covers what that rate shows, where test-based grading goes wrong, and what the rate cannot tell you.</description></item>
<item><title>Why agents fail: measuring reliability and cost</title><link>https://macanderson.com/research/why-agents-fail-measuring-reliability-and-cost</link><guid>https://macanderson.com/research/why-agents-fail-measuring-reliability-and-cost</guid><pubDate>Wed, 09 Sep 2026 16:00:00 GMT</pubDate><description>AgentBench, WebArena, GAIA, SWE-bench, and tau-bench each measure a different thing. None of them reports what a run costs. This post covers what that hides.</description></item>
</channel>
</rss>
