Mac Anderson

Part I: The argument

1Rented tokens

The token bill falls every year. The capability and the data never arrive. What a team owns after a year of renting.

Mac Anderson6 min read1,431 words
View markdown

A team that uses a frontier model through an API is renting. The rent has three parts, and only the first shows up on the invoice.

The first part is the token bill. It is real and it is falling. Epoch AI measured the price of reaching a fixed level of performance on a set of benchmarks and found it falling by somewhere between 9 and 900 times a year, depending on the task.1 That is good news for anyone who pays the bill, and it is the number vendors point to when someone asks why a team should not train its own model. The number is also beside the point. The question is not whether tokens will get cheaper. They will. The question is what the team owns at the end of the year.

The vendors in question are the frontier labs: Anthropic, OpenAI, Google, and the others whose models you reach only through an API and whose weights you never see. A black box is the right name for the product, whatever you think of the company. The second part of the rent is the capability you never keep. When a vendor ships a better model, your agents get better. When the vendor changes the model, deprecates it, raises the price, or changes the terms, your agents change with it. Nothing the model learned about your systems during a year of sessions stays with you, because the model learned nothing. It cannot. Inference does not update weights. Every session starts from the same checkpoint the vendor shipped, and every piece of context your team has written to steer it, every CLAUDE.md and every rule file, is a workaround for a model that does not know your code.

The third part is the data you discard. A session produces a record. The record shows which files the agent opened to understand a feature, which commands it ran to check its work, what it tried that failed, and what finally passed. If the agent was working in your repository on your task against your tests, that record is a worked example of your work being done. Nobody else has it. Nobody else can produce it. And in most teams, it is gone when the terminal closes.

Where the value of a session goes today

  1. TaskAn engineer or a work queue dispatches a task to a coding agent
  2. SessionThe agent reads, edits, runs commands, and reports done
  3. Tokens billedThe vendor bills every input and output token
  4. Record discardedThe transcript is deleted or buried in a local cache
The two emphasized steps are the rent. The money leaves and the record leaves. The model learned nothing, and neither did your organization's data.

What the open-weight trend changes#

The case for keeping traces rests on one bet: that open-weight models will stay close enough to the frontier that a model trained on your traces beats a rented model on your tasks. Three measurements say the bet is reasonable.

Epoch AI compared the best open-weight model to the best closed model across benchmarks and found the open model trailing by about a year, measured by the date at which a closed model first reached the same score.2 A year behind the frontier on general benchmarks is not a year behind on your codebase, because the frontier model is also starting from zero on your codebase.

Stanford's AI Index reported that the gap between the best open and best closed models on the Chatbot Arena leaderboard narrowed from 8.0 percent to 1.7 percent in one year.3 Chatbot Arena measures human preference on chat, not software engineering, but the direction holds across benchmarks the Index tracks.

On the benchmark that matters most for this book, SWE-bench Verified, open-weight models went from under 10 percent to above 40 percent in eighteen months. Meta's SWE-RL took a 70-billion-parameter Llama to 41.0 percent.4 Mistral's Devstral, a 24-billion-parameter model that runs on a single workstation card, reported 46.8 percent.5 The SWE-smith team took a 32-billion-parameter Qwen to 40.2 percent with about five thousand trajectories.6 The Skywork team took the same base model from 6.4 percent to 38.0 percent with about eight thousand, and found the gain still growing with each doubling of the data.7 These are not the frontier numbers. The frontier was above 70 percent at the time. They are the numbers that a team can run, inspect, and fine-tune.

Open-weight models on SWE-bench Verified

Model and methodResolved
Qwen2.5-Coder-32B, no agent training (SWE-Gym baseline)7%
Same model after fine-tuning on 491 SWE-Gym trajectories20.6%
Same model with a verifier and 16 samples per task32%
Skywork-SWE-32B, 8,209 trajectories38%
SWE-agent-LM-32B, 5,016 SWE-smith trajectories40.2%
Llama3-SWE-RL-70B, reinforcement learning41%
Devstral-Small-2505, 24B46.8%
DeepSWE-Preview, Qwen3-32B, RL with test-time scaling59%
Sources: Pan et al. (2024), Yang et al. (2025), Wei et al. (2025), Mistral (2025), Agentica and Together AI (2025). Each row is a different base model and training recipe, so the bars show the range open models reached, not a controlled comparison.

Two things about these numbers matter for the argument. First, the jump from 7.0 to 20.6 percent came from fine-tuning on 491 trajectories.8 That is not a large dataset. It is about a week of sessions for a mid-sized team. The further jump to 32.0 percent came from training a verifier on the same trajectories and letting it pick the best of 16 attempts, which is a preview of Chapter 9. Second, every model in the table was trained on public repositories solving public issues. None of them had seen the training team's own code. A model trained on your traces starts from these numbers and climbs on your distribution.

Licenses that permit it#

The models in the table ship under licenses that allow fine-tuning and commercial use. DeepSeek-R1 and its distilled variants are MIT licensed.9 The Qwen2.5 and Qwen3 families are Apache 2.0, with a small number of size variants under a Qwen license.10 The Llama 3 family uses Meta's community license, which permits commercial use below a very large monthly-user threshold.11 OpenAI's gpt-oss models are Apache 2.0.12 The point of listing them is not legal advice. It is that the permission to do what this book describes is ordinary and granted.

The asset that compounds#

Consider two teams of the same size doing the same work for one year.

Team A uses a rented frontier model. It spends on tokens, writes rule files, and gets better at prompting. At the end of the year it has a set of rule files, a bill, and a vendor relationship. If the vendor's next model is worse at the team's tasks, the team has no recourse. If a competitor uses the same vendor, the competitor has the same model.

Team B uses the same rented model, and also installs a trace collector and an oracle. It spends the same on tokens. Each session that flips an oracle from FAIL to PASS is saved as a verified trajectory. At the end of the year, Team B has thousands of worked examples of its own work being done correctly, each graded by a test the model did not write. It has fine-tuned an open model on them three or four times and has a model that resolves its own tasks at a rate the rented model cannot match on the same distribution, running on hardware it controls, at a marginal cost per token that is a fraction of the rent. If the vendor changes terms, Team B's model does not change. If a competitor wants the same model, the competitor needs Team B's traces.

The cost of being Team B instead of Team A, for the first several months, is close to zero. The collector is a plugin. The oracle is a test. Storage is cheap. The training runs come later and are cheaper than one month of a mid-sized team's token bill. The only thing Team A has to do to become Team B is to stop deleting the record.

Limits of the claim#

It does not claim that a fine-tuned 32-billion-parameter model will match the best frontier model on every task next quarter. It will not. It does not claim that the open-weight gap will close to zero. It may not. It does not claim that collecting traces is free of risk; Chapter 16 is about secrets, memorization, and the ways a pipeline like this goes wrong.

It claims that the traces are the asset, that the oracle is what makes them an asset, and that a team which starts collecting today will be in a different position in a year from a team that does not. The rest of the book is about how to do it carefully.

Footnotes#

  1. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI. https://epoch.ai/data-insights/llm-inference-price-trends ↩

  2. Cottier, B., You, J., Martemianova, N., & Owen, D. (2024). How far behind are open models? Epoch AI. https://epoch.ai/blog/open-models-report ↩

  3. Stanford Institute for Human-Centered Artificial Intelligence (2025). AI Index Report 2025, Chapter 2: Technical Performance. https://hai.stanford.edu/ai-index/2025-ai-index-report/technical-performance ↩

  4. Wei, Y., Duchenne, O., Copet, J., et al. (2025). SWE-RL: Advancing LLM Reasoning via Reinforcement Learning on Open Software Evolution. arXiv; NeurIPS 2025. https://arxiv.org/abs/2502.18449 ↩

  5. Mistral AI & All Hands AI (2025). Devstral. https://mistral.ai/news/devstral ↩

  6. Yang, J., Lieret, K., Jimenez, C. E., et al. (2025). SWE-smith: Scaling Data for Software Engineering Agents. arXiv. https://arxiv.org/abs/2504.21798 ↩

  7. Zeng, L., Li, Y., Xiao, Y., et al. (2025). Skywork-SWE: Unveiling Data Scaling Laws for Software Engineering in LLMs. arXiv. https://arxiv.org/abs/2506.19290 ↩

  8. Pan, J., Wang, X., Neubig, G., Jaitly, N., Ji, H., Suhr, A., & Zhang, Y. (2024). Training Software Engineering Agents and Verifiers with SWE-Gym. arXiv; ICML 2025. https://arxiv.org/abs/2412.21139 ↩

  9. DeepSeek-AI (2025). DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning. arXiv. https://arxiv.org/abs/2501.12948 ↩

  10. Qwen Team (2025). Qwen3 Technical Report. arXiv. https://arxiv.org/abs/2505.09388 ↩

  11. Grattafiori, A., et al. (2024). The Llama 3 Herd of Models. arXiv. https://arxiv.org/abs/2407.21783 ↩

  12. OpenAI (2025). gpt-oss-120b and gpt-oss-20b Model Card. arXiv. https://arxiv.org/abs/2508.10925 ↩

Cite this

Anderson, M. (2026). Rented tokens. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 1). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/rented-tokens

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Rented tokens},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {1},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/rented-tokens}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.