Mac Anderson

Part V: Delivery

15Serving and the economics

Routing between your model and the rented one, and when owning the weights becomes cheaper than the rent.

Mac Anderson4 min read918 words
View markdown

The model is trained. This chapter is about running it: where it serves, how requests are routed between it and the rented model, and the arithmetic that says when the pipeline has paid for itself.

Serving#

Open-weight models serve through a small number of mature engines. vLLM introduced paged attention, which manages the key-value cache in blocks the way an operating system manages memory, and made high-throughput serving of large models practical on commodity accelerators.1 Serving an adapter without merging it is supported directly; S-LoRA showed that thousands of adapters over one base model can be served from a single node with the base weights shared.2 For a team with one fine-tune per repository or per domain, that means one base model in memory and a set of small adapters, with the request choosing the adapter.

Quantization reduces memory and raises throughput at some cost in quality. A 32-billion-parameter model quantized to 4 bits fits on a single large accelerator. Re-evaluate on the frozen set after quantizing, as Chapter 13 said, because the cost is not uniform across models.

Routing#

The team's model does not have to handle everything on day one. A router sends each task to the team's model first and falls back to the rented frontier model when the team's model fails. The oracle makes the fallback decision cheap: a session that does not flip after the team's model's attempts is re-dispatched to the rented model.

The fallback sessions are traces. They are the hard tail of the distribution, solved by a stronger model, and graded by the oracle. They go into the trace store with their provenance and become the distillation examples from Chapter 9. Over time, the share of tasks that reach the fallback is the measure of how far the team's model has come, and it is also the training set for closing the gap.

Routing a task between the team's model and the rented model

  1. TaskAn oracle task is dispatched
  2. Team modelAttempts it; the gate asks the oracle
  3. FlipPASS: done. The trace is a self-generated example.
  4. FallbackNo flip after N attempts: re-dispatch to the rented model
  5. Rented modelAttempts it; the same gate, the same oracle
  6. TraceEither way, the trace goes to the store with its provenance
The fallback rate falls as the team's model learns. The fallback traces are what it learns from.

The arithmetic#

The cost side has four lines.

Oracle compute. Each oracle run is a container running a test suite. On a cloud runner, a suite that takes five minutes costs cents. At 100 oracle sessions a day with two attempts each and a baseline per task, that is a few hundred container-minutes a day.

Storage. A trace with redacted tool output is tens of kilobytes to a few megabytes. A year of a mid-sized team's traces is tens to hundreds of gigabytes. This is a rounding error.

Training. An adapter run on a 32-billion-parameter model over a few thousand trajectories is hours on one eight-accelerator node. At on-demand cloud prices, that is low hundreds to low thousands of dollars per run. Weekly, it is tens of thousands a year. Evaluation on a few hundred held-out tasks, each an agent session in a sandbox, costs about as much again.

Serving. One eight-accelerator node, or two for redundancy, serves a 32-billion-parameter model to a team of fifty with capacity to spare. Reserved, this is low six figures a year; owned, it is a capital cost amortized over several years.

The revenue side is the rent avoided. A team of fifty engineers running agents daily on a frontier model spends, at mid-2026 prices, in the range of several hundred thousand to a few million dollars a year, depending on how heavily they use it. The exact figure is the team's own bill, and the team should use it.

The comparison is not all-or-nothing. In the routing design above, the team's model handles the share of tasks it can flip, and the rented model handles the rest. If the team's model flips 40 percent of oracle tasks in its first quarter, 40 percent of the rent on those tasks is replaced by serving cost. As the share rises, the rent falls. The crossover, where the pipeline's total cost falls below the rent it replaces, depends on the team's bill, but for a team spending more than the cost of two accelerator nodes a year on tokens, it arrives within the first year of flips.

What the price decline does to this#

Token prices fall fast, and the rent will be lower next year.3 Three things keep the arithmetic in the pipeline's favor anyway.

Open-model serving costs fall on the same curve, because the hardware and the serving software improve for everyone. The ratio between rent and serving cost is more stable than either number.

The rented model's price per token is for a model that does not know the team's code. The team's model's cost per token is for one that does. If the team's model flips a task in fewer tokens, which a model trained on that codebase should, the comparison per task is better than the comparison per token.

And the rent buys no asset. The pipeline's cost buys weights, a trace store, and an oracle, all of which the team keeps. The question in Chapter 1 was what the team owns at the end of the year. The arithmetic here is about when owning it is also cheaper, and the answer for most teams of this size is: soon.

Footnotes#

  1. Kwon, W., Li, Z., Zhuang, S., et al. (2023). Efficient Memory Management for Large Language Model Serving with PagedAttention. SOSP 2023. https://arxiv.org/abs/2309.06180 ↩

  2. Sheng, Y., Cao, S., Li, D., et al. (2023). S-LoRA: Serving Thousands of Concurrent LoRA Adapters. arXiv; MLSys 2024. https://arxiv.org/abs/2311.03285 ↩

  3. Cottier, B., Snodin, B., Owen, D., & Adamczewski, T. (2025). LLM inference prices have fallen rapidly but unequally across tasks. Epoch AI. https://epoch.ai/data-insights/llm-inference-price-trends ↩

Cite this

Anderson, M. (2026). Serving and the economics. In Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces (Chapter 15). macanderson.com. https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/serving-and-the-economics

BibTeX
@incollection{anderson2026continuousdeliveryof,
  author    = {Anderson, Mac},
  title     = {Serving and the economics},
  booktitle = {Building a continuous delivery system for fine-tuned open-source, open-weight models trained on your organization's traces},
  chapter   = {15},
  year      = {2026},
  url       = {https://macanderson.com/research/continuous-delivery-of-fine-tuned-open-weight-models/serving-and-the-economics}
}

Updates by email

New research reaches subscribers first.

No spam. Unsubscribe any time. Read the privacy note.