Mac Anderson

Part 7: Context engineering

Context compression is worth more than a bigger model

Removing irrelevant information beats buying a larger model.

Mac Anderson4 min read995 words6 sources cited
View markdown

An agent underperforms on your repository, and the available levers look like this: a bigger model, a longer window, a larger reasoning budget. All three cost money and none of them change what you are sending. The question that predicts success is not how much the model can hold. It is what fraction of what it holds bears on the task, and that is a question you can answer with instrumentation you already own.

The Shannon view of a context window#

Shannon's framework18 gives the right vocabulary. The task has an intrinsic information requirement, the minimal set of facts needed to produce the correct patch. Everything else in the window is noise relative to that task, and attention is the finite channel the signal has to pass through. Two consequences follow. First, mutual information matters more than volume: a 4k-token context where every line bears on the task carries more usable information than 100k tokens where the answer is diluted 25 to 1. Second, compression is safe when it is lossless with respect to the task, which is what deterministic slicing (parts 2 to 4) gives you and what generic truncation does not. A stack-trace parser that keeps your frames and drops framework frames is a task-aware compressor with a retention property you can state. head -c has no such property.

Four results point the same way#

U-curve
Liu et al.: accuracy collapses for facts mid-context, long-context models included6
18 / 18
Chroma: every frontier model tested degrades as input grows, even on trivial tasks7
−dozens of pts
NoLiMa: when matches are semantic rather than literal, most models fall far below short-context accuracy by 32k8
Related > random
Power of Noise: semantically related distractors hurt accuracy most9

Read the four together for coding agents specifically. A repository dump is long, which triggers length degradation. It is full of near-duplicate distractors, which is the most harmful noise type the fourth result identifies. Using it requires semantic rather than lexical matching, the NoLiMa condition where the drop is steepest. A raw-context design sits at the intersection of all three.

Counterweight · long-context is improving

Be honest about the trend. Newer models post high scores on simple needle-in-a-haystack retrieval at extreme lengths, and vendors are attacking positional bias directly. The gains are strongest on literal retrieval, the easiest case, while semantic reasoning over long, distractor-dense context remains the weak point, which is the regime coding agents live in. And a model that holds its accuracy at length still bills you for every irrelevant token. Compression wins on cost even where it stops winning on accuracy.

Compression as an engineering discipline#

  • Retrieval precision over recall by default: start from verified anchors (part 3) and expand along graph edges (part 4) instead of dumping candidates.
  • Representation changes: signatures instead of bodies (a repo map is roughly a 100 to 1 compressor12), schemas instead of migrations, sliced traces instead of logs.
  • Budgets enforced in code: every context assembler takes a token budget and ranks content to fit it, and going over budget fails the build.
  • Context utilization as a metric: what fraction of the provided tokens does the final patch depend on? In the traces this book draws on it often sits below 10%, and it is the clearest single indicator of wasted spend (part 13, and part 17 for attributing that spend to the team that caused it).

A budget the assembler cannot talk its way out of#

def assemble(task, budget=8000):
    blocks = rank(candidates(task))        # anchors first, then graph neighbours
    kept, used = [], 0
    for b in blocks:
        if used + b.tokens > budget:
            break
        kept.append(b)
        used += b.tokens
    missing = required(task) - {b.id for b in kept}
    if missing:                            # a required anchor did not fit
        raise BudgetError(task, missing, used, budget)
    return kept
# Over budget is a failure with a name, not a truncation nobody sees.

The failure matters more than the cap. When a required anchor does not fit, you want a stack trace naming the task and the anchor, because that is a retrieval bug. Silent truncation turns the same bug into a wrong patch and a confusing review.

Do this week
  1. Record what you sent and what got used. For each run, log the ids and token counts of every block you provided, then log the files the accepted patch touched. One JSON line per run is enough to compute utilization.
  2. Replace one file dump with signatures. Parse with tree-sitter and emit declarations rather than bodies for anything outside the edit site. The artifact is a repo map you can diff against the old prompt, token for token.
  3. Give the assembler a budget. Add the cap and the named failure above. Done looks like a test that asks for an over-budget assembly and asserts BudgetError.
  4. Drop near-duplicates before they ship. Run minhash or simhash over retrieved chunks and collapse anything above your similarity threshold. Related distractors are the ones the fourth result says cost you most.
  5. Re-run last week's failures at the new budget. If accuracy holds with 40% fewer tokens, you have found the model upgrade you were about to buy.
Measure it
  • context_utilization: tokens of blocks the accepted patch depended on, divided by tokens provided. Higher is better. Below 10% means most of your bill is noise.
  • tokens_per_accepted_patch: total input tokens across the run, divided by patches a human accepted. Lower is better, and it is the number to quote when somebody proposes a larger window.
  • distractor_rate: share of retrieved chunks the final patch does not reference. Lower is better, and a rising value usually means recall crept back into your retriever.
  • budget_overflow_rate: assemblies that raised BudgetError, divided by assemblies. Lower is better, and each one is a retrieval bug with an address.
Takeaway

A bigger model reads your noise at a higher price. A better system removes the noise. Only one of the two compounds: every point of retrieval precision you engineer carries over to every model you run after this one.

Cite this

Anderson, M. (2026). Context compression is worth more than a bigger model. In Engineering Deterministic AI Coding Agents (2nd ed., Part 7). Oxagen Inc. https://macanderson.com/manual/compression-beats-a-bigger-model

BibTeX
@incollection{anderson2026compressionbeatsa,
  author    = {Anderson, Mac},
  title     = {Context compression is worth more than a bigger model},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  chapter   = {7},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/compression-beats-a-bigger-model}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.