Mac Anderson

Part 6: Context engineering

A rule written in prose cannot fail CI

Constraints belong in a representation a validator can check.

Mac Anderson4 min read1,025 words4 sources cited
View markdown

It starts with one file that tells the agent to stop repeating a mistake. Then a second for architecture, a third for conventions, a fourth for the things the first three left out. Six months later every run loads 30,000 tokens of prose, nobody can say which rules are still true, and the agent has started ignoring half of them. What you have is a second codebase written in a language with no compiler, no tests, and no dead-code detection.

Prompt bloat is measurable harm, not just cost#

The instinct says instructions are free, and that the worst case is a model skimming the irrelevant ones. The research points the other way. Performance degrades as input grows even when the added material is task-relevant in spirit. Liu et al. showed models lose information positioned in the middle of long contexts,6 and Chroma's 18-model study showed degradation with input length across every frontier model tested, at different rates.7 Anthropic's context-engineering guidance is direct about the mechanism: attention is a finite budget, every token in context draws it down, and the goal is the smallest set of high-signal tokens that achieves the outcome.19 Thirty thousand tokens of standing instructions are not a safety net. They are attention spent before the task begins, and billed per request.

The caching economics make it worse#

Prompt caching looks like it rescues the large standing prompt. Cache reads cost about 10% of base input price on Anthropic's API, with a 25% premium on writes, cutting cost by up to 90% and latency by up to 85% for long stable prefixes.16 The catch is that caching is prefix-exact. Edit one character of an early block and everything after it recomputes at full price, plus the write premium. A pile of frequently edited prose concatenated into the prompt is the worst shape for this, because each edit invalidates the cache across your whole fleet. Caching rewards small, stable, ordered context. That is an argument for structure, not for volume.

◌ Rules as prose

  • Loaded whole, in every run, relevant or not
  • No validation, so rules rot silently as the code moves
  • Conflicting rules resolved probabilistically by the model
  • Every edit invalidates the prompt cache prefix
  • Grows monotonically, and nobody dares delete a line

● Rules as data

  • Retrieved selectively, only the rules matching the touched domain
  • Machine-checkable, because rules reference real symbols and run in CI
  • Conflicts detected at build time
  • Stable structured core caches cleanly, variable tail stays small
  • Dead rules detected the way dead code is

What replaces the prose#

  • JSON Schema or typed config for anything that is a constraint: allowed dependencies, naming patterns, review requirements. A constraint in schema form is enforced by a validator rather than suggested to a model.
  • Tool metadata: capabilities, costs, and preconditions declared on the tools themselves, the direction MCP standardizes, so capability discovery is a lookup instead of a paragraph.
  • Rule objects with scopes: {rule, applies_to: "billing/**", verified_by: "lint:no-float-money", since: "a1b2c3"}. The retrieval layer injects the three rules that bear on this diff, not the three hundred that exist.
  • Generated docs: the part 5 dossiers, derived from code, carrying their source SHA.

One rule, both ways#

# before: a line in conventions.md, loaded in full in every run
# "Money is cents. Do not use floats for money."

# after: a rule object, injected when the diff touches billing/
{"id": "no-float-money",
 "applies_to": "billing/**",
 "verified_by": "lint:no-float-money",
 "since": "a1b2c3",
 "text": "Money columns and money variables are integer cents."}
# 40 tokens in the prompt, and CI fails on a float literal under billing/.

The second form costs less, arrives only when it applies, and has a witness. If the lint rule stops firing because the code moved, you learn that the rule is dead instead of carrying it for another year.

Order the prompt so the cache survives your edits#

prompt = [
  schema_snapshot,   # stable for days
  domain_dossier,    # stable until that domain merges
  matched_rules,     # 3 of 300, changes per task
  task_slice,        # changes per task
]
# Edits land in the tail. The cached prefix survives them.

Keep prose for what prose is good at: intent, history, and taste. Anything that functions as a rule belongs in a representation you can validate, scope, version, and retrieve selectively. If a rule cannot fail CI, it is a preference.

Do this week
  1. Count what you are loading. Run every standing prose file through a tokenizer (tiktoken, or your provider's token-count endpoint) and print a table sorted by tokens. The number at the bottom is what each run pays before it reads any code.
  2. Sort each line into constraint, context, or history. Constraints move to a rule file. Context moves to a generated dossier. History stays in prose and leaves the prompt.
  3. Give your three most repeated constraints a check. An ESLint or Ruff rule, a JSON Schema validator, or a dependency-cruiser rule. Done looks like the check failing on a violation you introduce on purpose, then passing when you revert it.
  4. Inject rules by glob. Match applies_to against the paths in the diff and pass only the matches. Done looks like a prompt carrying 3 rules on a billing change rather than the full registry.
  5. Reorder the prompt. Stable structured blocks first, task slice last, so an edit lands in the tail and the cached prefix holds.
Measure it
  • prompt_tokens_static: tokens loaded before the task-specific slice. Read it from the request log. Lower is better, and it should fall in steps as files move out.
  • rules_injected_per_task: rules in the prompt divided by rules in the registry. Lower is better, and a ratio near 1 means your scoping is not working.
  • rule_enforcement_rate: share of rules whose verified_by check runs in CI. Higher is better. Rules without a check are the ones that rot.
  • cache_read_ratio: cached input tokens divided by total input tokens, taken from the provider's usage fields. Higher is better, and a sudden drop means somebody edited an early block.
Takeaway

Natural-language configuration is configuration without a type system. Move constraints into structured, validated, selectively retrieved data, and let prose go back to being prose.

Cite this

Anderson, M. (2026). A rule written in prose cannot fail CI. In Engineering Deterministic AI Coding Agents (2nd ed., Part 6). Oxagen Inc. https://macanderson.com/manual/rules-belong-in-data-not-in-prose

BibTeX
@incollection{anderson2026rulesbelongin,
  author    = {Anderson, Mac},
  title     = {A rule written in prose cannot fail CI},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  chapter   = {6},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/rules-belong-in-data-not-in-prose}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.