Mac Anderson

Part 3: Deterministic retrieval

Parse the user prompt before the model ever sees it

Your prompt already carries structured information, so extract it before you pay for inference.

Mac Anderson4 min read916 words1 sources cited
View markdown

Someone files a bug against your agent and it reads like this: "The POST /api/v2/refunds endpoint throws DecimalConversionError in billing/processors.py after #4821 merged." That string holds an API route, an exception class, a file path, and an issue reference. Four retrieval keys, sitting in plain sight. Most agent stacks hand the raw string to the model and let it decide what to search for.

Deterministic extraction is a solved problem#

Long before you need a language model, thirty lines of parsing gets you:

  • File paths: anything matching path grammar, validated against the actual file tree, with near misses fuzzy-matched.
  • Symbols and signatures: CamelCase and snake_case identifiers, module.func() call syntax, function signatures, validated against your symbol index (part 4).
  • Stack traces: a pasted traceback has a rigid grammar in every language. Parse it fully rather than summarizing it.
  • Issue and PR references: #4821, JIRA-123, resolved through an API into titles, diffs, and linked commits.
  • URLs and API endpoints: a route pattern maps straight onto a router definition in the codebase.
  • Version and config literals: package names, versions, environment variable names, feature flags.

Every extracted entity is a verified anchor. It either resolves against your index or it does not. That check is binary, and it is the check a generative approach cannot give you, so an invented file path or a misremembered function name survives into the next call instead of being dropped at the door.

# Thirty lines, no model, runs before the first call.
PATH   = re.compile(r"\b[\w./]+\.(py|ts|go|rs|java)\b")
SYMBOL = re.compile(r"\b([A-Z][A-Za-z0-9]+|[a-z_][a-z0-9_]{2,})\(?\)?")
ISSUE  = re.compile(r"#(\d+)|\b([A-Z]{2,}\-\d+)\b")

def anchors(prompt, index):
    found, dropped = [], 0
    for raw in PATH.findall(prompt) + SYMBOL.findall(prompt):
        hit = index.resolve(raw)   # None when the tree has no such path or symbol
        if hit:
            found.append(hit)
        else:
            dropped += 1
    metrics.record("anchor_resolution_rate", len(found), dropped)
    return found

From entities to a retrieval plan#

The output of parsing is not decoration. It is an execution plan that runs before inference begins:

plan = build_retrieval_plan(prompt)
# {
#   anchors:   [File("billing/processors.py"), Symbol("DecimalConversionError"),
#               Route("POST /api/v2/refunds"), Issue(4821)]
#   expand:    graph_neighbors(anchors, hops=1)      ← Part 4
#   tests:     tests_covering(anchors)               ← Part 10
#   schema:    tables_touched(anchors)               ← Part 8
#   budget:    8_000 tokens, ranked by relevance
# }

The model's first call now holds the right file, its direct dependencies, the covering tests, and the issue diff, assembled by code, ranked by graph centrality, and cut to budget. Compare that to the common pattern, where the model spends its first five calls, each carrying a growing history, rediscovering what a regex had at millisecond zero.

Evidence · Structured localization beats free exploration

This is the design that let Agentless outperform autonomous agents. Its first phase is hierarchical localization, a fixed funnel from repository to files to classes and functions to edit locations, rather than a model wandering with tools. The paper's ablations show that staged narrowing keeps ground-truth edit locations in the candidate set while shrinking the code the model must read, and it is a large part of why the pipeline reached top open-source results at roughly a tenth of the cost of agent baselines.1 Constrained input, focused attention, fewer irrelevant tokens: the localization step is deterministic scaffolding doing the model's foraging for it.

Xia et al., "Agentless," FSE 2025

◌ Prompt as opaque string

  • The model infers what to look for, probabilistically, once per run
  • Search terms may be invented, and paths go unverified
  • 3 to 8 exploratory tool calls before the real work starts
  • Cost scales with how long the model explores

● Prompt as structured input

  • A parser extracts entities the same way on every run
  • Every anchor is validated against a real index
  • The retrieval plan executes before the first model call
  • Cost scales with the complexity of the task
Do this week
  1. Write the entity extractor. Start with four patterns: path grammar, identifiers, issue references, and route patterns. The artifact is a function from prompt string to a list of typed entities, with unit tests over ten real prompts from your own queue.
  2. Validate every entity against an index. Resolve paths against the file tree and symbols against your symbol index. Drop what does not resolve, fuzzy-match near misses, and count both.
  3. Emit the retrieval plan as an artifact. Write the plan to the task record before the first model call, with anchors, expansions, and the token budget. When a run goes wrong, the plan is the first thing you read.
  4. Seed the first call from the plan. Assemble the first message from resolved anchors instead of letting the model open with a search, then compare tool-call counts against the previous week.
Measure it
  • anchor_resolution_rate: entities that resolved against an index divided by entities extracted. Up is good. A drop usually means the index went stale, not that the prompts got worse.
  • first_call_hit_rate: share of tasks where the file the final diff edited was already in the first call's context. Up is good, and this is the metric that tells you whether localization works.
  • preflight_tool_calls: tool calls between the prompt and the first edit. Down is good.
  • unresolved_anchor_count: entities the model was handed that no index could confirm. Down is good, because each one is a chance to act on a path that does not exist.
Takeaway

Treat the prompt as the first document your system parses, not the first thing your model reads. Extraction is cheap, checkable against an index, and identical on every run. Inference is none of those things.

Cite this

Anderson, M. (2026). Parse the user prompt before the model ever sees it. In Engineering Deterministic AI Coding Agents (2nd ed., Part 3). Oxagen Inc. https://macanderson.com/manual/parse-the-prompt-before-the-model-sees-it

BibTeX
@incollection{anderson2026parsetheprompt,
  author    = {Anderson, Mac},
  title     = {Parse the user prompt before the model ever sees it},
  booktitle = {Engineering Deterministic AI Coding Agents},
  edition   = {Second},
  chapter   = {3},
  publisher = {Oxagen Inc.},
  address   = {Los Angeles, CA},
  year      = {2026},
  url       = {https://macanderson.com/manual/parse-the-prompt-before-the-model-sees-it}
}

Updates by email

Get the next edition of the field manual and new research when it is published.

No spam. Unsubscribe any time. Read the privacy note.