# Tests should be first-class retrieval objects

> Tests describe behavior in executable form, and a test runner's verdict is computed rather than inferred.

Part 10: Orchestration and memory. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/tests-as-retrieval-objects

Your agent writes a patch, then runs the suite to find out whether it broke something. The tests that would have told it what the function was supposed to do were sitting in the repository the whole time, and it read none of them before it started. Most stacks treat tests as a thing you run after generating a patch, not a thing you retrieve before. That is backwards. The test suite is the best-maintained behavioral specification your repository has.

### Behavioral indexing

Make tests queryable along the axes that matter:

-   **Coverage mapping.** Run the suite under coverage once per merge and invert the result: for every function, the exact tests that execute it. `tests_covering("process_refund")` becomes a deterministic lookup.
-   **Contract extraction.** Parse assertions into behavior statements. "A refund of a settled order returns `RefundResult(status=PENDING)`. A refund exceeding the order total raises `InvalidRefund`." Signatures told the model the shape (Part 4). Assertions tell it the meaning.
-   **Fixture graph.** Fixtures encode how valid domain objects are constructed. Link them into the code graph and the model gets canonical object construction for free.

Now the retrieval plan from Part 3 gains a behavioral layer. Touching `process_refund` pulls its three covering tests, a precise executable account of current behavior, usually far smaller than the docs that would attempt to describe it.

The inversion is a few lines in CI, and the artifact it produces is a map from symbol to tests.

```
# once per merge, in CI
pytest --cov=app --cov-context=test --cov-report=json

# invert contexts: {symbol: [tests that executed it]}
index = defaultdict(list)
for fn, arcs in coverage["files"].items():
    for line, contexts in arcs["contexts"].items():
        for test in contexts:
            index[symbol_at(fn, line)].append(test)

tests_covering("process_refund")
→ ["test_refund_settled_order",
   "test_refund_exceeds_total",
   "test_refund_is_idempotent"]
```

> **Evidence · Tests as selection machinery**
>
> Agentless does more than run tests at the end. Tests are load-bearing in its pipeline. It generates *reproduction tests* from the issue, uses existing *regression tests* to filter candidate patches, and re-ranks the survivors, a deterministic selection stage that the paper's ablations credit for a substantial share of final accuracy.[1](https://macanderson.com/manual/sources#r1 "Xia, Deng, Dunn & Zhang. \"Agentless: Demystifying LLM-based Software Engineering Agents.\" FSE 2025. arXiv:2407.01489 · github.com/OpenAutoCoder/Agentless") SWE-bench makes the broader point: the industry's shared yardstick for agent capability defines "solved" as nothing more or less than "the fail-to-pass tests now pass."[15](https://macanderson.com/manual/sources#r15 "Jimenez et al.. \"SWE-bench: Can Language Models Resolve Real-World GitHub Issues?\" ICLR 2024. arXiv:2310.06770") If tests are how you judge agents, they should be how agents see.
>
> Xia et al., FSE 2025 · Jimenez et al., ICLR 2024

### What an assertion says that a docstring does not

Contract extraction is worth a worked example, because the output is smaller than the input and carries more.

#### ◌ The docstring

-   "Processes a refund for an order."
-   Written once, at creation
-   Silent about the settled case, the over-total case, and idempotency
-   No signal when it drifts from the code

#### ● The extracted contract

-   settled order → `RefundResult(status=PENDING)`
-   amount > order.total → raises `InvalidRefund`
-   same idempotency key twice → one ledger entry
-   CI reruns it, so a drift turns the build red

### Execution-driven, contract-first loops

Tests also give the workflow skeleton (Part 9) its cheapest validators. A contract-first loop runs like this: retrieve the covering tests, have the model write or extend the failing test first so the intent becomes executable, have it patch, run the deterministic runner, then feed the structured failure (sliced per Part 2) into the next attempt. A compiler and a test runner grade each iteration. Their verdicts are computed from the code that ran, not inferred from a diff by another model emitting an opinion. The loop converges on green, and green is a fact about a process exit code.

> **Do this week**
>
> 1.  **Turn on per-test coverage contexts.** Run the suite once with coverage.py recording a context per test (or your language's equivalent) and write the JSON to a build artifact. Done looks like one file that records which test touched which line.
> 2.  **Invert it into a symbol index.** Map each covered line to its enclosing symbol with tree-sitter and store `{symbol: [tests]}` next to your code graph. Done looks like `tests_covering("process_refund")` answering in milliseconds.
> 3.  **Put covering tests in the retrieval bundle.** When a node targets a symbol, attach its covering tests ahead of any prose documentation, under the same token budget. Done looks like a bundle you can diff against the old one and see docs give way to tests.
> 4.  **Extract assertion contracts for your ten hottest symbols.** Parse the assert statements into one-line behavior statements and store them as graph properties. Done looks like ten symbols whose contracts a reviewer agrees with.
> 5.  **Make the loop write the test first.** Change the patch node to emit a failing test before it emits a patch, and reject attempts that skip it. Done looks like every merged agent patch carrying a test that failed before it.

> **Measure it**
>
> -   `symbol_test_coverage_index_freshness`: hours since the coverage inversion last ran, compared against commits merged since. Down is good. A stale index points an agent at tests that no longer exist.
> -   `covering_tests_in_bundle`: share of retrieval bundles for an edit task that include at least one covering test for the target symbol. Compute it from the bundle logs. Up is good.
> -   `first_attempt_green_rate`: patches that pass the suite on attempt 1, divided by all patches. Up is good, and it moves when retrieval improves rather than when the model changes.
> -   `test_first_compliance`: agent patches that shipped with a test that failed before the patch, divided by all agent patches. Up is good.

> **Takeaway**
>
> Index your tests the way you index your code. They are the documentation in your repository that executes, and the ones running in CI are current as of the last green run. That makes them the highest-signal retrieval objects you own.
