Part 10: Orchestration and memory
Tests should be first-class retrieval objects
Tests describe behavior in executable form, and a test runner's verdict is computed rather than inferred.
Your agent writes a patch, then runs the suite to find out whether it broke something. The tests that would have told it what the function was supposed to do were sitting in the repository the whole time, and it read none of them before it started. Most stacks treat tests as a thing you run after generating a patch, not a thing you retrieve before. That is backwards. The test suite is the best-maintained behavioral specification your repository has.
Behavioral indexing#
Make tests queryable along the axes that matter:
- Coverage mapping. Run the suite under coverage once per merge and invert the result: for every function, the exact tests that execute it.
tests_covering("process_refund")becomes a deterministic lookup. - Contract extraction. Parse assertions into behavior statements. "A refund of a settled order returns
RefundResult(status=PENDING). A refund exceeding the order total raisesInvalidRefund." Signatures told the model the shape (Part 4). Assertions tell it the meaning. - Fixture graph. Fixtures encode how valid domain objects are constructed. Link them into the code graph and the model gets canonical object construction for free.
Now the retrieval plan from Part 3 gains a behavioral layer. Touching process_refund pulls its three covering tests, a precise executable account of current behavior, usually far smaller than the docs that would attempt to describe it.
The inversion is a few lines in CI, and the artifact it produces is a map from symbol to tests.
# once per merge, in CI
pytest --cov=app --cov-context=test --cov-report=json
# invert contexts: {symbol: [tests that executed it]}
index = defaultdict(list)
for fn, arcs in coverage["files"].items():
for line, contexts in arcs["contexts"].items():
for test in contexts:
index[symbol_at(fn, line)].append(test)
tests_covering("process_refund")
→ ["test_refund_settled_order",
"test_refund_exceeds_total",
"test_refund_is_idempotent"]
Agentless does more than run tests at the end. Tests are load-bearing in its pipeline. It generates reproduction tests from the issue, uses existing regression tests to filter candidate patches, and re-ranks the survivors, a deterministic selection stage that the paper's ablations credit for a substantial share of final accuracy.1 SWE-bench makes the broader point: the industry's shared yardstick for agent capability defines "solved" as nothing more or less than "the fail-to-pass tests now pass."15 If tests are how you judge agents, they should be how agents see.
Xia et al., FSE 2025 · Jimenez et al., ICLR 2024
What an assertion says that a docstring does not#
Contract extraction is worth a worked example, because the output is smaller than the input and carries more.
◌ The docstring
- "Processes a refund for an order."
- Written once, at creation
- Silent about the settled case, the over-total case, and idempotency
- No signal when it drifts from the code
● The extracted contract
- settled order →
RefundResult(status=PENDING) - amount > order.total → raises
InvalidRefund - same idempotency key twice → one ledger entry
- CI reruns it, so a drift turns the build red
Execution-driven, contract-first loops#
Tests also give the workflow skeleton (Part 9) its cheapest validators. A contract-first loop runs like this: retrieve the covering tests, have the model write or extend the failing test first so the intent becomes executable, have it patch, run the deterministic runner, then feed the structured failure (sliced per Part 2) into the next attempt. A compiler and a test runner grade each iteration. Their verdicts are computed from the code that ran, not inferred from a diff by another model emitting an opinion. The loop converges on green, and green is a fact about a process exit code.
- Turn on per-test coverage contexts. Run the suite once with coverage.py recording a context per test (or your language's equivalent) and write the JSON to a build artifact. Done looks like one file that records which test touched which line.
- Invert it into a symbol index. Map each covered line to its enclosing symbol with tree-sitter and store
{symbol: [tests]}next to your code graph. Done looks liketests_covering("process_refund")answering in milliseconds. - Put covering tests in the retrieval bundle. When a node targets a symbol, attach its covering tests ahead of any prose documentation, under the same token budget. Done looks like a bundle you can diff against the old one and see docs give way to tests.
- Extract assertion contracts for your ten hottest symbols. Parse the assert statements into one-line behavior statements and store them as graph properties. Done looks like ten symbols whose contracts a reviewer agrees with.
- Make the loop write the test first. Change the patch node to emit a failing test before it emits a patch, and reject attempts that skip it. Done looks like every merged agent patch carrying a test that failed before it.
symbol_test_coverage_index_freshness: hours since the coverage inversion last ran, compared against commits merged since. Down is good. A stale index points an agent at tests that no longer exist.covering_tests_in_bundle: share of retrieval bundles for an edit task that include at least one covering test for the target symbol. Compute it from the bundle logs. Up is good.first_attempt_green_rate: patches that pass the suite on attempt 1, divided by all patches. Up is good, and it moves when retrieval improves rather than when the model changes.test_first_compliance: agent patches that shipped with a test that failed before the patch, divided by all agent patches. Up is good.
Index your tests the way you index your code. They are the documentation in your repository that executes, and the ones running in CI are current as of the last green run. That makes them the highest-signal retrieval objects you own.
Cite this
Anderson, M. (2026). Tests should be first-class retrieval objects. In Engineering Deterministic AI Coding Agents (2nd ed., Part 10). Oxagen Inc. https://macanderson.com/manual/tests-as-retrieval-objects
BibTeX
@incollection{anderson2026testsasretrieval,
author = {Anderson, Mac},
title = {Tests should be first-class retrieval objects},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
chapter = {10},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/tests-as-retrieval-objects}
}Updates by email
Get the next edition of the field manual and new research when it is published.