Part 2: Deterministic retrieval
Stop making agents read your entire logs
Logs are structured data, and treating them as prompt text costs tokens you do not have to spend.
Open a run where your agent debugged a failing test and you will find the same four steps every time: cat the log file, grep for "error", tail -n 200, paste the result into context. The most expensive component in the stack is now doing the job of a parser, and doing it probabilistically.
Signal to noise is a system property#
A production stack trace from a Python web request can run hundreds of frames, and most of them are framework internals: WSGI plumbing, middleware, object-relational mapper (ORM) dispatch. The frames that matter for the bug are usually the handful that touch your code. An irrelevant frame is not neutral padding. It costs tokens on the way in, and it competes for attention once it is there. Chroma's "Context Rot" study evaluated 18 frontier models and found that performance degrades as input length grows, even on simple tasks, and that it degrades faster when the context holds distractors that are related to the target without being the target.7 A raw log is a distractor-dense document by construction: hundreds of lines that look like the one line you need.
Retrieval research reports the same shape. Cuconasu et al. ("The Power of Noise," SIGIR 2024) found that adding semantically related but non-answer-bearing documents to a model's context degrades accuracy, and that related noise hurts more than unrelated noise, because the model cannot cheaply dismiss it.9 Framework frames in a stack trace are that kind of high-similarity distractor.
Execution slicing: preprocess, do not prompt#
The deterministic alternative treats program output as structured data and slices it before inference:
- Stack trace reduction: parse the trace, keep frames inside your package roots, collapse framework frames to one-line markers, keep the exception type, the message, and the innermost user frame's local variables.
- Dynamic trace extraction: run the failing test under a tracer (
sys.settrace,coverage.py, eBPF, OpenTelemetry) and extract the execution path that actually ran, with the values that flowed through it. - Log windowing: anchor on the failure timestamp or request id and take a bounded, correlated window instead of the file.
Reach for stack trace reduction when
The failure raises. You have a traceback, and what you need is the frames in your own packages, the exception type, and the locals at the innermost user frame.
Reach for dynamic trace extraction when
The test fails without raising, or a legal-looking wrong value arrives somewhere downstream. You need the path that ran and the values that moved along it.
Reach for log windowing when
The failure happened in another process or another service. You have a request id or a timestamp, and you need the correlated window around it, bounded in lines.
# Instead of handing the model a shell and hoping:
# grep -rn "error" logs/ | tail -500 ← the model pays for every candidate line
result = run_code_with_trace( # deterministic, runs in milliseconds
cmd="pytest tests/test_billing.py::test_refund -x",
package_roots=["app/"], # slice to code you own
capture=["exception", "locals", "executed_lines", "sql"],
max_tokens=1200, # hard budget, enforced by code
)
context = result.to_prompt_block() # 1.2k tokens of signal
The difference is architectural, not cosmetic. grep | cat | tail makes the model the filter. You pay tokens for every candidate line, and the filtering step itself is probabilistic. run_code_with_trace() makes code the filter. Filtering costs no tokens, returns the same slice on the same input, and you can test it like any other function. Bash is a fine escape hatch. It is a poor primary retrieval engine, because each round trip is another model call that carries the whole conversation history with it.
SWE-agent (Yang et al., Princeton) showed that agent performance depends heavily on the agent-computer interface (ACI). Replacing raw shell interaction with purpose-built commands that return compact, structured views of files and search results improved resolution rates on SWE-bench.14 The model did not get smarter. The interface got stricter about what reached the context window. Anthropic's guidance on tool design for agents makes the same point: a tool should return a concise, high-signal representation and enforce token-efficient defaults, because an agent inherits every inefficiency of its tools.20
Yang et al., "SWE-agent," NeurIPS 2024 · Anthropic engineering, 2025
- Write a stack trace reducer. Parse the traceback with your language's own grammar, keep frames whose file path sits under a configured package root, and collapse the rest to one line each. The artifact is a function that takes a traceback string and returns a block under 1,500 tokens.
- Put a token budget on every tool result. Each tool returns at most N tokens and says how much it dropped and why. Done looks like a truncation counter you can graph, not a silent cut.
- Trace the failing test instead of reading its output. Run it under
coverage.pyorsys.settrace, keep executed lines inside your package roots, and hand the model the path rather than the log. - Correlate logs by request id. Emit a request id on every log line, then serve a bounded window around a failure instead of a file. If the id is missing today, adding it is the whole task.
context_tokens_per_tool_call: tokens returned by a tool, at p50 and p99. Down is good. The p99 is where an unboundedcathides.frames_kept_ratio: frames sent divided by frames in the raw trace. Down is good, as long as the innermost user frame survives every time.tool_calls_per_task: tool calls between the prompt and the first edit. Down is good, because each one carries the full history with it.truncation_rate: share of tool results that hit the budget. Watch both directions. Zero means the budget is slack, and a high rate means the slicer is not slicing.
Every line your system filters in code is a line the model does not pay to read and cannot be distracted by. Deterministic preprocessing moves work off the model and onto a CPU you already own.
Cite this
Anderson, M. (2026). Stop making agents read your entire logs. In Engineering Deterministic AI Coding Agents (2nd ed., Part 2). Oxagen Inc. https://macanderson.com/manual/stop-making-agents-read-your-logs
BibTeX
@incollection{anderson2026stopmakingagents,
author = {Anderson, Mac},
title = {Stop making agents read your entire logs},
booktitle = {Engineering Deterministic AI Coding Agents},
edition = {Second},
chapter = {2},
publisher = {Oxagen Inc.},
address = {Los Angeles, CA},
year = {2026},
url = {https://macanderson.com/manual/stop-making-agents-read-your-logs}
}Updates by email
Get the next edition of the field manual and new research when it is published.