# Bounded tasks and ongoing work

> Some work has an endpoint. Other work continues. Manage each in its own way.

Part 20: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/bounded-tasks-and-ongoing-work

You operate two kinds of agent and they tend to get one kind of management. The first fixes a bug, runs a migration, or writes a report. It starts, works, and stops. The second triages an inbox, watches a queue, or keeps dependencies current. It has no last step. Ask "is it done?" of the first and the question has an answer. Ask it of the second and the honest reply is that the question does not apply.

### Tell the two apart first

|  | Bounded task | Ongoing responsibility |
| --- | --- | --- |
| Shape | Has a start and an end state you can describe | Runs until someone stops it |
| Examples | Fix issue 4821, migrate a table, draft the quarterly report | Triage refunds, keep dependencies patched, answer tier-1 tickets |
| The question | Did the checks hold? | Is it inside its authority and budget, and what is it waiting on? |
| Managed through | Completion criteria set before the work starts | Authority, budget, and review points |
| Reviewed | At the end of the run | On a schedule, and when a threshold trips |

The common mistake runs in one direction: inventing a finish line for ongoing work so that it fits a task tracker. An inbox agent with a "done" state will report done. That tells you nothing about the inbox.

### Bounded tasks: define completion before the work starts

Part 10 made tests the ground truth of behavior and had the agent write the failing test first. Generalize that. Before a bounded run begins, write down the checks that decide completion, and fix them so the run cannot revise them.

-   **A command that must pass.** `pytest tests/test_billing.py` exits 0.
-   **A file or a diff condition.** The migration exists. No file outside `billing/**` changed.
-   **A human decision.** The release owner approves the pull request.

Fixing the checks first matters for a plain reason. An agent that judges its own work can stop early or skip verification. The Berkeley taxonomy names this family of failures: premature termination, and no or incomplete verification, both recorded across the frameworks they studied.[5](https://macanderson.com/manual/sources#r5 "Cemri et al. (UC Berkeley). \"Why Do Multi-Agent LLM Systems Fail?\" (MAST). NeurIPS 2025 Datasets & Benchmarks. arXiv:2503.13657") A check the run cannot edit removes the agent's opinion from the verdict. The verdict becomes a function of recorded results, and another person can recompute it from the record (part 19).

```
# done/issue-4821.yaml     written before the run, hashed when it opens
task: "Fix DecimalConversionError on POST /api/v2/refunds"
checks:
  - run:   "pytest tests/test_billing.py::test_refund -x"
  - run:   "ruff check billing/"
  - diff:  { only_paths: ["billing/**", "tests/**"] }
  - human: { role: release-owner, question: "Safe to ship in 2026.09?" }
on_fail: { retries: 3, then: route_to_operator }
```

> **Counterweight · What a passing verdict means**
>
> A passing verdict means the specified checks held. It does not establish that every requirement was captured or that every action was correct. Your team decides whether those checks are enough for the task, and a thin set of checks passes thin work. When checks fail, the useful record is not just "failed". It is the reasons, and whether the agent continued, waited for a person, or reached its retry limit.

### Ongoing work: authority, budget, and review points

An ongoing agent is managed the way you manage a standing responsibility held by a person. You do not ask whether they are done. You look at what they are allowed to do, what they have spent, what they are waiting on, and what they did since you last looked.

-   **Authority.** The mandate from part 15, reviewed on a date, not only when something goes wrong.
-   **Budget.** The monthly limit and the per-run ceiling from part 17, with the `at_limit` action stated.
-   **Open requests.** The routed requests from part 16 that a person has not yet answered. This is the list an operator reads first each morning.
-   **Review points.** A schedule (weekly for a new agent, monthly once it is stable) and a few thresholds that trigger a review early.

| Review trigger | Reading | Usual cause |
| --- | --- | --- |
| Denials rise week over week | `requests_by_answer` | A tool or a task changed and the rules did not |
| Run cost p99 leaves its band | `cost_per_run_p99` | A loop, often on a failing check or a stale index |
| Routed requests wait longer | `time_to_answer_p95` | The named approver changed roles, or the rule routes too much |
| Tool use ratio falls | `tool_use_ratio` | The work drifted away from the equipment |
| The operator changes | Inventory | Nobody has re-read the mandate since |

### The operator's week

Put together, the operating job is small and regular. Daily, answer what is waiting. Weekly, read the spend and the denials for each agent, and review the runs that tripped a threshold. Monthly, re-read each mandate with its three owners, retire the agents nobody can justify, and turn repeated approvals into narrower rules. The field kit has this as a one-page checklist. None of it requires reading transcripts, which is the test of whether parts 14 to 19 are in place.

> **Where Oxagen fits**
>
> Oxagen is where that week happens. Every agent enrolled in Oxagen is on the Runs page with its mandate, its open requests, its spend, and its last run. An operator answers routed requests from the Access page, and an ongoing agent shows its authority, its open requests, and its spend this month with no finish line. Oxagen treats completion checks as an optional control for bounded tasks. They are not the definition of the product, and ongoing work does not need them.

> **Do this week**
>
> 1.  **Sort your agents into the two columns.** Use the inventory from part 14. Anything you cannot place is doing both jobs and should be split.
> 2.  **Write completion checks for one bounded task type.** Start with the one that reaches review most often. Three checks are enough: a command, a diff condition, and a human decision if the change ships.
> 3.  **Fix the checks before the run.** Hash the file when the run opens and record the digest as the first row. Reject a run whose checks changed midway.
> 4.  **Set two review triggers for one ongoing agent.** Pick from the table. Send the alert to the operator by name.
> 5.  **Put the monthly mandate review on a calendar.** Thirty minutes, three owners, one agent at a time.

> **Measure it**
>
> -   `checks_held_first_attempt`: bounded runs whose checks held without a retry, over all bounded runs. This is `first_attempt_rate` from part 13, against criteria the run could not edit.
> -   `runs_closed_without_checks`: bounded runs that ended with no completion criteria at all. Each is a verdict by assertion.
> -   `open_requests_age`: the oldest unanswered routed request, per operator.
> -   `days_since_mandate_review`: per agent. Over 90 is a review that is due.

> **Takeaway**
>
> Define completion when the work has an endpoint, and fix the checks before the run starts. When the work continues, do not invent a finish line. Manage its authority, its requests, and its spend on a schedule.
