# Equip agents on purpose

> A tool in the catalog is not a tool in the agent's hands. Assign equipment the way you assign access.

Part 18: Operating the workforce. From *Engineering Deterministic AI Coding Agents*, second edition, by Mac Anderson. Canonical page: https://macanderson.com/manual/equip-agents-on-purpose

Your platform team has built the tools this book describes: a trace slicer, a code graph, a schema index, a test lookup. They are registered once and every agent gets all of them, along with forty others that accumulated. The refund agent can see the deploy tool. The docs agent can see the database tool. Nobody decided that. It is what happens when the catalog and the assignment are the same list.

### The tool list is context, and it is the largest part

Every tool an agent can see is sent with every request: its name, its description, and its input schema. Part 7 showed what irrelevant context costs, and tool definitions are context like any other. One measurement from Oxagen's own in-app agent, taken by hand because nothing reported it: 271 tools came to 45,007 tokens, which was 92.4% of the cacheable prefix. Providers also cap the count. OpenAI's limit is 128 tools per request, and a request over the cap fails at the provider with an error about a payload nobody on your side can inspect.

Anthropic's guidance on writing tools makes the design point: more tools do not always lead to better outcomes, and it recommends building a few tools aimed at the workflows that matter most.[20](https://macanderson.com/manual/sources#r20 "Anthropic Engineering. \"Writing effective tools for agents.\" 2025. anthropic.com/engineering") The assignment is an engineering decision. Make it in the mandate, where it can be reviewed.

### Three kinds of equipment

| Kind | What it is | Assigned by |
| --- | --- | --- |
| Tools | Callable actions with typed inputs: `schema_slice`, `tests_covering`, `open_pull_request` | A list in the equipment clause |
| Skills | Packaged instructions for one kind of task, loaded when that task comes up, not on every request | A list in the equipment clause |
| Business context | The knowledge the work needs: domain dossiers from part 5, the graph from part 12, conventions, decisions | A scope, such as `domain:billing` |

Skills follow the argument of part 6. A rule that applies to billing migrations should reach the agent when it touches a billing migration, and stay out of the window otherwise. A skill is that rule with a trigger.

Context scope follows the argument of part 12. A graph makes retrieval exact, and it also makes permission exact. If knowledge is nodes with domains attached, then "this agent may read `domain:billing`" is a filter on a traversal. The same filter serves relevance and authority, and you maintain it once.

### Assignment is not permission

Equipment and access are different clauses for a reason. Assigning `open_pull_request` to an agent means the tool is in its hands. Whether a particular pull request may be opened is still a request that a rule answers (part 16). Keep both checks. The first keeps the window small and the agent on task. The second decides each action at the moment of use.

```
equipment:                       # owner: engineering
  tools:
    - run_code_with_trace        # part 2
    - schema_slice               # part 8
    - tests_covering             # part 10
    - open_pull_request          # each call still meets the rules
  skills:
    - billing-conventions        # loads when a diff touches billing/**
    - migration-review
  context_scope:
    - "domain:billing"
    - "adr:*"
  tool_budget:
    max_tools: 24                # a build failure over this, not a warning
    max_tokens: 6000
```

> **Evidence · Interfaces decide outcomes**
>
> SWE-agent's central result was that the agent-computer interface changes what the same model can do: purpose-built commands with compact, structured output raised resolution rates on SWE-bench over raw shell access.[14](https://macanderson.com/manual/sources#r14 "Yang, Jimenez et al. (Princeton). \"SWE-agent: Agent-Computer Interfaces Enable Automated Software Engineering.\" NeurIPS 2024. arXiv:2405.15793") Anthropic's tool-writing guidance adds the operating advice: choose tools deliberately, namespace them, return high-signal output, and keep token use low by default, because an agent inherits the inefficiency of its tools.[20](https://macanderson.com/manual/sources#r20 "Anthropic Engineering. \"Writing effective tools for agents.\" 2025. anthropic.com/engineering") The Model Context Protocol moves in the same direction by having each server declare its tools, their schemas, and their descriptions, so that discovery is a lookup.[25](https://macanderson.com/manual/sources#r25 "Model Context Protocol. Specification: servers, tools, and schemas. modelcontextprotocol.io/specification") A declared catalog is what makes a per-agent assignment possible.
>
> Yang et al., NeurIPS 2024 · Anthropic engineering, 2025 · Model Context Protocol specification

### Check that the equipment reached the work

An assignment is a claim until a run shows it. For each run, record which tools were offered, which were called, which skills loaded, and which context scopes were read. Two findings come out of that data quickly. Tools that are offered on every run and called on none are candidates to remove. Tools that an agent asks for and does not have show up as failed attempts or as workarounds through a shell.

> **Counterweight · Context is scoped, not universal**
>
> Giving an agent business context does not mean it recalls it correctly, and it does not mean every agent should have it. Storing a fact is not the same as retrieving it at the right moment, which is why parts 3 to 5 exist. Scope context to what the agent's work requires and what its mandate permits, and measure retrieval precision for the context layer the same way you measure it for code.

> **Where Oxagen fits**
>
> In Oxagen, tools and skills reach an agent through its mandate. A tool in the catalog does not grant every agent permission to use it, and each request still meets the decision rules. Business context is served within the scope the mandate permits. Oxagen wraps four harnesses: Claude Code, Codex, Cursor, and Stella. What a wrapper can gate and meter differs by harness, and the product says which.

> **Do this week**
>
> 1.  **Count the tools and the tokens.** For one agent, serialize the tool list exactly as it is sent and count both. Done is two numbers you did not have yesterday.
> 2.  **Log offered against called.** Thirty days of runs is enough. Sort tools by calls per run.
> 3.  **Remove the tools with no calls.** Take them out of that agent's assignment, not out of the catalog. Watch `first_attempt_rate` from part 13 for a week.
> 4.  **Move one always-loaded rule into a skill.** Pick the longest block of instructions that applies to one kind of task. Give it a trigger and load it on demand.
> 5.  **Add a tool budget to CI.** Fail the build when an agent's assignment exceeds a count or a token limit, or a provider's cap.

> **Measure it**
>
> -   `tool_tokens_per_request`: tokens spent on tool definitions, per agent. Compare it with the agent's median task context.
> -   `tool_use_ratio`: distinct tools called over tools offered, per run. A low ratio means the assignment is wider than the work.
> -   `skill_load_rate`: runs in which a skill loaded, over runs where its trigger matched. Under 1.0 means the trigger is wrong.
> -   `out_of_scope_reads`: context requests outside the agent's scope. Each one is either a scope to widen or a task the agent should not have.

> **Takeaway**
>
> Give each agent the tools and skills its work requires, and the business context its mandate permits. Keep the catalog and the assignment as separate lists, and check the assignment against what the runs used.
