Skip to content

How to approve AI agent tools as a set, not one at a time

Explainer 5 min read

Approving an agent's tools as a set means asking what the whole kit can do, and recording each approval so that it cannot be reused by accident.

Left: the ten tools in our test, each safe alone. Right: the three smallest sets that leaked the stored secret. One is a pair. In the other two, every pair is safe within our three-call search and only the three together leak.

A dotted underline marks a number read straight from a published file when this page was built.

In this post
  1. Start from what the whole kit can do
  2. Look for data that can travel
  3. Keep the review small and repeatable
  4. Test your own sets
  5. Record each approval
  6. What this means for you
  7. What this does not show

Start from what the whole kit can do

Approving a set of tools starts by asking what the agent could do with all of them, in any order.

The usual habit is to review tools one by one. Does this tool read files? Does that one post to the network? Each answer is fine, and each tool is signed off. The approval attaches to the tool, so the question about the whole kit never gets asked.

Our result shows why that matters. We took ten tools, each safe on its own, and ran 2,379 executed sessions of up to three calls. Our claim reports three smallest dangerous combinations, and only one is a pair. The first post in this series explains the finding.

Look for data that can travel

A useful question is whether anything the agent can read can also leave.

The pattern that Simon Willison calls the lethal trifecta is a good starting list: private data, untrusted content and a way to communicate outward. Our combinations share two of its parts, private data and a way out; our test supplied no untrusted content.

Keep the review small and repeatable

A set-level review does not need to be heavy. Three habits are a reasonable start (untested by us):

  • Write down, for each tool, what it can read and what it can send out.
  • Record, durably, when an approval has been used. In the run our evidence file records, one approval let three posting calls through; the lab’s later claim says a later version marks it used only while the program runs, and no run record of that version is published. If your platform relies on single-use approvals, check that each use is recorded and survives a restart.
  • Re-run the review whenever a tool is added, since a new tool changes every set it joins.

The review is only as good as the list of tools it covers, so keep that list current, and treat any tool that can be added at run time as part of the set.

Tool sets change, and a change can bite in two ways:

  • A team adds a tool to fix one problem and, without meaning to, completes a dangerous combination.
  • Because danger lives in sets, one new tool can create several new sets at once.

That is why the review has to be repeated, not done once. Treat a new tool the way you would treat a new dependency: look at what it reads, what it sends, and which existing tools it can hand data to. Then run the same short-sequence test on the new kit and compare it with the last one.

Test your own sets

One direct way to find out is to run short sequences against real systems and look.

We ran sequences of up to three calls, which is feasible when the set is small. For a larger set, a team can still sample sequences and watch for the signature we found: a group of tools where every pair passes the pair check but the group does not pass. Invariant Labs makes a related point in its toxic flow analysis, noting that any combination of an agent’s available tools may be used at run time.

A test world with a known secret, a known outlet and real files is small compared with the platform it tests. It turns a worry about what might happen into a count of what did.

Record each approval

If your platform asks a human to approve risky steps, record each approval where it will survive a restart.

In the run our evidence file records, an approval was not used up: one approval followed by three posting calls put three copies of the credential in the sink. The lab’s later claim says a later version consumes an approval at the first step it covers, but only while the program runs, so an approval rebuilt from its saved form, or shown to another process, is unused again. No run record of that later version is published.

Our limits, in plain words:

  • We chose the ten tools and the setup.
  • Sequences longer than three calls were not tried.
  • In our recorded run an approval was not used up by the step it covered.

That last gap matters in practice. A platform that restarts its agent process, or hands a saved approval to a second worker, would treat a used approval as fresh. One way to close it, which we have not tested, is to consume each approval atomically in a durable store, with one compare-and-set or transaction per use, so that two workers cannot both use it.

The lesson is plain. An approval that can be replayed is not an approval. Record when each one is used, and check that record on every step, not only in the process that first received it.

What this means for you

Treat tool approval as a property of the whole configuration.

  • Agent-platform teams can review the set and keep a record of which combinations were approved, not only which tools.
  • Security teams can build a small test world and run the combinations through it before an agent ships.
  • Buyers and diligence teams can ask a vendor to show how combinations of tools are reviewed and how approvals are stored.

What this does not show

The result does not show that any real platform has these combinations.

We chose the ten tools and the test for danger, so a different set could give different answers. The search covers sequences of up to three calls only. The advice in this post is a way of using the finding, not a result we measured, and we have not tested the review steps described here. The first thing to measure is the same search on a real tool set. The result page lists the claim, the limits and the evidence files.

Ask about a result, or check one yourself

Every result on this site links the published files behind it, and its page says what each file covers. Acquisition, licensing and partnership enquiries go to one address, and a person reads it.

Write to us Read the research results

How we show numbers

Every result figure links to the file it comes from. See our published files