Skip to content

Why a set of AI agent tools can be unsafe even when every pair is safe

Explainer 5 min read

Checking agent tools one at a time, or two at a time, can pass a combination of three that leaks a secret.

Left: the ten tools in our test, each safe alone. Right: the three smallest sets that leaked the stored secret. One is a pair. In the other two, every pair is safe within our three-call search and only the three together leak.

A dotted underline marks a number read straight from a published file when this page was built.

In this post
  1. Safe on its own is not the same as safe together
  2. What we searched
  3. Why pairs are not enough
  4. What it found
  5. Approvals and their gap
  6. What this means for you
  7. What this does not show

Safe on its own is not the same as safe together

A tool can be harmless by itself and still be one piece of something harmful.

Picture an AI agent with a handful of tools. One reads a file that holds a secret. One posts a message to the network. Alone, the first cannot send anything out, and the second has nothing worth sending. Together they can move the secret out of the building. Neither tool is the problem. The combination is.

Security people have named this pattern before. Simon Willison’s lethal trifecta describes three capabilities that are dangerous together: access to private data, exposure to untrusted content, and a way to communicate outward. Willison also points out that the Model Context Protocol encourages users to mix and match tools from different sources that can do different things. The idea is not new. What is useful is to measure how it behaves in a controlled setting.

What we searched

We built a test world and searched it for dangerous combinations of up to three calls.

Its result page gives the setup in plain words. The setup had a few plain parts:

  • Ten tools, each safe on its own.
  • Sequences of up to three calls.
  • Real files and a real SQLite database, not a simulation.

Our claim counts 2,379 executed sessions against real files and a real SQLite database. A session counts as dangerous when the stored secret, a credential, ends up sent out through the network-posting tool.

Using real files and a real database matters, and our claim file is published on this site. A simulation can hide the way data moves through storage.

We then looked for sets of tools that are dangerous together and where no smaller part of the set is dangerous. Such a set is called minimal. It is the smallest unit of danger, because removing any one tool makes it safe again.

Why pairs are not enough

The reason is that danger can need all three tools at once. For example, the tool that reads the secret and the tool that sends data out may have no link until a third tool connects them. No pair does harm alone; all three do.

A review that stops at pairs looks at the smaller groups and assumes the larger ones follow from them. They do not.

Put simply, safety is not inherited by a bigger set: the danger can live in the way the parts feed each other. A secret read by one tool becomes input to the next, and the second tool never needed to be dangerous to be useful to an attacker.

What it found

Our claim reports that the search found three minimal dangerous combinations, and that only one of them is a pair.

That sentence carries the whole point. Two of the three are triples. By the definition of minimal, every pair inside a triple is safe. A reviewer who checks tools two at a time, over sequences of up to three calls as we did, would therefore approve both triples, because each pair in them looks fine; we did not test pairs over longer sequences.

In plain words, a check that looks only at pairs has a blind spot, and we found it by trial in a concrete setting. This also matches an observation from Invariant Labs, which notes that any combination of an agent’s available tools may be used at run time. What an agent can do is a property of its whole set of tools, not of each tool in turn.

The idea also has a long history on phones. Kirin checked combinations of permissions when an app was installed, for the same reason: some permissions are only risky together.

Our limits, in plain words: we chose the ten tools and the setup, sequences longer than three calls were not tried, and in our recorded run an approval was not used up by the step it covered.

Approvals and their gap

Our own limits point at one more gap: an approval can be used more than once.

Many agent platforms ask a human to approve a risky step. In the run our evidence file records, an approval was not consumed: one approval followed by three posting calls put three copies of the credential in the sink. The lab’s later claim says a later version consumes an approval at the first step it covers, but only while the program runs, so an approval rebuilt from its saved form, or shown to another process, is unused again. No run record of that later version is published. Anyone who relies on approvals needs the use of each one recorded durably.

What this means for you

If you decide which tools an agent may combine, review sets, not only single tools.

  • Agent-platform teams can add a set-level review to their approval process, and treat a clean pair check as incomplete.
  • Runtime-security teams can test their own tool sets the way we did: run sequences of up to three calls against real files and a real database, and look for a combination whose pairs are all safe.
  • Buyers and diligence teams can ask a vendor how approvals are recorded, and whether a restart resets them.

What this does not show

The result does not show that the idea is new, and we do not claim it is.

It also does not show that any particular platform’s tools behave like these ten. We chose the tools and the setup, so a different set could give different combinations. The search covers sequences of up to three calls, so longer sequences are not covered. The strongest next step is to run the same search on a real platform’s tool set. If a triple with all-safe pairs appears there, the finding transfers. If none appears, the question stays open for that platform.

Results behind this post

All results from OrbitalProof · The same result on VerifyCore Labs, our parent lab

Keep reading

All posts

Sources and further reading

Where the numbers come from

Each number with a dotted underline was read from one of these published files, field by field, when the page was built.

Every file this site publishes

Ask about a result, or check one yourself

Every result on this site links the published files behind it, and its page says what each file covers. Acquisition, licensing and partnership enquiries go to one address, and a person reads it.

Write to us Read the research results

How we show numbers

Every result figure links to the file it comes from. See our published files