Skip to content

Three agent tools can leak a secret together even when every pair of them is safe

Left: the ten tools in our test, each safe alone. Right: the three smallest sets that leaked the stored secret. One is a pair. In the other two, every pair is safe within our three-call search and only the three together leak.

Three smallest sets of our ten tools leaked a stored secret together, and only one of them is a pair, so a pair-by-pair check over the same three-call sequences would pass the other two.

An AI agent can be given a set of tools, each safe on its own. We took ten such tools and ran sessions of up to three calls against real files and a real database.

Why now

Agent platforms now let users mix tools from different sources; Simon Willison names the Model Context Protocol as encouraging exactly that in his "lethal trifecta" warning, cited below. A platform that approves tools one or two at a time cannot see a danger that needs three, and this result shows such a danger on real files and a real database.

What it shows

An agent platform can check tools one at a time. A tool that reads a secret is fine if nothing can send the secret out. A tool that posts to the network is fine if it never sees a secret. Trouble starts when the right tools are used together.

The setup

Our setup has ten tools, each safe on its own. One of them can post to the network. A session is dangerous when the stored secret, a credential, ends up sent out through that tool. We ran sessions of up to three tool calls, each against real files and a real SQLite database rather than a simulation.

The smallest dangerous combinations

The finding is about the smallest dangerous combinations: sets of tools that are dangerous together, where no smaller part of the set is. We found three. One is a pair. The other two are triples, and by the definition of minimal, every pair inside them is safe. A check that looks at tools two at a time, over sequences of up to three calls as ours did, would therefore miss both triples; longer pair sequences were not tried.

Why it matters, and to whom

This is for teams that decide which tools an AI agent may use together: agent platforms, runtime security products and internal platform teams. If the approval model reasons about pairs, it can pass combinations of three in which every pair looks safe, as two such triples did in our ten-tool setup. The result is a concrete, executed example of that gap.

Approvals, and their gap

Many platforms ask a person to approve a risky step. In our test an approval is used up by the first step it covers, so after one approval the posting tool leaks the secret once, not on every call. But the use is tracked only while the program runs: an approval reloaded from storage, or seen by another program, can be used again. Anyone who relies on such approvals needs their use recorded durably.

How it was checked

Every session was executed against real files and a real database rather than simulated, and the search covered sequences of up to three calls. The tools and the setup are our choice, so a different set of tools could give different combinations.

Testing the checker

We tested our own checker by corrupting the result on purpose. We flipped the recorded premise that every single tool is safe, the premise the whole finding rests on, and the check then reported a failure. The evidence file records that run, the command that re-runs the check, and its run on a clean copy of the code.

What it does not claim

  • It does not claim that the idea is new. Simon Willison's "lethal trifecta" names a dangerous set of exactly three capabilities for agents: private data, untrusted content and external communication. Invariant Labs' toxic flow analysis notes that any combination of an agent's available tools may be used at run time. Much earlier, the Kirin security service for Android checked combinations of an app's permissions at install time, because some permissions are risky only together.
  • It does not claim that any particular platform's tools behave like these ten. The strongest next test is to run the same search on a real platform's tool set.
  • It does not claim that approvals stay used across a restart; the limits above describe that gap.

What the result adds is a measured instance on real files and a real database, and a clear demonstration that a pairwise check misses cases in this setup.

How to reproduce it

The evidence file names the command that re-runs the check. It needs our code at the commit named in the evidence file, and the check fails if any one of its steps fails.

To extend it, replace the ten tools with tools from a real agent platform and run the same search. If a triple appears whose pairs are all safe, the finding transfers to real tools. If none appears, the question stays open for that platform. Either answer is worth having before a platform chooses how to approve tools.

Formal statement

Minimal means that no smaller part of the combination is dangerous on its own; in particular, removing any one tool makes it safe again. We found three such combinations, and only one of them has two members.

Prior art

Evidence

Related results across the group

All results from OrbitalProof · The OrbitalProof home page

How we show numbers

Every result figure links to the file it comes from. See our published files