Three agent tools can leak a secret together even when every pair of them is safe
Three smallest sets of our ten tools leaked a stored secret together, and only one of them is a pair, so a pair-by-pair check over the same three-call sequences would pass the other two.
An AI agent can be given a set of tools, each safe on its own. We took ten such tools and ran sessions of up to three calls against real files and a real database.
Why now
Agent platforms now let users mix tools from different sources; Simon Willison names the Model Context Protocol as encouraging exactly that in his "lethal trifecta" warning, cited below. A platform that approves tools one or two at a time cannot see a danger that needs three, and this result shows such a danger on real files and a real database.
What it shows
An agent platform can check tools one at a time. A tool that reads a secret is fine if nothing can send the secret out. A tool that posts to the network is fine if it never sees a secret. Trouble starts when the right tools are used together.
The setup
Our setup has ten tools, each safe on its own. One of them can post to the network. A session is dangerous when the stored secret, a credential, ends up sent out through that tool. We ran sessions of up to three tool calls, each against real files and a real SQLite database rather than a simulation.
The smallest dangerous combinations
The finding is about the smallest dangerous combinations: sets of tools that are dangerous together, where no smaller part of the set is. We found three. One is a pair. The other two are triples, and by the definition of minimal, every pair inside them is safe. A check that looks at tools two at a time, over sequences of up to three calls as ours did, would therefore miss both triples; longer pair sequences were not tried.
Why it matters, and to whom
This is for teams that decide which tools an AI agent may use together: agent platforms, runtime security products and internal platform teams. If the approval model reasons about pairs, it can pass combinations of three in which every pair looks safe, as two such triples did in our ten-tool setup. The result is a concrete, executed example of that gap.
Approvals, and their gap
Many platforms ask a person to approve a risky step. In our test an approval is used up by the first step it covers, so after one approval the posting tool leaks the secret once, not on every call. But the use is tracked only while the program runs: an approval reloaded from storage, or seen by another program, can be used again. Anyone who relies on such approvals needs their use recorded durably.
How it was checked
Every session was executed against real files and a real database rather than simulated, and the search covered sequences of up to three calls. The tools and the setup are our choice, so a different set of tools could give different combinations.
Testing the checker
We tested our own checker by corrupting the result on purpose. We flipped the recorded premise that every single tool is safe, the premise the whole finding rests on, and the check then reported a failure. The evidence file records that run, the command that re-runs the check, and its run on a clean copy of the code.
What it does not claim
- It does not claim that the idea is new. Simon Willison's "lethal trifecta" names a dangerous set of exactly three capabilities for agents: private data, untrusted content and external communication. Invariant Labs' toxic flow analysis notes that any combination of an agent's available tools may be used at run time. Much earlier, the Kirin security service for Android checked combinations of an app's permissions at install time, because some permissions are risky only together.
- It does not claim that any particular platform's tools behave like these ten. The strongest next test is to run the same search on a real platform's tool set.
- It does not claim that approvals stay used across a restart; the limits above describe that gap.
What the result adds is a measured instance on real files and a real database, and a clear demonstration that a pairwise check misses cases in this setup.
How to reproduce it
The evidence file names the command that re-runs the check. It needs our code at the commit named in the evidence file, and the check fails if any one of its steps fails.
To extend it, replace the ten tools with tools from a real agent platform and run the same search. If a triple appears whose pairs are all safe, the finding transfers to real tools. If none appears, the question stays open for that platform. Either answer is worth having before a platform chooses how to approve tools.
Formal statement
Minimal means that no smaller part of the combination is dangerous on its own; in particular, removing any one tool makes it safe again. We found three such combinations, and only one of them has two members.
Prior art
- Simon Willison, "The lethal trifecta for AI agents: private data, untrusted content, and external communication": a published dangerous set of exactly three agent capabilities
- Beurer-Kellner, Milanta and Fischer, "Invariant Labs Exposes Novel Prompt Injection Attack Vulnerabilities, 'Toxic Flows,' in Agentic Systems & MCP Servers" (Invariant Labs): notes that any combination of an agent's available tools may be used at run time
- Enck, Ongtang and McDaniel, "On lightweight mobile phone application certification" (the Kirin system): install-time checking of combinations of permissions on a phone
Evidence
- The evidence file for this result
- Our claim file, with its limits
- All our published files, each with a fingerprint you can check
Related results across the group
- Random crashes found unclean recoveries in our AI agent's transaction layer; every crash point we chose recovered cleanly
- A permission grant cut to a minimum: for one workload every kept permission proved needed, for the other the result is partial
- Six slightly weakened bounds checks pass random fuzzing, and two coverage-guided fuzzers find all six
- The same result on VerifyCore Labs, our parent lab