Safe on its own is not the same as safe together
A tool can be harmless by itself and still be one piece of something harmful.
Picture an AI agent with a handful of tools. One reads a file that holds a secret. One posts a message to the network. Alone, the first cannot send anything out, and the second has nothing worth sending. Together they can move the secret out of the building. Neither tool is the problem. The combination is.
Security people have named this pattern before. Simon Willison’s lethal trifecta describes three capabilities that are dangerous together: access to private data, exposure to untrusted content, and a way to communicate outward. Willison also points out that the Model Context Protocol encourages users to mix and match tools from different sources that can do different things. The idea is not new. What is useful is to measure how it behaves in a controlled setting.
What we searched
We built a test world and searched it for dangerous combinations of up to three calls.
Its result page gives the setup in plain words. The setup had a few plain parts:
- Ten tools, each safe on its own.
- Sequences of up to three calls.
- Real files and a real SQLite database, not a simulation.
Our claim counts 2,379 executed sessions against real files and a real SQLite database. A session counts as dangerous when the stored secret, a credential, ends up sent out through the network-posting tool.
Using real files and a real database matters, and our claim file is published on this site. A simulation can hide the way data moves through storage.
We then looked for sets of tools that are dangerous together and where no smaller part of the set is dangerous. Such a set is called minimal. It is the smallest unit of danger, because removing any one tool makes it safe again.
Why pairs are not enough
The reason is that danger can need all three tools at once. For example, the tool that reads the secret and the tool that sends data out may have no link until a third tool connects them. No pair does harm alone; all three do.
A review that stops at pairs looks at the smaller groups and assumes the larger ones follow from them. They do not.
Put simply, safety is not inherited by a bigger set: the danger can live in the way the parts feed each other. A secret read by one tool becomes input to the next, and the second tool never needed to be dangerous to be useful to an attacker.
What it found
Our claim reports that the search found three minimal dangerous combinations, and that only one of them is a pair.
That sentence carries the whole point. Two of the three are triples. By the definition of minimal, every pair inside a triple is safe. A reviewer who checks tools two at a time, over sequences of up to three calls as we did, would therefore approve both triples, because each pair in them looks fine; we did not test pairs over longer sequences.
In plain words, a check that looks only at pairs has a blind spot, and we found it by trial in a concrete setting. This also matches an observation from Invariant Labs, which notes that any combination of an agent’s available tools may be used at run time. What an agent can do is a property of its whole set of tools, not of each tool in turn.
The idea also has a long history on phones. Kirin checked combinations of permissions when an app was installed, for the same reason: some permissions are only risky together.
Our limits, in plain words: we chose the ten tools and the setup, sequences longer than three calls were not tried, and in our recorded run an approval was not used up by the step it covered.
Approvals and their gap
Our own limits point at one more gap: an approval can be used more than once.
Many agent platforms ask a human to approve a risky step. In the run our evidence file records, an approval was not consumed: one approval followed by three posting calls put three copies of the credential in the sink. The lab’s later claim says a later version consumes an approval at the first step it covers, but only while the program runs, so an approval rebuilt from its saved form, or shown to another process, is unused again. No run record of that later version is published. Anyone who relies on approvals needs the use of each one recorded durably.
What this means for you
If you decide which tools an agent may combine, review sets, not only single tools.
- Agent-platform teams can add a set-level review to their approval process, and treat a clean pair check as incomplete.
- Runtime-security teams can test their own tool sets the way we did: run sequences of up to three calls against real files and a real database, and look for a combination whose pairs are all safe.
- Buyers and diligence teams can ask a vendor how approvals are recorded, and whether a restart resets them.
What this does not show
The result does not show that the idea is new, and we do not claim it is.
It also does not show that any particular platform’s tools behave like these ten. We chose the tools and the setup, so a different set could give different combinations. The search covers sequences of up to three calls, so longer sequences are not covered. The strongest next step is to run the same search on a real platform’s tool set. If a triple with all-safe pairs appears there, the finding transfers. If none appears, the question stays open for that platform.