A permission grant cut to a minimum: for one workload every kept permission proved needed, for the other the result is partial
A program can be handed every permission its tools declare. The lab cuts a larger declared grant down to a smaller one, then tests the result by removing each kept permission and running the workload again. For one workload the kept set passes that test in full. For the other the result is partial, because one place in the code that could ask for a permission was never exercised.
What we found
Starting from a declared grant of 32 permissions, the lab cut each workload's grant down, then removed every kept permission in turn and re-ran the workload. For one workload every kept permission proved needed. For the other the verdict is partial.
Why now
AI agents and automation tools are now given standing permissions to files, credentials and network endpoints. A grant that is larger than the work needs is risk with no benefit, and a method that tests each kept permission by removing it shows which ones the work really uses.
What it shows
A permission grant says what a program may touch. The lab's gate file, served beside the result, names six kinds of permission it controls: which tools may run, what kind of action each may take, which credentials it may use, which sources of input it may read, which network endpoints it may reach, and which file paths it may touch. The claim calls each single permission an atom.
The removal test
The method has two steps. First, it cuts a larger declared grant down to a smaller one for each workload. Second, it tries to break its own answer. It removes each kept permission in turn and runs the workload again. If the workload still succeeds without a permission, that permission was not needed, and the set was not minimal. The gate file records that a removed permission stops the effect in the world itself: a file is not written, a message is not sent, a row is not inserted.
Two workloads, two verdicts
The lab's claim, in our claim file, reports the outcome for two workloads. For the first, the set is certified minimal in the sense of the removal test: every kept permission was needed. For the second, the verdict is partial. The lab's limits, in the same file, say why: one place in that workload's code that could ask for a permission was never exercised by the runs, so the grant is minimal only over the behaviour that was tested.
Why it matters, and to whom
This is for teams that set permissions for programs that call tools, such as agent platforms and cloud identity teams. A grant cut to a minimal set is smaller, and for the certified workload the removal test is the evidence that each remaining permission is needed.
A field that could not fail
The limits also say that one field the lab used to mark the work as done could never come out false: it follows from how the grant is built, so it is not evidence on its own. The evidence is the removal test, not that field.
How it was checked
The gate file lists each check the lab runs and its result, including that every one of the six kinds of permission is enforced on its own and reduced by the cut. The evidence file records the command, a run on a clean copy of the code, and a note that the check asserts both verdicts, minimal and partial, instead of only the good one.
The lab also tested the check by editing its recorded audit to call a needed permission unnecessary. Because the removal step re-runs the workload and the workload still needs that permission, the check reported a failure.
What it does not claim
- It does not claim that the grant is enough for other inputs. It is checked against the lab's chosen workloads, and an input that takes a different path may need permissions those workloads never used. The Confine study of system-call policies for containers makes the same point: policies derived from observed runs do not capture all the code that can potentially be needed.
- It does not claim the smallest possible grant in every sense. For the certified workload, the removal test shows that no single kept permission can be dropped without breaking the workload; that is the sense of minimal the claim uses.
- It does not claim that the permission kinds were always independent. The limits box above repeats the earlier failure the lab disclosed, and what its current gate file records for that case.
- It does not claim that the method is new. Amazon's Identity and Access Management (IAM) Access Analyzer generates policies from recorded access activity. Progent applies privilege control to an AI agent's tool calls. What this result adds is the removal test across six kinds of permission, and a report that keeps the partial verdict in view.
How to reproduce it
The evidence file names the command that re-runs the check. It needs the lab's code at the commit named in the evidence file.
The most useful next test is to run each workload across every kind of input it is meant to handle, including rare ones, so that the place in the second workload's code that was never exercised is reached. If the grant then fails, it was enough only for the inputs the lab chose.
Prior art
- Mathangi Ramesh, "IAM Access Analyzer makes it easier to implement least privilege permissions by generating IAM policies based on access activity" (AWS Security Blog): generating least-privilege policies from recorded access activity
- Shi, He, Wang, Li, Wu, Guo and Song, "Progent: Securing AI Agents with Privilege Control": least-privilege policies over an AI agent's tool calls
- Ghavamnia, Palit, Benameur and Polychronakis, "Confine: Automated System Call Policy Generation for Container Attack Surface Reduction": why policies derived from observed runs miss rare conditions
Evidence
- The evidence file for this result
- Our claim file, with its limits
- All our published files, each with a fingerprint you can check
Related results across the group
- Three agent tools can leak a secret together even when every pair of them is safe
- Random crashes found unclean recoveries in our AI agent's transaction layer; every crash point we chose recovered cleanly
- Thousands of candidate code patches scored by proof: each one certified sound, proven unsound, or left undecided