Skip to content

Six slightly weakened bounds checks pass random fuzzing, and two coverage-guided fuzzers find all six

One square per weakened check. Top: uniform random fuzzing, 20,000 draws each, found none. Middle: libFuzzer found all six, in every one of 30 runs. Bottom: AFL++ found all six in every one of 30 runs.

Six weakened checks each passed 20,000 random fuzz draws with no hit, and coverage-guided fuzzers found all six in every run.

A bounds check that is wrong by one can let through so few bad inputs that random testing almost never finds them. The lab built six such weakened checks, modelled on checks in widely used open-source libraries, and measured how well different fuzzers catch them. Uniform random draws, 20,000 per check, missed all six. Two coverage-guided fuzzers found all six in every run they made.

Why now

Teams increasingly rely on automated testing to accept patches, including patches written by AI tools. A patch that is wrong by one can pass random testing; this measures that gap for six off-by-one checks the lab built, and whether two coverage-guided fuzzers find them.

What it shows

A bounds check decides whether an index or a length is safe to use. If a patch gets the check wrong by one, only inputs on the exact edge get through. Random testing draws inputs from a huge space, so it can run for a very long time without landing on that edge.

The random battery

The lab built six weakened checks, one each for models of checks in OpenSSL, libxml2, libpng, zlib, sudo and curl, each made wrong by a small slack so that very few inputs break it. Each passed a battery of 20,000 uniform random draws with no hit, 120,000 draws in all.

For three of the six, the claim counts the gap exactly: the chance that one random draw hits is a tiny fraction, given in the lab's claim in our claim file. For the other three, the gap is only estimated, and the claim gives no hit probability for them.

Coverage-guided fuzzers

Coverage-guided fuzzers do not draw blindly. They keep inputs that reach new code and change them a little at a time. At default settings, libFuzzer found all six weakened checks in 30 runs, within at most 27,898 executions. AFL++ found all six in all 30 of its runs. Every input they found replays in exact arithmetic as a case that breaks its weakened check.

Why it matters, and to whom

This is for teams that accept code changes on the strength of fuzz testing: library maintainers, security reviewers, and anyone running automated patch pipelines. A random fuzzer that reports no failures says little about an off-by-one check. A coverage-guided fuzzer found every weakened check here within a modest number of runs, so the choice of fuzzer matters. The comparison rests on the counted chance that one random draw hits, not on equal budgets: the random battery was smaller than the largest number of executions a coverage-guided fuzzer needed.

How it was checked

Every hit from the fuzzers was replayed in exact arithmetic and confirmed to break the check it was aimed at. The lab's limits, in our claim file, name the fuzzer versions and settings, and say that run times depend on the machine while execution counts do not.

What it does not claim

  • It does not claim anything about the libraries' own code, or about any copy of them in use. The weakened checks were built by the lab for this test, and the fuzzers ran on test programs generated from models of the checks, as the limits say.
  • It does not claim that random fuzzing is useless in general. The result is about checks that are wrong at a single edge, where random inputs almost never land.
  • It does not claim that the comparison is new. Coverage-guided fuzzing was designed to reach code paths that random inputs miss. What the lab adds is a measured gap on six concrete weakened checks, with every hit replayed.

How to reproduce it

The evidence file names the command that re-runs the check and the commit of the code it ran. It needs the lab's code, Python, and the two fuzzers named in the limits.

Prior art

Evidence

Related results across the group

All results from OrbitalProof · The OrbitalProof home page

How we show numbers

Every result figure links to the file it comes from. See our published files