{
 "_doc": "The lab's current claim and limits for one result, copied word for word from the portfolio's current-claims file at the line named here. A value this site does not print in full is replaced by its sha256 and byte count, with the reason; the excerpts beside it are copied word for word from the same value. Nothing here is written by hand.",
 "result": "code-admission:certbench",
 "title": "CertBench provable-soundness scoring",
 "claim": "Thousands of candidate guard patches are scored by the solver-free decision procedure: 2,673 certified, 3,077 proven unsound, 5,750 decided of 6,036 scored. z3 and cvc5 each agree with all 5,750 of those verdicts, though 333 of the counterexamples the scorer records (all for sudo) lie outside its own domain; SMT replaces each with an in-domain one that replays. The 286 the scorer leaves indeterminate were handed to SMT, the same SMT-LIB input to both solvers, which agree on every case: 38 are certified sound, 195 are proven unsound by a 32-bit counterexample that replays in exact integer arithmetic, and 53 stay undecided, because each is unsound in the scorer's integer model only through a value of 2^32 or more and sound once every variable is held to 32 bits. That brings the decided count to 5,983 of 6,036. The same scorer was pointed at 24 real, live GLM-5.2 candidate patches: 10 are proven unsound (0.4167, median over-acceptance 131,072), which lands inside the 16-46% band the DARPA SoK reconstructs for tests-plus-fuzzing pipelines (artifacts/certbench/summary.json :: live_glm_anchor).",
 "limits": "Every verdict is a fact about a linear integer model of each guard over the scorer's declared domain, not about the libraries' C code. The SMT leg ran z3 4.15.4 and cvc5 1.3.4 (artifacts/certbench/smt_leg.json :: solvers). Its 38 certifications rest on both solvers' unsat, and each also carries an integer-tightened Farkas certificate that replays. Its 195 refutations rest on counterexamples that replay. 53 of 6,036 (0.9%) stay undecided: their verdict turns on 32-bit wraparound, which the linear model does not represent. For 333 of the 392 sudo_set_cmnd patches the scorer calls unsound, the counterexample it records has su_size = 0, outside its declared domain (su_size >= 1). summary.json's every_decided_verdict_is_a_proof = true is therefore not true of those recorded points, though an in-domain counterexample from SMT stands behind every one of the 392 verdicts. The over-acceptance figures are counts over a declared 48-bit model domain.",
 "claim_excerpts": [],
 "limits_excerpts": [],
 "lab_commit": "b36a228c1b7ad32cc67717ac41d4419be13994a3",
 "source": {
  "file": "portfolio-control registry/CLAIMS_CURRENT.jsonl",
  "line": 3,
  "sha256": "81776df15a1498abec7228ea179c952d01c54e031ffd4144498a30660963074b"
 },
 "evidence": "/receipts/code-admission/results/certbench.json",
 "copied_utc": "2026-10-03T16:11:27Z"
}
