Skip to content

Random crashes found unclean recoveries in our AI agent's transaction layer; every crash point we chose recovered cleanly

Each square is one process kill. Above: 2,893 kills at random moments; the 31 red squares are recoveries that were not clean, gathered at the end here, not in the order the kills ran. Below: the 29 points we chose in advance, every one recovered cleanly.

At random moments, 31 of 2,893 recoveries were not clean, and some undid committed work; at all 29 points we chose in advance, recovery was clean.

An agent that makes a change in several steps can die halfway through it. Our transaction layer recovers such changes from a journal on disk. We killed its running process on purpose and checked what a new copy of the program, started after the crash, recovered.

Why now

AI agents are now being given the power to make multi-step changes to real files and databases, and research groups have only recently begun proposing transaction layers for them, such as SagaLLM and Cordon, both cited below. What such a layer leaves behind when its process dies is a question a buyer can now ask every vendor, and this test is one way to answer it.

What it shows

The setting is simple. An agent carries out a change in several steps. If the program dies between two steps, the change can be left half done, and nothing is left running to finish it or to undo it.

Our transaction layer keeps a journal on disk, and after a crash a new copy of the program recovers from it. A recovery is clean, or atomic, when it is all or nothing: the whole change or none of it.

Crashes at random moments

We killed the program at moments drawn from a seeded random source (the seed fixes the schedule of kills), many times. Here recovery was not always clean, and some of the unclean recoveries undid work that had already been committed.

Crashes at chosen points

We also chose a set of interruption points in advance and killed a real process at each one. A new copy of the program recovered every time, and every recovery was clean. The negative control, one case we built to be genuinely not clean, was caught as such, which shows the check does not pass everything.

Why it matters, and to whom

This is for teams that let agents change real systems. If an agent's changes are not all or nothing, a crash in the middle of one leaves damage that nobody asked for.

A test any buyer can ask for

The result is useful in two ways. First, it shows a test a buyer can ask any vendor to run: kill the process at random, recover, and count the half-done states. Second, it shows why a clean run at hand-picked points is not enough. We report both for the same layer: clean recovery at every chosen point, and unclean recoveries under random kills. A vendor who reports only hand-picked crash points is reporting the easy half.

How it was checked

Each kill at a chosen point ends the process at once, with no chance to tidy up. A new copy of the program then recovers from the journal. The rule is the usual one for transactions: a recovery passes only if the change is either complete or fully undone; a change left partly applied, or a recovery that undoes work already committed, counts as not clean.

Testing the checker

The negative control is genuinely not clean, and the check records it as such. We also edited the recorded crash results on purpose to see whether the check would notice, and it reported a failure.

What the files show

Our evidence file covers the chosen points; the random-kill counts are stated in our claim file, published beside it.

What it does not claim

  • It does not claim that the journal is safe against power loss. Killing a process is not the same as cutting power, which can lose writes or apply them out of order; Pillai and colleagues, cited below, explain why that matters for recovery.
  • It does not claim full coverage. The random kills sample the possible crash moments; they do not try every one. It does not claim the problems are fixed. The stretches of a change in which a crash leaves it half done, which the random kills found, are recorded, not repaired.
  • It does not claim that the method is new. SQLite's documentation describes the same test: run the operation in a child process, crash it in the middle of a write, then check that the transaction either finished or was fully rolled back. SQLite goes further and simulates power loss by reordering and damaging writes that never reached the disk. The work by Pillai and colleagues on crash consistency explains why that last step matters. Transaction and compensation layers for language-model agents are also published, for example SagaLLM and Cordon.
  • What this result adds is a measurement of one more system, our own agent layer, with the process-kill part of that published method, without its power-loss simulation. The measurement found real gaps.

How to reproduce it

The evidence file names the command that re-runs the check at the chosen points. It needs our code at the version named in the evidence file, which is private.

Without the code you can still confirm that the files are ours: download them and compare their fingerprints with our list of published files.

Going further

To go further, run the same plan under a tool that simulates power loss, or inside a virtual machine whose disk drops writes that never reached it. If recovery at the chosen points is still all or nothing there, that half of the result becomes much stronger; the unclean recoveries under random kills need a fix either way. If it is not, the journal needs a change before anyone relies on it.

Prior art

Evidence

Related results across the group

All results from OrbitalProof · The OrbitalProof home page

How we show numbers

Every result figure links to the file it comes from. See our published files