What crash recovery means
Crash recovery is the promise that a program which dies in the middle of a change will leave the system in a clean state, either with the whole change applied or with none of it.
Take an AI agent that is asked to rename a set of files and then record the new names in a database. That is several steps. The program running the agent can stop between any two of them, for several ordinary reasons:
- It might run out of memory.
- It might be shut down by a deploy.
- It might hit a bug.
Whatever the cause, nothing is left running to finish the job or to undo it.
A change that is all or nothing is called atomic.
Databases have offered this for decades, and the SQLite project explains its own approach in its description of atomic commit. Agents are newer at this. They are now being given the power to change real files and databases, and transaction layers for them have only recently been proposed in research, such as SagaLLM and Cordon below. So a crash in the middle of an agent’s change can now leave real files and databases half updated, which is why this matters now.
How a journal on disk helps
A journal is a record, written to disk before a change is made, of what the change will do and how far it has got.
The idea is simple, and it has three parts:
- Before touching anything, the program writes down the plan.
- As each step finishes, it records that too.
- If the program dies, a fresh copy starts, reads the journal, and works out what to do.
A step that was recorded as done stays done. A change that never reached its end is either finished or rolled back, depending on the design.
We built a transaction layer for agents along these lines. Its result page describes it in plain words and links our claim file. The layer keeps a journal on disk, and after a crash a new copy of the program, started fresh, recovers from that journal.
Two related research efforts point the same way. SagaLLM studies transaction and compensation guarantees for agent workflows. Cordon looks at staging and checking effects that cannot be undone. Our work is a measurement of one concrete layer, not a new theory.
How we tested our recovery
We killed real processes on purpose and then checked what a new copy of the program recovered.
The test followed a well-known method.
SQLite’s documentation on how it is tested describes running an operation in a child process, crashing it in the middle of a write, and checking that the transaction either finished or was fully rolled back. We did the same for our agent layer.
The test had three ingredients:
- Chosen points. We picked our interruption points in advance.
- Random points. We later added kills at random moments.
- A control. We included a case that is built to fail.
In the first part of the test, we chose our interruption points in advance. Our claim records that 29 such points were each recovered atomically. Every recovery was whole or absent.
The test also included a negative control, described on the result page. That is a case we built to be genuinely not atomic. Its job is to show that the checker does not pass everything. A checker that passes everything proves nothing, so a control that is expected to fail, and does, is part of the evidence.
We went one step further and edited the recorded crash results on purpose, to see whether the check would notice. It reported a failure.
What the random kills found
Choosing the crash points yourself is the easy half, so we also killed the program at moments drawn at random.
Here the picture changed. Our claim says that 2,893 further kills at seeded random moments found 31 recoveries that were not atomic, and that some of them undid work that had already been committed. A seeded random source fixes the kill schedule, so the same schedule could be rerun with our code, though timing can still change which crash states it reaches; the command we publish re-runs only the chosen points.
In plain words, the journal survived the crashes its authors chose and failed on some they did not. That is useful information, and the companion post on hand-picked tests takes it further. It is why testing at random moments matters at all, and the reason a buyer should ask a vendor for both numbers, not only the flattering one.
Our limits, in plain words: these were process kills, not power cuts; the random kills sample the possible crash moments rather than try every one; and the weak spots they found are recorded, not yet fixed.
What this means for you
If you run or buy an agent platform, ask what happens when the process dies mid-change, and ask how that was tested.
- Agent-platform teams that let assistants change files and databases can copy the method: kill the process at random, recover, and count the half-done states.
- Runtime-security and reliability teams can use the result as a concrete example of why hand-picked crash points flatter a design.
- Buyers and diligence teams can ask any vendor for both halves of the answer, the chosen points and the random ones, and treat a vendor who reports only the first half as having reported the easy part.
The full result page lists the evidence files, and you can hash them yourself.
What this does not show
The result does not show that the journal is safe against power loss, and we say so.
- Power cuts. A killed process and a power cut are different. With a power cut, writes can be lost or applied in a different order than the program expected, and the order of writes to disk is exactly what the journal relies on. Our own limits, above, call that order untested.
- Coverage. The random kills sample the possible crash moments rather than try every one.
- Fixes. The problems they found are recorded, not fixed.
- Novelty. The method is not new, since SQLite has used it for years.
- Other vendors. Nothing here shows that any other vendor’s agent layer behaves this way. It is one measurement of one layer, and its worth lies in showing what an honest test of recovery looks like, including the part that found problems.