Skip to content

What crash recovery means for AI agents that change real files and databases

Explainer 6 min read

An agent that dies halfway through a multi-step change can leave a mess, and a journal on disk is the usual way to make that change all or nothing.

Each square is one process kill. Above: 2,893 kills at random moments; the 31 red squares are recoveries that were not clean, gathered at the end here, not in the order the kills ran. Below: the 29 points we chose in advance, every one recovered cleanly.

A dotted underline marks a number read straight from a published file when this page was built.

In this post
  1. What crash recovery means
  2. How a journal on disk helps
  3. How we tested our recovery
  4. What the random kills found
  5. What this means for you
  6. What this does not show

What crash recovery means

Crash recovery is the promise that a program which dies in the middle of a change will leave the system in a clean state, either with the whole change applied or with none of it.

Take an AI agent that is asked to rename a set of files and then record the new names in a database. That is several steps. The program running the agent can stop between any two of them, for several ordinary reasons:

  • It might run out of memory.
  • It might be shut down by a deploy.
  • It might hit a bug.

Whatever the cause, nothing is left running to finish the job or to undo it.

A change that is all or nothing is called atomic.

Databases have offered this for decades, and the SQLite project explains its own approach in its description of atomic commit. Agents are newer at this. They are now being given the power to change real files and databases, and transaction layers for them have only recently been proposed in research, such as SagaLLM and Cordon below. So a crash in the middle of an agent’s change can now leave real files and databases half updated, which is why this matters now.

How a journal on disk helps

A journal is a record, written to disk before a change is made, of what the change will do and how far it has got.

The idea is simple, and it has three parts:

  • Before touching anything, the program writes down the plan.
  • As each step finishes, it records that too.
  • If the program dies, a fresh copy starts, reads the journal, and works out what to do.

A step that was recorded as done stays done. A change that never reached its end is either finished or rolled back, depending on the design.

We built a transaction layer for agents along these lines. Its result page describes it in plain words and links our claim file. The layer keeps a journal on disk, and after a crash a new copy of the program, started fresh, recovers from that journal.

Two related research efforts point the same way. SagaLLM studies transaction and compensation guarantees for agent workflows. Cordon looks at staging and checking effects that cannot be undone. Our work is a measurement of one concrete layer, not a new theory.

How we tested our recovery

We killed real processes on purpose and then checked what a new copy of the program recovered.

The test followed a well-known method.

SQLite’s documentation on how it is tested describes running an operation in a child process, crashing it in the middle of a write, and checking that the transaction either finished or was fully rolled back. We did the same for our agent layer.

The test had three ingredients:

  • Chosen points. We picked our interruption points in advance.
  • Random points. We later added kills at random moments.
  • A control. We included a case that is built to fail.

In the first part of the test, we chose our interruption points in advance. Our claim records that 29 such points were each recovered atomically. Every recovery was whole or absent.

The test also included a negative control, described on the result page. That is a case we built to be genuinely not atomic. Its job is to show that the checker does not pass everything. A checker that passes everything proves nothing, so a control that is expected to fail, and does, is part of the evidence.

We went one step further and edited the recorded crash results on purpose, to see whether the check would notice. It reported a failure.

What the random kills found

Choosing the crash points yourself is the easy half, so we also killed the program at moments drawn at random.

Here the picture changed. Our claim says that 2,893 further kills at seeded random moments found 31 recoveries that were not atomic, and that some of them undid work that had already been committed. A seeded random source fixes the kill schedule, so the same schedule could be rerun with our code, though timing can still change which crash states it reaches; the command we publish re-runs only the chosen points.

In plain words, the journal survived the crashes its authors chose and failed on some they did not. That is useful information, and the companion post on hand-picked tests takes it further. It is why testing at random moments matters at all, and the reason a buyer should ask a vendor for both numbers, not only the flattering one.

Our limits, in plain words: these were process kills, not power cuts; the random kills sample the possible crash moments rather than try every one; and the weak spots they found are recorded, not yet fixed.

What this means for you

If you run or buy an agent platform, ask what happens when the process dies mid-change, and ask how that was tested.

  • Agent-platform teams that let assistants change files and databases can copy the method: kill the process at random, recover, and count the half-done states.
  • Runtime-security and reliability teams can use the result as a concrete example of why hand-picked crash points flatter a design.
  • Buyers and diligence teams can ask any vendor for both halves of the answer, the chosen points and the random ones, and treat a vendor who reports only the first half as having reported the easy part.

The full result page lists the evidence files, and you can hash them yourself.

What this does not show

The result does not show that the journal is safe against power loss, and we say so.

  • Power cuts. A killed process and a power cut are different. With a power cut, writes can be lost or applied in a different order than the program expected, and the order of writes to disk is exactly what the journal relies on. Our own limits, above, call that order untested.
  • Coverage. The random kills sample the possible crash moments rather than try every one.
  • Fixes. The problems they found are recorded, not fixed.
  • Novelty. The method is not new, since SQLite has used it for years.
  • Other vendors. Nothing here shows that any other vendor’s agent layer behaves this way. It is one measurement of one layer, and its worth lies in showing what an honest test of recovery looks like, including the part that found problems.

Results behind this post

All results from OrbitalProof · The same result on VerifyCore Labs, our parent lab

Keep reading

All posts

Sources and further reading

Where the numbers come from

Each number with a dotted underline was read from one of these published files, field by field, when the page was built.

Every file this site publishes

Ask about a result, or check one yourself

Every result on this site links the published files behind it, and its page says what each file covers. Acquisition, licensing and partnership enquiries go to one address, and a person reads it.

Write to us Read the research results

How we show numbers

Every result figure links to the file it comes from. See our published files