Skip to content

Why testing a system at crash points you chose proves less than it seems

Explainer 5 min read

A recovery design that passes every crash point its authors chose can still fail at moments they did not think of.

Each square is one process kill. Above: 2,893 kills at random moments; the 31 red squares are recoveries that were not clean, gathered at the end here, not in the order the kills ran. Below: the 29 points we chose in advance, every one recovered cleanly.

A dotted underline marks a number read straight from a published file when this page was built.

In this post
  1. The easy half of a crash test
  2. What our own test saw
  3. Why the two answers differ
  4. What a stronger test would add
  5. A note on seeds
  6. What this means for you
  7. What this does not show

The easy half of a crash test

When the people who wrote a system also choose where to crash it, they tend to pick the places they already thought about.

That is not dishonest. It is human, and it is the reason outside tests matter. It is how testing works. The routine looks like this:

  • You list the moments that worry you, such as between writing a record and marking it done.
  • You crash the program at each one.
  • If the design handles all of them, the list is green.

Each point on the list is a case the authors understood well enough to name.

The trouble is the points nobody thought to list. A real crash does not consult the list. It lands wherever the process happens to be when something kills it. Our tests killed processes; they did not cut power.

What our own test saw

We tested our own agent transaction layer both ways, and the two answers differed.

The result page explains the setup. In the first part, we chose our interruption points and killed a real process at each one. A new copy of the program recovered from the journal on disk every time. Our claim records that 29 points came back atomic, meaning the change was whole or absent.

In the second part we drew kill moments from a seeded random source. Our claim says that 2,893 such kills found 31 recoveries that were not atomic, and that some undid work already committed.

The same layer gave a clean result under one test plan and an unclean result under the other. Neither number is wrong. They answer different questions.

Why the two answers differ

A hand-picked list covers the cases someone imagined, while random kills sample the whole space of moments.

Think of the program’s execution as a long line of instants. Any one of them could be the moment of a crash. The authors’ list names a few. Random draws land anywhere, including in short windows where the program is partway through two things at once. Our random kills sample those instants, reaching a limited number of distinct crash states, and do not try every one.

Two things follow from this:

  • A green list is a statement about the list, not about the program.
  • The gap between the list and the whole space is where surprises live.

A seeded random source has a practical benefit. It is random enough to reach places a person would not think of, and the seed fixes the kill schedule, so another team with the code can rerun the same schedule; the crash states it reaches can still vary with timing.

Our limits, in plain words: these were process kills, not power cuts; the random kills sample the possible crash moments rather than try every one; and the weak spots they found are recorded, not yet fixed.

What a stronger test would add

Killing a process is only one kind of failure, and the people who build databases test more.

SQLite’s documentation on how it is tested describes the same process-kill method and then goes further, simulating power loss by reordering and damaging writes that never reached the disk. The research paper by Pillai and colleagues explains why that matters: whether recovery works can depend on the order in which a file system makes writes durable.

Here is the gap in one line: our own limits say the order of writes to disk that our journal relies on is still untested. So there are two natural next steps. One is to rerun the same plan under a tool that simulates power loss. The other is to run it inside a virtual machine whose disk drops writes that never reached it. If recovery at the chosen points is still all or nothing there, that half of the result becomes much stronger. If it is not, the journal needs a change before anyone relies on it.

A note on seeds

A seed is the starting value of a random number generator, and the SQLite testing notes describe a harness that crashes a process at random points in a write. If you record it, the same sequence of random choices can be replayed. That lets another person with the code rerun the same schedule of kills, though timing can still change which crash states it reaches.

Two things change once a seed is recorded:

  • A failure can be replayed and studied.
  • A clean run can be repeated by someone else.

Without a seed, a failure found by a random test may never be reproduced. And a clean random run says little unless it reports how many kills ran and how many distinct crash states they reached. With one, you can say exactly which crash moments were drawn and ask for the same draw next time. Our kills were seeded. The command in our evidence file re-runs the chosen points; the random-kill count is the one our claim reports.

What this means for you

When a vendor tells you their system recovers from crashes, ask what kind of crash and who chose the moment.

  • Agent-platform and infrastructure teams can add a random-kill stage to their own test plan, with a recorded seed, and count the half-done states it finds.
  • Buyers and diligence teams can ask for both results, the chosen points and the random ones, and treat a report of only the first as the easy half.
  • Reliability engineers can use the pairing as a template: a clean result on a declared list, then a harsher result on a sampled one.

What this does not show

This is one measurement of one layer, and it does not show how often random crashes cause damage in the wild.

  • It does not show that the unclean recoveries are fixed. We record them, and they are not repaired yet.
  • It does not show behaviour under power loss.
  • It does not show that other vendors’ layers would give the same split between chosen and random points, although it is a reasonable thing to check. The result page links the files so you can hash them and compare.

Ask about a result, or check one yourself

Every result on this site links the published files behind it, and its page says what each file covers. Acquisition, licensing and partnership enquiries go to one address, and a person reads it.

Write to us Read the research results

How we show numbers

Every result figure links to the file it comes from. See our published files