The easy half of a crash test
When the people who wrote a system also choose where to crash it, they tend to pick the places they already thought about.
That is not dishonest. It is human, and it is the reason outside tests matter. It is how testing works. The routine looks like this:
- You list the moments that worry you, such as between writing a record and marking it done.
- You crash the program at each one.
- If the design handles all of them, the list is green.
Each point on the list is a case the authors understood well enough to name.
The trouble is the points nobody thought to list. A real crash does not consult the list. It lands wherever the process happens to be when something kills it. Our tests killed processes; they did not cut power.
What our own test saw
We tested our own agent transaction layer both ways, and the two answers differed.
The result page explains the setup. In the first part, we chose our interruption points and killed a real process at each one. A new copy of the program recovered from the journal on disk every time. Our claim records that 29 points came back atomic, meaning the change was whole or absent.
In the second part we drew kill moments from a seeded random source. Our claim says that 2,893 such kills found 31 recoveries that were not atomic, and that some undid work already committed.
The same layer gave a clean result under one test plan and an unclean result under the other. Neither number is wrong. They answer different questions.
Why the two answers differ
A hand-picked list covers the cases someone imagined, while random kills sample the whole space of moments.
Think of the program’s execution as a long line of instants. Any one of them could be the moment of a crash. The authors’ list names a few. Random draws land anywhere, including in short windows where the program is partway through two things at once. Our random kills sample those instants, reaching a limited number of distinct crash states, and do not try every one.
Two things follow from this:
- A green list is a statement about the list, not about the program.
- The gap between the list and the whole space is where surprises live.
A seeded random source has a practical benefit. It is random enough to reach places a person would not think of, and the seed fixes the kill schedule, so another team with the code can rerun the same schedule; the crash states it reaches can still vary with timing.
Our limits, in plain words: these were process kills, not power cuts; the random kills sample the possible crash moments rather than try every one; and the weak spots they found are recorded, not yet fixed.
What a stronger test would add
Killing a process is only one kind of failure, and the people who build databases test more.
SQLite’s documentation on how it is tested describes the same process-kill method and then goes further, simulating power loss by reordering and damaging writes that never reached the disk. The research paper by Pillai and colleagues explains why that matters: whether recovery works can depend on the order in which a file system makes writes durable.
Here is the gap in one line: our own limits say the order of writes to disk that our journal relies on is still untested. So there are two natural next steps. One is to rerun the same plan under a tool that simulates power loss. The other is to run it inside a virtual machine whose disk drops writes that never reached it. If recovery at the chosen points is still all or nothing there, that half of the result becomes much stronger. If it is not, the journal needs a change before anyone relies on it.
A note on seeds
A seed is the starting value of a random number generator, and the SQLite testing notes describe a harness that crashes a process at random points in a write. If you record it, the same sequence of random choices can be replayed. That lets another person with the code rerun the same schedule of kills, though timing can still change which crash states it reaches.
Two things change once a seed is recorded:
- A failure can be replayed and studied.
- A clean run can be repeated by someone else.
Without a seed, a failure found by a random test may never be reproduced. And a clean random run says little unless it reports how many kills ran and how many distinct crash states they reached. With one, you can say exactly which crash moments were drawn and ask for the same draw next time. Our kills were seeded. The command in our evidence file re-runs the chosen points; the random-kill count is the one our claim reports.
What this means for you
When a vendor tells you their system recovers from crashes, ask what kind of crash and who chose the moment.
- Agent-platform and infrastructure teams can add a random-kill stage to their own test plan, with a recorded seed, and count the half-done states it finds.
- Buyers and diligence teams can ask for both results, the chosen points and the random ones, and treat a report of only the first as the easy half.
- Reliability engineers can use the pairing as a template: a clean result on a declared list, then a harsher result on a sampled one.
What this does not show
This is one measurement of one layer, and it does not show how often random crashes cause damage in the wild.
- It does not show that the unclean recoveries are fixed. We record them, and they are not repaired yet.
- It does not show behaviour under power loss.
- It does not show that other vendors’ layers would give the same split between chosen and random points, although it is a reasonable thing to check. The result page links the files so you can hash them and compare.