Crash-only software: More than meets the eye (2006)
lwn.net
lwn.net
1. write ahead log before you do side effects/idempotent side effects
2. double writes to disk to prevent torn writes
3. checksums to make sure we don't make bad decisions based on bad data
4. redundancy/anti-entropy/other distributed system patterns which attempt to obviate the need to be overly concerned with a single process crashing
5. self-healing patterns when bad data is found
anyone have any other ideas?
Systematically testing that software is crash-proof isn’t exactly easier though.
Regarding exceptions, one technique, if you can hook the throw sites, is to increment a counter anytime a throwing condition could occur, and in the tests indiscriminately throw an exception when the counter reaches N, and repeat that for N=1…MAX.
In case someone wants to search for more recent stuff, the author changed name later: https://en.wikipedia.org/wiki/Valerie_Aurora
Especially when it crashes during a recovery or when there is encryption involved.
So more concretely, if system A crashes 1 % of the time, half of crashes might cause corruption. If system B crashes 100 % of the time, the developers of B will have to face the corruption issues and might reduce corruptions to affect only one crash in every thousand.
In the end, system A has a corruption rate of 0.5 %, whereas system B has a corruption rate of only 0.1 % – despite crashing more often.
So the point is changing the incentives around dealing with crashy situations in a way that ultimately results in higher stability. It's saying "It will crash anyway, so we have to deal with it. How can we force people to deal with it? We can crash it always."
The current usual advice for a safe shutdown is REISUB, with the mnemonic 'Reboot Even If System Utterly Broken', but the old mnemonic a lot of folks will still know is 'Raising Skinny Elephants Is Utterly Boring'. It seems to have been replaced because it flushes data to disk earlier in the sequence than killing off processes, potentially leaving data from running processes un-flushed.