Crash-Only Software and Recursive Microreboots
web.archive.org
web.archive.org
Interesting to see a name put to this. I used this technique (not having heard of it before) when developing a short but finicky and critical cloud-coordinated cross-datacenter disaster recovery consensus algorithm a few years ago. There was a point at which the algorithm recorded its state and a timestamp to local storage, so progress and correctness were guaranteed in the face of a crash. I came to the same conclusions as these researchers -- this state transition did not need to be fast, but did need to be absolutely robust and well tested. So rather than code and test both the "happy path" and the crash-recovery path -- I just called `_Exit(1)` and only coded and tested the crash-recovery path.
(It wasn't a perfect confluence of testing space -- my code exited in a way that did not trigger a core dump; whereas most failure modes did. But generally, core dumps bogging down an already-ill system were an issue we had to contend with.)
Would be interesting to see a DSL designed around this technique -- a semi-persisted program state of sorts -- with the happy-path automatically generated as an optimization.
In Erlang world, I think this is basically advocating that your GenServer "handle_call/3" "handle_cast/2" methods should only ever return "{stop,...}" and never "{reply...}" or "{ok...}"
> This cheap form of recovery engenders a new approach to high availability: microreboots can be employed at the slightest hint of failure, prior to node failover in multi-node clusters, even when mistakes in failure detection are likely; failure and recovery can be masked from end users through transparent call-level retries; and systems can be rejuvenated by parts, without ever being shut down.
https://www.usenix.org/legacy/event/osdi04/tech/full_papers/...
@dang et al, that link would be preferable to the web archive; more of its paper links are still alive.
Why are we not funding this!?