> Turning something off and then on again may be one of the worst ways to fix problems with certain large systems.
From an overall system-wide perspective, sure. But from an individual component perspective, it seems to work pretty darn reliably. There's a reason why Erlang/OTP (as one example) has such a reputation for fault-tolerance and robustness.
And on that note, Erlang's model of supervised preemptive processes seems like it'd be a good fit for unified-memory computing, especially if taken to an "everything is a process" extreme.