On rebooting: the unreasonable effectiveness of turning computers off and on
keunwoo.com
keunwoo.com
He said he developed that style because the lack of protection in DOS meant that any error (like a buffer overrun) could trash everything, right down to damaging the machine.
He said that early on those asserts were enabled in the shipping code, making it appear that his compiler was less reliable than competitors when he felt it was the opposite, but I think he wound up having to modify his approach.
* asserts are a compact form of documenting invariants both expected and produced, making code easier to reason about
* each assert serves as a tiny unit test with exactly the necessary cases
* most importantly, asserts catch logic bugs when code doesn't produce the intended results
P.S. that has me thinking, there should be a tool to transform asserts into testing. The tool would have another keyword "expect" in addition to "assert", for preconditions. With both preconditions and postconditions, individual functions can be automatically verified with a prover or a fuzzer! The tool could also check that preconditions are met at each call site, without having to write and rewrite unit tests for each call site!https://www.cs.columbia.edu/~junfeng/08fa-e6998/sched/readin...
Gray’s paper provides a logical reason why: If a bug can be reproduced reliably it’s much more likely to be found during testing and fixed, so appearing seemingly at random is a kind of Darwinian survival adaptation.
Just last week I was on the verge of literally ripping my hair out trying to figure out a frustrating bug that never occurred when running in the debugger. After a lot of frustration I had an idea: rather than starting the application in the debugger, I’d attach later. When doing this, I began to see a huge amount of interesting first-chance exceptions (this was managed C# code) that clued me into the root cause of the issue and I was finally able to solve the problem.
I came to learn that an underlying library I was using had some code that ran on startup that actually detected if a debugger was attached and would turn off the problematic code path in this case! I came to learn that this was an intentional optimization the initialization of this code caused a bunch of noise during regular debugging (lots of first-chance exceptions that were ignorable). Unfortunately, this also meant that the interesting failure case I was running into was now completely hidden.
Ideally there would be a way to get the debugger to lie to the process that a debugger is indeed not attached. Sure, maybe there's a reason today why you actually need to know a debugger isn't attached but at least give future-developers an escape hatch if they need it.
- the timing of execution so nothing works => hello USB driver!
- the scheduling of your threads
- the memory layout => a memory overflow won't crash your program the same way or not at all
I feel lucky I've enough hairs to pull them out on this kind of bugs. Fun time indeed.
which means in order to find as many bugs as possible during development, the dev environment must match the prod environment as exactly as possible.
Bugs don't "evolve" like a life form - unless you code is self-modifying! It is caused by a permutation of all possible inputs (where the environment is also an "input").
Limiting possible inputs (such as functional programming) means you limit the possible differences between dev and prod - so functional programming ought to produce less bugs!
If there is no competition in the evolution then systems can become arbitrarily baroque and fragile, which leads to rules like 'no software updates/patches' and 'all config changes via the CCB and pass mandatory tests' or 'your problem is handled manually', the wire wrap of systems development.
Rebooting (the process or the computer) is trimming the state space. Model checking driven design really helps tame the tendency to let program state evolve without bound. FP tends to follow a bottom up design which makes this easier too.
Bugs live in the codebase, and the codebase evolves (through human modifications).
Every change in the codebase has a probability of introducing a new bug. If the introduced bug is deterministic and easy to catch, the probability of its survival is lower. If the introduced bug is nondeterministic and hard to catch, the probability of its survival is higher. So eventually, as time approaches infinity, the probability of a random bug in the codebase being nondeterministic approaches 100%.
Which, it actually can't in the current state, but only because of the system itself changin, or evolving, over time.
For example lets imagine some embedded device in the ISP network, it is not very accessible so high reliability is required. You can overengineer it to be a super reliable Voyager-class computer, but that a) will cost much more money than it should, b) you will fail to achieve the target.
Or you can go crash-first approach. Many things can be simplified then, for example no need to have stateful config, no need to code any config management which saves it, checks it etc. You just rely on receiving new config every boot and correctly writing it to controllers. Less complexity.
Then if you are crash-first you expect to reboot more often. You then optimize boot time, which would be a much lower priority task otherwise. And suddenly you have x times better (lower) downtimes per device.
You can optimize on some hard stuff - e.g. any and all 3rd party controllers with 3rd party code blobs and weird APIs. Instead of writing a lot of health checks for each and every failure mode of the stuff you can't really influence, you write a bare minimum and then watchdog to reboot whole unit and hope it will recover. And this works very well in practice.
The list goes on. Instead of very complicated all-in-one device you have a lightweight code which has a good and predictable recovery mechanism. It is cheaper, and eventually even more reliable than overengineered alternative. Another example - network failure. Overengineered device will do a lot of smart things trying to recover network, re-initialize stuff, re-try different waits (and there will be a lot of re-tries), and may eventually get stuck without access. Lightweight device has a short simple wait with simple retry, or several, and then reboots. Statistically this is better than running some super complicated code, if the device is engineered to reboot from the start.
We made slight adjustments (outros and intros) so that it seemed natural to have a 10 second break for the program to restart. And, we built in a longer cycle of computer resets. It was an unreasonably stable system for years!
Can you elaborate on
>We built it in Flash [...] we just couldn’t stop the memory leaks
I thought Flash games was written in a high level JS-like language ? did it grant you enough access to raw memory that you can leak ? or did you mean a high level equivalent to memory leaks ?
Light Years Ahead | The 1969 Apollo Guidance Computer
https://www.youtube.com/watch?v=B1J2RMorJXM
34C3 - The Ultimate Apollo Guidance Computer Talk
Erlang seems to also follow this kind of philosophy, although on a more granular level. The point seems to be in separation of "worker" code and "supervisor" code - where "worker" represents a well-behaved function without any (unexpected-)error checks, and "supervisor" represents error-handling code that will catch and resolve any errors that happens in the worker code, expected or not.
Joe Armstrong's "Making reliable distributed systems in the presence of software errors" contains more information on the topic.
Have you lost your ever-loving mind? Of course you're going to reboot. Bit rot improvements have only ever been incremental. To a first order approximation we've expanded from a day or two to a couple of weeks/months for general purpose computers (I argue that servers are not GP, and so their spectacular uptimes don't translate).
In thirty years, that's barely more than an order of magnitude improvement. You're gonna need a couple more orders than that before taking away the reboot option sounds anything like sanity. If you can keep a machine from shitting the bed for 2-3 times longer than the expected life of the hardware, then we can talk, but I'm not making any promises. Until then I'm gonna reboot this thing sometimes even if it's just placebo.
Sure lots of domestic modems & routers have to be reset often, but it's common to find infrastructure routers that are only reset when they have to be moved! Says a lot about the software in domestic equipment.
I've worked with a mainframe that ran for a decade, and was only rebooted when a machine room fire required a full powerdown "right now"! But people are impressed when their Windows stays up for a few day %#%$#@#$!!
And remember Voyager, been ticking along since 1977.
But all these impressive systems are built to be impressive. Probably error correcting RAM, etc. Shows we can build reliable hardware. And if a substantial fraction of users wanted it, things like ECC RAM etc would be only fractionally more expensive than our error prone alternatives.
Of course,
Having worked in ISP security, IMHO a years-long uptime of such critical components is nothing to be proud of (anymore). Quite the contrary, those are complex components, so if you care about security you have to regularly patch them, including occasionally required reboots. Just look at the list of security advisories of relevant vendors (Cisco, Juniper, Nokia/Alcatel-Lucent, etc), you can find scary vulnerabilities! Granted, "rebooting" a core router is more nuanced than a regular PC (you can e.g. reboot one management engine of a pair; or just a line card; etc), so it does not always mean that the entire traffic stops because of it.
Oh and btw. your network design should be able to cope with such necessary reboots, otherwise you have a single point of failure.
Regards
This is running a single unthreaded process.
My Mac, which isn't doing anything fancy, has over 500 processes running on it. In fact, I just checked to see if anything bad was going on, and I recognize everything I look at - almost 100 processes from Chrome alone, for example.
How sure am I that all of these processes are running correct code? Chrome is running "101.0.4951.54 (Official Build) (x86_64)", which gives a hint of the disposable nature of that binary.
A computer used for a single task is a bit like a 4WD truck that only stays inside the city limits. It could do those things, sure, but it never does, so it hasn't really proven anything.
Not really. That's very different from a very application specific hardware+software being used for the exact purpose vs a very general purpose hardware+software being used for all sorts of things.
It's more like using a F1 car to race vs taking your average sporty car to a race track.
Sure that F1 car will race better, but at the end of the day you can't drive home in it or move kids/groceries around in it.
Generalization itself is hard, it gets MUCH harder when you have to care about back compat and random executables that can alter system state because previous versions allowed that behavior and it needs to be supported for the common cases moving forwards.
Channel ECC is the ECC type most directly relevant for high clock rates and signal integrity aspects. I agree with you that Channel ECC becomes a practical requirement to meet the interface transaction rates of DDR5. It is also true that channel ECC is not mandatory in DDR5 and is not implemented by mainstream CPU platforms (like previous DDR generations).
In fact it could very well be higher depending on how the physical module is designed.
I imagine some portion of bit-flip induced reboots are due to the actual DRAM chips, but also some portion will be due to everything else that can bit flip both on the memory module itself and in the interconnect.
I haven't seen anything yet to say that DRAM chip bit flips will be in the majority.
I started in the PC industry in 1988.
Then, they did. All IBM PC kit used 9-bit RAM, with a parity bit.
It was discarded during 1990s cost-cutting.
However, NVRAM enables power-efficient sleep. My old ThinkPad used almost all of its battery charge overnight in S3 standby. On the unreasonable efficiency of instant-on computers
Probably the FTC should have stepped up and ordered Microsoft to issue refunds. But they didn't, and here we are.
I run xorg, and requiring X11 to restart is equally rare. Thus, I often go for months or longer, without a restart at the gui level.
If I get a browser update, I restart that, and so on.
Microsoft has conditioned the world to accept absurdity. Just the lost productivity alone, due to reboots...
I think that's also reminiscent of Alan Kay's philosophy behind OO in its most original form, and probably most closely realized in Erlang:
"I thought of objects being like biological cells and/or individual computers on a network, only able to communicate with messages (so messaging came at the very beginning -- it took a while to see how to do messaging in a programming language efficiently enough to be useful). - I wanted to get rid of data. The B5000 almost did this via its almost unbelievable HW architecture. I realized that the cell/whole-computer metaphor would get rid of data[..]"
I wonder why so many of the most popular programming languages went into the opposite, very state and lock based direction given the strong theoretical foundation that computing had for systems that looked much more robust.
http://userpage.fu-berlin.de/~ram/pub/pub_jf47ht81Ht/doc_kay...
(Which is not to claim that OO has, as Kay insists, any universal merit.)
> In Smalltalk, every entity is called an object; every object belongs to a class (which is also an object). Objects can remember things about themselves and can communicate with each other by sending and receiving messages. The class handles this communication for every object which belongs to it; it receives messages and possibly produces a reply, typically a message to send to another object.
So far as I can tell, this means that Smalltalk-72 was what-Kay-considers-object-oriented.
I'm not sure whether this refutes the claim you're actually making, because apparently you're saying that Kay lies about what he used to think, and maybe e.g. you're saying that what Kay now says about object orientation is not what he used to say, and that in the 1970s he wouldn't have considered Smalltalk object-oriented. Or maybe you're saying that the right way to think about object-orientation is something different from Kay's, and that Smalltalk-72 was not what-you-consider-object-oriented. Or something.
Could you maybe be more explicit?
Could you give some examples of things that Kay now says he used to think, and explain why you believe he didn't actually think them? Could you explain in what sense you reckon object-orientation was absent in Smalltalk-72 but bolted on later, and why you think that indicates intellectual dishonesty on Kay's part?
Where can we find him saying "requires a constrained sort of variable run-time binding"?
To this day, all my systems which are intended for prolonged unattended operation reset themselves at least once a day.
[edit] More specifically there are a number of things that have to happen every 24 hours - some memory has to be zeroed, firmware integrity verified, etc. Most vendors like Verifone [1] implement this by rebooting the reader at least every 24 hours on a timer.
[1] https://developer.verifone.com/docs/verifone-documentation/e...
Shouldn't that be 'at most' or 'at least once'?
(My English skills are lacking so that was quite the mental stumble for me)
However, "at least every 24 hours" is also perfectly acceptable and idiomatic English. It is very common in English to use this construction, "at least every X time period," to mean "X often or more often." If you say, "I make sure to call my distant relatives at least every year," you mean with that frequency or more. If you say, "I make sure to stand up and stretch at least every hour," you mean with that frequency or more. If you say, "these machines are required to reboot at least every 24 hours," you mean with that frequency or more.
In other words, "at least" does not qualify the number of hours (such that more hours would also be acceptable), but the frequency (such that greater frequency would also be acceptable).
Hopefully this response helps you to understand why English, fickle language that it is, works this way in this case.
'at most every 24 hours' would imply reboots_per_day < 1 is bad and reboots_per_day >= 1 is good; i.e. that the standard was focused on preventing unnecessary reboots, but didn't care if the reader didn't reboot regularly, which I think is contrary to the comment's meaning.
'at least every 24 hours' is correct.
I understood as that being exactly the point of the GP, rather than contrary. Because of the 'most' referring to the 24 hrs, not amount of reboots.
But I guess this is the point where my lack of English skill comes into play.
> ... would imply reboots_per_day > 1 is bad and reboots_per_day <= 1 is good;
is what I meant to say!
But now I've taken a look at your phrasing ... I think I read your post as 'at most once every 24 hours', not 'at most every 24 hours' which is actually pretty ambiguous, and not as clear-cut as I was trying to say.
And when people whine about 9x being unstable they don't remember (or even never experienced it themselves) how awful was the hardware it worked on.
I "fondly" remember some combinations of hardware were a literally ticking time bombs, you never knew when it whould BSOD. Though by that time I had enough understanding what if I see CMSXXX.VXD failing it is the problem with a cheap ass sound card drivers, not Windows itself.
So, you should never deploy anything to production which won't be rebooted more frequently then you reboot it in testing. In practice, you should probably reboot much more frequently - as frequently as possible - to keep the delta between "known good" and "mutated" as small as possible.
[1] https://en.wikipedia.org/wiki/MIM-104_Patriot#Failure_at_Dha...
It might be acceptable for AWS to crash every few months, but it’s not acceptable for my systems to be out for the length of a reboot.
https://en.wikipedia.org/wiki/Robustness_principle
It's funny, I spent some time developing tools for CPU architects. Both the concept in the OP's anecdote, and the above principle don't really apply because logic doesn't break in the same way source code breaks. You don't run a program in HDL, you synthesize a logic flow. One could concievably test all possible combinations of that logic for errors, but it becomes 2^N combinations where N is the number of inputs and the number of state elements. Since this cannot be tested because the space is huge (excluding hierarchical designs and emulation), you generate targeted test patterns (and many many mutexes) to pare down the space, and perhaps randomize some of the targeted tests to verify you don't execute out of bounds. And even "out of bounds" is defined by however smart the microarchitect was when they wrote the spec, and that can be wrong too.
The only way to find and fix these bugs is to run trillions of test vectors (called "coverage") and hope you've passed all the known bugs, and not found any new ones.
There are four decades of papers written on hardware validation, so I'm barely scratching the surface, but I think it's a very different perspective compared to how programmers approach the world. I think a lot of the bugs that OP is talking about fall into the hardware logic domain. There isn't really a fallback "throw" or "return status" that you can even check for. Just fault handlers (for the most part).
This is all true of software too. Certainly, it’s easier to test software, but I don’t think it’s because software state is combinatorially simpler than hardware state.
Software fuzzing is more difficult as function parameters can be (literally) infinitely more complicated to define (e.g., variable length strings), whereas logic test vectors are 1's and 0's (a lot of them, but a finite number).
I've yet to see a fuzzing library that could handle all possible combinations of State, even with judicious mutex support.
Uhh, sure it does. Any non-trivial digital hardware design, be it for an ASIC or an FPGA, will contain a lot of state machines interacting with each other, sometimes asynchronously. Hardware isn't immune from reaching a system state which requires a reset of the device.
> Knight, seeing what the student was doing, spoke sternly: “You cannot fix a machine by just power-cycling it with no understanding of what is going wrong.”
> Knight turned the machine off and on.
> The machine worked.
In the novice's particular case, Knight understands what is actually wrong with the machine, and that in this case it will be fixed by turning it off and on again, so he does that.
The idea that in such a situation the machine wouldn't be fixed by power-cycling it when the novice does it, but would when an expert with deep understanding does it, is a joke.
The joke is mimicing the form of a Zen koan. Koans often play with contradiction, I think with the idea of shaking the reader out of simplistic black-and-white thinking into something more holistic and less dichotomizing. I don't think "you need to understand things deeply and not just make random easy changes in the hope of fixing them, but sometimes it happens that those random easy changes are what's actually required" really counts as the sort of holistic non-dichotomizing thing Zen is trying to teach, but it's kinda in the right ballpark.
Like, the “startup process” seems to be framed wrong to me. Why does it take tens of billions of clock cycles (seconds) to start up? Essentially all the system is doing is filling some memory with a known-good kernel image, and initialising some hardware devices. If we think about startup as “restoring a (deterministic) boot image”, then booting is much faster. And we might be able to use that to verify the kernel’s integrity periodically and maybe restore parts of that image at runtime if the system ends up in an invalid state. Sort of a “soft / partial reboot” process.
In my testing code I make heavy use of check() methods. These methods go through all my runtime invariants and checks that they all still hold. This is fantastically useful in testing - a check method and a fuzzer do the work of 10x their weight in unit tests. I wonder if a similar method in the kernel (run every hour or something) could be used to find and track down bugs, and automatically recover when they’re found. “Oh the USB device is in an invalid state. Log and reset it!”.
One of the main responsibilities of the OS is to manage hardware resources. During init, kernel drivers and modules perform health checks and determine availability and status of hardware resources, peripherals, I/O, etc. Since lots of things could go wrong with the various hardware components, at anytime, this process should be performed during each system boot.
Mind you, a once in a blue moon reboot of the network interface might cause more problems in a web application than a reboot of the entire machine.
> I offer the following argument that restarting from the initial state is a deeply principled technique for repairing a stateful system [emphasis original]
Whereas, Wigner's original article on the "Unreasonable Effectiveness of Mathematics in the Natural Sciences" has this statement:
> The miracle of the appropriateness of the language of mathematics for the formulation of the laws of physics is a wonderful gift which we neither understand nor deserve.
Maybe we can make a browser extension that rewrites bad titles.
Offhand I remember some discussion of how old dialup friendly multiplayer games would transfer state. Differential state would be transferred. There might or might not be a checksum. There would be global state refreshes (either periodically or as bandwidth allowed).
The global state refreshes are a different form of re-initilization. The current state discarded in favor of a canonically approved state.
Since the computer is a state machine, and all initial state is read from disk, restarting the computer reinitializes the state machine from initial state. Rebooting resets the state machine to before state mutation entropy led to conflicts. Rebooting is state machine time travel.
We could prevent this situation by designing systems whose state and algorithms don't mutate. But that would require changing software development culture, and that's a much harder problem than rebooting.
One mitigation is regularly snapshotting state, and on error allow the user to revert to a previous state. Your browser's Back button is a form of this.
Another is programmed responses to invalid state. If the program encounters an invalid state, it can attempt to resolve the issue. One method would be to bubble up a signal to the system, and the system can have pre-programmed responses. For example, if a program raises an exception of "Error: Out of disk space", then the larger system can be triggered to perform some disk space garbage collection. If a program raises an "Out of memory" error, the Linux kernel already has an "Out of memory killer" that attempts to relieve applications of their memory.
We can come up with much more sophisticated methods if we try. But again, this would require software development culture to think outside the box, and that is highly unlikely.
I highly recommend this talk on the subject: https://pron.github.io/posts/correctness-and-complexity
It is very strong: its corollary is there is no general algorithm to build Turing Machines with arbitrary non-trivial properties (if there was such an algorithm we could use it to build a TM that solves the halting problem)
So for any non-trivial property the only way to build a TM with this property is to use some ad-hoc method (as there is no general algorithm) - but how do we know the TM we've just built really holds the property we wanted it to? There is no way to verify it due to Rice's theorem.
> for any non-trivial property of partial functions, no general and effective method can decide whether an algorithm computes a partial function with that property
This basically translates to "for each non-trivial property P, there exists a program x for which it is undecidable if x has the property P", or "∀P ∃x such that it's undecidable if x has property P".
Notice that in this statement, programs are quantified using an "exists", not a "forall". Something similar holds for what you said here:
> its corollary is there is no general algorithm to build Turing Machines with arbitrary non-trivial properties
Translated to logic, this becomes "for every program y which takes as input a property and outputs a Turing Machine, there exists a property P such that the y(P) does not have property P", or "∀y ∃P y(P) does not have property P".
Because this statement uses "exists" on the property, it does not rule out the existence of a y which works for most properties P that we'd care about. It only rules out one that works for all properties.
As a simple example where it's possible to prove a non-trivial property, take the property "x is the identity function". This is a non-trivial property, because it is not true about all programs, nor is it false about all programs. And it is trivial to prove that the function "(n) => n" satisfies this property.
Now, I additionally think that most practical programs can have their correctness / incorrectness proven. My reason for this belief is that generally, I expect coders to understand why the code they're writing should work. If nobody understands why the code works, I'd generally consider that a bug, or at least a major code smell. But if the programmer actually understands why the code works, and hasn't made any mistakes, then in principle their understanding could be converted into a proof, because proofs are more or less just correct reasoning, but written out formally. Of course, it would still be very hard, but it's not "provably impossible".
Edit: minor text change
Translated: there exists a computational model weaker than TM but strong enough to implement programs we care about.
Unfortunately, we don't know such computational model and we know empirically surprisingly many problems require a TM (ie. surprisingly many languages are accidentally Turing complete)
It is in fact extremely common to build wholly deterministic machines out of the same parts as our unreliable computers, that are correct by construction. And they may implement a Turing-complete language and still be wholly deterministic, just by limiting the input to what can be proven deterministic in a strictly limited time. Again, this is normally by construction, where the proof is always trivial.
Such a proven-correct system can model literally any space-constrained algorithm.
It is neither common nor easy but actually provably impossible. Please read this piece carefully: https://pron.github.io/posts/correctness-and-complexity
You must be misunderstanding the work if you believe that what is being done routinely is impossible.
Quite possibly I don't understand what you mean by "work". If by "work" you mean creating provably correct software then no, it is not done routinely.
People create wholly deterministic state machines all the damn time.
What do you mean by "wholly deterministic"?
You can buy an 8-bit adder/accumulator from a catalog.
We still don't know if it is possible to build a provably correct multi-tasking OS or a web browser without security vulnerabilities. If it was easy we would already have one written in FStar or Idris.
While admiring all the work done by Project Everest - the fact is that it is a multi-year project (I first read about it about 3-4 years ago and it was already under way) with a very limited scope - it is more of an evidence of difficulties of writing correct software.
1. Even for as weak computational model as FSM the problem of verifying its correctness is intractable (NP-hard). And modularisation does not help.
2. We don't know any weaker computational model that is strong enough to be viable for the problems we want to solve and still weaker than TM. It is well known that it is surprisingly easy for a language to be (accidentally) Turing complete.
In practice, this means that the reductive take of Turing completeness, the halting problem and Rice's theorem is misleading in some way.
Yes, there is no general algorithm for constructing termination proofs for any arbitrary Turing machine. But there are clearly classes of problems for which a biological computer (a human) is able to construct termination proofs and proofs regarding other semantic properties.
See above - there is no reliable procedure (algorithm) that would allow you to do that.
Perhaps you interpreted my post to suggest that you could build a machine that verifies that any arbitrary software run on it never enters an invalid state? If so, I agree, that is not possible. I made a minor edit to make it clear what I’m talking about.
If you think about it - it is the same: how do you know the system you just built is flawless?
You have to use ad-hoc methods but then it is not possible to know your ad-hoc method is correct!
So by "flawless" you can only mean "having some property but we cannot say what property it is".
Edit: added "flawless" definition
This procedure (algorithm) is ill-defined (is not realisable by a TM) as it would require verifying two things:
a) the code you are about to write has a bug (ie. has a specific property)
b) the code you wrote does not have a bug (ie. has a specific property)
Rice's theorem says both are impossible.
So, to go back to the original point, you can't reliably determine the correctness of every possible program. But you could, if you were not prone to error like all humans, always write your own programs such that they are always correct.
From the fact that we know how to write programs that halt you cannot deduce we know how to write programs with arbitrary (or even just useful) property, because there is no general procedure to do that.
In other words: you cannot use the algorithm of writing programs that halt to write programs computing nth digit of Pi :)
In reality most interesting programs accept Turing complete languages - albeit not formally specified :)
Nobody told i3en's have "ephemeral storage" that gets wiped on shutdown.
Basically how can you know all the variables at one time? Start on initialization.
IMO the fact that everything "works" after reboots is simply because the state is 1. well defined and 2. available to all developers.
I do not think that this whole discussion about returning to a "known good initial state" has much merit, because computers are Turing machines: they modify their own state and then act upon that state. Which means that unless your boot drive is read-only, nobody can predict what your computer will do on the next reboot. (Unless someone solved The Halting Problem and I missed it.)
But the discussion about assertions and initially fragile software that quickly becomes robust did strike a chord with me, because that's exactly what I have been doing for decades now. (Also, I see other programmers using assertions so rarely, that I feel kind of lonely in this.) I have a long post about assertions (https://blog.michael.gr/2014/09/assertions-and-testing.html) to which I would like to add a paragraph or two about this aspect, and preferably cite Burrows.
I think that it might be something that you do occasionally (say, once a month or so) for each of your boxes, regardless of whether it's a HA system or not.
Though perhaps I can only say that because the HA systems/components (like API nodes or nodes serving files) won't really see much downtime due to clustering, whereas the non-HA systems/components (monoliths, databases, load balancers) that i've worked with have been small enough for it to be okay to do restarts in the middle of the night (or just weekends or whatever) with very little impact to anyone, even without clustering at the DB level/switchovers.
I mean, surely you do need to update the servers themselves, set aside a bit of time for checking that everything will be okay and don't try to patch a running kernel most of the time, right? I'd personally also want to restart even hypervisor or container host nodes occasionally to avoid weirdness (i've had Docker networks stop working after lots of uptime) and to make sure that restarts will result in everything actually starting back up afterwards successfully.
Then again, most of this probably doesn't apply to the PaaS folk out there, or the serverless crowd.
No. It's the best that can sometimes be done quickly.
Additionally, this doesn't mention the value of a good postmortem. Or the horrors of cloud computing, where restarting things is deemed a good enough mitigation because in the cloud these things happen, and nobody pushes for a good postmortem and repair items.
The last two sections do include the value of postmortems (or forensic analysis, as the author puts it) though from the perspective that this is more feasible and effective on a system that promptly crashes when things go wrong.
So the upshot was, a system crash just meant the audio would cut out for maybe 6 seconds, then start playing again. For the user, completely indistinguishable from poor radio reception, which is always going to happen from time to time anyway. Crashing wasn’t a problem at all as long as it was rare enough.
(I like making robust software and hate crashes, so I don’t like that approach, but I use that experience to remind myself that sometimes worse really is better, or at least good enough.)
Even when no unexpected exception occurs, sometimes frequently restarting the application can be a valid solution. For example when there are potential memory leaks (maybe in your code but maybe also in external library code) and the application should just run infinitely.
Btw, a simple method to restart an application: https://stackoverflow.com/questions/72335904/simple-way-to-r...
0. Cache invalidation
1. Naming things
2. Off-by-one errors
Index 0 is why reboots are so effective.The smaller the scale (amount of people affected), the more sense do business hours make. Ot times of non-operation. Who cares if a personal/smallbiz webserver sleeps for some hours at night?
LOL etc
I also don't know why I would, I can restart the broken components and get a less cluttered debug log when doing only that.