Very nice article, thanks. Reading with a slightly cooler head than yesterday, I still have strong disagreement with this.
I think it boils down to this:
> Here I'm classifying bohrbugs as repeatable, and heisenbugs as transient.
> If you have bohrbugs in your system's core features, they should usually be very easy to find before reaching production.
I don't know about "bohrbugs", but the vast majority of my bugs are "i-am-an-idiot-bugs", and boil down to "some combination of the handling one of the many possible inputs is going to trigger some code path where a function does not have the proper arity because I made a typo somewhere because I'm an idiot."
I claim that those are not "easy to find before reaching production" as a human being. (And I argue a large part of them are trivially found by a decent type static analyzer. Which, unfortunately, dialyzer was not last time I checked, 2 years ago.)
> By virtue of being repeatable, and often on a critical path, you should encounter them sooner or later, and fix them before shipping.
I disagree with that. My bugs happen on critical paths, but with non obvious shape of data. (I receive input from untrustworthy, aka "real-world", sources.)
I've spent enough time fighting bugs that were in the critical part of the code, were tested multiple times at the unit, integration and acceptance level, with different kind of simulators, but still broke on the first real-life usage because one function down a stack was not handling a nil or expected a string instead of a number, somewhere.
Encountering those issues "sooner" becomes a matter of "having enough imagination when writing automated tests".
I've heard property testing could be a tool to "automate" such imagination. But I'd rather have something telling me I just made a typo ;)
> So really, how efficient is restarting as a strategy?
> Well for repeatable bugs on core features, restarting is useless.
Amen.
> if the feature is a thing very important to a very small amount of users, restarting won't do much. If it's a side-feature used by everyone, but to a degree they don't care much about, then restarting or ignoring the failure altogether can work well.
It's hard to not read this as "Let it crash" being a good philosophy - provided it does not really matter, if your application works ;) (I know this is not the intended message, but still...)
I had great expectations about this section of the article:
> > I like statically typed language, and I restart my daemon after unhandled exception. What is Erlang gonna win me on fault tolerance
> This question was asked to me on a forum where I was discussing programming stuff and discussing the Erlang model. I copied it verbatim because it's a great example of a question a lot of people ask when they hear about restarting and Erlang's features.
The author follows with question (which is, basically, my opinion), with a very insightful explanation of supervision tree and OTP, but unfortunately, does not seem to talk about the "statically typed" part.
Which is a shame ! Message-passing, supervision and OTP are the great aspects of Erlang.
They're extremely powerful when used appropriately (which, in my experience, mean "sparingly, and as a plumbing-layer over purely functional code".)
My whole frustration with Elixir is that I want static typing "behind" the boundaries of GenServers.
Sure, I understand "GenServer.send" can not be typed, in general, and I'm ready to do some level of input validation at the boundaries, but beyond that point, please, If I tell you to add two apples, and I mispelled the name of one of them, don't even let me run, okay ?
I also know that there are efforts to do that (alpaca & co), but they're just too early-stages for my taste.
As said earlier, if you're solving the kind of problems where you never those kind of "i-am-an-idiot-bugs", or you're just not as idiots as I am, you may not find this problematic. I do.
But then, for dramatic purpose, the article ends with this gem:
> At some point, a team member got a very expensive credit card bill for the logging service we were using to aggregate exceptions. That's when we took a look at it and saw the horror on the leftmost side of the diagram: we were generating between 500,000 to 1,200,000 exceptions a day! Holy cow, that was a lot. But was it? If the issue was a heisenbug, and our system was seeing, say 100,000 requests a second, what were the odds of it happening? Something between 1/17000 and 1/7000. Somewhat very frequent, but because it had no impact on service, we didn't notice it until the bandwidth and storage bill came through.
Really ? Where those 500,000 to 1,200,000 exceptions not blocking anyone ? Not wasting someone's time ? Not causing improper or delayed results ?
Or, worst, just slowly making people lose faith into your software, or in software in general ? (Okay, I'm being slighly grandiloquant here, time to wrap up.)
Thanks for the link.