Real System Failures
c3.nasa.gov
c3.nasa.gov
"For more than 30 years, our design lab has seen that no IC greater than 16 pins (except memory) has worked according to its documentation"
That matches the experience of every single embedded engineer I have ever known.
Few lessons I took from it:
- "There is no such thing as digital circuitry. There is only analog circuitry driven to extremes." Digital is a pretty leaky abstraction. You can't safely ignore the physical world. In particular, be wary of your digital parts changing into other parts (or new "parts" appearing out of the blue) thanks to physics.
- There's so much that can go wrong. I'm in awe of people working on life-critical systems, and of challenges they deal with.
- What the fuck is going on with IC durability? The presentation quotes a text from 2013, which says "Commercial semiconductor road maps show component reliability timescales are being reduced to 5–7 years, more closely aligning with commercial product life cycles of 2–3 years." I.e. if your device has modern electronics on-board, it already won't last long because semiconductor devices themselves are expected to naturally fail after few years. This makes me really sad about the state of our technological civilization.
- Don't ignore specs you don't fully, 100% honest-to-god understand! Slide 38 is a damning enough description by itself. I'd add that this also applies to bureaucracies and laws - just because you think some rule is stupid, doesn't mean it is. "Move fast and break things" approach has no place where lives (or livelihoods) can be affected.
- Even adding a node to a linked list isn't a trivial thing, and has many places in which you can screw it up. This highlights just how much acummulated complexity we're dealing with here.
- Life always finds a way... to grow in your electronics and break it.
So, it looks like they'd spend a lot extra with effects on performance/watts/cost to make them last for NASA lengths of time. Those same companies are incentivized to sell new chips regularly in a price-competitive market. So, no reason for them to do aforementioned work since it's just throwing away money.
Again, just overview of a non-HW guy that reads a lot of HW industry's publications. HW people correct anything I missed please.
Yet, the slide goes on to argue this is a software problem? It was my impression that Byzantine tolerant systems required agreement among ⅔ of the nodes; if the system is split 50/50, how can even a tolerant system not fail? (Or rather, is it the difference between failing gracefully and failing spectacularly, and the slide fails to elaborate on exactly how the system failed? But I don't see how we can expect this to succeed.)
See the abstract in Lamport's original paper introducing the Byzantine generals problem [0].
We also see a similar issue in error correction -- an introductory undergrad course might teach this via Lagrange interpolation [1], where you need only n+k of the coefficients in the presence of erasure errors, but n+2k in the general case (where n is the size of the actual message, and k is the maximum number of errors to correct).
[0] https://www.microsoft.com/en-us/research/wp-content/uploads/...
In the situtation described in my link, there were 4 data sources, and 4 processors. A byzantine fault occurred with the link from one of the data sources to the processors. This cascaded into a failure of the entire 4 processor system when it fell into a 2:2 split. In concept, the processors could have communicated with each other and detected a disagreement in one of the four data sources.
The software bug is that a fault with one of the data sources cascaded into a fault with the entire processing unit. In a correctly functioning system, the fault should have been contained to just the one data source. That is, the problem should have been no worse then simply losing the bus entirely.
Assuming there was sufficient redundancy across the buses, the system could have continued functioning properly despite the fault. However, because the software did not correctly handle the fault, any other redundancy became useless.
[0] https://www.cs.indiana.edu/classes/p545-sjoh/post/lec/fault-...
Also, it's a great example of the Edward-Tufte-hating, horribly, hilariously bad PowerPoint NASA presentation style.
The NASA environment, which includes Honeywell in this case, is full of these sorts of things. This one is actually quite good.
This applies to the modern web as well.
On the other hand this presentation is not that bad, at least it does not have bullet lists nested five levels deep where text size does not match the nesting levels and such things.
NASA and its contractors (which, in my experience, do most of the work attributed to NASA) have a weird, self-reinforcing cycle of decision-making by PowerPoint such that slide decks are important, necessary, and ubiquitous and therefore almost universally bad. Tufte's example of of a presentation burying the lede that the Shuttle will get blowed up is just one consequence.
The amount of confidence people have in their ability to plan for contingencies seems to go down in proportion to their exposure to hardware. Complex systems are endlessly inventive when it comes to finding ways to fail.
"If only we'd had the human, time, money and organizational support resources to plan ahead more accurately, we wouldn't have made this particular mistake!" That's called the benefit of hindsight, and it's the project manager's classic "told you so". To management it sounds like "give me more budget and a slacker timeline", and to engineering it sounds like "someone wants to use a different one-true-solves-all-problems-solution".
Experienced system designers know that the real art is knowing that out in the real world, things will fail no matter how careful you are, so anticipating and detecting both known and unknown failure modes and recovering from them is really the critical need.
For an accessible, real world study of how this can be achieved with arbitrarily complex software systems, I can highly recommend reading about Erlang, or alternatively deploying a nontrivial pacemaker/corosync cluster. Most engineers never build a system this resilient in their lifetime, but once you have, you can never look back.
Further, instead of the "build the perfect system" philosophy put forward in the presentation (ie. formal analysis), both solutions use the alternative "tolerate and control for failure". This is a significant philosophical and practical distinction.
[1] https://mars.jpl.nasa.gov/mer/newsroom/pressreleases/2004012...
But how do you formally verify/analyze a system for fault and failure tolerance if the methods of detecting failure and other faults are themselves not enough?
e.g. The slide about COM/MON, which I admit I didn't fully understand, seems to be that the solution picked wasn't the very best possible one due to constraints and that failures were not detected that the point they were expected to.
I guess you at least would know those are failure/fault points which can not be tolerated or handled somehow and should be watched.
The only place I have encountered something like this was on an Arduino board where the use of a buzzer was causing a voltage drop that affected the logic of the code. (It appeared that a delay function returned immediately instead of taking 250ms, which sped up the loop.)
Question:
How do you actually implement Byzantine Fault Tolerance?
I found this in Wikipedia:
Byzantine fault tolerance mechanisms use components that repeat an incoming message (or just its signature) to other recipients of that incoming message. All these mechanisms make the assumption that the act of repeating a message blocks the propagation of Byzantine symptoms.
Is verifying the interpreted input value the primary way to design for Byzantine Fault Tolerance?
Practical BFT: http://pmg.csail.mit.edu/papers/osdi99.pdf The Night Watch: https://www.usenix.org/system/files/1311_05-08_mickens.pdf
Generally the idea is to assume that there will be fewer than k failures out of the n nodes you have.
[0] https://embdev.net/topic/118781?page=single [1] https://aerocontent.honeywell.com/aero/common/documents/myae...
It's a great bit of hacker lore, if you haven't yet read it.
It shows how lucky an average programmer is. We have to deal with relatively easy issues; we can modify code, recompile, debug and repeat until success. :)