Note I wrote when X fails, not if X fails. It's a different way of thinking.
Note I wrote when X fails, not if X fails. It's a different way of thinking.
This to me is the biggest difference between writing code for the software industry vs. an industrial industry.
Software is all about the happy path ("move fast and break things") because the consequences typically range from a minor inconvenience to a major financial loss.
Industrial control is all about sad paths ("what happens if someone drives a forklift into your favorite junction box during the most critical, exothermic phase of some reaction") because the consequences usually start at a major financial loss and top out in "Modern Marvels - Engineering Disasters" territory.
I have worked on projects where in retrospect the LOC generated per day, if spread out across the whole project, were between 1 and 3.
But typically, writing of the code does not even commence in the first year, sometimes two.
Then there is the test cases and test coverage etc etc.
This is the difference between engineering code and just producing it - all the effort that goes into understanding all the unwanted code behaviour that may occur and how to detect, manage and/or avoid it.
Implicit state is the enemy, therefore the best code has all states explicitly defined.
Fukushima:
badthink - the seawall is high enough that it will stop tidal waves
goodthink - what happens when the seawall is overtopped? Answer: the backup generators drown. Solution: put the backup generators on a platform.
Deepwater Horizon:
badthink - the pipe is strong enough to never break
goodthink - what happens when there's enough force to bust the pipe off? Answer: the pipe flow cannot be shut off. Solution: put a fuse (a weak spot) above the valve, so when the pipe busts off, it breaks above the valve, and the valve can be turned to shut off the flow. (The valve was located on the sea floor.)
badthink: the Fukushima backup generators must be placed on a platform to keep them out of the range of a once in a millenium tsunami
goodthink: what happens when a typhoon comes and damages the generator on an exposed platform; an event which happens predictably and far more often than tsunamis. Answer: put the backup generators in the basement of a reactor building behind a large seawall. What catastrophe could put the reactor building completely underwater, and still have the reactor survive?
Yeah, trivial changes to the design can prevent all sorts of disasters, but you have to know what you are trying to prevent in a world of infinite complexity
Our customers are in a cutthroat market with low margins. We can't spend a ton on pre-analysis, redundancies and so on.
Instead we've focused reduced the impact of failures.
We've made it trivial to switch to an older build in case the new one has an issue. Thus if they hit a bug they can almost always work around it by going to an older build.
This of course requires us to be careful about database changes, but that's relatively easy.
It's like buying a thermometer from Home Depot vs a highly accurate, calibrated lab thermometer. Sometimes you just don't need that quality and it's a waste paying for it.
I find that I can learn a ton from those industries, and as a software engineer I have the added advantage of being able to come up with zero-cost (or low cost), self-documenting abstractions, testing patterns, and ergonomic interfaces that improve the safety of my software.
In software, a lot of safety is embodied in how you structure your interfaces and tests. The biggest cost is your time, but there are economies of scale everywhere. It really pays to think through your interfaces and test plan and systems behavior, and that's where lessons from these other industries can be applied.
So yeah, if you think of these lessons as "do tons of manual QA", you'll run into trouble resourcing it. But you can also think of them as "build systems that continuously self-test, produce telemetry, fail gracefully in legible ways and have multiple redundancies".
An airliner is a lot of lives, a lot of money, a lot of fuel, and a lot of energy. Which is why a lot has been invested in training, procedure, and safety systems.
Cars operates in an environment which is in most ways a lot more forgiving, they’re controlled by (on average) low-training low-skill non-redundant crews, they’re much more at risk of “enemy action”, the material stresses are in a different realm, and they’re much, much more sensitive to price pressure.
Hell, the difference is already visible in aviation alone, crop dusters and other small planes are a lot less regulated amongst every axis than airliners are.
A whole lot more people die from car accidents, yet there are few reports on national news on accidents. So fewer people care. Meanwhile each time there is an aviation disaster, 100s of people die and it's all over the news for weeks. Similarly with train accidents and nuclear accidents. There where only 2 very large ones but they still haunt the field to this day, while (for example) the deaths from solar installations by people falling from roofs are mostly ignored.
Large accidents have to be avoided, a lot of small ones are more acceptable.
But that is cost/benefit analysis. When any accident can kill hundreds and do millions to billions in damage besides (to say nothing of the image damage to both the sector and the specific brand), the benefit of trying to prevent every accident is significant, so acceptable costs are commensurate.
Which is actually counterproductive! This makes it harder to compete as a bus service, bus lines shut down, and more people drive. I wrote more about this at https://www.jefftk.com/p/make-buses-dangerous and https://www.jefftk.com/p/in-light-of-crashes-we-should-not-m...
I mean you're implying that there are more accidents with autopilot than without it, right? Seems like quite the claim...
Example: https://www.theguardian.com/technology/2023/nov/22/tesla-aut...
The fact is, there’s a lot of history and best practice around building safety critical systems that Tesla doesn’t follow.
Additionally, even with the practices they follow, they call a consumer facing product that isn’t really an autopilot “autopilot”, while focusing outbound comms on a beta product that is more like an autopilot, but not available to them.
A plane pilot knows very well what the limits of the autopilot are and what the passenger believes is irrelevant.
Conversely if too many/most car “autopilot” users believe it does more than what it really does then it’s dangerous.
In electrical engineering 600V is still “low voltage”. Any engineer in the field knows that so that’s fine right? But if someone sells “low voltage” electric toothbrush or hand warmer no normal person will think “it’s 600V, it will probably kill me”. When you sell something, what your target audience takes away from your advertisement matters. If they’re clearly confused and you aren’t clearing it up after so many years then “confusion” and misleading advertising are part of your sales strategy.
Nobody here on HN, because we're really into tech. Outside the tech world, I would guess that 50% of the population thinks that "autopilot" (on any device) means that no human is needed.
I like the idea of thinking 'when' instead of 'if', but the verdict should be even harder when it comes to software engineering because it has this rare material at its disposal, which doesn't degrade over time.
On the 757, one set of control cables runs under the floor. The backup set runs in the ceiling.
As it was pointed out to me, airplanes sitting on the ground are a black hole sucking up money. Airplanes in the air carrying payload (note the "pay" in payload) are making money. Boeing understands this very well, and is very focused on getting that airplane in the air making money as much as possible.
crickets, let's just randomise which sensor we use during boot, that ought to do it!
let's just build a system that pushes the nose down under those conditions, have it accept potentially unreliable AoA data, and not tell pilots about it!
And the reference is presumably to 737 MAX accident. https://www.afacwa.org/the_inside_story_of_mcas_seattle_time...
(Airbus is not Boeing.)
Boeing planes (before MCAS): we have detected a problem with your engines, would you like to shut down?
Airbus planes: we have detected a problem with your engines, we have shut them down for you.
Like what’s the secret sauce of nvidia vs radeon or AMD vs intel? Reliable execution, seemingly - and this is an environment where failures are supposed to be contained to very specific rates at given levels of severity.
The FAA has gotten into a mode where they let boeing sign off on their own deviations from the rules, the engine changes forced the introduction of the nose-pusher-down system which really should have required training, but Boeing didn't want to do that, because the whole point of doing the weird engine thing was having ostensible "airframe compatibility" despite the changes in flight characteristics. And they have become so large (like intel) that they don’t have to care anymore, because they know there’s no chance of actual regulatory consequences, nor can the EAA kick them out without causing a diplomatic incident and massively disrupting air travel, so they are no longer rigorous, and we simply have to deal with Boeing’s “meltdown”.
And yes they should be doing better but in the abstract, certification processes always need to be dealing with “uncooperative” participants who may want to conceal derogatory information or pencil-whip certification. You need to build processes that don’t let that happen and nowadays there’s so much of a revolving door that they can just get away with it. Like none of this would have happened with the classified personnel certification process etc - it is fundamentally a problem of a corrupted and ineffective certification process.
This decline in certification led to an inevitable decline in quality. When companies figure out it’s a paper tiger then there’s no reason to spend the money to do good engineering.
The FAA’s processes are both too strict and too lax - we have moved into the regulatory capture phase where they purely serve the interests of the industry giants who are already established and consolidated, and they now serve primarily to exclude any competitors rather than ensure consistent quality of engineering.
The specifics are less interesting than that high-level problem - there obviously eventually would be some form of engineering malfeasance that resulted from regulatory capture, the specific form is less important than the forces that produced it. And that regulatory capture problem exists across basically the whole American system. Why do we have forced arbitration on everything, why are our trains dumping poison into our towns? Because from 1980-2020 we basically handed control of legislative policy over to corporate interests and then allowed a massive degree of consolidation. Not that airbus is small, but the EAA isn’t regulatory capture to the extent of most American bureaus.
Most of what was written about the MAX crashes in the mass media is utter garbage and misinformation. No surprise there, as journalists have zero expertise in how airplanes work.
Both crashes could have been easily averted if the crews had followed well-known procedures. There was also nothing wrong with the aerodynamics of the MAX, nor the concept of the MCAS system. The flaw was in the way the MCAS system was implemented, and the way the pilots responded to it.
For example, rarely mentioned is the third MAX incident, where the airplane continued normally to their destination. The crew simply turned off the stab trim system.
BTW, I had a nice conversation with a 737 pilot a few months ago. He told me what I had already concluded - the crashed crews did not follow the procedures. I've also had unsolicited emails from pilots who told me what I'd written about it was true.
The EA crew oversped the airplane (you can hear the overspeed warning horn on the CVR) and did nothing to correct it. This made things worse. They were also given an Emergency Airworthiness Directive which said to restore normal trim switches, then turn off the trim system. They did not.
That's it.
I'd say half the fault was Boeing's, the other half the flight crews'.
The MCAS is not a bad concept, note that MCAS is still there in the MAX.
Pilots are a brotherhood, and they don't care to criticize other pilots in public. But they will in private.
Which is why the entire worldwide MAX fleet was grounded for more than a year, and the regulators didn't just mandate a bit of extra training.
Coming up with this narrative about how it's the crew's fault because they failed to disable Boeing's quietly introduced little self-destruct system fast enough to save their own lives was a particularly despicable move from their PR department and I lost a lot of respect for them over that.
As for training, turning off the stab trim system to stop runaway trim is a "memory item", which means the pilots must know it without needing to consult a checklist. Additionally, after the first crash, all MAX crews received an EMERGENCY AIRWORTHINESS DIRECTIVE with a two-step procedure:
1. restore normal trim with the electric trim switches
2. turn off the trim system
I expect a MAX pilot to read, understand, and remember an EMERGENCY AIRWORTHINESS DIRECTIVE, especially as it contains instructions on how not to crash like the previous crew. Don't you?
> might have been able to save the aircraft
It's a certainty. Remember the first LA MAX incident, the airplane did not crash because after restoring normal trim a couple times, the crew turned off the trim system, and continued the flight normally. They apparently didn't even think it was a big deal, as the aircraft was handed over to the next crew, who crashed.
> a bit of extra training
They are already required to know all "memory items".
> Coming up with this narrative about how it's the crew's fault because they failed to disable Boeing's quietly introduced little self-destruct system fast enough to save their own lives was a particularly despicable move from their PR department
AFAIK Boeing never did say it was the crew's fault. The "have to respond within 5 seconds" is a fantasy invented by the media. It is not factual.
Both Boeing and the crews share responsibility for the crashes.
And I never said they didn't. I just choose to assign Boeing the lion's share of the blame, as they should never have let that rush-job, cost-cutting death trap of a machine take to the skies in the first place.
Anyway, I see you have your mind made up, so there's not much point in arguing further. If you feel like continuing, why don't you take it up with - let's see - every single global aviation regulator, who also somehow came to the conclusion that there was maybe something a little bit wrong with the type.
I thought that the majority of the problems was that Boeing wanted the same type-rating, so that airlines could avoid paying for training. This resulted in crews not getting proper training and so not knowing the proper procedures ... which was by decision.
Both the airlines and Boeing should take the blame; I don't really see how it would be the pilots fault, if you lie and say "it's the same plane, it flies the same, you don't need conversion training".
I am not in aviation, most of this is from YouTube sources, so y'know ...