A Matter of Millimeters: The story of Qantas flight 32
admiralcloudberg.medium.com
admiralcloudberg.medium.com
Being an SRE at a FAANG and generally spending a lot of my life dealing with reliability, I am consistently in awe of the aviation industry. I can only hope (and do my small contribution) that the software/tech industry can one day be an equal in this regard.
And finally, the biggest of kudos to the Kyra Dempsey the writer. What an approachable article despite being (necessarily) heavy on the engineering content.
Try drawing the software monstrosity you work on / with as an airplane. 100 wings sticking out all different directions, covered with instruments and fins, totally asymmetrical and 5 miles long. Propellers, jets, balloons, helicopter blades.
Yep, it flies.
When it crashes, just take off again.
Every single model is somewhat bespoke. There’s common components but each ends up having its own special problems in a way I assume different car models in a common platform (or two small SUVs from competing manufacturers) just don’t.
Two 767 made few months apart will have initial difference, like two different versions of java 8 SDK.
1: https://en.wikipedia.org/wiki/List_of_Boeing_customer_codes
However, I have been told by an insider that supply chain integrity is an underappreciated issue. Someone has been caught selling fake plane parts through an elaborate scheme, and there are other suspicious suppliers, which is a bit unsettling:
"Safran confirmed the fraudulent documentation, launching an investigation that found thousands of parts across at least 126 CFM56 engines were sold without a legitimate airworthiness certificate."
https://www.businessinsider.com/scammer-fooled-us-airlines-b...
https://admiralcloudberg.medium.com/riven-by-deceit-the-cras...
And to be able to reconstruct the chain of events after the components in question have exploded and been scattered throughout south-east Asia is incredible.
Or at least, I assume the turbine parts weren’t defective, although given what seems to be quite a happy-go-lucky approach to manufacturing defects in Hucknall, maybe my assumption is not made on solid grounds…
Not exactly. The idea is not not making mistakes, it's whatcha gonna do about X when (not if) it fails.
Note I wrote when X fails, not if X fails. It's a different way of thinking.
crickets, let's just randomise which sensor we use during boot, that ought to do it!
And the reference is presumably to 737 MAX accident. https://www.afacwa.org/the_inside_story_of_mcas_seattle_time...
let's just build a system that pushes the nose down under those conditions, have it accept potentially unreliable AoA data, and not tell pilots about it!
An airliner is a lot of lives, a lot of money, a lot of fuel, and a lot of energy. Which is why a lot has been invested in training, procedure, and safety systems.
Cars operates in an environment which is in most ways a lot more forgiving, they’re controlled by (on average) low-training low-skill non-redundant crews, they’re much more at risk of “enemy action”, the material stresses are in a different realm, and they’re much, much more sensitive to price pressure.
Hell, the difference is already visible in aviation alone, crop dusters and other small planes are a lot less regulated amongst every axis than airliners are.
A whole lot more people die from car accidents, yet there are few reports on national news on accidents. So fewer people care. Meanwhile each time there is an aviation disaster, 100s of people die and it's all over the news for weeks. Similarly with train accidents and nuclear accidents. There where only 2 very large ones but they still haunt the field to this day, while (for example) the deaths from solar installations by people falling from roofs are mostly ignored.
Large accidents have to be avoided, a lot of small ones are more acceptable.
But that is cost/benefit analysis. When any accident can kill hundreds and do millions to billions in damage besides (to say nothing of the image damage to both the sector and the specific brand), the benefit of trying to prevent every accident is significant, so acceptable costs are commensurate.
Which is actually counterproductive! This makes it harder to compete as a bus service, bus lines shut down, and more people drive. I wrote more about this at https://www.jefftk.com/p/make-buses-dangerous and https://www.jefftk.com/p/in-light-of-crashes-we-should-not-m...
I mean you're implying that there are more accidents with autopilot than without it, right? Seems like quite the claim...
Example: https://www.theguardian.com/technology/2023/nov/22/tesla-aut...
The fact is, there’s a lot of history and best practice around building safety critical systems that Tesla doesn’t follow.
Additionally, even with the practices they follow, they call a consumer facing product that isn’t really an autopilot “autopilot”, while focusing outbound comms on a beta product that is more like an autopilot, but not available to them.
A plane pilot knows very well what the limits of the autopilot are and what the passenger believes is irrelevant.
Conversely if too many/most car “autopilot” users believe it does more than what it really does then it’s dangerous.
In electrical engineering 600V is still “low voltage”. Any engineer in the field knows that so that’s fine right? But if someone sells “low voltage” electric toothbrush or hand warmer no normal person will think “it’s 600V, it will probably kill me”. When you sell something, what your target audience takes away from your advertisement matters. If they’re clearly confused and you aren’t clearing it up after so many years then “confusion” and misleading advertising are part of your sales strategy.
Nobody here on HN, because we're really into tech. Outside the tech world, I would guess that 50% of the population thinks that "autopilot" (on any device) means that no human is needed.
Our customers are in a cutthroat market with low margins. We can't spend a ton on pre-analysis, redundancies and so on.
Instead we've focused reduced the impact of failures.
We've made it trivial to switch to an older build in case the new one has an issue. Thus if they hit a bug they can almost always work around it by going to an older build.
This of course requires us to be careful about database changes, but that's relatively easy.
It's like buying a thermometer from Home Depot vs a highly accurate, calibrated lab thermometer. Sometimes you just don't need that quality and it's a waste paying for it.
I find that I can learn a ton from those industries, and as a software engineer I have the added advantage of being able to come up with zero-cost (or low cost), self-documenting abstractions, testing patterns, and ergonomic interfaces that improve the safety of my software.
In software, a lot of safety is embodied in how you structure your interfaces and tests. The biggest cost is your time, but there are economies of scale everywhere. It really pays to think through your interfaces and test plan and systems behavior, and that's where lessons from these other industries can be applied.
So yeah, if you think of these lessons as "do tons of manual QA", you'll run into trouble resourcing it. But you can also think of them as "build systems that continuously self-test, produce telemetry, fail gracefully in legible ways and have multiple redundancies".
Fukushima:
badthink - the seawall is high enough that it will stop tidal waves
goodthink - what happens when the seawall is overtopped? Answer: the backup generators drown. Solution: put the backup generators on a platform.
Deepwater Horizon:
badthink - the pipe is strong enough to never break
goodthink - what happens when there's enough force to bust the pipe off? Answer: the pipe flow cannot be shut off. Solution: put a fuse (a weak spot) above the valve, so when the pipe busts off, it breaks above the valve, and the valve can be turned to shut off the flow. (The valve was located on the sea floor.)
badthink: the Fukushima backup generators must be placed on a platform to keep them out of the range of a once in a millenium tsunami
goodthink: what happens when a typhoon comes and damages the generator on an exposed platform; an event which happens predictably and far more often than tsunamis. Answer: put the backup generators in the basement of a reactor building behind a large seawall. What catastrophe could put the reactor building completely underwater, and still have the reactor survive?
Yeah, trivial changes to the design can prevent all sorts of disasters, but you have to know what you are trying to prevent in a world of infinite complexity
As it was pointed out to me, airplanes sitting on the ground are a black hole sucking up money. Airplanes in the air carrying payload (note the "pay" in payload) are making money. Boeing understands this very well, and is very focused on getting that airplane in the air making money as much as possible.
I like the idea of thinking 'when' instead of 'if', but the verdict should be even harder when it comes to software engineering because it has this rare material at its disposal, which doesn't degrade over time.
This to me is the biggest difference between writing code for the software industry vs. an industrial industry.
Software is all about the happy path ("move fast and break things") because the consequences typically range from a minor inconvenience to a major financial loss.
Industrial control is all about sad paths ("what happens if someone drives a forklift into your favorite junction box during the most critical, exothermic phase of some reaction") because the consequences usually start at a major financial loss and top out in "Modern Marvels - Engineering Disasters" territory.
I have worked on projects where in retrospect the LOC generated per day, if spread out across the whole project, were between 1 and 3.
But typically, writing of the code does not even commence in the first year, sometimes two.
Then there is the test cases and test coverage etc etc.
This is the difference between engineering code and just producing it - all the effort that goes into understanding all the unwanted code behaviour that may occur and how to detect, manage and/or avoid it.
Implicit state is the enemy, therefore the best code has all states explicitly defined.
(Airbus is not Boeing.)
Boeing planes (before MCAS): we have detected a problem with your engines, would you like to shut down?
Airbus planes: we have detected a problem with your engines, we have shut them down for you.
Like what’s the secret sauce of nvidia vs radeon or AMD vs intel? Reliable execution, seemingly - and this is an environment where failures are supposed to be contained to very specific rates at given levels of severity.
The FAA has gotten into a mode where they let boeing sign off on their own deviations from the rules, the engine changes forced the introduction of the nose-pusher-down system which really should have required training, but Boeing didn't want to do that, because the whole point of doing the weird engine thing was having ostensible "airframe compatibility" despite the changes in flight characteristics. And they have become so large (like intel) that they don’t have to care anymore, because they know there’s no chance of actual regulatory consequences, nor can the EAA kick them out without causing a diplomatic incident and massively disrupting air travel, so they are no longer rigorous, and we simply have to deal with Boeing’s “meltdown”.
And yes they should be doing better but in the abstract, certification processes always need to be dealing with “uncooperative” participants who may want to conceal derogatory information or pencil-whip certification. You need to build processes that don’t let that happen and nowadays there’s so much of a revolving door that they can just get away with it. Like none of this would have happened with the classified personnel certification process etc - it is fundamentally a problem of a corrupted and ineffective certification process.
This decline in certification led to an inevitable decline in quality. When companies figure out it’s a paper tiger then there’s no reason to spend the money to do good engineering.
The FAA’s processes are both too strict and too lax - we have moved into the regulatory capture phase where they purely serve the interests of the industry giants who are already established and consolidated, and they now serve primarily to exclude any competitors rather than ensure consistent quality of engineering.
The specifics are less interesting than that high-level problem - there obviously eventually would be some form of engineering malfeasance that resulted from regulatory capture, the specific form is less important than the forces that produced it. And that regulatory capture problem exists across basically the whole American system. Why do we have forced arbitration on everything, why are our trains dumping poison into our towns? Because from 1980-2020 we basically handed control of legislative policy over to corporate interests and then allowed a massive degree of consolidation. Not that airbus is small, but the EAA isn’t regulatory capture to the extent of most American bureaus.
Most of what was written about the MAX crashes in the mass media is utter garbage and misinformation. No surprise there, as journalists have zero expertise in how airplanes work.
Both crashes could have been easily averted if the crews had followed well-known procedures. There was also nothing wrong with the aerodynamics of the MAX, nor the concept of the MCAS system. The flaw was in the way the MCAS system was implemented, and the way the pilots responded to it.
For example, rarely mentioned is the third MAX incident, where the airplane continued normally to their destination. The crew simply turned off the stab trim system.
BTW, I had a nice conversation with a 737 pilot a few months ago. He told me what I had already concluded - the crashed crews did not follow the procedures. I've also had unsolicited emails from pilots who told me what I'd written about it was true.
The EA crew oversped the airplane (you can hear the overspeed warning horn on the CVR) and did nothing to correct it. This made things worse. They were also given an Emergency Airworthiness Directive which said to restore normal trim switches, then turn off the trim system. They did not.
That's it.
I'd say half the fault was Boeing's, the other half the flight crews'.
The MCAS is not a bad concept, note that MCAS is still there in the MAX.
Pilots are a brotherhood, and they don't care to criticize other pilots in public. But they will in private.
Which is why the entire worldwide MAX fleet was grounded for more than a year, and the regulators didn't just mandate a bit of extra training.
Coming up with this narrative about how it's the crew's fault because they failed to disable Boeing's quietly introduced little self-destruct system fast enough to save their own lives was a particularly despicable move from their PR department and I lost a lot of respect for them over that.
As for training, turning off the stab trim system to stop runaway trim is a "memory item", which means the pilots must know it without needing to consult a checklist. Additionally, after the first crash, all MAX crews received an EMERGENCY AIRWORTHINESS DIRECTIVE with a two-step procedure:
1. restore normal trim with the electric trim switches
2. turn off the trim system
I expect a MAX pilot to read, understand, and remember an EMERGENCY AIRWORTHINESS DIRECTIVE, especially as it contains instructions on how not to crash like the previous crew. Don't you?
> might have been able to save the aircraft
It's a certainty. Remember the first LA MAX incident, the airplane did not crash because after restoring normal trim a couple times, the crew turned off the trim system, and continued the flight normally. They apparently didn't even think it was a big deal, as the aircraft was handed over to the next crew, who crashed.
> a bit of extra training
They are already required to know all "memory items".
> Coming up with this narrative about how it's the crew's fault because they failed to disable Boeing's quietly introduced little self-destruct system fast enough to save their own lives was a particularly despicable move from their PR department
AFAIK Boeing never did say it was the crew's fault. The "have to respond within 5 seconds" is a fantasy invented by the media. It is not factual.
Both Boeing and the crews share responsibility for the crashes.
And I never said they didn't. I just choose to assign Boeing the lion's share of the blame, as they should never have let that rush-job, cost-cutting death trap of a machine take to the skies in the first place.
Anyway, I see you have your mind made up, so there's not much point in arguing further. If you feel like continuing, why don't you take it up with - let's see - every single global aviation regulator, who also somehow came to the conclusion that there was maybe something a little bit wrong with the type.
I thought that the majority of the problems was that Boeing wanted the same type-rating, so that airlines could avoid paying for training. This resulted in crews not getting proper training and so not knowing the proper procedures ... which was by decision.
Both the airlines and Boeing should take the blame; I don't really see how it would be the pilots fault, if you lie and say "it's the same plane, it flies the same, you don't need conversion training".
I am not in aviation, most of this is from YouTube sources, so y'know ...
On the 757, one set of control cables runs under the floor. The backup set runs in the ceiling.
Here at HN we want a post mortem for a cloud failure in a matter of hours.
I'll go one further - I've yet to finish writing a postmortem on one incident before the next one happens. I also have my doubts that folks wanting a PM in O(hours) actually care about its contents/findings/remediations - its just a tick box in the process of day-to-day ops.
And then, I saw an endless stream of aggrieved comments from people who were personally outraged that the outcome, whatever it might be, hadn't been finalized yet at the late, late date of... late February.
They may want a mitigation or RCA in hours, but even AWS gives us NDA restricted PMs in > 24 hours.
I’d love to be an engineer with unlimited time budget to worry about “when, not if, X happens” (to quote a sibling comment).
But people don’t tend to die when we mess up, so we don’t get that budget.
It must have something to do with the number of mistakes, otherwise it's all a waste of time!
It's all well and good responding to mistakes as thoroughly as possible, but if it's not reducing the number of mistakes, what's it all for?
Not really. Imagine two systems with the same amount of mistakes. (Here the mistakes can be either bugs, or operator mistakes.)
One is designed such that every mistake brings the whole system down for a day with millions of dollars of lost revenue each time.
The other is designed such that when a mistake happens it is caught early, and when it is not caught it only impacts some limited parts of the system and recovering from the mistake is fast and reliable.
They both have the same amount of mistakes, yet one of these two systems is wastly more reliable.
> if it's not reducing the number of mistakes, what's it all for
For reducing their impact.
Checklists of course are not the same as detailed post-mortems but they belong to the same way of thinking. And they would cost pretty much nothing to implement.
Also CRM: it's very important to have a culture where underlings feel they can speak up when something doesn't look right -- or when a checklist item is overlooked, for that matter.
I was a submarine nuclear reactor operator, and one of my Commanding Officers once ordered that we stop using checklists during routine operations for precisely this reason. Instead, we had to fully read and parse the source documentation for every step. Before, while we of course had them open, they served as more of a backstop.
His argument – which I to some extent agree with – was that by reading the source documentation every time, we would better engage our critical thinking and assess plant conditions, rather than skimming a simplified version. To be clear, the checklists had been generated and approved by our Engineering Officer, but they were still simplifications.
Even then, you have to get people to read them, which is somehow a monumental task. Docs? Nah, lemme read this Medium blog instead.
On a side note, that's also why there's all the nomsense security theater at airports.
There's a slight difference in terms of what kind of damage an airplane malfunctioning causes compared to a button on an e-commerce shop rendering improperly for one of the browsers. My point is that the level of investment in reliability and process should be proportional to the potential damage of any incidents.
See the excellent To Engineer is Human in just this topic of analyzed failures in civil engineering.
This part of the article really leapt out at me:
The tolerance for this bore was supposed to be Ø 0.05 mm according to the design drawings, but was changed to Ø 0.5 mm in the manufacturing drawings without explanation. Even so, the non-conformance on the accident hub was between Ø 0.90 and Ø 0.98 (an offset of 0.45–0.49 mm), which should have been flagged by the machine. The CMM records from the accident hub were not retained, so it was not possible for investigators to confirm that the error was actually registered.
The meaning might not be obvious if you've never worked in a machine shop, but it's crystal clear if you have. Many people at that plant knew that they were delivering out-of-spec parts. Everyone who handled that part could have told you at a glance that the counterbore was badly off-centre. Rather than going back to remake the parts, rather than figuring out why the parts were bad, they just went through the motions of QC, shipped them anyway, falsified documentation and discarded evidence. For all the complexity of the analysis, the root cause is blindingly simple - flagrant negligence, concealed by flagrant deceit.
Closing comment: damn I thought people paid more attention when building turbines.
More recently I’ve seen pictures of people evacuating down the slides with their luggage! Seems incredibly dangerous, not just for the slide experience but in slowing down evacuation. We had no fire in the cabin but what if we had?
Oh yeah, you know the stereotype of the press sticking their camera in your face to see how freaked out you are? It does happen in real life.
But it is ignored. Which is sad, people could really get hurt.
Your right though the fact as many people comply as they do is kind of incredible given how people act in other situations.
You’re right. What I said above used to be true. That seems to have been questioned in the 90s and in 2000 the FAA finalized a rule changing it.
The current recommendation (https://www.faa.gov/travelers/fly_safe/information) say you can keep your shoes on but to remove high heels, as you said.
A bit of googling says it was changed because of passengers injuring their feet on the terrain/debris after crashes. Additionally modern slides are much tougher than they used to be and won’t tear from shoes and probably even high heels.
But I bet high heels are probably not a smart thing to be wearing on possibly uneven debris covered terrain in an emergency when you need to move fast and safely.
Learn something new every day.
* shoes | boots with sharp objects embedded in soles (glass, bent nails)
* extra spikey high heels,
* work boots with hard edged metal hooks for laces,
(etc) causing damage to both inflatable slipways and to other passengers.
How often has a passenger going down an emergancy slide caused a rip that deflated that slide?
Not very often .. and aircrew are taught to issue instructions that make that as an unlikely occurence as possible.
Case in point when you provide an example with Russia. Other example is Covid.
Fire creates smoke. Smoke quickly makes people unconscious, then kills them. It doesn't matter whether you trust the government or not, whether you follow the rules or not, whether you hold your luggage or not, if you can't get out of smoke, your fate is pretty much set in stone. Putting the blame on people who never had the time in the first place, and couldn't supernaturally turn to liquid and go through the door all at once is like telling that those who were robbed could've trained themselves to run faster.
Although you can argue that Titanic could survive, if only had it blasted bass-boosted <anthem of a proper country>.
People aren't taking their luggage with them during an airplane fire because of their distrust of The Deep State.
You irritation (to the extreme, I mean you just linked the deep state with assholes) is probably due to powerlessness in fixing the issue, let alone interpreting it properly.
However, it is still a widely-acknowledged idea that an unhinged state like the former-USSR-newly-Putin-land makes people afraid of not being able to feed themselves the next day.
I don’t suppose you oppose the idea that when the state organize something, civil servants’ interests are so badly aligned with the proper execution of the plan that it’s a shitshow every single time. And with entirely incompetent people in power, one’s better off doing the opposite of what regularly-proven liars tell them to do.
If you’re upset, is it because you envision lying as a normal form of citizen management? How’s that going for your camp? Do you feel that people trust you?
I'm saying that's not the common case. If you pick 100 random people out there who are disobeying some safety rule, where the outcome is greater convenience for themselves, I'd be willing to bet money that at least 90 of them are doing it out of selfishness and not some principled distrust of the government.
That's just one reason why it's important to listen to the safety briefing, even if you've heard it before. The repeated drill helps us to remember what to do, even when there's added stress.
I built qa software to digitize the forms and signature process like what’s mentioned in the article as having not correctly been signed off on.
I ate lunch with repair engineers that had dark wells of knowledge about the engines they worked on. They could talk so deep on a subject that lunch break was over and we’d resume conversation over weeks.
There’s a paragraph in this post that hits a few points that are very subtle. The missing sign offs and engineers not knowing the process and and and. I think the criticism of RR is valid here. The qa manager at the mro I worked at was a force of nature. He was feared and uncompromising. He was also the signature that could cause an engine shutdown in flight. I admired this person and still do.
There’s small issues like this that go on every day on every engine model all over the world. There’s thousands of engines flying right now that have little defects that could cause a shutdown. There’s issues that have been identified, signed off as low risk and will be checked next time the engine comes in for overhaul.
There’s engineers out there that see the same fault, a premature cracked pipe, carbon buildup, abnormal corrosion, after a while of seeing this problem, they’ll raise the paperwork which will go up the chain and sit. It may be ignored, taken for information for future designs, identified as something that should be fixed or monitored or the frequency of monitoring increased. Maybe the part life will be reduced or you will be forced to NDT the part at each overhaul.
The cheese wheel concept is great as these systems are so complex there’s always going to be some issues.
As for Qantas, near the end it mentions the plane was repaired at great cost. It’s a source of company pride that they’ve never lost an airframe. They repair planes which are BER (beyond economic repair) just to keep this record.
Indeed. Qantas has been ranked the safest airline int he world almost every year since forever [1]
I clearly remember when QF32 happened and everyone was utterly shocked. That simply DOES NOT happen to Qantas.
[1] https://www.forbes.com/sites/laurabegleybloom/2023/01/03/ran...
It is a situation very similar to the downfall of Boeing.
Financial engineers should be banned from operating businesses. They are not focused on the quality of the business, from which profits are derived. They work backwards from their financially engineered results to drive down "costs", even if those "costs" are entirely essential to the operation of the business.
Qantas (and its subsidiary Jetstar) are having to recover their engineering, customer service, and other "costs" to actually achieve the operating business that their expensive tickets require. Currently they are being priced out of operating in Asia, not because they have too expensive operations, but because their board and CxOs were entirely driven by shareholders, not the ongoing operation of the business.
The things I saw left me question how any innovation could happen at all in there or why we did not have a much higher rate of fuck-you-shima per year or how the hell plane engines are not exploding daily.
IIRC the B777 engine controllers are still m68k. Discontinued in 1995.
That seems sensible? You’d need a really compelling reason to rewrite the entire control software and recertify the engine to match. Especially for an engine which has seen no order in 15 years.
What I heard was that there was quite a scramble to buy up all existing supply and also talk some alternate manufacturers into continuing production at a low rate.
B777 was introduced in 1995. Having an engine controller that is obsolete and not available any more at the moment it is launched, seems a bit shortsighted to me.
Then again it works, the planes are flying in the end it's fine.
First, in 1995 Motorola stopped development of the ISA, that says nothing about chip manufacturing which is what RR or airlines would care for. Ti launched the 68k-powered 89 three years later, and only switched away with the N-Spire CAS in 2007. Pilots launched 68k-powered Pilots in 1997.
Second, the early 90s were a time of flux for ISA and you could not necessarily know the plans of your provider, the 68k probably looked quite reasonable when RR started developing the Trents in the mid 80s. RR launched the 777’s 800 in 1991. And even after that, 68ks powered much of the early 80s and early 90s.
That would then have produced major failings in the audit, if not the outright revocation of the quality accreditation, which I would then expect to be followed up on by an audit from the customer (which in the case of TFA would be Rolls Royce), asking some rather uncomfortable questions of the management, examining whether the inter-company concession process was being adhered to, and perhaps reflecting internally (i.e. within RR) - "Do we think these folks are the right people to be making these parts for us?"
From what I've read here it seems to me that Rolls Royce were astonishingly lax in not riding their subcontractors nearly hard enough, quality wise.
Kudos to the Qantas crew on board as well as Captain de Crespigny and his co-pilots and two check captains. We happened to have a lot of experienced pilot power on board.
A video from that time: https://youtu.be/U8Un2boLZD8
- The main one was that I had a flight from Vancouver to Victoria and the weather was too bad for the helicopter to fly. So we took a prop. On takeoff, some cross-wind hit the plane and we tipped over. My colleague and I who were sitting across from each other thought that was it.
- The other one was my plane was reported crashed when I was visiting my parents for some holiday or other. I got panicked call on drive back from airport.
I had a near tip-over coming out of a DIA years ago. DIA gets very windy. We were nearing speed to lift off and a gust of cross wind hit the plane. Looking out the window I thought for sure the wing was going to hit the ground, but in that moment the pilot seemed to shift from a standard take off to something that felt much more vertical. Once we were airborne the flight attendant came by who looked a little shaken and offered me a free drink.
If you are like me, you've probably said “hmm…” to yourself multiple times when certain things were mentioned, because those were things that actually didn't work (that they were left intact really boosts the credibility of the author). From calculation software that had never ever been tested with out-of-ordinary data to the computer keeping the broken engine running. From pure luck with fuel tanks being almost full and unable to explode to absence of any physical kill switch to stop the engine. An hour being generously available to go through ALL the checklists to clear the notifications. An hour of passengers and crew staying on top of the poodle of fuel hoping that nothing would ignite it. Finally, pure randomness in debris flying the way it did. It's not a story of “layers of safety” overlapping, it's a story of “layers of randomness” overlapping.
What would be really interesting is a distribution of outcomes for all possible trajectories of debris, i. e., how (un)lucky they actually were. I guess corporations don't release models like those to the public.
Also, that special chamber for oil filter requiring precise drilling of a perfectly fine pipe seems “ewww” to me. It is not serviceable anyway without reinstalling everything from scratch, as far as I understand, why not make it a single piece?
I do agree though it did not spend enough effort focusing on the areas to improve:
- A computer controlled engine that runs for 60 seconds while on fire, and lets a dangerous part spin too fast. It seems like something that should of been covered ahead of time.
- An engine manufacturing process that is so complex it’s almost impossible to validate.
- A fault management system that only shows you 1 or 2 at a time when you have 40.
As long as the system prioritizes the warnings/cautions with the most pressing ones shown first, this is a very good thing. In a high-stress situation, you don't want the pilots to have to deal with figuring out which of the 40 warnings need to be taken care of first.
Also, as the article says, the pilots did their job following the aviate, navigate, communicate mantra. They first made sure they had the appropriate time to follow the checklists, and only then did they proceed to follow them.
There's over 100 years of aviation experience backing many of these procedures and approaches to dealing with problems. Many are hard-won with literal blood and lives.
Did we read the same article ogurchik? This situation was not simple, and the computer and manuals were as helpful as they could be given the unknown situation.
All I was trying to point out that your assumption about the ECAM system may be ignoring some reasons for why it works that way. No need to be a salty pickle about it.
Certainly there was a fair bit of luck involved as well.
There has been very systematic and deliberate effort to better aviation safety DESPITE commercial pressures.
The swiss cheese means that there are many more layers of randomness that have to line up. Many of those layers came from previous accidents. Those layers are not random at all. Also none of those layers are hole free.
If that disk had disintegrated differently a potentially different set of layers would have applied. Would it have meant fatalities? Possibly. Would it have instantly blown up the plane? We don't know.
But it is pretty obvious that had many of those layers not existed then the chances of a much more disastrous outcome would have been much higher.
[0] https://upload.wikimedia.org/wikipedia/commons/e/ef/Fataliti...
This flight can be seen as an expensive (thrilling, entertaining, newsworthy, etc.) experiment on live subjects whose outcome was not controlled by existing tools and procedures.
The same for everything before to which it is compared so lightheartedly.
Please don't forget that your image shows a giant graveyard.
Have a mice day.
There a whole field of Fault Tree Analysis that looks at how adjacent faults can propagate into unrelated components, then Event Tree Analysis to determine what will happen next. Models that assess robustness against failures even when we have no idea how the failure will occur.
Reliability of cyber physical systems is a constantly evolving field, lots of recent work on concepts like probabilistic model checking, ML for anomaly detection, resistance to cyber attacks, and so on.
The beauty of it is that everyone in aviation seems eager to learn and build on errors. This event prompted new actions that makes future flying even safer, despite having no victims.
Sheer dumb luck was certainly involved. Those discs could have cleaved the plane in half to say nothing of the humans in its way but somehow missed most of the plane entirely. We definitely need to count every single one of those blessings. It's hard not to be positive when such an episode ended with zero fatalities, zero injuries even.
That’s on purpose, you don’t want an automation decide such a drastic move as shutting down an engine. That’s the pilot’s decision.
> absence of any physical kill switch to stop the engine
There is, you shut down the fuel flow with a valve. But that “kill switch” was damaged.
> An hour being generously available to go through ALL the checklists to clear the notifications
Again, pilot decision to do it if time is available. Isn’t it safer that way?
> pure randomness in debris flying the way it did
Well that’s the nature of the failure. It’s like complaining that which HDD fails in a datacenter is random.
> outcomes for all possible trajectories of debris,
Yes it’s not public data, but all positive trajectories are analyzed at the design stage, and structural and systems components are kept segregated accordingly.
It seems to me that fixing one complex problem creates 10 other complex problems. They can be rare, but it's ignorant to shift focus from them.
Kyra has written so many great articles under her nom de cloud. Trust me, just pick any of them and you will learn something.
https://news.ycombinator.com/from?site=admiralcloudberg.medi...
https://en.wikipedia.org/wiki/United_Airlines_Flight_232
>Despite the fatalities, the accident is considered a good example of successful crew resource management. A majority of those aboard survived; experienced test pilots in simulators were unable to reproduce a survivable landing. It has been termed "The Impossible Landing" as it is considered one of the most impressive landings ever performed in the history of aviation
plane lost all hydraulics and had to be steered and crash landed using only the engines
> Rescuers did not identify the debris that was the remains of the cockpit, with the four crew members alive inside, until 35 minutes after the crash.
I can't imagine spending a half hour waiting to be rescued, not knowing whether any of your passengers had survived.
https://en.m.wikipedia.org/wiki/2003_Baghdad_DHL_attempted_s...
If you like this article you’ll also likely like the show Air Disasters too (also known as Air Crash Investigations and Mayday, depending on where you are). It goes into a lot of detail based on crash reports without sensationalizing things too, though not quite as far as this article.
Admiral Cloudberg's articles are more like Columbo: they start with what happened, and then go back in time to find out and explain all the little details that caused it. In a way it's much more logical that way.
Mentour Pilot constantly has to say "remember this, it will prove important later". But we don't know why it's important, and so we don't remember, and as a result the narrative is much less clear.
Amazing article. So well written. Kudos to the Qantas flight team, especially the pilot - they know their stuff for sure. And also kudos to the Airbus engineering team, that was such an epic win for redundant systems.
(It was interesting to see how stopping calculations were improved as part of the post mortem, for one.)
Worth noting that there was an unusual flight crew: 3 captains (one to check the captain's proficiency, and another to check the checker's proficiency) plus the first and second officers.
One of my favourite things about the A380 is that in-flight live feed from the tail, surprised more planes don't do it. Offers visual detail of the entire topside and a lot of information that might not otherwise be available.
Slightly lower, yes, but since, as the GP pointed out, the stall warning sounded just before touchdown, it looks like their calculations already took that into account.
2 datums might need to be very similar but not quite the same, so checking for it might present a lot of hard to handle false positives and make the system very complex.
. Furthermore, initial inspections at the start of the production run were supposed to verify that the manufacturing process was creating products that satisfied the “design intent,” but the initial products were checked against the manufacturing drawings, not the design drawings.
Information flies so fast in the modern world. There is a classic XKCD about learning about an earthquake via Twitter moments before the ground starts shaking.
Hacking overflows in an emergency, topnotch.
Playing devil's advocate: you don't have to. It just has to be better than a pair of experienced airplane pilots working together. Which is still very hard, and there's still a good chance we won't see it in our lifetimes, but at least it's not impossible.
Surely, it depends on the nature of the emergency. As I understand it, in this Qantas example, the pilots did not need to fly the plane with real-time responses, just to make good decisions.
Let's not completely discount remote pilots, while recognising they are not a universal panacea.
A partial mitigation of these issues could be high bandwidth / low latency networks just in take-off / landing corridors?
They may still need help at landing, or by then it could be relatively normal.
But if you can’t provide that level of help everywhere (including over oceans) the design of the system is choosing to lose planes in a trade off for needing fewer human pilots.
Let's do completely discount remote pilots, please.
For AI to replace pilots, you don’t need to prove that sometimes humans fuck up. You need to demonstrate that AI would fuck up less often and in a more acceptable way. This requires looking at the big picture, not only bad cases.
But I reckon single pilot operation with emergency autoland will happen. The tech already exists for general aviation.
> Emergency landing.
> You look like you've been through it.
> The engine... exploded.
> So is he a pass or a fail?
> He's a pass.
Both this captain and the Sullenberger of thMiraclenon the Hudson were Air Force (RAF and USAF respectively). Since, you will be going against an enemy who may damage your aircraft, there is likely more training on how to assess and recover from damage as well as how to handle these types of situations.
However such pilots being very authoritarian or having bad crew resource management and not listening/refusing to let the copilot help has caused a number of accidents (or been a contributing factor) numerous times too.
But so many companies (including mine) still work on more of the “heroic” model where it’s up to individuals to just learn the hard way through helping lots of customers and noticing patterns.
Getting people to understand that SRE is a code of practise and an overall approach has been very difficult, even with the so-called "QA" team, who think their job ends when the latest upgrade is deployed.
We do work in public transport, and the best solution I've found so far, is when they say they're "done", I ask them whether they are willing to stand at the railway station at peak hour and explain to passengers why they can't get home on time (or to work).
The usual result is that they go away and think about it and there is more testing done. But getting that to be a standard approach and way of thinking is very difficult, especially when product owners and project managers are only focussed on the next milestone/payment.
Turbulence is an obvious one. Downdrafts another. You can have a perfectly functional aircraft, but if the whole air column it's in goes down faster than the aircraft can climb, the aircraft will go down with the air column no matter what.
Reminds me of an Air Crash Investigation episode: some volcano had erupted, ash was high up in the air, air traffic control wasn't aware of this, and iirc it didn't show up on weather radar or similar systems (or on the planes' systems).
So it looked all clear. Meanwhile the whole plane was getting ash-blasted. To the point that paint was stripped, cockpit windows went from clear to matte, and ash attached itself to engine fan blades. Obviously trouble followed...
Bottom line: the environment a vehicle moves through, is always a factor. Sometimes an unpredictable, uncontrollable and/or hazardous one.
> From a complete loss of power to all engines to on the ground with zero deaths, zero injuries. That's exactly the kind of story I'm talking about that gives me such faith in flying!
Understood (and agreed). But you missed my point: fate of that flight didn't result from safety engineering. It depended entirely upon the ash-laden air it flew into, and its effect on the aircraft & its engines. No amount of systems redundancy could have made it a safe flight.
So yes: flying is very safe these days. But there are limits to what safety engineering can provide.
But it did! One obvious example being the redundancy that allowed the plane to fly safely despite one of the engines not restarting.
The plane encountered an entirely unpredicted situation that caused damage, but thanks to its design was still able to land safely.
To RetroTechie’s point, they got lucky. No design decision saved them. Without that they’d have been a glider until they hit 0ft and it likely would have been far worse.
We’ve clearly gotten very good at flying, managing most weather conditions we’re likely to fly through, the mechanics/maintenance of the planes, and pilot training.
I’ve gained a ton of appreciation for how detailed our preparations are from watching Air Disasters. But we just can’t control everything, some danger is inherent.
Once I got into our regular flight profile and following our flight plan, I just sat back in my seat and let out a hysterical laughter. I am the calmest person when I fly now :)
https://www.reuters.com/business/aerospace-defense/engine-ma...
Doesn't make me feel comfortable about flying.
Mentour Pilot did a video on that: https://youtu.be/JSMe1wAdMdg?si=YSgbqFpR_EBe-FvX
I'm not afraid of flight. I'm afraid of fall.
(Sounds much better in Greek: δεν φοβάμαι την πτήση, φοβάμαι την πτώση).
Understatement of the century. Somehow those spinning discs completely missed those passengers!! They were in such a state that when they disintegrated they went up and through the wing, towards the ground and one nearly missed the plane itself, striking its ventral section instead of passenger cabins. That near miss took out a huge number of aircraft systems so how catastrophic would the damage have been if they had gone through the plane??
I read the entire article and my conclusion is their luck was immeasurable from the very beginning. They were blessed with such tremendous luck even before the flight crew got the chance to demonstrate their badassery and heroism. One of those discs destroyed part of a building.
That seems like an odd message to include in an emergency action system, which by definition is only active in unexpected situations. Is there really no system to confirm if a fuel leak is happening?
Nope, there is no system to confirm a leak apart from a camera around the tail if you're lucky enough to have one, my previous airline had a flight where an engine leak was detected this way. Think about it, how would you design such a system? So this falls on the crew.
The procedure to determine if you have a leak is pretty much the same across types: add the fuel on board (FOB) to the fuel used (FU) and make sure that the number you get is the same as what you started the flight with. If it's less by some margin then you probably have a leak. You can confirm further by looking at tank quantities (but they take time to reduce depending on the size of the hole). If you get an engine or pylon leak then you might also see increased fuel flow on that engine. If the leak is elsewhere in the system then you might notice a smell. If you can't work it out then the procedure (at least on Airbus types) usually involves turning an engine off to see if the leak stops (yep, really).
As for the ECAM "open fuel transfer valves" message, I don't know for sure on the 380 but all the other Airbus types I've flown have something like:
.IF NO FUEL LEAK
FUEL IMBALANCE....MONITOR
So it doesn't really instruct you to open the transfer valves but leads you into the fuel imbalance procedure if you think you need it. The very first line of the fuel imbalance procedure says something like "Don't apply this procedure if fuel leak is suspected".
At its simplest you measure estimated volume delivered to the engines against estimated volume remaining in the tank. Both are things that should be digitally measurable.
The problem seems to be that the only case it really matters is in a catastrophic accident where such measurements are going to be broken anyways.
E.g. the A330 has an inner tank in each wing (which itself can be split into two compartments if damaged), an outer tank in each wing and fuel in the horizontal stabiliser which is used for CG control in the cruise. All of that plumbing can leak too. You’d be adding significant weight and complexity implementing leak detection across all that.
Regardless of all of this, the aircraft is still fully controllable even with a total asymmetry (one side empty the other full) so balancing the tanks isn’t a massive priority.
The engines have predictable fuel consumption patterns. Even if fuel move across a bunch of tanks, you can still calculate total onboard fuels and detect a leak.
The challenge is knowing where the leak is.
Given that the aircraft can be landed over max landing weight (needs a maintenance inspection) and is still controllable with total imbalance I’d say that balancing just wasn’t as pressing of a concern.
Also, with that much damage you never really know where else it could be leaking. Leaking fuel into critical spaces of the aircraft could be bad so turning on the fuel crossfeed might add extra issues.
A true miracle.
A thoroughly enjoyable write-up this was.
There's actually another failure - (6) - which is the failure to perform visual inspection of the assembling mechanic which should have generated a query or note at the very least. While at this point the mechanic is probably given specific points to check, and an aggressive timeline for assembly, and has little motivation to extend critical thinking beyond the assembly task as they cannot possibly intuit all of the design intent for every part and subassembly, a better run assembly process with a culture of observation would have had this flagged for verification with designers.
It’s very interesting that there were not wall thickness measurements. That would have solved this whole issue.
The failures were more with the whole process (like the reference points with different tolerances and the inadequate paperwork) rather than machinist incompetence. They are just the guys at the bottom.
There was no correcting mechanism and nothing to catch the issue. This is a system failure. The person who set the second datum for a different operation did nothing wrong: a new reference was needed for production reasons and that new reference did not need the same tight tolerance. They were not responsible for what happened later. Then, as the author mentions in the article, the software should have flagged the tolerance mismatch when that datum was used for something else.
> The machinist would have seen with the naked eye very easily that the hole was not close to center, an old salt would have raised it up.
No machinist is ever going to eyeball tolerance violations by a fraction of a millimetre on all the measures of every piece they build. That’s science fiction. Checking should have been in the manufacturing check list. Again, a system failure.
> The machinist not being aware that moving the jig was ruining the setpoint. Incompetence number 3.
Moving it was more or less required for another operation, which is the reason why there were different reference points for seemingly the same thing. Again, that is not the problem. The fundamental problem was that the second datum was used instead of the first.
Besides, if a machinist being aware that a bit of metal must not move at all is required to keep aircrafts flying, the real failure is not to specify that. I don’t know where you live, but in most of the world humans do what they can but are not perfect. That’s why there are checks and procedures to correct mistakes. No single decision or action should result in a crashed aircraft. Otherwise the whole system is just creating death traps.
> There were clearly incompetent individuals working at the facility.
You call them incompetent without seemingly understanding the actual problems, even though they were explained in detail in the article. There will always be out-of-specs pieces and random issues everywhere. If your system depends on humans being perfect, then your system is the problem.
Even great people are bound to make a mistake sometimes. You need to reduce it, sure, and I hope that this specific failure never happens again, but we need to take a broader view.
The article mentioned some of these tubes being rejected, and yet this one made it through.
That said it seems like did have a poor process where a part could be out of spec and they had no good way to check it. As they mentioned about Swiss cheese, you want as many layers as possible, and checks like that are needed.
Why not counterbore the pipe before installation, so it’s a trivial process?
Would it then not survive welding perhaps?
Laserscan it or capacitor probe measure it and encounters the lack of material.
In car part companies there is usually a qm-lab drawing samples from production. Does airplane turbine production not have this step?
> For engineering purposes, disk fragments are assumed to have infinite energy at the moment of release; they will cut through any reasonable material and cannot be contained.