The Reliability Trap: The crash of Emirates flight 521 (2020)
admiralcloudberg.medium.com
admiralcloudberg.medium.com
The article isn't attempting to make a point here, but for anyone who is unaware, it is this way on purpose. It's not an accident or an oversight or a bad policy.
It used to be that pilots had to justify doing a go-around. So in that critical moment, when the pilot is making the go-no-go decision, they're thinking about whether or not they can justify the go-around, not about whether or not the go-around is the right thing to do. There was a tendency for pilots to try to force a landing that just wasn't gonna happen and it resulted in more than a few disasters. So now the policy is to just do the right thing and the pilot doesn't have to explain, not to tower, not to management, not to anyone why they decided to go-around.
Yes...
> In both cases, the pilots assumed that the autothrottle would increase thrust during a go-around, but were unaware that they had run up against an edge case where it would not.
Also yes, but...
> The common denominator between the two crashes was an overreliance on automation.
That doesn't follow. The pilots were absolutely right to assume that pressing the go-around button would indeed perform a go-around. The core issue is that the button silently failed. It's like a function that doesn't throw an exception despite failing and instead chugs along in a broken state.
It's a simple case of bad UI/UX. If a button that's supposed to do X is unable to do it, the failure should be indicated to the pilots in an obvious way.
The recommendations say as much: "The GCAA also issued recommendations intended to make the Boeing 777 a safer aircraft, including that the configuration alarm go off when the throttles are not advanced during a go-around;".
Take-off/go-around is a complicated maneuver involving quite a few tasks that need to be orchestrated to preform it correctly, in short you need to quickly transition the aircraft from it's landing configuration to it's take-off configuration. an ideal situation for automation. However according to the article.
"However, the TOGA switches are inhibited on landing after touchdown, because if a pilot were to accidentally press it during rollout, it could cause the plane to run off the runway. The autothrottle system remains active, but if sensors detect that there is weight on the wheels, the TOGA switches simply won’t do anything."
Just show a bright red warning or a voice alert when the inhibition activates.
So we look to find what could have been done differently, There are many small things that perhaps could have prevented this, basicly this is the article. But one of those things was to notice that the huge lever in the center console was not moving. and their hand was right there to pull the togo switches.
The engines are also very far from the cockpit, and behind them. They hear them, but not well.
So it’s not that surprising they didn’t immediately notice the problem.
That they did notice, and relatively quickly, is a sign of how well they did. But seconds matter in this scenario, and by then the lag to spin up resulted in no significant change of thrust before impact.
If a pilot doesn't pay attention to where his thrust levers are when appropriate thrust is critical to safe flight, that is pilot error.
A go-around isn't complicated enough to merit reliance on automation anyway: Max thrust (or whatever is appropriate), pitch up and climb according to ATC orders or go-around procedure at airport, landing gears up, and come back for another landing attempt.
In this scenario, the auto throttle doesn’t auto throttle up because unexpectedly, they were partially wheels down for a moment, and that disables auto throttle up.
Oops.
On multi-engine jets, common training is to remove your hand from the throttle at V1 (takeoff decision speed). Most problems after that speed are to be taken airborne and dealt with there.
Definitely one of those ‘if the PIC had been a little more paranoid, and a bunch of other rare/weird things had not happened, it would have been fine’ type accidents. The first officer was supposed to also verify things, but the specific step wasn’t in the training either, and wasn’t normally applicable because of the TO/GA automation.
Luckily no loss of life from the crash directly. If the firefighters had listened to prior crash issues and fixed them in their own response, likely no firefighters would have died either.
(In fact, this idea of gravity-relative-to-butt is very strong in people and when that fails to be a useful model is often when we see accidents happen.)
It's the same in a car, if you are coasting in neutral and encounter a hill, even though the car will be slowing down there is no "perceptible" longitudinal acceleration because it's caused by gravity (equally acting on the car and the occupant), just like in space or one of those zero-G flights.
There would be an increase in drag when pitching up, that would also cause deceleration.
So it's both correct to recommend that the Boeing 777 try to alert the pilots to the misconfiguration, and to reinforce pilot behavior to reduce the risk of being in a position where that misconfiguration happens.
even basic consumer devices bark at you when a button press is invalid, and the people who work at boeing are obviously not idiots.
my guess is that there are a series of guidelines, solid reasoning or standardizations that lead to this nonsensical result and that decision-making and design process itself would be the interesting thing to understand in what went wrong here...
well... knowing what we know of MCAS, and their whole approach to user interaction ... that bar is very shaky.
anyway, the answer seems to be that the button has a clear feedback in the form of physical thrust set-point indicators. basically the pilot (if I remember the post correctly, the first officer) should have noticed that despite pressing the button the set-point indicators did not move.
i suspect the story behind most decisions in flight control ui is a very long one, with a long list of lessons learned, expectations set and human factors. these things arise in any complicated design, where some new feature violates the existing design principles and a compromise is made that on its face makes no sense, but in context is totally understandable.
that context would be interesting.
As far as I know - please correct me if write this wall of text based on inadequate information - the problem with MCAS was that they wanted to hide it, but that's not necessarily bad UX. what's bad is that they obviously failed in two distinct, but connected and together critical steps/aspects.
1) MCAS did its thing in 10 sec bursts. which was just crazy, never before seen madness. (I agree local design context is important, and I would like to know WTF was the context for this.)
2) it was undocumented because Boeing argued that it was technically just a runaway stabilizer failure, already covered by the manual/training
I would argue that the 1st failure was the (more) fundamental one, the UX one. because if the fucking thing looks like what a typical runaway stabilizer malfunction looks like, then yeah, it's "okay", pilots should be able to recognize/remember the correct remediation method for it.
...
of course it's debatable how close the new airframe + engines (without MCAS) were to the old one in terms of flight characteristics. it's possible that it really flies like an old one with too much weight in the back, and that's routine for pilots.
...
and just to be clear, I think it's just inexcusably dumb to try to cheap out on proper training.
This seems related to the deep belief in risk compensation and risk homeostasis, despite weak evidence for the former and nearly none for the latter. We can be cynical about technology all we want, but to actually improve that needs to be driven by evidence.
https://admiralcloudberg.medium.com/blind-to-the-problem-the...
https://admiralcloudberg.medium.com/the-long-way-down-the-cr...
Automation is convenient, and in many cases (but not all!) it can making flying safer. But an aircraft with an idiot at the controls is still an unsafe aircraft, regardless how good the automation is.
MCAS was bad because Boeing and the airlines wanted to retrofit without additional training. It wasn't the automation that was bad per se.
1. The automation is designed and implemented in a reliable, known, safe way.
2. The pilot are competent at flying the aircraft completely manually if necessary.
If either of those two assumptions are violated, the flight is not safe.
Automation cannot replace incompetent pilots, no matter how good. Pilots can't overcome bad automation if it is not designed and implemented properly (eg: MCAS), no matter how competent.
This accident was caused by a violation of #3. I've no doubt that the pilots in question were completely capable of performing these maneuvers manually - they failed at identifying which particular mixture of automation and manual control they were in, and in taking the correct steps during them.
And because it worked inconsistently due to terrible implementation.
The implementation was extremely terrible (no redundancy on the inputs), and it was Boeing that decided to hide its existence and overrides from everyone. Automation isn't bad, but when implemented with criminal negligence, it can be.
Covers advice on when to go from maximum automation towards more manual control of the aircraft when the cockpit gets busy.
This is why self driving cars of today, that claim that humans should take over in case of failure are nonsense. We have a huge body of knowledge showing that when faced with a system that almost always works, humans suck at overseeing it and taking over when needed.
The point of the article isn't that the automation just needs to be made more reliable and that will solve the problem, it's that systems like this are becoming too complicated to understand all aspects about how they work.
You would have thought overriding the TOGA switches should have sounded an immediate alarm though. It seems perfectly reasonable that a very dangerous command or input would be inhibited by the automatic systems, but it seems completely crazy that any command or input would ever be inhibited silently. I'm actually flabbergasted that this is not a fundamental rule of aircraft control systems design at Boeing.
On Boeing planes, many automations such as TOGO typically have a corresponding physical counterpart - in this case, autothrottle slides the throttle controls upwards to indicate that the engines were asked to go to full power. There’s also the engine instruments which show the engines reaction to the requested throttle level.
It’s not hard to know that it didn’t set the throttle to full power, but the pilots didn’t check or didn’t know to check. An alarm would help, but would also be another fallible automation to be relied upon.
I still think it's crazy that there aren't alarms for any situation where the airplane disregards an input like this.
Idea: Add a few "Robocide" buttons to the cockpit. If pressed, they deliver a figurative bullet to the brains of the autopilot, dropping the plane into a far simpler "full manual flying mode". Pilots regularly train in doing that, and flying the plane when suddenly dropped to manual.
I wish...
In fact this is how AF447 crashed... a plane under direct control of pilots...
Read: https://www.vanityfair.com/news/business/2014/10/air-france-...
Also, when they went very far into stall, the stall warning disabled under the assumption it was a sensor problem rather than the airplane actually getting into that bad of a stall. So, paradoxically, starting to improve the situation caused the stall warning to come on, while doing the wrong thing caused the warning to go away.
Both pilots were trying to debug with information intentionally being hidden from them. (Honestly, the stick should vibrate or something if your inputs are being ignored or counter-commanded by the other pilot. Better yet, mechanically link the controls as in Boeing planes... it's not great that the stronger pilot wins, but at least both know what's going on.) Granted, there was a lot of pilot error in AF447, but there were multiple user interface issues that greatly contributed to the problem.
Edit: Also, as I remember, the start of the incident was that the pitot tubes iced up and the autopilot disengaged itself because it had no idea what to do. It's hard to point to a case where the automation has explicitly given up as a case where we should rely on more automation. Clearly a world where the automation was better would have been better, but just letting the autopilot do its thing wasn't an option. The autopilot disengaged itself.
Loss of pitot tubes only implies loss of air velocity indicator. The attitude indicator, the altimeter and the thrust levers worked fine. I still can't believe they intentionally reduced thrust and pitched up for an extended amount of time, and not only thought this was a good idea, but didn't crosscheck the most fundamental instruments to confirm that the plane was flying level. And then proceeded to ignore stall warnings and stick shaker.
The human error was so severe that I'm not convinced that better alarms could substantially mitigate. At the end of the day, if a pilot forgets how to fly a plane, then their peer (presumably still capable of independent thought) needs to have the presence of mind to take over.
Not necessarily. They are mechanically linked but there is a breakaway mechanisms designed to be failsafe against one of the sticks jamming. You could have a scenario of them fighting each other so hard the breakaway trips and then only one of them has a working stick, at least until both sticks are aligned, allowing the clutches to engage again.
(From Vanity Fair, it sounds like the Air France pilot's union is both extremely powerful, and profoundly hostile to the idea that pilots must actually be competent. Also like opaque layers of both poorly-understood automation and infernally-clever instruments repeatedly got in the way of the pilots doing plausible things to "manually" recover.)
See https://en.m.wikipedia.org/wiki/Qantas_Flight_72 for a flight computer failure. The Mayday episode lists states some time after that accident, another Qantas flight had the same failure, but the crew knew of the potential issue, so they cut power to all 3 flight computers.
Other similarities are poor and incomplete documentation and/or withholding of crucial technical information.
Airplanes have never been safer, despite many more planes flying. Crashes of airliners is rare. If anything, that points to automation greatly improving safety.
Thing of it this way, many more pilots accidentally hit the TOGA button on the ground. The automation has surely prevented more accidents than it has caused.
To be fair, that's what Boeing's MCAS was also trying to do. It's just that Airbus aren't incompetent and/or criminally negligent, and don't hide the existence of such automations.
The issue was they decided to use automation to avoid training pilots on how the plane actually behaves. That kills people.
Far worse than the occasional silent failure are too many false alarms. The medical profession has this problem.
If you get too many alerts, people ignore them. There was a series of incidents where pilots were routinely pulling circuit breakers to silence takeoff config warnings going off while taxiing. Eventually a plane took off without flaps and a lot of people were killed.
Automation is very complex and the solutions aren’t always so straight forward. People forget there is a serious risk of adding an alarm causing more accidents.
Should there be an alarm here? Seems obvious there should be and there better be a damned good reason there isn’t.
But yes, this stuff is very complex and I'm (we're?) only judging this from the outside.
unless it has changed in the last few months, or if I somehow missed it
Many of the criticisms in the comments below that I've seen are deeply uncharitable in a way that's so typical of this site, where man y people post dense opinions with all kinds of viewpoints that mis-attribute blame and say mistaken things for all sorts of reasons on numerous subjects.
How ridiculously typical of the self-congratulatory audience on HN. Have none of you ever written code that has a flaw or two?
The engines were at idle during an extended float; so it could take longer to spool up. The crew knew this, but that's five really important seconds.
Definitely should have waited for the kick in the pants before raising the gear.
In a heavy aircraft, things happen slower. The airspeed will decay slowly after raising the nose, but you will see some initial climb until gravity has time to assert itself. Increasing headwind will introduce transient increased airspeed masking the decaying energy situation.
I wonder why spoilers were not used, but don't know how that fits in with 777 procedures.
This is it. Humans can understand any system, if you have pilots failing to understand they weren’t trained properly. If you can’t train your pilots properly, use different planes. Nearly every plane crash for decades has been a human failure not a technological one. The way airliners behave is well known and not a mystery at all, if pilots don’t know what to expect its their fault.
Designing aircraft for human training needs is indeed important and constantly under review, but I cannot really blame aircraft for the failures of airline training programs.
a nice airframe and a complete shitshow of software and interaction with the machine ensemble
I theorize that since software is so invisible/intangible, these complexities are not understood by many.