Crisis over The Atlantic: The near crash of Air Transat flight 236
admiralcloudberg.medium.com
admiralcloudberg.medium.com
One of my pet peeves: the first 7 "causes and contributing factors" effectively blames the engine technicians for taking shortcuts. Nowhere does it say anything about the scheduling pressure from management that made the technicians consider shortcuts in the first place. If you have hiring and firing power and you tell someone they need to get the plane back in the air a few days later, they will do things to get the plane back in the air a few days later.
This is considered simply a "risk" by the report, and not a direct cause. As a consequence, the actual actions taken basically all belong to the category of "fix the engineers", "fix the flight crew", or "fix the plane". Not a single "fix the stressful maintenance environment".
[0]: https://en.wikipedia.org/wiki/Human_Factors_Analysis_and_Cla...
https://news.ycombinator.com/item?id=31315234
> There are way too many jobs where management doesn't give the front-line worker an appropriate amount of time and/or resources, and when things inevitably go wrong, the front-line worker is blamed, and the problem is studied as if the only factor was the front-line worker.
On the contrary, the evidence strongly suggests that the maintenance crew did not feel they were under any special pressure: in their minds, they had found the correct solution, one which would allow them to complete the job on time, so any risk of flights being delayed or canceled had been solved.
You do not even have to work in the airline industry to see that unscheduled maintenance risks creating delays, and that is less than optimal for everyone involved - no-one has to be told it. If that knowledge creates too much stress and conflict of interest to be safe, commercial aviation would simply be untenable.
Now as a senior developer it's hard for me see why I was so terrified. Some managers yell and bluster. (Or, they used to, back in the day.) Some managers shame and insinuate. Some managers obliquely threaten. But they all, no matter how they express their authority, have to face up to the unpredictability of project planning. The question is not whether your manager is nice or mean. The question is, do they really think they're better off firing you and replacing you? Whether because they think they can get an upgrade or simply to pin the blame on you for something. After you've been around the block once or twice, you understand that a manager who is chill to your face is ultimately subject to the same motivations as a manager who uses stress as a weapon.
I wonder what it's like for airline mechanics? They are represented by unions. I would think that safety infractions would be the real potential career killers, but that's just me guessing.
Is this really the defense we're using?
Somebody who has been in the industry for a while and seen people fired and laid off has learned that someone's job is no more or less safe with a friendly, empathetic manager. (Your personal stress levels are another question, and an important one, but complying with a toxic manager doesn't decrease your stress levels, so it isn't leverage to make you do or not do anything besides change jobs.)
That's why I mentioned aviation mechanics being unionized, and I wondered if a worker's history of safety infractions follows them. Those would seem to be very important factors. A safety-critical worker should be able to say, "Look boss, you're telling me if I don't break this rule I could lose this job, but 1) the union has my back, and 2) if I do break this rule I might not get my next job."
That's important, because threatening in a loud and obviously abusive way is not the only way to make someone afraid for their job. You could even say that it's the amateur hour way to manipulate an employee. If management wants to scare people into skimping on safety, they can do it without doing anything you could point at specifically as "abusive" behavior, so you need protections that don't rely on identifying obvious misbehavior by management.
EDIT: Edited to point out and clarify, I didn't say anything about the behavior of the manager in my story, and I don't remember the manager being angry or threatening. Actually I can barely remember his face, and I doubt it had anything to do with him. I think it's something that junior people bring with them, just like they're scared of disappointing teachers, scared of their high school principal, etc., in a way that's hard to understand when you're older.
This should be a major point for a maintenance shop. Documentation should always be available, both online and in offline versions.
" The Rolls Royce SBs were also listed in the Trent Illustrated Parts Catalogue, accessible from any computer at the facility, but he was apparently unaware of this, so he instead switched to plan B and called the Air Transat Maintenance Control Center for help."
TechPubs teams often have no control over the compliance (Optional, Mandatory, etc) of the SBs - that, at least in my experience, always came from way up high in the chain, sometimes from a VP who wasn't even tangentially involved in MRO. The required procedures - MX or ACRW - also are often coming from a costs concern, rather than criticality. The whole process of figuring out what's critical and what's not - that's a problem still being smacked into, every day.
It's not hard to make a Sanity Check for pubs vs ERP/PDM/CMMIS, but often people involved REALLY DON'T WANT TO SEE THAT. Because it's horrifying.
The commenter who mentioned HFACS is right on the money. Unless you implement something equivalent, you're gonna be chasing your own tail.
What a great article.
These are the kinds of failures you get when you try to proceduralize everything with foolproof procedures, checklists, run-books, processes, etc. The maintenance people didn't even know they were making a judgement call.
Had the technicians sleeved the offending tubes with appropriate chafing gear and secured them to reduce movement as would be common place in more earthly situations instead of trying to get it "technically in spec" without thinking about the implications of what they were actually doing this problem may very well have been averted. But maintenance people in aviation aren't really ever exercising those creative problem solving skills since it's such a book driven workflow they're always dealing with so of course they missed the forest for the trees (not that they would have been allowed to implement a one-off fix like that).
This is the opposite of what happened. The engineers were confronted with a novel and unexpected problem, and instead of following procedure (which they lacked access to) they applied "creative problem solving skills". From TFA:
> Although it was possible to install the post-SB fuel lines and hydraulic pump with a pre-SB hydraulic tube, the tube would rest against one of the fuel lines at a point where it rounded a 90-degree bend close to the pump. Aware that the plane could not be dispatched unless there was clearance between the tubes, the technicians torqued a nut on the end of the hydraulic tube until it rose approximately 0.635 mm off the face of the fuel line.
Torquing a nut so that the hydraulic tube isn't technically rubbing against the fuel line is pretty creative and seems to solve the problem. But it didn't take into account pressure and vibration while in flight, and almost caused the death of 306 people.
If 0.635mm (an odd number itself which happens to be exactly 1/40th of an inch, or "25 thous" in American sizes) and they let it go that sounds really odd. That's literally the size of a grain of sand. If the concern is rubbing, sure that implies it will move during flight.
> These are the kinds of failures you get when you try to proceduralize everything with foolproof procedures
> Had the technicians
The report makes it clear that the problem was their lack of access to the relevant details, not that the process was defective. If they had seen the need to replace the hydraulic hose, they would have presumably changed it or flagged the problem to someone higher up.
> The maintenance people didn't even know they were making a judgement call.
It seems like they did. If you replace a part and something is touching, you are not trusted in aviation to try and make some space with a bit of fiddling for exactly this reason. They did make a judgment call. If they had relied on the process, again, they could have queried why the new pump was touching when the old one wasn't.
"Read and follow the directions exactly" is a rule that is written in blood.
Aviation safety is utterly dependent on procedures, checklists, run-books and processes. The fact that things sometimes fall through the cracks is no argument against them.
Now if management had for example knowingly procured inferior replacement parts or hired non-certified personnel to work on the aircraft, they would be at fault.
The NTSB report called it a military airbase, and other reports say “Military air traffic controllers guided the aircraft”. Civilian aircraft can use the military side with a permit.
Plenty more info on the air field at https://military-history.fandom.com/wiki/Lajes_Field
Lajes’s other claim to fame is that Bush and Blair met there just days before the Iraq invasion to hammer out whatever deal they thought had to be dealt.
The tale of the stricken aircraft was fascinating. I’ve always wondered, over the decades as I’ve flown slight doglegs over Earth’s oceans, if the strategy to route over remote airfields in case of emergency has ever had a “cost” estimated. How many hours and gallons of fuel has been spent for this safety blanket? The infrastructural insurance spent to prevent lives lost in air travel appears to be much higher than lives on road travel. Could it be that the insurance is for the aircraft, not the lives? Hmm.
The island, and the others of the Azores, are a wonderful place to visit by the way.
Here's where it is: https://goo.gl/maps/fpwxuPzGVDtDEKmB8
and what it looks like when you drive on it: https://www.youtube.com/watch?v=3C5LSrRkFyU
See where the grass divider ends, and gets replaced by a jersey barrier for 2 km? That part is the runway - it would require some hours of preparation before it would be usable, it's not quite "declare mayday and dump the plane there."
The thing with the road in Uruguay is that as you're driving, there's basically a yellow sign with a picture of an airplane and an arrow pointing up, and you're like "what the fuck did that mean?" And then suddenly you understand exactly what it meant because you're driving on a fucking runway...
I googled and found a list on https://www.mil-airfields.de/se/list.htm but IIRC there must be something further northwest, not on the list.
That said, it’s not the first road doubling up as a landing strip I’ve found myself on - there’s a stretch of the main road between Donetsk and Mariupol that was, in soviet times, used as an emergency field for strategic bombers. It’s this relatively quiet provincial highway that suddenly turns ridiculously wide, and goes dead straight and level for what feels like an eternity - 11km looking at google.
Oh, and RAF manston used to have a level crossing across a taxiway.
https://www.google.com/maps/@13.7720949,-89.3716038,3a,40y,1...
https://goo.gl/maps/5JE2YpEh75V52coR6
And Runway 21 (the same runway in the opposite direction):
This is accomplished by having air strips which can take massive planes like A350s at remote places which can’t justify planes of that size normally like Newfoundland, Greenland and the Azores
While these are mainly there for military and temporary emergency use, in 2001 they were used far more - with dozens of planes landing at tiny towns like Gander and St. John’s in Newfoundland, places which only really expect a single large plane to land and refuel rather than 30 with passengers to accommodate for days.
A modern 777 at most airlines for example has an ETOPS rating of 180 minutes.
That's 180 minutes of flying with a failed engine. So you have to take into account whether you can maintain altitude on 1 engine. Most of the time you can't, you'll be flying lower. Which means slower speed, so the calculation is 180 minutes at that speed.
I enjoyed the history of ETOPS-138 [0]. ETOPS-120 leaves some small triangles of the Atlantic inaccessible, particularly if one emergency airfield is unavailable. A 15% increase to 138 minutes covers it.
Anecdotally, though, my experience as a commercial passenger feels like airlines fly almost directly over pacific islands rather than within X minutes of diversion. My memory may be anchored in bygone years.
[0] https://aviation.stackexchange.com/questions/30979/why-etops...
> pacific islands
Depending on the jet streams, the flight can take the great circle route [0] or taking advantage of the jet stream [1] and flying more southernly
[0] https://flightaware.com/live/flight/JAL6/history/20230108/02... [1] https://flightaware.com/live/flight/JAL6/history/20230105/02...
The Japan to Europe flights are even worse. The great circle route cuts west, straight across Russia [0]. These days, they fly east to the Bering Strait then more or less directly over the north pole [1] (though probably easier to see on great-circle: [2])
[0] http://www.gcmap.com/mapui?P=HND-FRA
[1] https://flightaware.com/live/flight/DLH717/history/20230109/...
Back in the Soviet era when USSR airspace was closed to most Western flights, planes couldn't do that route non stop from Europe, so had to stop over in Anchorage for refueling.
http://www.gcmap.com/mapui?P=BINP%2CFRA-KUTAL%0D%0A&MS=wls&D...
Need to find an excuse to go to Japan
Playing around, a theoretical NUE (Nuremberg) to KUTAL flight comes within 1 mile of the North Pole
Well into the 24 hour polar night/day period though (depending on time of year)
I would be surprised if the cost of equipment and training pilots and whatelse was not factored in.
But the major driver might be that you have to pay for more expensive insurance to get regular people to strap themselves into something going many miles per hour many feet above the ground. Not because it's inherently less safe (though I would argue it is) but because the prospect is scarier!
If you haven't looked into Gibraltar airport before, it's a fascinating thing - the land it's on is reclaimed using rock dug out from the siege tunnels, there's a level crossing on it as the only route into the town crosses the runway; if a plane is coming into land, or taking off, you have to wait at the barriers. It's also the city airport closest to the city centre it serves at 500m.
What us failure analysis nerds need, what we deserve, is a 90s era Discovery Channel show written by Cloudberg. Either that or a a United States Chemical Safety Board[1] style series.
What I enjoy most is that frequently, the crash appears to have a clear “culprit” early on, then cloud berg points out 5-7 other factors, any one of which might have prevented the accident. If the series had a theme, I think it would be that airplane crashes are less about human failings than systemic ones.
Because if your ass is being repeatedly saved by the equivalent of a $0.50 zip tie, you probably want to know that.
You know the fuel level in each tank, and have a good estimate or actual flow data on engine fuel usage. Simple math says "leak" or "no leak".
If that would work, it is surprising it isn't in the software.
In monitoring IT systems seeing metrics like this over time is super valuable.
Also, like most engineers, you don't necessarily add systems for every conceivable thing that could go wrong. If you think about how many times this has happened against the millions of plane-miles travelled, you can see why it wouldn't necessarily have been high on the list.
If the number does not reach within expected margins, then show an error.
This protects against leaks, but also miscalculation of fuel needed by the pilots, misfueling, or efficiency loss somehow.
Perhaps the big innovation needed is accurate fuel quantity measurement.
Any sensor has a failure rate. If the probability that a sensor has failed isn't dramatically lower than the probability of a leak, then the pilots will do just what they did in the incident and assume a bad fuel sensor reading rather than a leak.
Another failure mode of fuel tank sensors is short-circuiting and blowing up the plane: https://en.wikipedia.org/wiki/TWA_Flight_800
It's not as simple as "fuel per time" because fuel usage and speed change massively with altitude.
What was missing here was a big warning to the pilots (that has since been added). But it's also standard procedure at all airlines to monitor fuel and required fuel calculations, which could have helped this crew if they did it earlier.
> Meanwhile, Airbus and the French Directorate General of Civil Aviation worked together to produce a recommended service bulletin modifying the Flight Warning Computers on A330 and A340 aircraft, allowing them to warn of possible fuel leaks by continuously comparing the planned fuel with the actual fuel on board.
If a human can do it, I would imagine a computer could automate the process. Likely wouldn't remove the need for the pilots to do it anyway, though; redundancy and reliability and all that.
FTFY.
It's a common misconception in the general population that the Captain always flies the plane and the FO supports them.
There's always a pilot flying, and pilot monitoring. The roles are usually decided before the flight by the Captain and have nothing to do with who is Captain and who is First Officer.
On the driver side, my immediate thought was: How did a fuel leak in one engine cause a total loss of fuel? I would hope that there are now checklists that prioritise shutting off the cross-feed valve if a leak is suspected. Weight and balance are subordinate to the risk of running out of fuel and starving both engines.
On a side note, we would have regular "Safety Days" where incidents of ours, and major incidents like these that have happened to other aircraft were presented to the squadron (both aircrew and maintenance) in detail and discussed. Having the potential consequences in your mind at all times really helped fortify the necessity of following the correct procedures during the regular conduct of your work.
If something goes wrong, they are the ones managing the checklist.
In any case, the automation handles the vast majority of flight engineer duties. And fuel leak warnings have been added now.
While this is in no way scientific, it seems that what most catastrophes have in common are bad CRM (crew intra-communication) and lower-than-average airmanship. In most cases, for a catastrophe to occur there must be an initial problem, sometimes very small, that is made worse and worse by bad communication and a series of misunderstandings. And ultimately, in some cases, pilots fail to fly the plane.
For poor airmanship, AF 447 comes to mind (very minor initial incident, plane working perfectly fine, disaster caused by incompetence), as well as a comparable problem on AirAsia 8501 (the captain decided to reboot the computers in flight (!) and the first officer was unable to fly the plane properly without autopilot). Disasters caused by poor CRM include the Tenerife Airport disaster (deadliest accident in aviation history).
But I find incidents that don't result in a catastrophe often more interesting. Incidents and problems are impossible to avoid when using machines as complex as modern airplanes; what's fascinating is how humans, working together and using all of their skill, are able to save the day.
This is the case here. The problem was caused by the maintenance team who didn't use the proper hydraulic tube when replacing an engine, because they didn't have immediate access to the text of the service bulletin that described the procedure, and they thought they could do without. (The wrong hydraulic tube didn't have the proper clearance; when under pressure, it wore away at the fuel line beneath it, until it cracked and leaked.) The actions of the pilots in flight compounded the problem but this was completely excusable given the information and training they had, and the checklists as they were written. Ultimately, disaster was averted by excellent airmanship.
There were other events like that. The most famous one is of course US Airways 1549 ("miracle on the Hudson"), but a less well-known, very similar, and possibly more impressive one, happened in 1988. On TACA Flight 110, the pilot landed the plane on a patch of grass with no engines, with no injuries to the passengers or damage to the plane.
That pilot, Carlos Dardano, had only one eye, having been shot in the face while landing in El Salvador a few years prior and being caught in a cross fire. He had started flying when he was a toddler, on the knees of his father. At the time of the incident, he was 29 but he had already amassed 13,410 flight hours, with almost 11,000 of these as pilot in command.
Another amazing story is the Air Astana flight 1388, in 2018. A heavy maintenance operation had resulted in the inversion of the aileron cables, but not the spoilers, making the plane utterly impossible to control. This problem escaped detection until the plane was in flight. The pilots, after calmly discussing how to ditch the plane in the sea to minimize the number of victims on the ground (it was a test flight with only 6 people on board, all pilots or engineers), managed to regain a semblance of control and landed safely, after 90 minutes of the wildest roller coaster imaginable. It looks like Kazakhs can fly.
It seems the modern world doesn't hold skills in high esteem (what used to be called tradecraft before the word became associated with espionnage); the MBA culture essentially regards workers as interchangeable; and of course we have all these machines!
But the opposite is true. What will save your life is not the machine, it's the experience, professionalism and competence of the people using the machine.
For example https://en.wikipedia.org/wiki/Emirates_Flight_521 - a go around was commanded by the pilots to abort a landing, but they weren't aware that the automated go around system - specifically the auto-throttles - would not engage if the wheels had touched the ground (if only briefly). Thus engine power remained at idle/landing speed, and the plane fell out of the sky.
The pilots should have checked the throttles of course, but they would not necessarily have been aware that the wheels had touched down, may not have known/remembered that the automatic go around system would not function in that case, and they were concentrating on flying the plane in presumably unusual circumstances that led to the go around in the first place - they assumed the automation would function as it usually does.
A better human-machine interface (e.g. with a warning that go around automation was disabled) would have prevented the crash. I think anyone designing human-machine interfaces, especially safety critical systems, should read up on these cases.
https://admiralcloudberg.medium.com/a-mathematical-miracle-t...
This story is more ambiguous. The plane should never have left the ground and the captain committed a serious error when he decided to depart with no fuel gauge working. According to airline procedures at the time, no plane could be dispatched without working fuel gauges. That wasn't up for debate, it was a hard rule and it was disregarded.
Then, contrary to legend, nobody really mixed up imperial and metric units, but they did the math wrong (twice!) when converting from one to the other. They weren't surprised when they found that 8,000 liters of fuel would weight 14,000 kg. One liter of water weights 1 kg; fuel is less dense than water; without knowing anything else, one should question a calculation that results in the weight of 8,000 liters of fuel being more than 8 metric tons.
So yes, ultimately, the plane landed safely thanks to outstanding airmanship; but it was improper following of guidelines, and math incompetence, that caused the problem in the first place.
Also known as the Swiss cheese model[1].
I wonder how many (significantly) lower than average airmanship days happen everyday without incident. I suspect it’s a lot, and the overall system catches/corrects many and is resilient to the effects of many others.
I consider the lack of respect for experience and skill one of the biggest failures of western culture. It also applies to things like mathematics where nowadays some people are even proud about their lack of skills. But the end result is that most people are amateurs in their own job.
1) Design error. Ok I am a total nobody in this area but still fuel leak detection not being included seems like a major design fault. I mean something that measures difference between whatever fuel injectors deliver into the engine and whatever leaves the tank is quite possible I think.
2) How many times have we read when a fucking management pushed for event happen / to be on time no matter what and the following disaster. Challenger, Columbia etc. etc.
3) Service procedures, manuals etc. I guess those are combinations of (1) and (2)
4) Errors by service and flight crews. I can imagine that the person who bent the line had enough reason to raise concern. The rest can't really be blamed.
As for Robert Piche - he is a hero in my book
I've often found that means his reports conflict in various ways regarding either what happened, or what caused what happened, with what is found in so many other reconstructions, many of which seem to goes as deep as a Wikipedia article and not a whole lot further.
I find he provides a ton of insight then...
There's only a few YouTubers I have time for like ones that don't hide complexity, such as Dave Jones from EEVBlog.
A case in point would be this article, as it is clearly written, balanced and very digestible.
YouTube’s algorithm is pretty good at doing this for you!
"Oh, you're watching a video on foo? Here's a list of video clickbait trash, with sometimes one more video on foo thrown in.
"Oh, you either never open or always bail out of 'shorts'? Here are more of them!"
The "Shorts" section can be hidden pretty easily.
For me, YouTube hasn't been able to suggest properly for years; ages ago, suggestions were all clearly related to the current video in some way. The old behaviour lead me to spend lots of time on YouTube.
When "Shorts" started appearing on the main page, it somehow seemed like I couldn't get rid of it. I somehow don't want to see it ever, on any computer so it being default is as evil to me as blinking ads.
The addition of forced-autoplay-unless-you-sign-in[1] is definitely a case of Making It Suck[2].
So, I increasingly use other clients, or exclude YouTube results from searches. This pretty well gets me away from clickbait rubbish like "$profession doesn't want you to know this!!", grossly exagerated faces/postures, and Shorts.
[1] Welll… actually I guess it's "default to autoplay whether signed in or not, and keep resetting it if you're not signed in".
[2] Meaning, moving people towards what another behaviour by making something else more difficult.
I also get annoyed by the Shorts section (they're almost pure clickbait/teasers, addictive but no real content, waste of my time!). But I can just click the "X" and it disappears for 30 days. If they didn't have the "close for 30 days" option, I agree, this would really suck.
[0]: https://aircraft.airbus.com/en/services/enhance/skywise/skyw...