737 Max: 1960s Design, 1990s Computing Power and Paper Manuals
nytimes.com
nytimes.com
There is a large degree of coupling between avionics hardware and software amd updating that software is prohibitively expensive. There's just no point in doing it unless you get a tangible benefit.
Furthermore, avionics systems take a long time to develop are not often rebuilt from scratch, so the hardware naturally lags far behind the computing power we're used to.
The recycling of avionics systems is definitely fueled by a desire to reduce costs. However, it's also important to note that older, mature systems with a history of service inherently offer a degree of safety compared to something which is untested.
The issue this article complains about can easily be seen as a safety feature.
To be fair, the author does spend most of the article talking about missing features and design issues rather than harping on "processing power". It is a poor title.
To be fair the article doesn't critique using 1990s (or earlier) hardware/software. It's merely a clickbait headline. The only reference is dismissed by the very next sentence as insignificant [Boeing could be replaced with any ~~expert~~ somewhat-knowledgable person on the subject]:
> The flight-control computers have roughly the processing power of 1990s home computers. A Boeing spokesman said the aircraft was designed with an appropriate level of technology to ensure safety.
This only adds to my lowered standards when approaching NYT articles over the years. They've long been abandoning being correct for being entertaining at a seemingly accelerating rate. Which is too bad considering plenty of people continue to hold them to a higher credibility over the typical clickbaiter sites. Considering they are a subscription service I'm not convinced this is justified.
I actually edited in an extra paragraph at the end of my original post because I wanted to at least credit the author for talking more about design issues than "processing power" in the body of the article. The title is very bad, though.
We don't need to pay for content. The world is full of people who will write for nothing more than getting their name out there. What we need are good editors. The news media are firing the very people who could get them out of this mess they ended up in.
What we used to have that we don't have any more is a forced-delay, quality-control, gatekeeping function. This was fulfilled by the slowness of manufacturing printed content and good editors. Turns out that this is a place where cutting out the middleman, the wisdom of crowds, and automation worked against our best interests.
What are the correct numbers?
They've also hired younger, cheaper writers.
As a result, the overall quality of the product has fallen, noticeable.
[1] https://www.google.com/amp/s/deadline.com/2017/06/new-york-t...
It was designed in the 1960s using technology they were completely confident would work and keep working - so it was basically 1940s electro-mechanical technology: relays, worm gears etc.
I don't think they really trusted electronics at that point, let alone computing devices for something so super critical.
No idea how that is done now though!
The "1990s home computers" power was simply an illustration of Boeing just trying to do the minimum work to move forward. Other examples that, while not bad per se, seem to indicate shortcuts:
> If one of two sensors malfunctioned, the system could struggle to know which was right. Airbus addressed this potential problem on some of its planes by installing three or more such sensors.
> Most new Boeing jets have electronic systems that take pilots through their preflight checklists, ensuring they don’t skip a step and potentially miss a malfunctioning part. On the Max, pilots still complete those checklists manually in a book.
> A second electronic system found on other Boeing jets also alerts pilots to unusual or hazardous situations during flight and lays out recommended steps to resolve them. On 737s, a light typically indicates the problem and pilots have to flip through their paper manuals to find next steps. [...] Boeing decided against adding it to the Max because it could have prompted regulators to require new pilot training, according to two former Boeing employees involved in the decision.
All of these things still get the job done, but they seem to indicate that Boeing simply wanted to do one last turn of the crank with the 737 design to compete with the A320neo, and may have rushed thing too much.
The only thing a more modern processor buys you is more memory, a faster tick rate for your RTOS, potentially an MMU if it doesn't cause substantial timing skew, and high-precision floating point. In the case of an avionics system I would imagine the requirements for fixed-point arithmetic are very well-understood at this point, so floating point doesn't buy you a whole lot and potentially creates new and exciting problems.
So I would say that 90's-era computers are plenty powerful for what they're used for.
I would not be surprised if they still use 68k in the A320 NEO.
Between specification, implementation, validation and testing, the cycles in the aircraft industry can be quite long. 5 to 10 years usually. And it's also quite costly to validate a brand new system. That being said, given the safety requirement, it's for the better.
For example if a UI display showed:
"Warning: MCAS overriding pitch up. Disable?"
Is better than a LED #234 and LED 412 are on. Consult manual for instructions.
I would imagine that a lot of the fly by wire or control systems nowadays are either analog systems or very very simple digital systems (with extreme redundancy).
Even if they could have changed the UI they still needed to hide the existence of the MCAS system, again to avoid retraining, so it's doubtful such a clear an helpful error message would be allowable.
I think the way you criticize the article is about as faulty as the article itself.
This piece is about detailing how the 737 is an outdated design that got patched and patched till it more or less broke. The processing power is just /one/ aspect and is given as a point of reference, like the 1968 photo.
I am no expert in aviation, but I guess it's quite a safe guess that fly-by-wire systems need more processing power than mechanical/analog designs. And that a glass cockpit with a trouble database needs more than a blinking LED and a paper handbook.
Nowhere in the article does the author claim the 737 MAX crashed because it had too weak CPUs. It is, however, one indication of an obsolete design - newer jets like the A320 or the 787 have more beefy processors for a reason.
I know I worked in 1-dimensional transportation and their 486-based onboard-processors felt cramped, the software engineers had to hack around their limitations and they wished they could switch to something newer. I doubt 3-dimensional transportation is easier to implement.
I address that in my post and I do not fault the author for pointing out design issues. My issue is with the implication that processing power is an issue.
> Nowhere in the article does the author claim the 737 MAX crashed because it had too weak CPUs. It is, however, one indication of an obsolete design - newer jets like the A320 or the 787 have more beefy processors for a reason.
Newer jets have faster processors because they're built on a newer system. If you're developing a new system, you might as well throw in some faster hardware. The cost of the hardware itself is nothing compared to the overhead development costs for avionics.
From my experience, this is not how the development of transportation systems work. They throw in faster hardware if they need it, as it is insanely expensive to do so. And I still believe glass cockpits and fly-by-wire need more computing power.
Any sane person would prefer the solid well-proven choice over the bleeding edge. People used to make similar comments about the AGC and the Space Shuttle GPCs e.g. that such-and-such-a-PC-was-faster. Yeah, so what?
There exist possibilities for far more safety with more compute.
For example, when a mechanical failure occurs (for example an engine explodes and partially falls off) the flight characteristics of the plane change dramatically. Current systems fall back to human control, and hope that human will be able to figure out how to control a now very "unique" machine.
Future systems could dynamically create new flight models based on collected data to be able to fly the machine in entirely new ways.
For an example of this, did you know it's possible to fully control a quadcopter even with 3 out of 4 rotors broken and no working fins/flaps? [1]. That sort of 'fly it how it has never been done before' is out of reach of human control, and might save lives.
[1]:. https://www.ethz.ch/en/news-and-events/eth-news/news/2013/12...
You randomly deform the airframe in a physics simulator, then get a real pilot to try to fly the simulated damaged airframe to the nearest airport.
You then get an AI to try doing the same.
As soon as the AI manages to get to the airport more often than the human, the tech will save lives, on average.
My guess is a PhD student could come up with an AI passing the above test inside a year. Yet aviation standards are stringent enough AI will be decades away from production use, if ever.
There's some weird psychology behind this: human pilots/drivers/etc are seen by passengers as a personal proxy, with agency over any situation.
If you take away the proxy you take away the illusion of agency. Humans really do not like being put into situations where they believe they have no agency at all.
As AI gets smarter, this will become more and more of a problem, until there some kind of cultural shift because AI is obviously safer it's not even a question any more.
This may or may not happen.
Keeping those with enough hubris to say "yeah, shouldn't take long if we've got some sort of sort of simulator to devise scenarios for a neural network to overfit to" about a wide range of failure modes that an entire field has spent decades studying well away from safety critical systems might be one of the more underappreciated aspects of tight regulation.
The linked article does not support your point, it shows a quadcopter losing one rotor which then goes into a kind of 'controlled fall', two (or even three) rotors is never mentioned and even the one rotor missing scenario involves the whole craft spinning (see video on the page).
I don't think so. You build a plane that's expected to run for decades -- if you put a 2010 computer in there now, it'll be 20 years old in 2030.
That 1990's computer was put there in the 90's, does that really make it more reliable?
The CG change from the new engines specifically, being masked by new software, MCAS. The only old part culpable is the legacy 737 limit of 2 AoA sensors (with only 1 used for MCAS input (a new thing)).
And if they weren't trying so hard to purposely avoid modernizing the cockpit, the issue would have been avoided altogether.
Also, planes are extremely expensive so they don’t get thrown away, they get constant upgrades. The U2 spy plane is a good example: a 1950/60s airframe running with more modern instruments. So if the need arise they will upgrade those CPUs on the 737s.
The whole point of this issue that the technology and cockpit design is purposely 30 years old to prevent the need for re-certification not for safety issues.
I'm not suggesting technology be used that isn't needed; I'm just suggesting using the technology to make air travel safer and easier for pilots.
According to wikipedia, the AMD29000 processor was first released in 1988 and "In late 1995 AMD dropped development of the 29k..."
I'm not commenting on whether that's appropriate to be running your aircraft or not, but it does very much seem to be a 1990's CPU running approx 50mhz.
And paper manuals? Are they expecting to use iPads for in flight documents or something? Shall we compare failure states of paper vs tablets?
Though I think most of them that are flying are C-47's from the war era.
The less software is in something(and less cloudy/iot) the better. And the more I program the more I prefer stupid and mechanical things when reliability is on the line.
A paper manual is always there for you.
Magic is great. Until it's not.
Then have the paper manual as a choice and/or backup.
If the programmers working on aircraft are so incompetent that they can't get a the digital manual correct, than they have no business programming any of the other multitude of systems that are critical to fly-by-wire systems.
The manuals have never been an issue. In fact, having a physical, rigorous checklist has been shown to improve the likely hood of a successful outcome.
That cat is already out of the bag. "According to the FAA, Class 1, Class 2 and Class 3 EFB may act as a substitute for the paper manuals that pilots are otherwise required to carry with them. "
And are them in your opinion also as reliable as paper manuals?
I think it's fair to say that MAX versions present some compromise solutions (like the now-infamous MCAS, which is there to compensate for the "unnatural" bigger engines). But I think that is not the main point. They would be good solutions it they worked as intended.
There are some other more fundamental and more daunting, afaik unanswered questions.
Like, why does such a critical system like MCAS take only a single AoA sensor as input, when there are two sensors available? Specially considering that the inputs from both are hardware-available to MCAS (the new software version is going to take data from both).
Boeing affirmed in its manuals that the elevators would be able to compensate for the trimmed vertical stabilizers. Now the preliminary report in the Ethiopian's crash shows that the pilots wheren't able to perform such compensation, even by pulling the control columns all the way back.
Those and some other issues are much more critical.
The classic approach is to have three sensors, so in case one fails you can know which one. Having two only indicates something is wrong but is not useful on the fly.
If true, it's possible signal averaging isnt necessarily the best choice
Air France 447 [1] three independent air data systems, two of them failed due to environmental conditions
XL Airways 888T [2] three independent AOA sensors, two failed because the plane was washed without the right covers in place
US Airways 1549 [3] two independent engines, both disabled by bird strike at the same time (No fatalities)
Qantas Flight 72 [4] three independent inertial reference units, bug in voting system if a single sensor's output had multiple spikes 1.2 seconds apart (no fatalities)
An in the data centre, no amount of power-supply redundancy will save you if a technician pulls out the power cables on the wrong server :)
[1] https://en.wikipedia.org/wiki/Air_France_Flight_447#cite_ref... [2] https://en.wikipedia.org/wiki/XL_Airways_Germany_Flight_888T [3] https://en.wikipedia.org/wiki/US_Airways_Flight_1549 [4] https://en.wikipedia.org/wiki/Qantas_Flight_72
Besides, those AoA sensors are EXTREMELY reliable. So reliable that some have raised the hypothesis that the real problem is not in the sensors themselves, but in some piece of hardware or software between them and the flight computers.
It seems plausible to me since failures in those sensors are too rare in the other planes but, despite that, they allegedly failed in two 737 Max 8s and in a really short timespan.
Which can also fail:
In 2008, on a customer-acceptance flight of an Airbus A320, two of the angle-of-attack sensors froze and those two sensors then outvoted the third. When the pilots went to demonstrate the stall-prevention system, they were not aware of the malfunctioning sensors. The plane crashed, killing the seven people on board.
The same problem arose again on a 2014 Airbus A321 Lufthansa flight leaving Spain. Eight minutes after takeoff, two of the angle-of-attack sensors froze at the same pitch. This time, after a drop in altitude, the pilots were able to regain control and complete the flight. [1]
I don't think the fundamental problem with MCAS was the number of sensors, but that it was too difficult for the pilots to override MCAS when it faulted.
1. https://www.seattletimes.com/business/boeing-aerospace/a-lac...
Are you saying the answer in the article didn't give enough information?
"Airbus addressed this potential problem on some of its planes by installing three or more such sensors. Former Max engineers, including one who worked on the sensors, said adding a third sensor to the Max was a nonstarter. Previous 737s, they said, had used two and managers wanted to limit changes.
The angle of attack sensor, bottom, on a Boeing 737 Max 8.CreditRuth Fremson/The New York Times “They wanted to A, save money and B, to minimize the certification and flight-test costs,” said Mike Renzelmann, an engineer who worked on the Max’s flight controls. “Any changes are going to require recertification.” Mr. Renzelmann was not involved in discussions about the sensors."
Clearly these signals can be cross connected, because that's the solution Boeing are testing at the moment, but it's outside the normal design of aircraft systems.
In the MCAS case, however, we are speaking about a computer which not only interferes in the flight controls, but also does that in a way impossible to override and it's too difficult for the pilots to spot the root cause.
However, for the autopilot example, I don't think you are quite right. Typically there are multiple autopilots (2 or 3 is quite common), and each is driven by a different set of flight data. In some scenarios multiple autopilots are engaged at same time, for example during CATIII auto-landings but as far as I'm aware that is the exception rather than the rule.
It's true that the autopilot may disconnect if there is a warning like 'IAS Disagree', I suspect this is driven by a separate monitoring process though, rather than being an integral part of the system.
The root cause is without doubt relying on a single sensor, and then downplaying the importance of the system so that nobody opted for the additional expense of the extra sensor. Boeing also have to answer for their lack of transparency; their flight control logic has always left the pilot fully in control of the plane, and can override any automatic system. This sets them apart from Airbus, which under almost all circumstances will defer to the computer.
In ways, the 737 MAX crashes are the antithesis of the 447 crash - the pilots thought they were in full command of the plane, whereas an automatic system designed to protect them malfunctioned, versus the pilots in the Air France plane believed the computer would protect them from exceeding the plane's capabilities, whereas the plane's computers could not get reliable data and passed full control to the pilots.
Where the risk analysis seems to have gone most wrong is that Boeing apparently grossly underestimated the difficulty of both figuring out what actions were needed to respond to the symptoms of MCAS failure, and to perform them. I don't know whether it was a significant factor in the former, but when the AofA sensor failed, it caused the stick shaker, as well as MCAS, to kick in.
The other mistake in analysis seems to be that when the power of MCAS was increased after initial flight testing, the additional risk it created was not properly taken into account. In particular, the ability of MCAS to drive the trim all the way forward appears to have been an unintended and overlooked side-effect of one design change.
Full investigation is hopefully going to tell us what has really happened.
> Pilots start some new Boeing planes by turning a knob and flipping two switches.
> The Boeing 737 Max, the newest passenger jet on the market, works differently. Pilots follow roughly the same seven steps used on the first 737 nearly 52 years ago: Shut off the cabin’s air-conditioning, redirect the air flow, switch on the engine, start the flow of fuel, revert the air flow, turn back on the air conditioning, and turn on a generator.
So? What does this have to do with anything? Is the goal to produce an airplane where pilots press a button "fly to destination", and the plane does it?
> The strategy, to keep updating the plane rather than starting from scratch, offered competitive advantages. Pilots were comfortable flying it, while airlines didn’t have to invest in costly new training for their pilots and mechanics. For Boeing, it was also faster and cheaper to redesign and recertify than starting anew.
> But the strategy has now left the company in crisis, following two deadly crashes in less than five months.
How was it the strategy to keep updating the plane that left to this crisis? The strategy itself is not to blame here, and I very much like the idea to gradually improve a proven model. It was a bad execution of this strategy that left the company in crisis.
In Germany, the national train agency Deutsche Bahn (and its predecessors) basically had a policy for nearly a century to order rolling stock that was designed to be produced for around 40 years. During this production run, the model was gradually improved. Some of the rolling stock designed in the 50ies is still in use, and quite reliable at that [0]. During the 90ies, agency and industry switched to a policy where basically every train generation was newly developed from scratch (for example, the ICE high speed trains). Guess what - you can channel a lot of public money into private hands that way, but bleeding edge technology is not what you want or need when you are trying to build a reliable transportation system.
When you're nose down heading into the ground, taking your hand off the stick to adjust the trim can get panic-y. Especially since the MCAS may be working against your trim adjustments.
The design of the Max line seems to be just enough compete with the A320neo, and just enough to not "need" re-certification. But these two justs, together, may have cost several hundred people their lives.
Those other Boeing jets required costly new training. Boeing were trying to build an update to the 737 that did not require this.
Sometimes this is justified -certainly the MCAS incident is a example. Changes like the MCAS should never have been allowed to change a fundamental part of the airplane interface without considerably more processes being met. That said, sometimes opposition to enhancements is trade unionists being trade unionists.
Apologies for not finding an exact link, but here is the wikipedia article for the flight: https://www.wikiwand.com/en/US_Airways_Flight_1549
Even on aircraft without fully integrated checklists, many operators have switched to using iPads instead of paper manuals as well.
> A second electronic system found on other Boeing jets also alerts pilots to unusual or hazardous situations during flight and lays out recommended steps to resolve them.
The 737 is stuck in the past because given the choice, nobody wants to have to retrain their pilots on more modern systems.
Ding ding ding!
Now, in the Air France flight, part of the problem was the system trying harder than it should have: it would have sufficed to alert the pilots to the lack of air speed indication and let them figure out what to do, but instead the computers cascade alarm after alarm.
You can see the problem: various pieces of the system were engineered separately, and there was no single system that could have suppressed the downstream alarms so that the pilots could focus on first on understanding the first alarm. Then again, that too might not have been a good design: perhaps a stall alarm should take precedence, say, over a frozen pitot alarm.
First Premise is that there re not many avionics expers - I agree so having more of those experts able to look at review and learn from different implementations would be a good thing.
Second premise is that you need to be an avionics expert to review and improve avionics software, I'm not convinced, there are many clever people out there.
Third premise you can't test or run the software becuase there is no emulator; maybe if the software was open source someone would start writing an emulator - maybe an emulator could be open sourced?
Fourth point; nothing would change if it was open source - I think the quality of code that would be released would go up and that's the main thing I'm talking about, embarrassment would be the optimal solution to something like this - I'd be very suprised if Boeing would release software that used a single sensor as input to a critical system like this.
Now the other problem that I have this in this thread. The Boeing issue is not a software issue at all. Even if we chose the best case scenario for the software it would be still a physical design & lack of pilot education problem.
Of course, that will never happen.
Within days of the Lion Air crash, MCAS was strongly suspected to be the problem. And yet Boeing management allowed 157 more people to die in the Ethiopian Airlines crash. And then the Boeing CEO still tried to tell everyone that it was a safe airplane.
The entire top management of Boeing, and the entire board of directors, should be perp-walked out the door.
Of course, that will never happen.
Fact checking at the NYT seems to be dead. The MCAS moves the entire stabilizer -- which is not a "flap" in any sense. Can they not find a pilot or some other knowledgeable person to read through an article like this before publishing it?
Paper manuals actually instill confidence. Should pilots contact Stackoverflow in an emergency?
Because as we are literally discovering, not knowing the details of how and why the automatic stabilization software is driving your trim wheel can kill you. The biggest single part of the failure chain here was the fact that this should have been an invasive change requiring recertification as a different aircraft and retraining for pilots instead of pretending that it was just like a 737-600.
I'm not familiar with pilot training and certification but it seems to me that even if the 737 MAX was type certified certifying a pilot with previous 737 experience would have been just supplemental training. Ultimately a nothing burger.
I've seen a lot of cases where people get bent thinking of all the bad things that will happen because they think customers won't like some change. Always in the end if it's not a show stopper the customers just lump it.
> Boeing also designed the system to rely on a single sensor — a rarity in aviation, where redundancy is common. Several former Boeing engineers who were not directly involved in the system’s design said their colleagues most likely opted for such an approach since relying on two sensors could still create issues. If one of two sensors malfunctioned, the system could struggle to know which was right.
> Airbus addressed this potential problem on some of its planes by installing three or more such sensors. Former Max engineers, including one who worked on the sensors, said adding a third sensor to the Max was a nonstarter. Previous 737s, they said, had used two and managers wanted to limit changes.
This seems to be the root cause of the problem. Competition is not always good.
One thing I remember is Microsoft tried repeatedly with Win95, 98 and ME to get USB support to work correctly. Failing each time and having to start over. Much to the horror of peripheral manufacturers. With Windows XP that finally mostly worked.