After Alaska Airlines planes bump runway, a scramble to ‘pull the plug’
seattletimes.com
seattletimes.com
* There were two unusual events in very short order
* Someone quickly noticed and gave the order to go to ground stop
* The problem was figured out quickly and a work-around was developed
* Flights Resumed after successfully deploying the work around to the 'production process'
* A patch was quickly developed and deployed once the underlying bug was uncovered.
I think everyone (but perhaps the developers) comes out looking like a champ.
Now, do I think telling pilots what thrust they should use to take off is wise? I dont know - I'm not a pilot, and I'd want pilots feelings on the matter before commenting further, but it feels to me like a loss of authority that would make me uncomfortable - particularly if I was held to account for undetected failures.
The other part that I'd note here, is process is just as important as software could process alone have caught this failure without the tail strikes? possibly, and its something worth looking into further - but only if it doesn't add an unacceptable workload burden to pilots workload that could otherwise compromise safety further.
>“The goal is to lower the power used on takeoff,” he said. “That reduces engine wear and saves money” on fuel and maintenance.
>Flights to Hawaii are typically full, with lots of baggage and a full load of fuel for the trip across the ocean. The planes are heavy.
> That morning, a software bug in an update to the DynamicSource tool caused it to provide seriously undervalued weights for the airplanes.
>Peyton added that even though the update to the DynamicSource software had been tested over an extended period, the bug was missed because it only presented when many aircraft at the same time were using the system.
> Both planes headed down the runway with less power and at lower speed than they should have. And with the jets judged lighter than they actually were, the pilots rotated too early.
>The data “confirms that the airplane was safely airborne with runway remaining and at an altitude by the end of the runway that was well within regulatory safety margins. Both aircraft got airborne well within safety limits despite the lower thrust.
It’s also probably lucky that two Hawai’i bound flights both didn’t notice the discrepancy, or two bumps wouldn’t have bee as suspicious. Interesting that someone cited divorce and sleep examples as emotional pleas, reassuring us they are only human.
Like all complex failures at least 5 things had to happen. Software update that morning, some kind of hard to simulate resource issue in processing, overweight flights, pilot not pricing plane data appeared light, pilot turning early despite low speed, two in a row.
Huh? FAA spends a huge amount of time and energy focusing on human factors, tasks saturation rates and crew resource management. Nothing to do with with argumentum ad passiones, just prudent risk mitigation.
But another class of mitigations is to make the human less error-prone in the first place. This is especially important in preventing failures that were caused by improper management of the automation. Hence the emphasis on rest periods, training, and other mitigations on the human side.
In the "swiss cheese" model, a disaster needs to slip through both the human's defenses and the automation's defenses to actually happen. It makes sense to improve both of those layers of defense against disaster, while acknowledging that humans will never be perfect. But at least they can be awake and knowledgable.
The full quote:
> …if there’s a glitch, naturally some pilot somewhere is going to miss it.
“Not everyone gets eight hours sleep the night before. Someone is going through a divorce. Someone is not so sharp that morning,” he said. “The sanity check isn’t perfect every day of the week.”
No, it's more that using lower thrust during take off saves on engine wear and noise levels around the airport. It doesn't really have that big an impact on fuel use, at least that's not the primary purpose.
The article concentrates far too much on the thrust setting. The important bit is the speed where the plane should be rotated to take off, which is known as V_r. That depends on the weight of the plane. On the takeoff roll, the pilots watch the airspeed indicators and pull back to lift off when that speed has been reached. If they have been told a V_r that is lower than it should be, then pulling back will rotate the aircraft without it taking off, and that's when you have this danger of the tail hitting the ground. When that happens, the plane must be inspected, because a tail strike has the potential to weaken the pressurised container that is the fuselage, and this could lead to an explosive burst and depressurisation when flying at high altitude. See https://en.wikipedia.org/wiki/China_Airlines_Flight_611 for why this is a Bad Thing.
Which is saving on money
There's a difference between dumping toxic waste into a river to save money and keeping a plane engine in safe operating conditions to save money.
A somewhat less direct concern but a related one is ground effect lift. the aircraft can remain a short distance off the ground (roughly a wingspan of altitude) at lower speeds than it can actually "fly," due to ground effect. Smaller aircraft might routinely spend some time in ground effect gaining additional speed before they begin climbing, but airliners have so much thrust they usually rotate pretty directly to their climb speed. This makes it more of an issue though that if rotation occurs too far before Vy climb speed the ground effect period will be prolonged and increase the amount of time the aircraft spends at risk, flying but slow with poor control authority and too close to the ground to have much of a recovery opportunity. The ground is a pretty safe place to be, the sky is a pretty safe place to be, but that first couple thousand feet between the ground and the sky is rather hazardous and jetliners get out of it as quickly as possible.
Then the problem here is trading safety for something - cost apparently.
BTW, in a small plane I found soft field takeoffs quite fun. Nose up early, wheels up early, fly in ground effect to gain speed. I suppose that is a safety tradeoff and might not want to do it in some conditions.
Edit: Actually, that probably doesn't even come into it. Not a pilot, but don't imagine "barely rotate so the nosewheel is in the air but still have negative angle of attack so no lift" is very easy, or very stable if possible. Even if you could do it, you'd risk slamming the nosewheel back into the runway.
So, once you get to the stage where you're ready to rotate, or even late in rotating, you're committed to getting the plane off the ground.
But it could also just be statistically true. Being in a position of power, surrounded by beautiful coworkers, lots of overnight hotel stays..
Without further investigation a conclusion can't be drawn though.
That there were no tests for this system under high load indicates a poor engineering practice at DynamicSource. They don’t deserve any kudos for a post-hoc patch that a test in production discovered the bug for when the system is a safety critical system like this.
Also process, procedure and more importantly people are always the last line of defense against failure.
Like why would the increased system load change the results. The calculations should be independent and designed in a way where they can't interact.
Bad weight estimates have resulted in crashes in the past.
A few modern pickup trucks have a system for measuring how much weight is in the bed or being transferred through the tongue of a trailer because of how common it is for people to overload their trucks.
AFAIK these systems just use the same sorts of suspension height sensors that have been used for auto-levelling for years and are factory calibrated to know that if the rear end squats by X distance then there must be Y load in it.
Seems like something along those lines would be technically easy to retrofit to large aircraft, possibly even as a software update if they already have some relevant sensors.
Pilot here :) the journalist got it a bit wrong, performance data isn't primarily about how much thrust to use. It is part of the calculations, but crews are free to deviate from reduced thrust calculations and will do so.
This incident is much more about the calculated speeds to use and aircraft trim settings. That's what caused the tail strikes, trying to lift off at a speed that is too low for the weight.
Those numbers are required by every airliner taking off. And getting them from software is standard everywhere. The alternative is doing the calculations manually using paper charts, every pilot is required to know how to do that, but nobody does it that way for commercial operations.
I thought I was missing something in the way the story read, I'm glad to know I was.
> “Still, the mishaps point to the need for more vigilance by pilots in checking automated data.”
In commercial aviation there can never be a single point of failure that could lead to catastrophic outcome. The performance calculations are such a thing, so there must be validation steps.
In many airlines for example if the performance calculations are done manually (or by an app based on manual input) they need to be done by both pilots independently and the results cross-checked. Otherwise a single typo could have a big impact.
That seems horribly wrong to me. I can understand software being slow under load, but being wrong under load sounds like a horrible internal architecture problem.
"the bug was missed because it only presented when many aircraft at the same time were using the system"
Obviously sessions should be independent and not sharing data, but that's why it was a bug.
https://www.bloomberg.com/news/articles/2019-06-28/boeing-s-... (paywalled)
https://archive.is/vdg9S (paywall workaround)
Note that there were several articles about it around the time, the above one is just the first reasonable seeming one from a quick search.
Such as weight=plane+fuel+people+luggage. If one of the variables became 0 you’d still have the other 3.
It was a low enough value to have made multiple pilots question its value that day, like fuel or passenger load was missing completely.
Tons of software can be "wrong under load" - things like race conditions and memory leaks are common problems, and they don't necessarily point to a huge architectural defect. E.g. Cloudflare had a bug years ago that caused highly sensitive data to leak, but only for a teeny (relatively) portion of their site. Similar issue happened to GitHub: https://github.blog/2021-03-18-how-we-found-and-fixed-a-rare...
For example, one way to use SQL is to escape your parameters and then concat them straight into your SQL string. However, if someone leaves out an ‘escape’ somewhere, you have a major security incident
The other way is to pass parameters separately as data to your SQL driver. It completely negates the problem
If you’ve chosen the first way in your project already, you’ve committed to a major internal architectural problem. In the same vein, maybe if your code requires sprinkling ‘synchronized’ everywhere, you did it wrong
But that said, we don’t know what actually is the problem since it’s just a marketing statement saying that high load was the cause and that doesn’t tell us anything
If you're writing software for giant metal flying cylinder carrying hundreds of people, we can't just brush mistakes off the same way we might with web-based software (and yes, privacy is important, but it's not life or death).
FTFY. IT shops on the whole aren’t any better here.
Software's accuracy being affected by load is completely unacceptable.
There's no such thing as perfect software, but there is definitely software that doesn't make up a false value if it can't work.
To be fair to the other commenters, it's conceivable that the practical difference between the current unsound architecture and a sound replacement might come down to to some small defect. But an attempt to fix the problem by fixing just the defect, even some kind of RCA process on it, will fail if the architecture itself is not fixed, and that requires a change in attitudes and understanding by the devs responsible for the system, not just a software change. That is where the real flaw lies. But some people don't believe that it's possible for the flaw to be in these places, outside the code. All the RCAs in the world can't help them.
That said they have a pretty explicit cap on maximum activity (number of planes) so it is weird they didn't test around this. It isn't like 400,000 aircraft suddenly DDOS their system.
All the other arguments seem to assume a large consumer type of load - tens of thousands of users, etc...
I just can't see an undue strain being placed on a well designed system from < 300 data points. And I haven't even accounted for the distribution of needing to compute takeoff data over the course of a day nor how many planes are NOT taking off at the same time, etc...
Also, to somewhat change the topic, didn't Alaska Airlines disband their QA org a few years ago as part of cost cutting? IIRC, they did this to model the software company models (that ship bugs regularly to consumers) and seem to be getting some data that they need to bring back that org...
Follow up thought: When software results are this critical, I wonder if a totally separate program should be used as well and result compared. An independent implementation from another vendor.
An absolute pro. There's a hundred variations of this story, to varying degrees of criticality and impact; seeing a pattern out of two data points, connecting the dots, making the tremendous call to immediately pull the plug, to stop the world and give engineering time, then diagnosing and triaging the problem in less than a half hour; that's world class reliability engineering.
Such a rule likely saved an SR-71 once: https://theaviationgeekclub.com/the-story-of-belmont-86-the-...
The natural tendency for military air crews is to complete the mission if humanly possible. To counter this inclination, the Wing Commander had designated certain emergencies sufficiently critical to require immediate landing. This was one of those emergencies.
The system reports data on number of passengers, weight of cargo, plane balance, etc., to the pilots. The calculation is done by the plane’s flight computer. How can it be off by 20,000 pounds, but only under heavy server load?
The explanation that comes to mind is that DynamicSource has a subservice for each source of weight and one of those subservices crashed under heavy usage. So the top-level aggregate-and-report service got an error from one subservice and said “well, guess it’s zero, lol”?
Not exactly true.
The dispatch release has a flight & performance plan based on the load manifest. It's calculated out of band (here, by the software vendor) and punched into the airplane configuration for departure.
If the plan is wrong, the plane flies wrong. GIGO at its finest.
> the update to the DynamicSource software had been tested over an extended period, the bug was missed because it only presented when many aircraft at the same time were using the system
> the data was on the order of 20,000 to 30,000 pounds light. With the total weight of those jets at 150,000 to 170,000 pounds, the error was enough to skew the engine thrust and speed settings.
Multithreading/contention issue? But how would that alter the weights?
> The software code was permanently repaired about five hours later
That's surprisingly fast, isn't it?
Off the top of my head: failure to sum values correctly, misreporting the weight to the target plane (reading values from another plane’s weight data).
If there are multiple data streams feeding in weight data per aircraft, under a low aircraft load (the scenarios they apparently tested) the odds of interruption during an operation like += can be low enough to not see the issue. Under high load, though, the odds of interruption and incorrect tallying increases substantially.
"multiple data streams feeding in weight data per aircraft"? What does that word salad even mean?
Even if what any of what you said was true: these developers are writing critical safety software. If they can't manage to write code without generating race conditions and testing their system for such conditions, they are grossly unqualified to be writing this sort of software.
As others have said: this is a huge fuck-up.
And that's not word salad, though maybe not as clear as I could have made it. But to clarify for this hypothetical (if this is the issue): Each luggage weigh station submits data to the application or database which then has to tally the weight. Those are multiple data streams, not a hard concept. An airport has many luggage weigh stations all potentially submitting data at the same time.
If those are being processed concurrently and the database or however the data is stored is not properly locked then the tally can be incorrect for the same reason non-atomic increments/decrements can become incorrect in a multithreaded application. But just like the problem with non-atomic (unlocked) increments/decrements the problem may not manifest without high enough load on the system. Low-load testing (what the article seems to describe) means that updates can happen fast enough (server or application is under loaded, less likely to interrupt any update) or submitted slow enough that the issue never appears. So you need to test with a higher load to detect the problem (if you're going to be able detect it in testing at all).
"multiple data streams feeding in weight data per aircraft"
it's literally some integers (floats at the worst?).
Maybe they stored something in a global variable, and with two planes querying at the same time values for the wrong plane could be retrieved.
A trivial error could be "list all active planes", "list current luggage weight for all these plane IDs", and have the latter give responses in a different order than the input list. Or perhaps "fuel load" returns in a different order. Or something.
Concurrency seems like the obvious one, but it's not the only source I can imagine here. And ultimately, the max is under 300 planes. You can do that serially unless the integration layer is slow.
Actually there were such guardrails, so the system had multiple manual checks built in.
> “At that point, two in a row like that, that’s when I said, ‘No, we’re done,'” said Peyton. “That’s when I stopped things.”
Hats off to Peyton, and possibly Alaska. I would also love to hear about times when something similar happened, and it was discovered that nothing was wrong.
It’s also fragile, and the fact that it strikes many as an uncommon level of integrity for a business decision is precisely why aviation needs to be shielded from having the prevailing business culture sneak in, like we have seen happen at Boeing in a few highly notable instances.
[0] https://en.wikipedia.org/wiki/1983_Soviet_nuclear_false_alar...
That's no small mistake. It's not a rounding error on the weight of the soda cans loaded into the galley for the flight.
That's most of the passengers. It's most of the freight. It's heavy enough to be a spare engine in the hold. It's a goodsize fraction of the fuel needed to fly from Seattle to Honolulu. It's certainly enough to foul up the mandatory weight-and-balance computation the pilot in command is required to do.
Somehow the input to this software package missed something big. It would be interesting to know exactly what was missed.
No. You don’t get to build complex configurations that only work with supporting software and then BLAME THE USER when the software fails.
“Ah, but you didn’t follow rule 137 in the 2000-page manual that nobody reads because computers made it obsolete” is a deflection of responsibility from the owners to the operators
and it’s not okay.
> Alaska’s Peyton said “several crews noticed the error and notified dispatch.”
> The pilot at American Airlines said “requesting manual data is not standard” and that if there’s a glitch, naturally some pilot somewhere is going to miss it.
> “Not everyone gets eight hours sleep the night before. Someone is going through a divorce. Someone is not so sharp that morning,” he said. “The sanity check isn’t perfect every day of the week.”
So they already knew (or should have known) about this glitch, and did nothing about it until it started causing damage to aircraft. I would have thought the pilots who reported it earlier would have also predicted the outcome and got it fixed then. From my lay perspective, it's obvious that if aircraft is heavier than the flight computer was told it is, then something bad is going to happen (e.g. could have been complete failure to take off and crash at the end of the runway).
Well, maybe such incidents where pilots do make the prediction, and get it fixed before it causes the problem they predicted, never make it to the news...
>"It delivers a message to the cockpit with crucial weight and balance data, including how many people are on board, the jet’s empty and gross weight and the position of its center of gravity.
>In a cockpit check before takeoff, this data is entered into the flight computer to determine how much thrust the engines will provide and at what speed the jet will be ready to lift off."
Given that overhead bins are regularly maxed out with carry-on luggage now since airlines began charging for checked bags, how are they able to accurately account for the weight and balance? Airlines seem to almost never weigh customers carry-on at check in.
They'll have a decent idea of the average weight of the average passenger (https://airinsight.com/the-pending-new-faa-weight-balance-ru...). If you're chartered by an anorexia treatment facility to move 300 patients, you may put a little extra attention to the exact values.
> The new FAA standards will increase an average adult passenger and carry-on bag weight to 190 pounds in the summer and 195 pounds in the winter. Up 12% from 170 pounds and 175 pounds, respectively. This includes an extra ten pounds for winter and five pounds for summer. This also includes 16 pounds for personal items, up from ten. Airlines must increase the average weight of female passengers and their carry-ons from 145 pounds to 179 pounds in summer, and from 150 pounds to 184 pounds in winter. The average weight for males with carry-ons is increased from 185 pounds in summer to 200 pounds, and from 190 pounds to 205 pounds in winter.
What I do know is that with something like this, a little could go a long way. I wonder what the inspection and repair for a tail strike is, and whether that cancels out the money saved by minimum viable thrust across the fleet. I’m sure someone is punching the calculator on this to determine that.
Aviation's LGTM
Make no mistake. The Exec/finance class will cut into any margin that you aren't willing to make them directly legally culpable for.
Note that these are manufacturer-sanctioned procedures, not applicable to small piston aircraft, which should be using full power/RPM if the AFM/POH calls for that.
Note also that the full rated power (and sometimes more) is available in the case of an abnormal or emergency.
As long as you save more than $1.5 billion ($5 million x 300 passengers) before a plane crashes, no problem.
Until we completely bankrupt a company for killing people under its jurisdiction, this will all continue.
And based on the inferred mass, can’t you directly calculate what the takeoff velocity should be?
When they start the 'rotate' maneuver, the nose of the aircraft doesn't immediately go to (for instance) 10 degrees. It's a progression.
As soon as the nose starts pitching up (by 1 degree for instance), the aircraft is already generating lift to pull the aircraft up.
If the speed is set right (so you're fast), when you get to that 10 degrees angle of attack, your tail has already cleared the runway by a very safe margin, because you've been producing X lift from 0 to 10 degrees of rotation.
However, if your speed is set too slow, you're not generating as much lift, so the aircraft will be pulled up slower (you have been producing Y% less lift from 0 to 10 degrees of rotation). This means that by the time you reach your designated 10 degrees of angle of attack, you haven't cleared the tail yet, and it will hit the runway.
(Note that you can't keep increasing AoA indefinitely if you need more lift; at a critical speed, the airflow will separate from the wing and stop generating any lift. That's an aerodynamic stall.)
Now that's strange. Anyone have more details? Why should there be any connection between the calculations for different aircraft?
It basically boils down to “this flight started nosediving, almost hit the ocean, and we don’t know why”.
https://www.nytimes.com/2023/02/13/us/united-maui-flight.htm...
Interesting that baggage weight changes by destination. I'm sure passenger weight does too, but for different reasons.
\> an ad (or whatever, didn't read it) popup appears
\> close popup, try to start reading again
\> another popup appears and blurs the article
\> close website, I wasn't that interested anyway
Also shocking that the system gives an incorrect answer when it is under load. Lack of load testing is a problem, sure, but an architecture that allows the answer to just be wrong is fundamentally flawed.
This is turning into the same sort of high drama news site that most of us are trying to get away from.
In this instance, at least it was a software bug that was the cause. So quite suitable for HN.
> the computer then calculates just the right amount of engine thrust so the pilots don’t use more than necessary. “The goal is to lower the power used on takeoff,” he said. “That reduces engine wear and saves money” on fuel and maintenance.
This is not an accident, but rather a feature of capitalism, this time it is human lives that are commodified and have costs externalized unto.
There’s no need to have market based competition to keep costs down. Look at USPS for example. That operation is run without a profit motive, serves literally every zip code in the US (by definition), and runs healthily year after year (except for the part where Congress imposed arbitrary pension funding requirements[1] and private capital interests have been chomping at the bit to buy up all the real-estate holdings).
They also don’t have UPS’s problem of not putting AC in their delivery vehicles.
[1] https://ips-dc.org/how-congress-manufactured-a-postal-crisis...
This feature of capitalism caused the money to be invested to achieve this efficiency, and also the systems and processes in place to ensure it was implemented properly; with sanity checks, authoritative oversight to halt ops when something was determined to be off, manual backup systems in place and well trained on, a fix rapidly deployed, and new testing put in place to detect any similar bugs in the future.
Yes indeed the stench of capitalism is strong here! </s>
Comment OP isn’t wrong entirely - this is the result of an economic decision. Not without it’s good effects - both can be true, but this is not a safety optimization by any means.
Many aviation accidents can be traced back to capitalistic pressure - cost saving on inspections, time saving for flow, etc. It’s appropriate to ask whether this is true here as well.
There are many considerations to setting takeoff thrust like climb speed limits, tire speed limits, elevation, temperature, engine-out safety margins, and many more I’m sure I don’t know about with my limited flight knowledge.
This is before we get into discussing weight load, which is primarily a factor of fuel load, which is also not just filled to 100% for every flight, but carefully calculated based on distance, weight, wind speed, etc. to not carry unnecessary fuel, which is also a huge efficiency boost.
In short there are safety reasons to use the appropriate (non-maximum) thrust on these types of planes just as much as there are efficiency reasons.
Capitalistic market forces are a big reason we have such an amazing airline industry with planes that are technological marvels of comfort, speed, and efficiency. Capitalism drove 120 years of relentless R&D into commercial flight systems, from the very first flight in 1903.
Not paying for necessary safety inspections is not “capitalistic” in any way. It seems like an ideological battle to use the term that way.
This means they are also leaving a safety margin of fuel, even though that costs extra money.
The influence of capitalism would be on the requirements. But assuming the requirements are well-considered trying to minimize fuel and engine usage to meet those requirements is a good thing.
If the engineering requirement is “do something that brings you close to the margin of safety in order to save resources”, is that not an influence of capitalism?
How to hit that line precisely and with confidence is where engineering takes over.
If efficiency is the curse capitalism brings us, I’m feeling a little better about it.
https://www.theguardian.com/business/2012/feb/28/ryanair-sta...
You’re trolling, right? That seems like a really implausible world view.
Software to optimize resource usage while providing a service => blame capitalism is che-guevara-tshirt-edgelord level of inane.
This story is about the system working well. Detected immediately via well-trained and alert humans, someone had and used their individual authority to ground the entire airline for safety reasons, mitigated with a workaround 20 minutes later, permanently fixed within five hours, and new tests implemented to account for the root cause.
Sure. And there are bugs in their code, too.
https://www.theverge.com/2019/9/3/20847243/spacex-starlink-s...
"SpaceX acknowledges that it failed to communicate due to the bug and missed the emails about a higher probability of collision. Finally, on Monday morning in Europe, ESA made the call and used the thrusters on Aeolus to raise the satellite’s orbit by about 984 feet (300 meters) without waiting for SpaceX to take corrective action."
The first failure was the bug itself - without knowing more about the details it's hard to say more. But one should ask what kind of practices exist at DynamicSource - e.g., compile time thread safety checks, code review practices, testing requirements, rollout procedures - that would prevent bugs like this from getting out into production.
The second failure was that the software failed to validate its own output and fail noisily or generate a warning of some kind. It's safety-critical software - bugs may be excusable, but not giving at least a warning for an unusual output is poor engineering.
The third failure was that the flight staff failed to notice the error until tail strikes happened. Human errors do occur, but we need better systems to assist human operators. E.g., perhaps the flight software itself could act as another line of defense to prompt the pilot when values appear out of normal range.
Yes, in this case the planes weren't anywhere close to crashing. But if the assumption is that bugs are unavoidable, then there needs to be better systems in place to catch those bugs before they cause a major accident. Because the next time we might not get so lucky with detecting the issue and having someone with integrity calling the shots.
Also, no need to start off your comment with an ad hominem - a simple "software errors are unavoidable" would've said just as much.
It's literally a safety critical system, not move-fast-and-break-things front end web dev code.
Remember the extraordinarily in depth review of their software while their entire fleet was grounded, where each and every individual bug found was reported in mainstream media?
Where they weren't allowed back in the sky until we were all assured everything was fixed?
The bug they're reporting here sounds like they didn't do reasonable testing. And that's the kind of thing that shouldn't be happening with Boeing 737* aircraft especially, after their recent problems. :/
So was the Space Shuttle, whose software team is widely regarded as the absolute gold standard in this regard.
https://www.fastcompany.com/28121/they-write-right-stuff
Still, bugs:
"This software is bug-free. It is perfect, as perfect as human beings have achieved. Consider these stats : the last three versions of the program — each 420,000 lines long-had just one error each. The last 11 versions of this software had a total of 17 errors. Commercial programs of equivalent complexity would have 5,000 errors."
> Where they weren't allowed back in the sky until we were all assured everything was fixed?
Because no aircraft would fly again with that standard.
> And that's the kind of thing that shouldn't be happening with Boeing 737* aircraft especially, after their recent problems.
The software in question isn't made by Boeing, and it's not just for Boeing aircraft. It's a third-party thing, picked by individual airlines. https://www.dynamicsource.se/ lists support for 13 aircraft across 7 manufacturers.