Boeing 787s must be turned off and on every 51 days to prevent 'misleading data'
theregister.co.uk
theregister.co.uk
But seriously, this is clickbait and nothing to see here. Many things on the aircraft are checked, cycled, etc before every flight, let alone on a 51 day mx schedule.
Have you ever had to do this mid-flight?
Edit: Well, if this was a normal month.
This article is turning a routine checklist/maintenance item into scary sounding clickbait.
It takes a certain type of person to fly a plane and resilience in face of unknowns and following checklists are some of the qualities they have.
In the same vein, these kind of failures are known to programmers for tens of years, just like metallurgists are aware of metal fatigue _and plan for it_.
Failure of software professionals to plan and mitigate for this kind of foreseeable problems is inexcusable and I liken them to some incompetent metallurgists in an alternate Univers who brushed off De Havilland Comet's like there was nothing to learn from.
Yes software bugs happen, but they are fixing this with documentation rather than root causing the problem.
That should worry people, because until its been root caused the actual implications are also unknown. For all we know its a symptom of a bug that will cause a more severe problem somewhere else.
This is a bug with ""several potentially catastrophic failure scenarios". Yet, its not been fixed in the ~10 years since it first flew. Nor is it the first, there have been a number of fairly critical bugs on this airplane that took a long time to diagnose before changes were submitted for certification.
So, In ~10 years and many major revisions of the aircraft, multiple re-certifications, etc none of them have bothered to fix it. One might argue they are afraid of changing the software because it might cause other catastrophic failures, but that leads down a thought process just as severe.
The off-the-charts arrogance of programmers allows creation of unbelievable tools and at the same time enslaves the person to the blind ego.
That's nothing, I fix "failures of the mind" by shutting off my brain for eight hours every day. That's a reboot!
In software, it's what we call an "ugly hack." An "ugly hack" meant that 737s didn't rely on both sensors, and people died. An ugly hack meant that the Ariane 5 rocket exploded in mid-air.
Ugly hacks should not be a part of any project where lives are at stake.
To be fair, this is somewhat of an over-exaggeration of the requirements since not all systems are critical and not all errors cause critical problems. In addition, the risk must be balanced against the alternative, in this case the risk caused by making sure a reboot is done every 51 days, so you would need to do an analysis of the failure probability and possible consequences of the status quo and compare that against the possible error modes of a software fix.
As an addendum to the risk analysis, the above analysis was only for one year error and on a per-flight basis. If you expect the 787 to fly for ~30 years then the fix must not cause two crashes over 30 years so a 1 in 7,500,000 chance. The average flight is ~5,000 KM which is ~4-5 hours per flight for a total flight time of ~60,000,000 hours. A plane takes ~3 minutes to fall from cruising altitude, so we need fleet downtime of 6 minutes per 60,000,000 hours which is 1 in 600,000,000 downtime. That is 99.9999998% uptime, 8 9s, 6,000x the holy grail of 5 9's availability in the cloud industry, 60,000x the availability guaranteed by the AWS SLA (again, somewhat of an over-exaggeration since you need correlated failures to amount to 3 minutes of continuous-ish failure, but that depends on an analysis of mean-failure time and mean-time-to-resolution which I do not have access to).
To then circle back to airplane software, the standards of original development are even higher than the standard I stated above. There are ~10,000,000 flights in the US per year according to the FAA and for at least the 10 years before the 737-MAX problems (I believe closer to 20), software was not been implicated in a single passenger air fatality. That means that in over 100,000,000 flights there were only two fatal crashes due to software (not even in the US, so we would actually need to include global flight data, but I will not bother with that since I am unaware of the count of software-related fatalities in other countries) for a total fleet-wide reliability of 1 in 50,000,000, 7 9s on a per-flight basis. If we use the per-time basis I used above, there are 25,000,000 flight-hours per year, so over 10 years, 6 minutes in 250,000,000 hours is 1 in 2,500,000,000 or 99.99999996% uptime, 9 9s, 25,000x gold standard server availability, 250,000x the availability guaranteed by the AWS SLA. Also note that with servers, we can use independent replicated servers to gain redundancy allowing uptime to multiply (1 in 100 failure for each server means chance of failure of both at the same time, assuming independence, is 1 in 10,000), but the same does not apply to airplanes since every airplane must succeed.
The thing to understand is that the software problems we are seeing are not necessarily an indication that their standards are lower than the prevailing software industry and that they should adopt their practices. It could be, and likely is, that the OBJECTIVE quality level we require is extremely high, and they have not been able to achieve it as of late. This obviously does not excuse their problems since, as I stated above, they must reach the OBJECTIVE quality level we require; it is just an observation that maybe it is not because they are incompetent, maybe they are really, really good and the problem is just really, really, really hard.
The cost of recertifying software shouldn't be the sole reason not to rewrite something, but I imagine you've been part of a rewrite where not quite everything worked as intended, even after having a lot of tests.
"The devil you know" very much applies to software that controls such life-critical functions and flying airplanes. If a pilot knows how to work around something, introducing something they may not know how to work around could be the difference between life and death.
Why did two 737-MAXes crash and the fleet grounded? New systems were introduced that pilots didn't know how to address - and seemingly could not workaround, even in spite of the engineers who designed them not wanting that outcome.
A rewrite, even with the most rigorous architectural and QA standards, is not a panacea.
That's an improper calculation. Slight nuance but orders of magnitude difference. A better approximation is:
If a fix had a 1 in 250,000 chance of causing a different error that would result in a fatal crash then it would be worse than the 737-MAX problems.
And converse:
If a fix had a 1 in 250,000 chance of causing a different error that could result in a fatal crash then it could be worse than the 737-MAX problems.
MAX's MCAS generates way more errors than 1 in 250k. A few of them resulted in a crash.
If I'd know that it were every "(2^32 - 1) ms to days" -> "49 days 17 hours 2 minutes 47.3 seconds" (millis stored in an unsigned long) then I'd be at ease, but 51 days doesn't say anything to me.
I just hope that they know why it's max 51 days.
These things are often driven with a 32.768 MHz crystal, in case anyone's wondering why not just a nice even 1.000ms.
On another forum I talked about way long ago, when I had to debug a huge system that reset after 248.x days of uptime. Yep, run the math...
I don't know for a fact this is the case here for the 787, but I think there are far better things to worry about when it comes to technical security in airliners than how often they need to be rebooted. For example, whether the on-board WiFi is sufficiently separated from the in-flight systems, and (as discussed recently here on HN) whether the advent of touchscreens for critical flight systems is sufficiently durable, tested and redundant.
Not necessarily. If the computer will never actually need to run for 51 days continuously, it may be a reasonable trade-off to require the reboot instead of writing (potentially buggy) code to handle a scenario that can be easily prevented from happening.
It reminds me of this story:
https://devblogs.microsoft.com/oldnewthing/20180228-00/?p=98...:
> I was once working with a customer who was producing on-board software for a missile. In my analysis of the code, I pointed out that they had a number of problems with storage leaks. Imagine my surprise when the customers chief software engineer said "Of course it leaks". He went on to point out that they had calculated the amount of memory the application would leak in the total possible flight time for the missile and then doubled that number. They added this much additional memory to the hardware to "support" the leaks. Since the missile will explode when it hits its target or at the end of its flight, the ultimate in garbage collection is performed without programmer intervention.
But it's not an issue, because servers are constantly redeployed for each code deploy.
How did they manage to create leaks? Or was it bugs in their custom PHP interpreter?
If the maintenance schedule does document system reboots, this is boring and business as expected... no different than periodically reinflating tires or replacing oil. I'd have no concerns flying on such a plane.
Things like Corona virus not being that bad and how surgical masks are useless. I genuinely question whether or not what you're saying is actually true or just following the trend.
I honestly can't tell these days.
This type of thinking is proving to be more dangerous than the ideas expressed.
Power cycling as a maintenance requirement is absurd. You don't sell someone a garage door opener for example and say "oh yeah, by the way, part of the ownership maintenance is that you need to unplug it at least every month, otherwise it might just start opening and closing randomly. No biggie".
This is an indication of a serious problem with the engineering, and after everything that has been failing with Boeings products lately, makes me want to avoid getting on a Boeing aircraft.
This has to be a joke. I can't fathom that this sort of issue is in production. It means that there is some sort of unaccounted saturation of buffers, or worse, memory leaks. In a deterministic system, these are the obvious things you need to get right.
https://ad.easa.europa.eu/blob/EASA_AD_2017_0129_R1.pdf/AD_2...
It would also be absurd to say that part of the maintenance for your garage door opener is tearing the whole thing apart after x amount of use, but that's absolutely standard for aircraft engines.
It was funny in Windows 95 days (and Unixes knew how to handle those) now it's just sad.
Of course the problem might be a bit more complex and it might be a combination of issues. Still not good though
Ummm, this says otherwise:)
Waxing anecdotal from a randos unverified, admittedly fuzzy memory is a pretty banal application of my agency. “People aren’t sharing info correctly! Allow me to fix this while admitting to a fuzzy understanding myself!”
Imagine what could be done if we curved that emotional sense to think and know into agency towards learning it and not bloviating online, bending readers mood to our position, disregarding literal effort at verification.
Add in behind the back character assassination of the author, and yeah really buying into your expertise. At least I’m being dismissive to your face.
You don't know that he didn't do a quick search and nothing came up.
Why not treat it like security? Yes, there are other layers of defense, but any given layer needs to be measured on its own. If I find that some large government website allows for javascript to be inserted for an XSS, but prevents it from running because it only allows javascript executed from a specific javascript origin, it is still a security flaw because some user might use a browser that does not implement content security policy. Yes, the user shouldn't be using such an insecure browser, but the website itself should not allow for scripts to be injected and not properly encoded.
I don't know about the case here, but any time I've hit an issue in my work where "Thing X needs to be done every Y or bugs start happening" it's a pretty clear sign of some deeper issues and likely a lot of underlying bad dev processes.
This issue might be as "simple" as a memory leak that will suddenly require reboots every N minutes when a seemingly unrelated patch exacerbates an issue.
Aircraft firmware requiring mandatory reboots in alignment with maintenance schedule, but working reliably otherwise, inspires more confidence than firmware advertised to run bug-free forever.
When lives are on the line software should be tested for reliability beyond 51 days. Having to restart is a symptom of reckless disregard for safety IMO.
Avionics software is written a world of verifiable requirements.
For how many days should the software be required to operate?
Is it acceptable to add that many [more] days to the software verification schedule in order to verifiably demonstrate that it works according to requirements?
Why is 51 days not long enough?
That this test could have been performed does not mean that all possible tests could have been performed.
Which is really what we're talking about here.
Is "able to run 51 days without reboot" a requirement, or not? If not, and it's not a use case, then it shouldn't have been tested for.
Instead, the limited time and resources available should have been spent on more important things.
I don't have as much a problem with that. There's always been tension between regulator access and proprietary corporate information. In order to get the former, you have to design a scheme to protect the latter.
I'm not saying that this is even acceptable or a great trade-off, but the way you worded your comment is presumptuous.
If this issue is clearly identified and tested around that's alright, it isn't a huge deal to have to reboot periodically... I'm more concerned this issue is one of those "Oh well, it just... gets a bit off after fifty days - try rebooting it, that seems to fix it."
Airliners are very long-lived equipment. So in fact they ship new releases. New releases have features that may be really valuable to safety, as well as features that are nice quality of life improvements. They're not shipping once per hour like a web startup, or even once per day like the NT internal team, but they do need to ship more than "once per new model of aircraft".
I've written before about an accident I spent a bunch of time looking at. No fatalities, just a smashed runway light but still reportable because of the "But for..." rationale. Two of the easiest things that would have prevented that from occurring were firmware tweaks. One was a recommended (but not mandatory) change in a newer build and the other exists only in Airbus planes so far.
Specifically the newer build does OAT disagree meaning if you tell the plane "It is -20°C outside" thus automatic takeoff thrust is much lower, the plane considers the temperature sensor at the engine inlet and it says to itself, this reads +15°C which is 35K different, that's the difference between flying and crashing into the fence at the end of the runway. I disagree with your guess about the temperature and so I refuse to try to figure out what to do next. You can realise you entered it wrong and type a more realistic value in, or you can set the thrust yourself manually if my sensors are broken.
The fancier Airbus approach was not to focus on the result of air temperature calculations. If the plane isn't accelerating enough, it can't fly, we don't care why it isn't accelerating, maybe the wheels are square - we need to abort takeoff so we don't crash. So teach the plane how long runways are, it can use GPS to figure out which runway it's using, and then it can tell pilots if they aren't getting enough acceleration and they'll abort because they don't care why it's not enough acceleration either, they don't want to die in a fireball.
I had an A380 flight that was slightly delayed due to a “software update” taking longer than expected.
It was at the SIN layover for QF1 LHR to SIN, so it was kind of worrying/amusing to have your plane need a software update halfway through your journey
No it's a symptom of having bugs in your code.
And they can be there for a host of reasons ranging from "this is a once off accident" to "systematic failure in the software engineering process".
The requirement of a reboot on its own, though, would not strike me as a blatant disregard for safety, as long as the period between reboots is long enough to exceed the maximum possible length of flight (taking any emergencies into account) with leeway to spare.
(Franky, I am surprised airplane software runs for days—somehow I assumed they get shut down completely every time they refuel.)
[0] https://en.wikipedia.org/wiki/Aircraft_maintenance_checks#A_...
Probably this is a counter that rolls over if it's not reset, the predictability of needing to reset it before time T is an indicator it's a sequence number that's driven by a hard real-time trigger with extremely predictable cadence.
Hey, maybe <51 was just a off-by-one error... or maybe the actual advisory is to be <50 and some PM decided that number was too round or violated an SLA.
1. 4,294,967,296 or 4.29e9
Context[1] for anyone not familiar with the old Win32 APIs.
[1]: https://docs.microsoft.com/en-us/windows/win32/api/sysinfoap...
"for reasons the directive did not go into, the 787's common core system (CCS) – a Wind River VxWorks realtime OS product..."
Early in development we ran everything (real-time embedded system) on a system interval whose finest level of granularity was a 1ms tick. The system scheduler used a 32-bit accumulator that we knew would roll over after 50 days. However, we were given assurances that the system would have to be powered down for maintenance weekly so it didn't matter. Since proper maintenance is a hard requirement (or the instrument will start reporting failures) that was OK.
Eventually, some time after release we started getting feedback that the system was shutting down for no apparent reason. We investigated and found that those failures were all due to not having been powered down in months.
Apparently, since it could take up to 30 minutes after powering on the instrument before it was ready to run, some labs were performing maintenance with the power still on, so the machines were hitting much higher than expected uptimes. In many cases it wasn't an issue, but if time rolled over in the middle of a test, the instrument would flag the "impossible" time change as a fatal error and immediately shut everything down.
Next release moved to a 64-bit timer. I think we're good :-)
This is where FMEA (Failure Modes Effects Analysis) is useful: the likelihood of critical failures is assessed and the ones that are both dangerous and unacceptably likely to occur are removed by design. The rest are assigned specific ways of being handled.
In this particular case, the severity of the failure (not completing a test) is relatively, but not unacceptably high, but the fault is very unlikely to occur since it requires that (a) a lab violate the maintenance protocol we specified and (b) it happens during the time period between a test starting and ending. In all other cases it's a non-issue.
If we were to continue running in this scenario, the outcome could be far worse than shutting down since we would now have the possibility of providing incorrect diagnostic data to a physician. Again, the FMEA would say that although shutting down is bad, continuing to run is far worse.
However, they're also very difficult and time consuming to perform and keep updated throughout the development lifecycle. They're also necessarily sparse in terms of real coverage of a system's operational/behavioral domain for complex systems.
That said, I think way more software engineering organizations should be doing them as a matter of course even outside safety-critical systems. They're a very useful procedural tool to highlight blindspots at the very least.
But then the debugging started. Turns out the service discovery component would health check backend destinations once a second. This was fine as it made sure we would never try to call against a server that was long gone. The bug was that it never stopped health checking a backend. Even if the service discovery had removed a host from the pool long ago. Gmail had stopped deploying while it got ready for Christmas, and Calendar was doing a ton of small stability improvement deploys. We created the perfect storm for this specific bug.
The most alarming part? This bug existed in the shared code that did RPC calls/health checking for all services across Google and had existed for quite a long time. In the end though, Gmail almost took Google offline by not deploying. =)
The
https://en.wikipedia.org/wiki/Alaska_Airlines_Flight_261#Ext...
Basically, adding more code to make software "smarter" for those edge-cases was judged (rightly or wrongly) as having an even higher risk of introducing new bugs and creating new test procedures or invalidating previous testing.
* Patriot Missile launchers must be rebooted periodically to account for accumulated errors from truncation: http://www-users.math.umn.edu/~arnold//disasters/patriot.htm...
* Cruise missiles with memory leaks that do not matter because they'll reach the target before they allocate all memory: https://groups.google.com/forum/message/raw?msg=comp.lang.ad...
The issue appears after 51 days. I can't see in that AD a specific recommendation to reboot at predetermined intervals, just that it requires reboots before 51 days of continuous operation.
https://www.engadget.com/2015-05-01-boeing-787-dreamliner-so...
At one company I worked for, all of our National Instruments test equipment would start to fail with communication problems after about two months on our Windows XP computers. Being familiar with GetTickCount, I rebooted the computers, recorded the date, verified the next failure was 49 days later, then emailed National Instruments with a link to the GetTickCount documentation. They pushed out an update with a fix 3 days later. Oops.
I'm not an aerospace or embedded engineer.
We need an Elon Musk of flight, Boeing has gotten away with too much mediocrity
However, since Airbus != Boeing, nobody around here cares. Only stories that are pointing out problems with Boeing are allowed (or upvoted) apparently.
[1] https://www.theregister.co.uk/2019/07/25/a350_power_cycle_so...
http://www-users.math.umn.edu/~arnold//disasters/patriot.htm...
Seriously, Windows 10 does tend to get dodgy if you don't reboot for a few weeks. I'm not the only one who's noticed. Granted, it's less likely to outright crash, but it acts increasingly drunk.
I made an excercize of rebooting my work computer every Monday, but it really wasn't enough. Now I restart it every time I can.
https://sites.google.com/site/edmarkovich2/whywindows95andwi...
Still, it's remarkable that two separate Seattle-based companies have produced a similarly short time bomb on very expensive and highly visible product development projects.
The joke was that nobody had ever had a Win95 system stay up for 49 days. Mwah-hah-hah.
If the government deems that Boeing much be saved, it should also deem that the that prior management was negligent and cause for this situation and seize their personal assets and hold them criminally accountable.
Not published April 1.