Boeing 787s must be turned off and on every 51 days (2020)
theregister.com
theregister.com
> The Patriot missile battery at Dhahran had been in operation for 100 hours, by which time the system's internal clock had drifted by one-third of a second. Due to the missile's speed this was equivalent to a miss distance of 600 meters.
> The radar system had successfully detected the Scud and predicted where to look for it next. However, the timestamps of the two radar pulses being compared were converted to floating point differently: one correctly, the other introducing an error proportionate to the operation time so far (100 hours) caused by the truncation in a 24-bit fixed-point register. As a result, the difference between the pulses was wrong, so the system looked in the wrong part of the sky and found no target. With no target, the initial detection was assumed to be a spurious track and the missile was removed from the system. No interception was attempted, and the Scud impacted on a makeshift barracks in an Al Khobar warehouse, killing 28 soldiers, the first Americans to be killed from the Scuds that Iraq had launched against Saudi Arabia and Israel.
So such things happen also in military (although this was early nineties)
More detail in 1 link from the article: https://web.archive.org/web/20100702180720/http://mate.uprh....
> I was once working with a customer who was producing on-board software for a missile. In my analysis of the code, I pointed out that they had a number of problems with storage leaks. Imagine my surprise when the customers chief software engineer said "Of course it leaks". He went on to point out that they had calculated the amount of memory the application would leak in the total possible flight time for the missile and then doubled that number. They added this much additional memory to the hardware to "support" the leaks. Since the missile will explode when it hits it's target or at the end of it's flight, the ultimate in garbage collection is performed without programmer intervention.
https://groups.google.com/forum/message/raw?msg=comp.lang.ad...
> This alarming-sounding situation comes about because, for reasons the directive did not go into, the 787's common core system (CCS) stops filtering out stale data from key flight control displays. That stale data-monitoring function going down in turn "could lead to undetected or unannunciated loss of common data network (CDN) message age validation, combined with a CDN switch failure".
Excuse me? Clearing the queue is not a "solution" if you have to remind a human to do it.
It's actually fairly depressing and is part of the reason I transitioned to consumer software.
We have a tendancy in software to think that the software should do literally everything even with a user actively working against it. A lot of times that is even more harmful.
Exactly my point. Trust no humans is not a philosophy that applies in aviation software.
It shouldn't apply quite as much in some other software either.
People who want to design reliable systems and processes look to airplanes for how to do that. That does not mean the same techniques are applicable or cost-effective in a different context, but at the very least the historical processes have been empirically demonstrated to produce extremely high reliability far beyond what nearly every other industry can attain and what many industries, such as commercial software, do not even attempt to achieve and may not even think is possible in their environment.
[1] https://www.sciencedirect.com/science/article/pii/S100093612...
[2] https://en.wikipedia.org/wiki/Transportation_safety_in_the_U...
These are the kinds of things that separate commercial software from open source software: in many organizations, commercial programmers typically write to a spec, and the specs usually come managers asking for specific, needed things without both a proper understanding of the larger picture and a nuanced understanding about how those specific pieces intimately fit in to the larger picture.
Open source is usually written by people who take pride in what they write, so even though a 32 bit timer for millisecond events almost certainly won't still be running 50 days after you've started a program, they'd still consider, "What if it overflows? Should I handle that event, or should use 64 bits instead?" instead of, "Not in the spec. Why should I care?"
And with bad programming practices comes unmaintainability. You want the department who wrote that code several years ago to open it up again, fix some things, then give you a new version? Well, that's going to take months and hundreds of thousands of dollars because they're not set up like that.
It's a stupid game with stupid results, yet we have so many apologists who think that if companies with money do it, it must be right.
The actual larger cost is not the maintenance or updates but the verification. There are definitely issues, especially when it's a code base you've been maintaining for 15-20 years, but generally it's the verification that takes the most effort, since often you're not just verifying the new functionality but downstream functionality as well.
Airbus A350s had a similar type of bug where the fix was to reboot them at least every 149 hours: https://it.slashdot.org/story/19/07/25/1932229/airbus-a350-s...
Towards the extreme ends of the spectrum of complexity, humans need sleep. On the other end, my pocket calculator likely doesn't need to be switched off and on to ensure that numbers add correctly. I guess complex operating systems sit closer to humans than a calculator on this "spectrum". I do remember reading that the space shuttle's computer systems were close to perfect in design, but they're not operated as frequently as a 787.
As the number of possible states of the system explodes, and as you layer levels of abstraction onto each other, the probability that you somehow reach a state the designers didn't think about increases. Then you need to reset.
Add to that limited bits for representation of numbers, and people not thinking about what do to when one overflows (because this is really really hard in many cases), and you get to "better reset it every n days" scenario.
> Any interesting writings on the subject?
None that I know of.
WTF? Is this standard operating procedure?
https://news.ycombinator.com/item?id=22761395
or even 16 hrs ago some pts and talk: https://news.ycombinator.com/item?id=27111650
Fail safe and mitigate harm as much as possible. A power cycle is cheap and easy way to mitigate leaks and overflows.
That was my first thought as well. The Arduino millis() function wraps back around to zero at around 50 days of uptime. https://www.arduino.cc/reference/en/language/functions/time/...
I don't have a copy of the 787 maintenance manual, but I have a feeling it covers that concern.
This article is still critical, but they are pandering to FAANG etc. (see for instance the recent Stallman coverage).