It doesn't seem much worse then memory leaks in missle guidance tracking systems that exceed flight time. We have finite resources, if the effort to correct is minimal what's the harm?
It doesn't seem much worse then memory leaks in missle guidance tracking systems that exceed flight time. We have finite resources, if the effort to correct is minimal what's the harm?
That mentality at Boeing as, literally, costed many lives.
The harm is that nobody knows why there's a memory leak requiring a reboot (or if it's even a memory leak). What happens when that very same issue is combined with a rare case and causes the death of hundreds of people?
"Have you tried to turn it off and on again" may be fine for a $20 Internet-of-insecure-and-shitty-Thing bought on alibaba. Not so much when lives are at play.
> This condition is caused by a software counter internal to the GCUs that will overflow after 248 days of continuous power. We are issuing this AD to prevent loss of all AC electrical power, which could result in loss of control of the airplane.
> A simple guess suggests the the problem is a signed 32-bit overflow as 231 is the number of seconds in 248 days multiplied by 100, i.e. a counter in hundredths of of a second.
So design systems that does not exceed those resources.
This seems sane given that the planes don't operate for 51 days constantly (I'm not in aerospace so please correct me, it seems a reboot could occur with refueling without issue)
Commercial aircraft need continual software updates to operate. They are, in a sense, living, breathing machines. Things like navigation and terrain databases are updated inside of 30 days.
Adding a scheduled reboot is one more item on a checklist that was already being run through.
It's counterintuitive, but performing a reboot as a scheduled maintenance item is far more risk averse than going in and touching code that has been otherwise thoroughly tested and signed off by regulatory authorities.
The chances of introducing a new bug when attempting to repair the former presents additional risk to what amounts to a convenience issue.
Also in case of emergency, eg after a power loss or whatever, you might have to do a reboot anyway. So you might as well make sure that this code path is well exercised.
I'd rather deal with a ground hog day of the system being for the millionth time in its first day of operation, than dealing for the first time with the system being in its millionth day of operation.
Having had to migrate a 12 year old dying server this weekend, yeah, I was 24/7 strongly cursing the idiot who didn't document anything[0]. On the plus side I did get to update a bunch of stuff to more modern practices.
[0] You will not be surprised to learn that idiot was me.[1]
[1] My other servers are much better - anything that hasn't yet been properly service'd has its own `RUNME.sh` which runs whatever it is in the correct way.
But equally, they dont do this 24x7 - if only because airport curfews and maintainence schedules won't let them.
Rebooting the computer when doing regular maintenance is no big deal.
> Usually planes are turned around too fast to be waiting for them to fully reboot every time they fuel.
To be clear, this affected the Boeing 787, a plane usually focused on long-haul between medium sized cities. It is incredibly rare to see a long-haul flight turned-around immediately. Normally, they have max two flights per days, and for longer routes, just one route per day. There was plenty of time to reboot. I don't think anyone was ever in danger.Also, I am starting to grow tired of "anything Boeing does is bad" on HN in the last 6-12 months. The Boeing 787 was a huge hit, both technically and commercially. (I would say the same for the Airbus A350.) I certainly never worked anything as important or cool in my career. The endless booing from the HN peanut gallery adds little new and/or useful information to the discussion. Yes, I expect to be downvoted for this last paragraph.
"Let's build something that we KNOW will catastrophically fail, because we deliberately ignore to take account limited resource availability of that system."
For a critical systems, that's just lazy and unacceptable.