Uptime 15,364 days – The Computers of Voyager [video]
youtube.com
youtube.com
Is the uptime really technically true? Sure, the Voyager has been operating for 40+ years, but all embedded systems must have watchdog timers. And given how hostile the space environment is, I'll be surprised if the main system hasn't been reset by a watchdog timer for a couple of times to recover from fault conditions, thus the actual uptime must be much less than that.
"Practically all of Voyager's redundancy is gone now, either because something broke along the way or it was turned off to conserve power. Of the 11 original instruments on Voyager 1, only five remain"
Though I did wonder if they updated the firmware/code in all this time and als no definitive answer stood out: https://www.quora.com/Was-the-opportunity-to-update-the-Voya...
https://voyager.jpl.nasa.gov/news/details.php?article_id=16
It suddenly started returning corrupted frames. So engineers had it slowly transmit a core dump, and found a single flipped bit. To fix the problem, they "reset" the computer, which I assume means they rebooted it.
> Most spacecraft have more than one Command Loss Timer Reset for subsystem level safety reasons, with the Voyager craft using at least 7 of these timers.
I was not implying that "the fact that the system has been rebooted lessens the achievement", I said none of it. I was just wondering about details of the system and technical accuracy of the statement, isn't it the point of posting on HN?
Running a probe for 15,364 days without even a single bit flip or poweroff would be an extraordinary miracle that exceeded all reasonable expectations, not simply the grestest accomplishment.
Please don't assume that every technical statements/questions imply undervaluation, criticism or attack, regardless of how common they are in tech.
Thanks.
To make it more difficult, what would be the safest uptime for a box that allowed remote logins? SSH flaws don't count, since you can always upgrade that on the fly, but kernel-level privilege escalation weaknesses would count as critical.
I would imagine something small and stripped down serving a particular purpose would fare better. And OpenBSD prides itself in security for general purpose computing, but it still gets regular security fixes.
I’m not entirely sure how you define uptime for these machines if none of the original parts are still there.
The question is then where you define the boundaries of a system and it's uptime. At least from my recollection for mainframes they defined uptime based on the execution of batch jobs and availability of services not the OS/Hardware which if it crashed often involved Big Blue coming to investigate WTF happened and how it happened since System Z machines are designed with so much redundancy that you can swam RAM modules without interrupting the workflow.
Today with RAIM (RAID for Memory) IBM System Z machines even support an entire memory channel dying without interruption.
I think the 'allows user access' bit is the thing that would limit safe uptimes the most, since once you're logged in to a box, the kernel surface area for attacks is much larger.
The hardware might be sleeping for 99% of it's service life, and is typically over-engineered. They could run into the next century if corrosion doesn't get them.
The lack of listening ports might shrink the attack surface, but a malicious endpoint that it connects to might be able to confuse it. (or perhaps some evil MITM attacker)
1. Wiznet produces chips where the entire network stack is implemented in hardware, there is no code. Malformed packets will never touch my code. http://shop.wiznet.eu/chips/w5500.html However, this is a minority of devices.
2. These devices typically have no dynamic memory allocation at all, they will work with a circular buffer of fixed size. Normally they can't run out of memory.
3. You might be able to clog up their CPU, however they typically use an RTOS, those operating systems will place hard limits on amount of CPU time different components of the system could take. At best, you will prevent the device from sending out anything over the network while you are attacking it.
4. Of-course, if there are mistakes in the TCP/IP stack, you might cause issues. Most of them are using lwIP stack, I am not qualified to comment on it's security. https://savannah.nongnu.org/projects/lwip/
Although, in-place upgrades haven't been around all that long in the grand scheme of things. Perhaps there's a box out there that predates in-place upgrades and is still running securely?
In a smallish embedded system there are not too many (software) difficulties in keeping the software running virtually forever.
Getting the hardware to keep running without any fault for 42 years, on the other hand...
> The Viking CCS had two of everything: power supplies, processors, buffers, inputs, and outputs. Each element of the CCS was cross strapped which allowed for “single fault tolerance” redundancy so that if one part of one CCS failed, it could make use of the remaining operational one in the other. [1]
Modern systems like that of curiosity rover also use hardware redundancy, (triple redundancy even), but I believe this happens at a much higher level, i.e whole computer.
[1] https://www.allaboutcircuits.com/news/voyager-mission-annive...
Seriously, at this point I have multiple browsers open, multiple tabs, multiple programs split over multiple desktops with my workflows.
But more seriously an encrypted drive that has a lot of data on it that I've forgotten the password for, and can't remember the system I used to encrypt it / set it up, and figuring out how to change it is always pushed to future LandRs problem. A restart and I'm screwed!
I'd basically just give up and go live in the woods.
"We should forget about small efficiencies, say about 97% of the time: premature optimization is the root of all evil. Yet we should not pass up our opportunities in that critical 3%. A good programmer will not be lulled into complacency by such reasoning, he will be wise to look carefully at the critical code; but only after that code has been identified. It is often a mistake to make a priori judgments about what parts of a program are really critical, since the universal experience of programmers who have been using measurement tools has been that their intuitive guesses fail." - Donald Knuth
It's not "optimization is root of all evil". The key is "premature optimization". Maybe people gloss over that part, but it is right there.
Yes, Knuth goes into more detail on what he considers premature optimization in the context of programming computers. However the short sentence applies much more broadly in my experience.
For example, "premature optimization" of BOM costs in a hardware project can cost you dearly down the road when it turns out that leaving in some extra flexibility in the design would be mighty useful.
Premature X is bad.
Overusing X is bad.
These are true for most X. If it's not bad, then you didn't do it prematurely or overuse it!