Perseverance rover uses same PowerPC chipset found in 1998 G3 iMac
macrumors.com
macrumors.com
The CPUs used in space applications typically lag ~20 years behind the consumer market due to the need to be very sure that the kinks have been ironed out, and the need to produce variants that can withstand the environmental conditions outside earth's atmosphere.
EDIT: Also worth pointing out - the price of one of these specialized chips is in the six digits, due to the higher tolerances and low production volumes.
The single highest factor in all space design is cost per kg lifted. The only thing that is more costly - lifting it successfully and then having it break.
Heavier rover? You'll need more fuel for the skycrane's engines. And that fuel adds weight, so don't forget that you'll need more fuel to lift the extra fuel.
Now your rover and skycrane are bigger and heavier. Do you need a bigger parachute? Does everything still fit in the aeroshell? Is the fuel capacity for orienting and stabilizing the aeroshell still sufficient with the increased mass?
It's actually a chronic problem with Voyagers 1/2. The teams that maintain them regularly experience weird, unplanned behavior because a cosmic ray bit-flipped some register or other, and then they have to manually figure out what erroneous command got executed and how to recover from it, and those are extremely simple systems compared to what's running on Perseverance.
So yeah, you need radiation hardened parts in space or you're just begging for a dead probe.
Case in point: https://www.jpl.nasa.gov/news/engineers-diagnosing-voyager-2...
"Engineers successfully reset a computer onboard Voyager 2 that caused an unexpected data pattern shift, and the spacecraft resumed sending properly formatted science data back to Earth on Sunday, May 23. Mission managers at NASA's Jet Propulsion Laboratory in Pasadena, Calif., had been operating the spacecraft in engineering mode since May 6. They took this action as they traced the source of the pattern shift to the flip of a single bit in the flight data system computer that packages data to transmit back to Earth. In the next week, engineers will be checking the science data with Voyager team scientists to make sure instruments onboard the spacecraft are processing data correctly. "
Detecting complete blackouts is fairly easy, after a while when a watchdog tickle is absent for too long, but bad output due to flipped memory? Consensus-style systems can be used, but they are also a complexity in itself, and the point here was to replace a hardened single point of failure with redundancy.
This. Worked in physics department in the 90s that built satellites, and all the chips were late 70s or early 80s chips. Programming them was "fun".
Those old macs run Ubuntu just fine :)
Feature size is super important in radiation exposed environments because resistance to bit flips from a stray proton (or other charged particles flying around in space) is inversely proportional to the size of your transistors. So small transistors mean they’re more likely to be activated by stray charged particles causing all kinds of interesting and exciting problems (which you don’t want to be debugging from one end of a 12 light minute comms channel).
Yup. And while the transistor sizes haven't really shrunk by the 35x that this suggests, they've still shrunk by a lot. And, of course, the area has shrunk by this amount squared.
> causing all kinds of interesting and exciting problems
Sometimes super exciting. Latchup can turn the whole package into a white-hot parasitic transistor connecting Vcore to ground at low impedance. Rad-hard variants can largely eliminate this possibility.
Compare the PPC 740, which was ~300,000 transistors/mm^2 at 260nm. Apple M1 is ~130,000,000 transistors/mm^2 at 5nm. (260/5)^2 =~ 2700x, but the actual difference is ~430x.
So it's not as bad as line widths, but it still isn't quite keeping pace.
In plain English: an internal short-circuit.
Keeping the working data synchronized would be the real trick. One could imagine all of these CPUs are hooked to a single bank of redundant and ECC stabilized memory, and all access goes through the watchdog processes which will only let through the ones that are in agreement.
The end result could be a system that is lighter and faster than the traditional RAD hardened system simply because it's built on a smaller process. The downside is the enormous complexity of the watchdog systems, they would be very expensive to get right. Also, synchronization is one of the hardest problems in computer science. It's basically cache invalidation on steroids.
This is a common set of techniques in critical systems. The smallsat built by students that I mentored didn't have voting, but it had processors with moderately sophisticated watchdogs performing mutual-power-monitoring and simple hardware failsafes. Aerospace control systems often run in lockstep and have voting, etc.
> Keeping the working data synchronized would be the real trick. One could imagine all of these CPUs are hooked to a single bank of redundant and ECC stabilized memory, and all access goes through the watchdog processes which will only let through the ones that are in agreement.
Typically you make the memory redundant too, and just ensure that input and outputs are common and all code is deterministic. e.g. There's Tandem / HP NonStop which use these techniques.
> The end result could be a system that is lighter and faster than the traditional RAD hardened system simply because it's built on a smaller process.
Handling upsets by voting makes a lot of sense. But radiation can cause permanent damage of small geometry circuits. And even with fast mechanisms to crowbar power in the event of a latchup, you're not really sure that you'll save the day.
It's a great way to handle things on the cheap for cutting edge payloads and research projects, though.
Tandem was software fault tolerant - the Stratus machines were fully hardware fault tolerant and ran code in parallel on seperate logical cpus to test for faults. Truly neat machines. Also, it was a PL/1 based machine - even the OS.
12(?) Motorola 68K chips were wired into multiple 'logical' cpu's, and programs executed in parallel on each. If a discrepancy arose, majority won and 'wrong' logical cpu's were taken out of service. It would also automatically call Stratus's remote service, report the failure and order a replacement board :-)
The node sizes used in marketing today haven't actually had any relation to feature size for at least 10 years. You can read some of the actual feature details for 7nm on wikipedia.
https://en.wikipedia.org/wiki/7_nm_process#7_nm_process_node...
If "feature size" is the width of a fin, a 7nm feature will typically be ~15nm across and spaced ~15nm apart (ie the center of one transistor is 30nm away from the center of the transistor next to it).
It's worth noting that the actual thickness of a transistor isn't strictly the most important factor. Which is why that wiki page shows fin pitch (the width of the transistor plus the space between transistors), because that's the metric that you get transistor density from.
The first version of the rPI was pretty close in specs to computers from that time and it works fine - though the G3 mac shipped with way less RAM by default. The first gen rPI on my desk even boots into a usable desktop.
There's a lot of bloat in modern software, such as browsers, but most is in the form of code/features that are rarely used.
Maybe someday we'll see electron versions that can be stripped to stuff you actually need.
Yeah, you'd probably want mosh;)
I've done embedded work on MCUs with less horsepower than you probably have in your headphones.
Sure it's not going to run windows 10, but it'll reliably control that device until the mechanical parts wear down. Same applies here, where the advantage of huge transistors means much greater resistance to cosmic rays, and a 20 year old design means all the bugs and quirks are well understood.
The latest and greatest is nice, but not always the right choice for the problem at hand when you have the option of custom firmware and you aren't connected to a hostile internet.
MCUs are awesome. Given a 100mhz single core M3, half a meg of SRAM, some serious fun can be head.
(honestly 512K SRAM would be insane, but doable now days)
People don't realize that with a purpose built runtime and well crafted code, 100mhz is a LOT of power.
Modern OSes, complicated software stacks, preemptive multi-tasking, and the need for flexible general purpose usage, eat up a huge percent (I'd say well over ~70%) of the usable computing capacity in the devices we deal with every day.
Now I write web services that spend more CPU and RAM decoding JSON than entire runtimes I helped design. Go figure.
Of course productivity is a bit higher, having to design and code the most optimal binary serialization formats for every single task takes a lot of time and effort, but wow is the end result fast and memory efficient!
Working with a 50 MHz ASIC on an old process with 256K of SRAM shared between two different companies and codebases, now that was fun.
I think a big part of it is security. Having to isolate address spaces is expensive.
(Sometimes I like to imagine being able to run a unikernel OS where applications are written in something like safe Rust, everything runs in the same address space, and the OS delegates enforcing memory separation to a trusted compiler. So, things like writing to disk or the network are just regular function calls with no context switch overhead. In this model, unsafe code is treated sort of the way we treat kernel code or setuid binaries in Linux: regular users aren't allowed to run arbitrary unsafe code. With Spectre, I don't know if this sort of approach is still feasible on modern processors, or if it's possible for a compiler to guarantee that the compiled code is not susceptible to any known side-channel attacks.)
Better than that, static allocation. Embedded typically doesn't allow malloc anyway, so give me a platform that has the world's simplest MMU, static bounds checks that are put into write once read many memory upon boot. :-D
(Of course, the rules are totally different with BSD and Linux, where you generally want the newest software that supports the architecture and it works great as long as you choose your packages wisely.)
Which, in this case, include radiation resistance, which is much harder, if not impossible, to achieve with modern chips with very tiny transistors. Large transistors can much more easily shrug it off.
The old chip is not just "good enough". It's the only thing that works.
So far it worked fine for their near Earth applications and it will be interesting to see how they will handle long term deep space in this regard with starship. I guess they could still fly a rack encased in lead given the scale and performance of Starship though. ;-)
http://wordpress.mrreid.org/2012/08/25/curiositys-rad750-rad...
The key risk vector in semiconductors in a space environment is things like bitflips. So, when a high energy particle collides with your transistor in DRAM and a 0 becomes a 1, that is a problem. There are deep architecture changes made to the RAD 750 specific to deep space missions, for example SRAM is used instead of DRAM, as SRAm is less succesptable to bitflips from high energy particles.
There are other alternative RAD CPUs, such as the Leon. But ultimatley, a space vehicle does not need a lot of compute to have a successful mission. It runs VxWorks, a real time operating system. And a far, far, high risk of failure comes from bit flips, rather than having a bunch of compute for no reason
https://spectrum.ieee.org/automaton/aerospace/robotic-explor...
It's a great 2 for 1 thing that NASA typically does. They'll win even when they fail because they learned something and they won't fail that way again.
I remember reading that it only takes a couple milliseconds to reboot so they're not worried about it moving very far while they do.
For secondary purposes they have no problem running more modern silicon. E.g. the Ingenuity helicopter, as a last-minute addition (and secondary mission) uses much newer hardware and software: https://spectrum.ieee.org/automaton/aerospace/robotic-explor...
> There are some avionics components that are very tough and radiation resistant, but much of the technology is commercial grade. The processor board that we used, for instance, is a Snapdragon 801, which is manufactured by Qualcomm. It’s essentially a cell phone class processor, and the board is very small. But ironically, because it’s relatively modern technology, it’s vastly more powerful than the processors that are flying on the rover.
EDIT: Found the article again: https://spectrum.ieee.org/automaton/aerospace/robotic-explor...
I'm not sure where I got the reboot thing from, I think HN in a thread after the Perseverance landing
Does it try to keep state across such a reboot? Is the first course of action to try to deduce whether it's about to crash? Is every sensor reading done multiple times to reduce risk of reading wrong (typically, such sensors are basic and doesn't employ error checks or such)? Does it try to understand why the OS crashed and do any graceful degradation, eg "fine, don't use that sensor then"?
That might miss bit flips that don't affect the flight profile, but those do not need a real time response. Bit flips in collected data for example could probably be found afterwards when the data is analyzed.
Heck, it might even make sense to figure out what the minimum time, T, is such that something can go wrong that a reboot will fix fast enough to save the mission but only if the reboot happens within T seconds of the onset of the problem. Then just reboot every T seconds regardless of whether you know of something actually wrong.
Neither. The proper question to ask: "is it the right tool for the job?" The answer here is yes. Power is a vetted architecture and has been used in critical processes for quite some time.
Just because something is old does not mean it is bad or some sort of burden.
Also, spacecraft a little distributed systems. There are _lots more_ processors on these things. The instruments and mechanisms all have their own processors, from what I see ... which is limited.
Don't think too hard about how Percy is self-driving on its limited compute, that will make your head hurt to think about trying to make that work.
Think of it this way: technology is not a straight-line tech tree from "less advanced" to "more advanced." It's more about tradeoffs.
So these processors might have less processing speed compared to modern consumer ones, but they're more reliable and resistant to radiation (and that's partially a consequence of the reduced speed).
It's the only thing that works, so it's good by definition.
Modern chips with tiny transistors can easily get fried (or make gross mistakes) because of cosmic radiation. Large transistors are much more robust.
1: http://mirror.macintosharchive.org/developer.apple.com/perfo...
https://www.extremetech.com/extreme/278160-nasa-switches-cur...
You can start poking around here: https://www.nasa.gov/content/commercial-lunar-payload-servic...
There's actually quite a few options in the "next gen" space, as many observed the need for updated compute. The one I've been closest to is HPSC: https://www.nasa.gov/directorates/spacetech/game_changing_de...
The CRS-10 that fly to the ISS seem to use RadHArd Cortex M4 or M0, according to this ARM PR document:
https://www.arm.com/blogs/blueprint/arm-rad-hard-chips-space
WARNING: this was a lot longer and 'scattered' than I recall, apologies in advance.
DISCLAIMER: This was a casual talk and I casually took some notes, and it happened over three years ago as of this writing. In short, don't assume anything below is an accurate representation of what Jinnah said. As a long-time SpaceX fan, and as a much longer-time software engineer, I was super pumped during the whole talk, and was definitely not focused on accurate recording.
https://www.realms.org/spacex-talk-notes.html
PS: I know most of this veers way off topic, but for some reason I decided to take this opportunity to share this material.
PPS: The comment was too long to post, so you'll need to click the above link. Sorry about that. If someone would like to distill the on-topic parts into another comment, that would be fine.
> They saw that at NASA. Because of the handoff, it caused them to push reliability ratings way higher than needed, which added complexity and customization.
Very interesting!
also
> it took Elon six weeks to go from "OK let's land on a boat" to actually having a boat. The software VP guy didn't think it 'was a 2014 problem...definitely 2015'...people who had been there longer knew Elon would be able to get a boat quick. Definitely a 2014 problem.
then this
> So then they'd fly balloons off of the boat so they'd get low level wind data, said data would go to the vehicle and be used for control for the last few km.
So many gold nuggets. Thank you!
You're welcome! I'd forgotten most of this and so it was really fascinating reading over it again just now, especially in light of all that SpaceX has been doing over the past 3.5 years.
>Facebook got jammed up because they had to re-write a lot of software under DO189B/C.
I typed this up 3.5 years ago, and didn't give a crap about Facebook one way or another at that time, so whatever was said there didn't stick with me at all.
Hopefully somebody else where can somewhat decode mystery string (that might have typos) based on context.