The team had recently seen similar failures in simulation, and was able to quickly decide it was okay to proceed.
I believe this is incorrect. From an earlier HN discussion:
"The 1202s were also a lot less benign than is often reported. They occurred because of the fixed two-second guidance cycle in the landing software. That is, once every two seconds, a job called the SERVICER would start. SERVICER had many tasks during the landing. In order: navigation, guidance, commanding throttle, commanding attitude, and updating displays. With an excessive load as caused by the CDU, new SERVICERs were starting before old ones could finish. Eventually there would be two many old SERVICERs hanging around, and when the time came to start a new one, there would be no slots for new jobs available. When this happened, the EXECUTIVE (job scheduler) would issue a 1201 or 1202 alarm and cause a soft restart of the computer. Every job and task was flushed, and the computer started up fresh, resuming from its last checkpoint. It was essentially a full-on crash and restart, rather than a graceful cancellation of a few jobs. And unlike is often said, the computer wasn't dropping low-priority things; it was failing to complete the most critical job of the landing, the SERVICER.
Luckily, the load was light enough that of the SERVICER's duties, the old SERVICER was usually in the final display updating code when it got preempted by a new SERVICER. This caused times in the descent when the display stopped updating entirely, but the flight proceeded mostly as usual. However, with slightly more load, it was fully possible that the SERVICER could have been preempted in the attitude control portion of the code, or worse yet, the throttle control portion. Since each SERVICER shared the same memory location as the last one (since there was only ever supposed to be one running at a time), this could lead to violent attitude or throttle excursions, which would have certainly called for an abort. Luckily, this didn't happen -- and the flight controllers didn't abort the mission not because 1202s were always safe, but because they didn't understand just how bad it could be, were the load just a tiny bit higher."
I believe the volume is so low it does not warrant the investment to make radiation hardened versions more often.
How much more likely? Are the odds all that high over the mission duration?
Most of my knowledge of the subject is from a friend who did his PhD dissertation on it; specifically, triply-redundant adder circuits for single-bit operations, allowing for some rad-hard chip designs using regular foundry processes.
Is it like the situation high in earth's atmosphere (like using your ipad on a commercial flight), or would it be more like on the moon with no protection?
The software stack is probably not ready for that since the plan is to abandon it after its mission. And the Zigbee protocol is rather limmited.
Great talk about it from the FSW Workshop in 2019.
I hope they installed the Home Assistant Core docker container on Perseverance. Gotta get those sweet dashboards.
https://en.m.wikipedia.org/wiki/Radiation_hardening
That's why they don't use the latest and greatest in the space.
Therefore boot time becomes critical - if you end up rebooting due to bit errors multiple times per second, you can't afford to wait for Linux to start up each time...
I wonder how one could use micro kernels to further improve startup time and have a mini distributed OS/kernel for each component.
You still have 10% the surface area, power usage and weight and 10 times the speed of the radiation hardened ones.
- Single Event Burnout, SEB
- Single Event Gate Rupture, SEGR
- Single Event Latch-up, SEL these can be recoverable
In addition there are also Total Ionizing Dose (TID) Effects https://radhome.gsfc.nasa.gov/radhome/tid.htm
Just an example: consider you were to add some extra electrical charge to the gate of a transistor (from an electron or ion beam, I don't know).
A larger transistor has a higher gate capacitance, it is therefore quite immune against a few extra charges. On smaller transistors, though, it could dramatically increase the voltage, leading to a bit error, or destroying the transistor. Capacitance in this case is proportional to the area.
Higher density also has some inherent drawbacks against particles, since a damaged part is proportionately more damaged if it is smaller.
More ancient processes are also higher-voltage, and higher-current, so they can handle a lot more noise on these signals.
Surely we could still make a chip today with the same transistor size as one from 2001, but better in other ways.
In microelectronics, costs are directly proportional to the die area. Actually, they might rise faster due to yield issues.
Making masks [1] is expensive. The bigger the mask, the more expensive. A mask set for a modern CPU can easily be in the millions, I think. And it gets more expensive with size.
Then you have yield. The bigger the chip, the more likely it has some defects (due to dust or other issues during fabrication). Often, processes work more or less well across a wafer: temperature higher in the center, etc. That can affect performance.
Due to yields, bigger chips have to be scrapped more often, and are generally less performant. Binning (selecting the fastest, slowest, more efficient, or chips with specific intact features across a wafer) is less effective. You might have to add redundancy or mechanisms to cut power to damaged areas to avoid short circuits.
Now, that's why we don't generally make bigger integrated circuits. Now, we could make bigger integrated circuits with today's latest clean rooms and equipment to try to raise yields. I don't know if that's being done already, but it would likely raise costs. On the other hand, progress is being made on bigger chips as well [2].
Another more promising direction (IMO) is to use chiplets like AMD does it. You could use more of these for a bigger virtual size.
Now, like I wrote, a lot of the performance improvements actually come down from physically scaling down the transistors: if the gate is smaller, the transistor needs less electrons to charge up. That means faster transistors, and less energy. Also, transistors are closer, so signals reach the next one faster [3].
If you want bigger chips at a previous technological node, you are going to need a huge heatsink, or disable part of the chip ("dark silicon") [4].
The real answer might come from completely different architectures, based on light or spin, or more power-efficient circuit/computer architectures like with adiabatic computing [5] (or non von neumann based, closer to what I do).
Power efficiency is key, since that's the limiting factor for performance nowadays (ask any overclocker: you don't want to melt your CPU. Also, rovers have a small energy budget). With better efficiency, you have room to grow performance again.
Paradoxally, software seems headed in the other direction, generally speaking.
[1] https://en.wikipedia.org/wiki/Photomask
[2] https://news.ycombinator.com/item?id=20739408
[3] https://en.wikipedia.org/wiki/Dennard_scaling
Will be interesting when missions are being launched utilising the tech of today, like rad hardened neuromorphic chips https://www.businesswire.com/news/home/20200902005406/en/Bra...
Also Linus finally got his progeny to Mars. That’s a pretty cool accomplishment:
“What have you built?”
Well, I invented one of the most prolific operating systems the world has ever seen, but not just this world - there is a helicopter on Mars that is flying due to the seeds which I planted that day...”
What have you built?
Good rundown of how it was built: https://www.youtube.com/watch?v=GhsZUZmJvaM
On top of that, I'm sure JPL would love to move with the times, which is why the helicopter is using the more modern processor.
Source: https://orbitalindex.com/archive/2021-02-24-Issue-105/#ingen...
I've been told LinuxBIOS is able to get you a text console login prompt faster than hdd platters can spin up, but it takes Ubuntu tens of seconds on my SSD laptop to get me a login prompt.
I'm surprised they don't have an FPGA MEMS autopilot with a simple degree 2 or 3 polynomial model of the flight dynamics , with the CPU and Linux only being involved in making adjustments to the autopilot. Or, maybe that's what they're doing, and by "falling out of the sky" what they really mean is the autopilot drifting.
absolutely performance. Yes to the other three for sure, but the engineers reported that there was no way they were running flight control using image tracking on a 200MHz CPU.
I just don't get in what world you think military munitions are using CV for targeting bombs.
In more modern weapons, imaging IR sensors are well-established for terminal guidance on missiles like LRASM, JASSM, or NSM to distinguish targets from clutter and identify specific target features (specific parts of a ship, for example). Of course "traditional" "IR-homing" SAMs and AAMs now use imaging sensors (often with multiple modes like IR+UV) to distinguish between the target and decoys/jammers. Even your basic shoulder-fired anti-tank missile like Javelin requires some amount of CV to identify and track a moving target.
aka edge detection :) I don't remember if it was the Sidewinder or Walleye that eventually dropped in a CCD (or both), but I know that the Maverick (which is technically older than Walleye) got along without a CCD until the GWOT - when it finally upgraded. The Javelin actually beat Maverick in that regard, having a 64x64 sensor 10 years earlier - able to handle scaling and perspective change for the 2-d designated target pattern.
Would be interesting to know cost allocation for that Sony/Samsung chip. Project manager and scientists/engineers could have listed design challenges for industry player like TSMC/Apple go beat with an offering, a short run of 20 chips specifically for this helicopter.
Reliability is very important and space is harsh. I assume on the surface the radiation levels are low enough for Earth systems to work (maybe playing a bit with voltage/clock frequency helps, not sure how much shielding they can add, probably not too much).
NASA should be about long term and expensive projects, SpaceX is just a tool to achieve that goal which is to do new science regardless of whether it is in the air or in space.