Compiling an application for use in highly radio-active environments
stackoverflow.com
stackoverflow.com
So far this has luckily remained a research area in practice, because the vendors tend to do a good job of hardening their machines against errors as they get larger (most of the gloom and doom predictions take the error rates of current hardware and extrapolate). It will be interesting to see if it remains that way.
Relevant article: http://superfri.org/superfri/article/view/14
Given enough bits, random errors that individually are almost impossible become almost a certainty.
If the drives can't be trusted, then you'd need some kind of error correction in software, I'd presume. I thought the drives should auto-correct any read errors.
I'm not talking about in a radioactive environment, just normal usage.
In the lab, I accounted for every possible valid state with a transition out to another state.
I left my microcontroller running when I went to get lunch, came back and it was frozen. I did a dump, and discovered it was idling in a state that has no entry into it, and because I had no exit, it would just sit there.
I 'fixed' the problem by adding a transition to reset the controller from every state that shouldn't be reached.
Silently jumping back to a default state is sometimes useful, but could result in unpredictability or perhaps silent data corruption for some critical applications.
Most of the answers are talking about spaceflight, because it is one of the only radioactive environments that has a particle distribution that can be effectively fought.
Most earthbound environments have particle distributions that are practically unsolvable in software, and should instead be solved by using mechanical and noise tolerant analogue methods of sensing and manipulating and then massively shielding your computation (normally by controlling from a different building).
*Note. By particle distribution I am refferring to both rate and charge
Another place to look is in safety critical systems. IEC 61508 in general and the automotive variant xxx26262. There are several dual core processors available now that run two copies of the code and check for any errors. That doesn't tell you which one failed, but it does catch the error. There are methods defined for creating fault tolerant systems too.
We've been doing this shit for years - your electric steering system, brakes, even throttle control (with certain exceptions of course ;-)
The funny thing in cars is that most of those systems have mechanical backups or consider "shut down" an undesirable but reasonable failure (vs steering left or dumping brake fluid). Your car can coast to the side of the road under driver control without any power. All this self driving stuff will require fully redundant components for some of these systems (including LIDAR or cameras) to really be safe.
All of that redundency makes it possible to build some really robust flight envelope systems that keep the airplane within safe margins.
* I should note that all of this applies to the A320 family, the systems have been developed even further in recent years. For example with the A350 Airbus made some steps towards allowing the flight computers to be used in Simulators so that the same software/hardware as on the real plane can be used.
> they use different software and hardware
Does this mean different architectures? If it does my respect for the redundant-hardware approach just went through the window.
Also, how is the watchdog redundant? I can't imagine there's only one; how does this work? Are both watchdogs somehow wired in parallel, are they cross-connected to each other, or...?
Also visible in the diagram is that each side has its own watchdog, both connected to each other. The way this whole thing work is fail safe, so if one computer fails the backup can jump in, and if that fails too the flight controls will either retract or stop in their current position depending on what makes the most sense. It’s also mirrored, so for example if spoiler 2 on the right wing fails and is retracted, spoiler 2 on the left wing will also retract.
Here is a description about the flight controls and how pilot input gets passed through to the control surfaces: http://www.smartcockpit.com/docs/A320-Flight_Controls.pdf
And here is a general overview about the architecture: http://www.skybrary.aero/bookshelf/books/2313.pdf
(I wonder if there are any systems built on multiple architectures where each unit is itself a redundant system with CPUs in lockstep.....)
Am I to intuit from this diagram that the watchdog watches all the components - power, I/O, memory, and CPU? That's very impressive. Or does it watch a central bus/backplane everything is connected to?
Also, how does either side decide/figure out the other side has failed? Simply deciding that the other half is wrong if it doesn't match this half's output could fail catastrophically if one of the sides reaches this conclusion after entering an invalid state (ie, it's the other side that is correct, and this side is wrong).
I'm also mildly curious as to why the I/O on one side has two connections to the actuators, while the other has only one.
The fourth computer's software was only sufficient to abort the mission and return the shuttle to the Earth. IIRC, they never had to use it.
"A fourth AI is purposefully kept dormant. He's a bit single minded and his only purpose is to find a planet and touch down. If the other three ever can't agree or have a moment of clarity in which they realize they've become unstable, they deactivate and activate him. He regularly reloads from scratch, forgetting his previous incarnations and spends a majority of time validating the coordinates his previous incarnation left for him, making course adjustments, and ensuring the humans don't prevent him from saving them..."
I would read that book...
The crew are just caretakers: the ship is controlled by a disembodied human brain, called "Organic Mental Core" or "OMC", that runs the complex operations of the vessel and keeps it moving in space. But the first two OMC's (Myrtle and Little Joe) become catatonic, while the third OMC goes insane and kills two of the umbilicus crew members. The crew are left with only one choice: to build an artificial consciousness that will enable the ship to continue. The crew knows that if they attempt to turn back they will be ordered to abort (self destruct).
The odds of a simultaneous (within one hour) double failure is the square of that, or 1/763760000 per hour. This corresponds to a MTBF of roughly 3836880000 hours or 438000 years.
438000/10000 is 43.8, so 2.3% chance over 20 years.
It's a longshot but it's not never.
[1] http://web.mit.edu/airlinedata/www/2014%2012%20Month%20Docum...
[2] https://en.wikipedia.org/wiki/List_of_most-produced_aircraft
So yes the odds of any one plane experiencing that problem are absolutely tiny but across the fleet not so much.
CERN for example has a lot of material, a quick search leads to e.g. https://lhcb-elec.web.cern.ch/lhcb-elec/html/rad_hard_links.... or the Radiation to Electronics workgroup https://r2e.web.cern.ch/R2E/
It's sponsored by IBM, A large Israeli college, A technological high school chain - And Several Israeli Technological Military Units (Air force, Communications, Cyber). So yeah, that's probably a really good indicator you might be on to something.
The microwave clock knows it's a microwave because it controls the microwave. It would be wasteful to put extra layers of abstraction in such a small system.
Likewise, a system designed for radiation tolerance will expect a close coupling between software and hardware; the whole point of the software in that case is the hardware.
My perspective on this is informed by work on ABFT (for more on that, see https://www.computer.org/csdl/trans/tc/1984/06/01676475.pdf, and the literally 1000 later papers citing it) -- you can design a version of the basic linear algebra subroutines that have fault-tolerance built in, and use them without changing program source code.
For codes that spend a lot of time doing numerical computations (e.g., preliminary data reduction in a spacecraft on-board computer), ABFT is an interesting option.
The fist thing I learned is that any plan that I come up with to help mitigate risks always ends up being proven less effective than I initially thought. The idea that the hardware can flip bits or latch up behind your back is rarely a consideration for most software design work, so it really requires an entire new way of thinking. In the end, it's all about reducing probabilities, but sometimes your mitigation techniques can be borderline useless. This could be because of my relative lack of experience though. I'm sure there are people who have been doing this a while and have converged on a design process that works well.
Couple this with the fact that code in the failure path very rarely gets actually tested with _real_ failures during design, and you end up with situations where your entire test/precaution/recovery technique is rendered useless by a single oversight.
It kind of reminds me of how some bootloaders perform a DRAM test during power up. The only problem is that these tests are usually executing after the bootloader is already running from RAM itself. Sure, it's better than nothing and may help to find some obscure errors like a bad address line pin or something, but the issue here should be obvious - if there was a real problem with the DRAM, it's likely you wouldn't even get to that test, so now what?
Another example: suppose you have a PI bus connecting to an FPGA, and you are relying on the PI_done line to pace transfers. Now, a common test to verify the bus is working would be to read some magic word from the FPGA or maybe use an XOR test register, but what happens when there's an actual failure like a bad connection to the bus or bad power to the FPGA? Well, the CPU throws an exception and resets because the timeout for the PI_done expires. This all happens before your test has a change to even print anything to the console.
I run into these kinds of issues all the time when designing software routines to self-test the hardware on our devices.
Everything is all good when the hardware works, but as soon as hardware starts actually failing, all of those diagnostic tests are suddenly useless.
I've gotten kind of off topic here, but the thought process is what I'm getting at.
It usually goes a little something like this:
Possible solution: Always have a backup OS image. Caveat: There's only one copy of the bootloader, so if that area of the flash gets screwed, we're done.
Possible solution: Scrub the contents of RAM to check for errors. Caveat: The scrub routine is running from RAM.
Possible solution: Use a watchdog timer to force a reset if the CPU runs away. Caveat: What if the watchdog timer latches up and the CPU hard locks? (Unlikely, I know).
In the end, you can do a lot in SW to make a system more rubust, but the ultimate solution is probably on the mechanical side. You're better off spending your efforts modifying the enclosure and submerging the thing under water than you are spending hours designing in software techniques to try to recover from HW errors.
TLDR: Most people agree that software techniques for mitigating risks in a rad environment are all about reducing probabilities, but I feel that such techniques have a good chance of being less effective in practice than people may think. I don't know what the point of this comment really was though, other than being negative. :)
It's my understanding that you would need a very thick layer of metal to actually protect a device. Generally, lead works, but the thickness required would mean that your device would be extremely heavy. However, like you said, this isn't a problem for ground applications.
Also, apparently water works well too.
This is all honestly beyond my knowledge though. Our customers do their own qualification since these are COTS devices, so any failures may be handled on their end by making separate enclosures or using another vendor's system as a backup. They haven't actually asked us to implement any software based risk mitigation, but I investigated it on my own.
BTW thank you for excellent information in your comment.
I believe most of the studies were conducted by NASA, so incorporating that into a search should help.
Some experiments have used 2000 year old Roman lead to avoid this http://www.nature.com/news/2010/100415/full/news.2010.186.ht...
1) You'd need a lot of lead to totally shield the equipment, like meters thick. Of course partial shielding reduces the radiation load but unless you can lower it to a negligible amount, you still need to have radiation tolerant code.
2) We generally want to minimize the amount of material in a particle detector, especially at layers nearest the collisions (where radiation is most intense). These inner layers (called trackers) are used to track light charged particles, while the outer layers (called calorimeters) are used to absorb and measure energy. So we need less material throughout to ensure we get good energy measurements by the time particles get to the calorimeter.
In addition to using radiation hardened components, a lot of this is mitigated by doing as little computation as possible on the front-end readout electronics. Most of it is switched capacitor arrays, asics, and the occasional FPGA, which can be made fairly reliable since they're basically just responding to clock signals and measuring voltage (or something), and buffering these measurements. The raw data is then piped away from the detector at absurd bitrates to another part of the cavern (called the "counting room") that is much less radioactive and has lots of standard computers/electronics to do further processing.
So for mfgr mistake code testing you find a problem you give up or set the board on fire or anything to make sure it never gets shipped to a customer. That's not good for a spaceship or a nuclear reactor.
For outside opposition you can assume a non-zero BER on the telecom link either due to signal issues or radiation (acute or long term degradation) in the analog circuitry. However anti-attacker code is just a place to get flipped bits and crash. On the other hand a really naive string copier looking for a null with a flipped bit will never terminate, but you're watching memory address bounds anyway, so ...
The right thing for a desktop computer to do when the stack pointer or PC gets lost is to run off into outerspace and get powned or panic the system because who cares that should never happen on a shipped system on earth in a non-critical application. The right thing for an outerspace vehicle is populate "unused" memory with a jump to a recovery routine.
going off on a tangent related to your post:
Simplicate and add lightness. Interrupts exist because slow 1970s era processor tech couldn't keep up with 1970s IO tech or UI. In 2016 you poll. Your interrupt service routine can't get corrupt if you don't use one. Your "interrupts enabled" bit can't get cleared if you're not using it. Your stack pointer can't get lost if you don't use one. If you need to repoint the antenna every couple minutes and your program crashes every couple hours on random average, you power it up every couple minutes for a few seconds and then it can't crash or lock up when there's no power. Loops exist because in 1977 when we had 128 bytes of memory and wanted to load 8 bytes of data its cheaper in terms of memory to write a loop that reads one byte 8 times, but in 2016 you unroll all loops so if there's a single bit error in one read then "most of" the process works and a single bit error in the loop code can't crash it because there's no loop code. You're paying the mfgr for great gobs of memory in 2016, use it. From an engineering standpoint its like replacing crane mounting bolts the size of bratwursts with 6-32 sized hardware from walmart because thats how we rolled in 1977 and ain't nothin gonna change that and at least sometimes in desktop simulation it works.
The stackoverflow article displays perfectly why I don't use stackoverflow. 90% of the comments are people talking about how the problem is no fun, weird assumptions about the asker, blind cargo-cult style devotion to paperwork and tradition and passing the buck to other agencies, or complaining (incorrectly) that if its not perfectly solvable then no effort should be made. If you want to know the option flag to sort the output of GNU ls by modification time, then SO works, but SO is useless for all serious questions. SO is great at answering the questions I don't need to ask and sucks at answering the questions I need answered.
Stupid question: would it be possible to burn the OS + deployment in a ROM, kinda like the old 680x0 macs? Or are those just as vulnerable to radiation as RAM?
At any rate, it seems like designing in this environment would require a level of redundancy not unlike what's in the human brain, to store & correlate information. Very fascinating stuff!
If you're working out of a lab, once the testing is complete, what happens to the hardware? Is it irradiated and does it need to be disposed of safely?
In today's hardware you can run a CRC check on your bootloader prior to starting it using only registers. When you start it, you can check a portion of memory from the bootloader, then using that memory, check the rest of memory. If you use ECC memory and initialize it properly, it will do a machine check if multiple bits get flipped. Multiple systems doing the same calculations should get the same answer, using three systems and flagging any system that disagrees with the consensus will identify failing hardware. I/O channels can use ECC as well, and software that is running in ECC protected I/O and memory can use data protocols to detect data corruption in flight (channel errors) or in situ (storage errors).
As with most things you can never be perfect (aka the halting problem) but you can make failure statistically rare enough that you are willing to risk it. It doesn't mean that bad things will never happen. We live with risk every day and still drive to work.
And the trick is it is easier to just say "Whoops, something isn't right here, we're resetting to a known state." but that can make functioning really difficult (see Kepler's safe mode, or the challenges Armstrong had landing the LEM when it was constantly resetting the flight computer). Sometimes you just break. And in those situations the best you can do is to break in such a way that you don't cause any additional harm.
One should be able to design a watchdog system to handle the all kind of HW failures.
----------------
The KISS design principle is : "If I am a working SW/HW, I always assume the Master has failed and ready to take over the overall control unless I got Absolutely Positive confirmation that the Master system is working perfectly within the in N milliseconds windows."
10 years ago, I designed the system HA software for such system. It was only with two completely identical redundant system. Each system had 19 Xilinx FPGAs.
It worked very well. One can trigger any type of failures on any of one the system. Within 250 milliseconds, the other system automatically can / did take over the control whenever for whatever failure that was detected.
The system software design was very simple and worked very well in the Comcast and other cable network deployments. The input was 10GBits MPEG2 streams. The output 48 channels of Analog NTSC TV pictures. If you blink your eyes, you won't notices the complete System HW / Data Flow had been switched over.
It also made very cool demo. Unplugged power, network, boards on any one system, the other one automatically took over the complete control within 250 ms. If you watch very carefully, you only see slight flicker on the output TV signals for all 48 channels.
One other benefit, once the system was in place, it was trivial to add in-service SW Upgrade to the system. One can upgrade any SW (uboot, kernel, FPGA code, HW, etc) with only 250 ms impact on the end customer's service.
Once you got the idea on how it worked, it is actually not hard to expand for system with N - redundant HW/SW and let the "master working" HW to own / generate the correct solution within set N-millseconds time limit.
After which, all the configure and mirrors to standby in C code.
A LOT more code was written to test the config + HA switch over process. The test script was continually configure the system on the master side via snmp and SW force switch over every 30 seconds and run over night, over weekend, and forever to see if there was any unexpected configure errors and failures.
http://learnyousomeerlang.com/the-hitchhikers-guide-to-concu...
EDIT: Thinking about this a bit more, for messing with registers, you probably need specialized hardware or you need to intercept the data on the way to the register.
Edit: if you had the resources maybe you could even fuzz simulations of the RTL. Our systems take about a week to boot Linux on simulations but I guess that would drop dramatically for simpler MCU systems.
Would be amazing if the method used for implementing it could be generalised somehow.
Have 3 or 5 or a dozen copies of the thing running in parallel. As many as it takes.
Every action calls for a vote among the copies. Majority rules.
Periodically validate the data. Again, majority determines what's valid.
The bug webpage mentioned in the article is dead, and clicking the link subtly redirects you. Here's the archived version: http://web.archive.org/web/20140824005145/http://www.guildwa...
Sometimes software needs to be able to work around bad hardware.
The answers there are very insightful
If you have something in front of you, there's really no reason you can't at least try to come to some kind of understanding of it. Then do your best to communicate your findings.
Write 5 lines of perl that run logical AND and logical OR on a randomly generated bitmap emulating flash memory error rates at 1e-3, 1e-5, 1e-7, then run a billion emulation cycles (depending on code complexity this might take a long time, or maybe not).
If you're doing it right, you have the code for the emulator, never use closed source for this kind of work unless you trust it with your life (LOL) so implement statistical error injection in your emulator. Guess what, your CPU adder now has a 1e-5 bit error rate.
This is before the dreaded "programmer with a soldering iron" is unleashed. A white noise source is attached to pin D0 of the CPU and levels are increased while behavior of the system is analyzed in the test jig.
Simplistically, radiation is linear, so running at 100x intensity for a month is the same as running at design intensity for 100 months. Talk to your local nuclear physicist or radiologist. Especially about neutron activated decay products, LOL. In more detail, if it takes xenon 12 hours to build up to a problem level, you can't just test at 100x intensity for an hour and call it good because 11 hours after the test ends you'll finally have 100x the xenon decay product as a contaminant, but that's a weird situation.
I take your basic point--and it's a good one--that engineering is actually a very creative pursuit, when done well. But like any other high-expertise field, it's difficult for people on the outside to understand how engineers solve problems and, thus, to appreciate the sorts of skills required to do it well.
But isn't one of the lessons here that it's probably not a good idea to make gross generalizations about the attitudes and dispositions of people in other disciplines? :) I think "Liberal arts types say engineers are not creative" is probably about as true as "engineers are not creative." There may be some kernel of truth to each, but neither is especially accurate. (The kernel of truth may be that each statement could often be true when relativised to dumb liberal arts types and dumb engineers.)
Your larger point is right: some commenters are giving up on problem analysis or solutions "because the problem is hardware". That implication does not follow. Perhaps people are too used to highly abstracted software systems?
If you want to build robust systems, picking the right solution is paramount. In this case the consensus seems to be to fix it in hardware rather than messing out with software (it takes longer to implement a verifiable solution, harder to test, etc).
I worked and published in the area for a few years, so I have a definite POV. I find it appealing that software can, in some cases, make up for the failings of hardware.
The topic area is fascinating. It has some of the intellectual aspects of security research: you can get somewhere by breaking through abstraction barriers in surprising ways.
Still, for things that are really mission-critical, I'm not sure there's any substitute for full-blown hardware redundancy, e.g. having 3 computers and requiring 2 of them to agree on an action (via an interlock mechanism that is itself robust and redundant, which is a whole other can of worms...).
Think of COTS servers, space-craft, or various people-management processes.
I don't know how these corruptions manifest themselves, but if for example a single bit gets flipped every now and then somewhere (probably a naive assumption), a CPU that had a 0b11111111 nop instruction and 8 reset instructions at 0b01111111, 0b10111111, 0b11011111 etc., you could interleave your code with tons of nop instructions to minimize the possibility of executing corrupted code. You'd then be more likely to reset than to execute something that will lead to further corruption.
As for data, I like the idea of storing/reading it redundantly in 3 or 5 places and assume that the majority of a read to all of them (bit by bit) is correct and rewrite the supposedly corrupt data. Again, not sure how well a solution like this actually applies to radiation corruption. You can of course use any existing techniques for data error correction, like Hamming coding (used in for example teletext to correct errors).
I like this. It's functional, yet wildly impractical.