I always knew about the theoretical cosmic ray bit flips. Before listening to this episode, I did not stop to think how often they actually cause problems.
I always knew about the theoretical cosmic ray bit flips. Before listening to this episode, I did not stop to think how often they actually cause problems.
My methodology was simple. On a Linux home server that had plenty of spare memory (non ECC RAM) I ran a process that simply alloced a large buffer and filled it with a pattern. It would then periodically scan through the buffer looking for changes to the pattern.
I ran this for over a year which should have been long enough given the amount of RAM I was using and the rates that I found in the literature for cosmic ray induced bit flips resulted in several flips.
My method would have missed a flip if the page it happened on had been paged out sometime in the past and that paged out copy still existed, and the bit flip happens between the time of the last scan and the time the kernel decides to discard that page. On the next scan it would page fault and load the good page.
But the system was very lightly loaded and almost never actually had to page things out, so most of the time if a bit flipped it should have still been there by the next scan and so I don't think this explains why I saw no flips.
A few years later I got a Mac Pro, which had ECC memory. I used that mostly at work from 2008-2017. I got another Mac Pro in 2009 which I used at home from 2009-2017. I'd occasionally look at the memory status in System Report which should say "ECC Errors" if the ECC had to fix any errors and only ever saw "OK". I'm not sure if that resets on boot and I only looked occasionally so if it does reset than it is quite likely I would have missed an error statuses.
If you ran that machine in a hotbox at 85°C for a year I think you probably would have experienced higher error rates. Bonus points if you also force the data to swap out a lot so it gets transferred through as many data paths as possible.
Also, there was an expectation that comsic ray induced bit errors would grow as ram circuit features shrunk, but it ended up not happening; reasons unknown or at least I never saw anything suggesting a reason.
Getting errors in a small sample of RAM is unlikely, unless you specifically induce them by using debug features or misconfiguring your system (but some systems conspire to misconfigure themselves, making it easier to observe! a couple years ago, retail motherboards really liked setting the ram voltage too low)
It's not 3.6 roentgen...
Joking aside that's incredibly fascinating, I never thought that ECC memory has that much of a performance impact. Might be more optimal to just get a large jerry can of water and put that over the server as radiation shielding lol.
I had a similar issue one time where a Pentium III era server rebooted and came up with a comically small amount of memory, maybe 2-4 mb instead of 128 mb. That wasn't too bad, because it was a very lightly service; important to be on its own machine for reasons, but didn't need much. Just ran a little slow when it was running from swap. I think it did trigger a swap usage alert, and then it was like why is it swapping, why is it so slow, wait why doesn't it have any memory!?
I'm told it's the equivalent of a chest x-ray...
I see you, brother! :D
Jokes aside, that it seems really unlikely that cosmic rays would be clustered past, like, a couple hours, right? It isn’t like some neutron star is, like, tracking your server as the world spins (well, I hope not, I mean who’d you piss off for that to happen?).
Anyway, this fits my totally unscientific expectation that cosmic rays are just sort of like an informal description of hardware bugs that nobody can reasonably find beforehand. A server that seems to be hit by lots of cosmic rays probably has a dodgy connection somewhere inside it, but I mean maybe it’s the RAM, swap that out or replace it… but maybe inside the chip SOC, so what are we going to do, bust out the electron microscope to check all those connections?
In terms of diagnostics, we pretty much just asked for ram replacement, if that didn't work, cpu replacement, if that didn't work, motherboard replacement. If that didn't work, the chassis / rack position is clearly cursed, don't give us anything there again, please. :D I don't know what they did with the hardware we didn't like, maybe send it to the manufacturer, maybe give it to customers they don't like, maybe surplus it.
So, at scale, they were getting failures constantly. (It didn't help that the QIC tape backup was basically write-only.)
As a rule of thumb, you get 2 muons through your head a minute - but of course your head has a huge volume compared to memory chips.
In the case of a scientific simulation, we can find algorithmic speedups that are "lossy"- an example would be approximations to n-body systems, where you need to calculate n-squared interactions (between all pairs). You can calculate all n-squared interactions which produces teh completely correct result. But since atoms that are far away don't interact strongly (falls off as 1 over r squared or more). So you can maintain a neighborlist- all atoms within a distance R- cheaper than you can calculate n-squared interactions, with some tiny error that is unknown. It's assumed in many cases the errror is neglible and the speedup is huge.
Or you can switch to using a particle-mesh method which involves taking a fourier transform, doing some work in fourier space, then an inverse transform. The results are nlogn, and the error is small (and known). Speedup is signfiicant but takes much, much more computational skill and infrastructure.
Next, the case I'm referring to of an accelerator running a tensorflow job, it's a totally different scenario- here, some random subset of machines will repeatedly return the wrong result- say, for a matrix-vector operation. Maybe garbage numbers, maybe all zero, maybe some infs. When that gets summed into your gradient, it often causes blow-up and the entire job terminates. It's not clear whether it makes sense to make high-performance jobs have to be robust- I see them as special cases where you're working hard to make sure the computing substrate is effectively 100% reliable (by sending/fixing those machines).
Other folks have observed that some amount of small noise injection to the gradient can help training, but the sorts of errors I've seen almost immediately terminate the training job. I don't mind intentional noise addition, but noise due to hardware that is provably, reliably, and repeatedly miscalculating results? Not so much.
The best was a one-off error log along the lines of "unknown type System.DateTime". Huh? That's a system defined type that just went missing. Never saw it again.
Another at a different employer was a crash that occurred after a check condition that absolutely should have gated the crash from being reached. Single threaded. Simple microcontroller. Had to reflash it to flip the bit back. After doing the math on how much RAM we had in the wild vs. cosmic bit flip rates reported in super computers, we had to expect one flip per year.
If it's a safety critical system, server or not, use ECC RAM!!
[0] - https://podcastaddict.com/wissenschaft-auf-die-ohren/episode...