Monitoring ECC memory on Linux with rasdaemon
setphaserstostun.org
setphaserstostun.org
FWIW I’m a believer in ECC. But give it some time. Maybe you’ll see errors. Or maybe you’ll get lucky and see none. Either way, you’re covered!
Google's results show that the error rate is strongly driven by bad modules rather than random errors. A module that throws one error is up to 224x more likely to throw another error, and corrected errors are a strong predictor of future uncorrectable errors.
The major exception is cosmic rays, which can flip bits even in a good module. But unless you're at high altitude they don't really dominate the error rate.
http://www.cs.toronto.edu/%7Ebianca/papers/sigmetrics09.pdf
Obviously without ECC you don't know if you have bad modules, but, if a machine passes a memtest scrub for 24+ hours then it's probably fine.
On-die ECC should also catch the majority of these errors, especially if you are not right at the limits of your memory training. JEDEC type timings should be very forgiving to transmission errors in general.
People are very down on the "on-die ECC" on DDR5, they act like ECC is completely useless unless you protect the whole datapath, and of course it would be preferable to do so. But actually the on-die ECC should catch a large number of the errors, because bad modules (dies) are where the errors mostly happen.
Not sure why, maybe it's the different memory access patterns, maybe bad dimms are more sensitive to repeated accesses below the rowhammer level, that healthy dimms have no problems with.
I find things like make -j for large builds MUCH better at finding memory errors than memtest.
I was just hoping to see some errors on a regular basis, so I could feel vindicated and rave to everyone about how useful ECC is and point to my console and say "look at how many errors you might be missing!". ;)
(I choose not to worry about flips of three bits or more. YMMV.)
This is dependent on the system configuration. For Linux, default behavior is to halt only if the uncorrected error occurs in memory used by the kernel.
It's more accurate to say that double bit flips will be detected.
Search for "edac_mc_panic_on_ue"
Although I was much more interested in ability to find dying hardware. Whenever weird (file corruptions, processes dying, reboots, things not working like they did yesterday) could be from numerous causes. Voltage dips, problems in the power supply, software bugs, OS bugs, security compromise, defective CPU/Ram/motherboard/dimms, etc. It can be very hard to track down, you might well spend weeks and significant $$$ playing the game of swap the part and wait for the problem to reappear ... or not. Much like what recently happened to LinusT. I'll happy pay another 10% system price for a system that has a few less crashes and will tell me when it starts seeing memory errors (while fixing the small ones).
I'm typing this from a Xeon E3-1230v5 bought in 2015, that cost LESS then the similar i7. I did spent a bit more on the motherboard ($50) and ram ($75), but it's been rock solid since and if I had a dimm problem it would likely be 10 minutes to troubleshoot and then I'd RMA or buy replacement ram. ECC reduces the chance of writing the wrong data to the right place on disk or, even worse, the right data to the wrong place.
There's a reason why every datacenter does it. You must be able to trust your machine.
In addition, rasdaemon doesn't seem to have received any commits since April of last year, where mcelog appears to be under active development. Can someone help me understand the recent uptick in interest for rasdaemon, when mcelog seems to remain the superior option?
[0] https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=890301
It's frustrating that 4.12 is five years old and this hole in functionality remains. I wonder why kernel devs don't consider userspace support for their features when swinging the deprecation axe.
See eg https://github.com/qemu/qemu/blob/48804eebd4a327e4b11f902ba8...