ECC Memory and AMD's Ryzen – A Deep Dive
hardwarecanucks.com
hardwarecanucks.com
Also I have no proof that any crazy thing can not happen, but there is no reason for single bit errors not to be corrected regardless of the OS. The worse that should happen for them is to not be reported.
IMO if you have the opportunity (the category of HW you want supports it) you would be crazy not to use ECC RAM. Non-ECC RAM is basically the only component in a PC that is not protected. Obvious weak point. I've been beaten at least twice (two defective components, way more than 2 errors before I figured out what was happening) only on computers I was directly owning or using at work (among a total of a dozen of computers). Now I don't want to loose my time anymore, I always use ECC memory when possible (I'm not going to pay a computer twice the price just for that, so it is a "little" difficult with laptops which also have a plethora of other choice criteria, but it is very easy to get affordable workstation desktop computers with ECC)
No modern digital communication bus will be designed without any form of protection, so this make not much sense to have computers without ECC RAM. I would even like to have it on smartphones, but unfortunately I doubt this will happen soon.
I thought perhaps video editing and rendering would be a more likely area where ECC could be valuable to the non-scientific user. I have gone googling for examples of artifacting in videos caused by non-ECC ram errors, but I found nothing.
The sole benefit is correcting random bitflips in memory. These bitflips don't occur often enough per byte of RAM in workstation environments to matter: It's far more likely you'll have an issue with corruption due to poorly developed software than due to EMI.
However if you're building a home NAS on commodity hardware, and you're concerned about things like data integrity, ECC is definitely preferred.
How often do you believe these random bitflips could happen in a year of continuous use?
Unless you have some funny notion of how you want to spend your time, you don't want to risk to fall in that situation...
The last thing you want if you are doing any kind of creative work is more or less random bitflips that happen very rarely but still more that lets say once in a year (YMMV). If your bitflips are too frequent you will soon notice and "fix" the problem (if you don't switch to ECC you actually fix nothing, but if you have a defective chip and switch to a good one at least you mitigate your issue)
For information the defect rate in consumer electronic is in the order of 1% (not specifically for RAM, I don't have any exact figure, but as a general order of magnitude figure it should be something like 1%). I know there are too much software bug here and there, but I don't want to add extra ultra-non-deterministic data destructing bugs just for fun.
I would far more prefer that the affected program(s) have a chance to react, or be killed as a subset of the system. If the error occurred in a filesystem context there may be other ways of correcting the issue (particularly if it's merely in read cache instead of write cache).
Obviously unhandeled exceptions should cascade until they are either contained or until the entire system halts.
There's no reason a fault while in the ECC error handler shouldn't have the same progression.
I really want to see someone get some radioisotopes and place them next to both ECC and non-ECC RAM (while forcing reads and writes to the affected memory) to see what sort of soft errors / SEUs happen.
https://www.cs.utexas.edu/users/skeckler/pubs/SELSE_2014_Rel...
http://users.nccs.gov/~vazhkuda/hpca.pdf (section 6)