The parent comment saying that they never see ECC errors in the wild is missing a few things:
- Server memory tends to be clocked lower than consumer memory, so errors are less frequent to begin with.
- The errors are not evenly distributed. Some memory sticks have a high error rate, others are virtually zero. There's batch-to-batch variations.
- I've done my own tests on hundreds of servers. We run burn-in tests for about 24-48 hours. About 95% have zero errors of any kind, but 5% have a high enough rate that putting them into production would be a mistake. ECC allows us to catch those bit errors instead of silently accepting them and allowing data corruption to creep in.
- I've personally had 3 different personal computers experience high memory error rates, to the point of multiple BSODs per day and data corruption. They all started off "good" and slowly turned "bad". The only reason I knew to look for memory corruption as the root cause is because of my extensive industry experience. A grandma using the same PC would have just blamed Windows for being unstable.
- Vendors like Microsoft simply ignore all crash error reports sent back by telemetry with only 1 or 2 samples, because those are virtually guaranteed to be caused by memory corruption, not programmer error. I've done similar memory dump collection and found that easily 30% of all crashes were unique in this way, suggesting that ECC memory could improve PC stability significantly.
- Suggesting that ECC memory is not needed because "good" memory doesn't need it and only "bad" memory is a problem is missing the point. All memory is bad, it's just that the bit error rates are different!
Eventually I got a corruption in a large Git repository with photos. First I suspected a disk error but reruns of "git fsck" reported different bad objects, run to run. How odd, so I ran memtest86 and it reported a bad 32 MB memory region at offset 4 GB. Never saw any kernel issues or other instability. Booting up the computer does not use up 4 GB of memory (Linux) but starting a bloated web browser does.
The computer is stable after mapping away the bad memory region by using GRUB_BADRAM. That was an new takeaway for me, a machine with ECC memory can do this bad blocks mapping automatically. It's not just beneficial in correcting single bit errors.
I would love to have ECC memory in my machine. I used it for my previous build in 2012 but I think the situation has gotten worse since then. Bigger price difference and ECC memory is not even available at the same clocks as non-ECC memory.
It's true that ECC will catch errors caused by cosmic rays, but in practice most memory errors are just from faulty memory. Large studies in Google datacenters showed that the errors were heavily concentrated in a small number of DIMMs: http://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf
The catch is that you don't really know if you have one of these faulty memory sticks. I always run memtest86 overnight to check for obviously faulty memory, but some of these errors could take months to manifest.
I had a laptop with a RAM cell that failed. It would show up in memtest86 in a matter of minutes, yet surprisingly I didn't notice the issue for a very long time in day to day usage. I always wonder if there were random bit flips in anything I worked with during that period, but I'll never know.
With ECC, it's just one less thing to worry about.
Without ECC any of the above will crash an app, or even the entire machine. With ECC you'll get a log message, and if it's a single bit error it will be automatically repaired. If it's not fixable and it's in userspace that application is killed (and the kernel logs the error), if it's in kernel space the kernel will panic.
So generally it makes your machine more reliable, and easier to debug. Generally repeated errors are a fault of some kind and you can easily tell which dimm it is. Without ECC you end up troubleshooting all causes of crashes, and even if you know it's memory, you can't be sure which dimm it is.
So I think it's well worth the minimal premium to make your machine more reliable, and if it's unreliable it's much easier to track and fix the problem.
Amusingly without ECC heavy CPU use with parallel gcc compiles causes a particular error. Not sure if it's in the FAQ, but it's well know that the particular error means you have a CPU/RAM problem, common with memory errors, CPUs that are too hot, or overclocked CPUs.
Sidenote I seem to have experienced Cunningham's Law for the first time semi-consciously, as I was aware my claim is most probably wrong. Still, I wonder if data corruption is a more common scenario rather than crashing, because ever since I use Linux I don't recall experiencing many crashes I couldn't attribute to a driver issue or known instability in the software I used, especially on compatible hardware. Windows days were another story... But maybe the memory modules were worse back then and perhaps having less memory made it more likely to crash?
There are some bugs that get raised on linux once every few years that are so obscure they are suspected to be random memory errors.