If you're Google (you're not Google) then the scale of errors becomes a choice of optimization.
If you're Google (you're not Google) then the scale of errors becomes a choice of optimization.
The frequency of DRAM errors is proportional with the amount of memory installed in a computer. For smaller quantities of memory, e.g. 16 GB or 32 GB, the frequency might be of one error every few months. With ECC, the mean time between errors is likely to become longer than the lifetime of the computer.
Intel, who has initiated this absurdity of convincing the naive customers that it is OK to buy computers that may make mistakes from time to time, has succeeded to get away with it only due to the huge amount of software bugs, which have made the computer users believe that whenever a computer crashes, or some corrupted data is discovered, it is more likely that the cause has been a software bug and not a hardware defect.
The integrated circuits made with modern CMOS processes, with very small devices, have a non-negligible aging rate.
Because of that, memory modules that have been used for many years start at some point to have much more frequent errors than when new.
When you have ECC, such memory modules are immediately detected and they can be replaced before causing some irremediable data corruption. This helped me a lot a few times.
It's some sort of variation of the Toupee Fallacy, you aren't aware of the frequency of bit errors because you have absolutely no way of knowing when they happen, and what "random" errors were caused by them.
I don’t care about triggering an occasional bug or a bit flip in a video stream, these things are inconsequential. If a bit gets flipped in a photo it’ll be a very slightly different photo. A piece of text will be slightly different… if either had the luck to be in memory ready to write when hit, in dozens of gigabytes of memory.
The most likely consequence of a bit flip is nothing at all happening because it hit unallocated memory. The next most likely consequence is nothing as it hit a piece of program memory that has no effect on operation and gets overwritten in moments.
ECC zealots are really excited about preventing something that just doesn’t happen all that often: actual data corruption with permanent significant effects.
This is really weird. How come even routers have ECC memory while our actual computers don't?
Why can't all memory just be ECC memory? I don't even see the point of non-ECC memory. Everything has error correction built-in: hard drives, SSDs, even optical media. Yet RAM doesn't?
(The 1st generation of Pentium, the 60/66 MHz models, had been introduced in 1993, while the 2nd generation, the 90/100 MHz models, introduced in 1994, used a different socket and different motherboards.)
Unlike with the previous Intel CPUs, Intel had the idea to introduce a market segmentation between Pentium and Pentium Pro (for the successors of Pentium Pro, Intel created the Xeon brand, a couple of years later).
Therefore, Intel has retained the ability to detect and or correct memory errors only in the Pentium Pro chipsets (at that time the memory controllers were in the so called northbridge, a part of the chipset, and not in the CPU), while removing this feature from the Triton chipset intended for Pentium.
While Intel has initiated this market segmentation policy, between "amateurs" and "professionals" (unlike IBM, which since the first IBM PC had always included memory error detection via parity), both the memory module vendors and the makers of other CPUs have been also happy to follow this policy, because it allows them to extract much more money from the knowledgeable users than the cost of adding ECC, while also having larger profit margins when selling to naive users, because the elimination of ECC/parity has never caused any price reduction.
When this feature was removed, all the new motherboards and modules without parity/ECC had the same prices as the previous models with error detection (for many years, the memory errors were detected via parity, but then the memory controllers were improved, so that when using the same number of extra bits, i.e. the same memory modules and motherboards as for parity, single errors could be corrected, not only detected).