It's seriously annoying that ECC memory is hard to get and expensive, but memory with useless LEDs attached is cheap.
It's seriously annoying that ECC memory is hard to get and expensive, but memory with useless LEDs attached is cheap.
I do hope the nuclear powerplant next door uses more fault tolerant hardware, though.
> I much prefer cheaper hardware.
The cost savings are modest; order of magnitude 12% for the DIMMs, and less elsewhere. Computers are already extremely cheap commodities.
Assuming that's more due to intentional market segmentation than actual cost, yeah I would pay 12% more for ECC. But I'm with the other guy on not valuing it a ton. I have backups which are needed regardless of bitrot, and even if those don't help, losing a photo isn't a huge deal for me.
That was me. It isn't "officially" supported by AMD, but it should work. You can enable EDAC monitoring in Linux and observe detected correction events happening.
> Assuming that's more due to intentional market segmentation than actual cost
That's the argument, yeah.
You mean like compilers and test suites ? Very few professional workloads don't parallelize well these days.
(Though games these days scale better than they used to, but only up to a to a point.)
I find that most tools I write for my own use can be made to scale with cores, or run so fast that the overhead of starting threads is longer than the program runtime. But I write that in Rust which makes parallelism easy. If I wrote that code in C++ I would probably not bother with trying to parallelize.
Imagine ECC was free -- would you rather have free ECC and no bitflips, or no ECC and bitflips? It's hard to imagine choosing bitflips.
Ironically, overclocking ECC memory is much easier than overclocking non-ECC DIMMs, because you know exactly at which point you start encountering instability and need to dial back, instead of relying on.. application crashes and BSOD's to know that you're running way too optimistic clocks/timings.
Meanwhile I overclocked 'low clock / loose timing' ECC DIMMs on Ryzen 7 platform with no issues at all – kept increasing clocks and lowering timings until ECC started reporting errors, then dialed it back a couple notches, and now it is not just stable, but I also have exact reporting of it being stable.
(For those out there following along with PCs, if you aren’t tuning with MBIST maxed out in your BIOS, you might want to revisit that.)
Ironically, that's around the time Intel started making it difficult to get ECC on desktop machines using their CPUs. The Pentium 3 and 440BX chipset, maxing out at 1GB, were probably the last combo where it pretty commonly worked with a normal desktop board and normal desktop processor.
I'm not really sure if this makes it overall more or less reliable than DDR2/3/4 without ECC though.
it's "ECC" but not the ecc you want, marketing garbage.
Because you can't track on-die ECC errors, you have no way of knowing how "faulty" a particular DRAM chip is. And if there's an uncorrected error, you can't detect it.
I think this sort of reporting is a pretty basic feature that should come standard on all hardware. No idea why it's an "enterprise" feature. This market segmentation is extremely annoying and shouldn't exist.
I would definitely like to have a laptop with ECC, because obviously I don't want things to crash and I don't want corrupted data or anything like that, but I don't really use desktop computers anymore.
And there's non-random bit errors that can hit you at any speed, so it's not like going slow guarantees safety.
The main reason ECC RAM is slower is because it's not (by default) overclocked to the point of stability - the JEDEC standard speeds are used.
The other much smaller factors are:
* The tREFi parameter (refresh interval) is usually double the frequency on ECC RAM, so that it handles high-temperature operation. * Register chip buffers the command/address/control/clock signals, adding a clock of latency the every command (<1ns, much smaller than the typical memory latency you'd measure from the memory controller) * ECC calculation (AMD states 2 UMC cycles, <1ns).
The main overhead is simply the extra RAM required to store the extra bits of ECC.
That said, memory DIMM capacity increases with even a small chance of bit-flips means lots of people will still be affected.
However, there are still gaps. For one thing, the OS has to be configured to listen for + act on machine check exceptions.
On the hardware level, there's an optional spec to checksum the link between the CPU and the memory. Since it's optional, many consumer machines do not implement it, so then they flip bits not in RAM, but on the lines between the RAM and the CPU.
It's frustrating that they didn't mandate error detection / correction there, but I guess the industry runs on price discrimination, so most people can't have nice things.