On-die ECC was added in DDR5 because the error frequency has increased too much.
Its only effect was that now DDR5, as seen externally, has a reliability similar to that of the older memory generations. On-die ECChas not brought any serious improvement in reliability. It has just prevented the degradation of the reliability.
The errors generated internally are caused mostly by the ionizing radiation from the environment, which discharges the storage capacitors.
However there are many errors that appear during the communication between memories and the CPU, which are caused by the electrical noise from the environment.
These external errors are much less likely to appear for soldered DRAM. Because of this, for soldered LPDDRx memory it is less important to have end-to-end ECC than for socketed modules, which are much more vulnerable to noise.
The DDR5 standard has some options for using some error detection for the communication link, but it is impossible to know whether a given computer or motherboard has implemented such options or if they are enabled by the firmware when the motherboard has the physical support.
The only way to be certain that the memory works fine is to use ECC that covers completely the circuit from the CPU memory controller to the values stored in the DRAM and back to the CPU memory controller and your operating system has the appropriate device driver for obtaining the error reports.
The imperfect contacts are equivalent to adding series resistors, possibly in parallel with parasitic capacitors, on the link traces, allowing a greater amplitude for the noise pulses.
Caveat that I was building a gaming PC at that time, I might have gone differently for a home server. And, this was before the memory market went nuts for unrelated reasons.
The "gaming" DIMMs are available in much greater speeds, which are "overclocked" speeds, i.e. there exists absolutely no guarantee from Intel or AMD that they work at the advertised speed, even if they frequently do work.
However, with such overclocked DIMMs, you do not know how well they work. They might work most of the time, but e.g. in a hotter room they might have an error or two per day without you noticing, especially if the computer is used mainly for games.
For someone doing anything professional on a computer, overclocked DIMMs are not an acceptable choice.
This is not true[1]. These DIMMs are not guaranteed to work with all Threadripper CPUs, but they are guaranteed themselves to be stable at their advertised frequencies, which are higher than JEDEC. (And from my experience, 9000-series threadripper is perfectly happy to drive 8-channel 1Rx8 at 7200MT/s.)
AMD (and possibly Intel, I don’t have any experience with their recent offerings) also lets you overclock memory completely on WRX90. So if you buy JEDEC sticks, you’re welcome to try to push them as hard as you can.
[1] https://v-color.net/products/ddr5-ocrdimm-amd-wrx90-workstat...
Without ECC, nobody knows if their memory is working or not. Regardless whether it is running a standard JEDEC speed or an overclocked one.
If you buy a complete gaming computer which includes overclocked memory, then hopefully the vendor has done a burn-in and has tested for some time the computer, as sold.
However, if you buy the CPU yourself, neither Intel nor AMD provides any guarantee that the CPU will work with overclocked memory.
It is upon you to test your assembled computer, but true tests would require a very long time and a climatic chamber, to offer any kind of certainty that the combination CPU-DIMMs works at the desired speed.
If you have ECC memory, you can try to overclock it and at least in this case you will know for sure if the memory works at the higher speed, or not.
Overclocking ECC memory is actually a method frequently used to verify whether the hardware ECC support and also the operating system EDAC driver work OK.
There's no guarantee that it works with standard clocked memory either.
If the CPU does not work at that speed, you are entitled to a replacement or a refund.
For instance, AMD Ryzen™ 9 9950X3D is guaranteed to work with DDR5-5600, but with nothing faster. No AMD non-server CPU goes above this limit for DDR5.
Some of the Intel CPUs are guaranteed to work with faster memories, e.g. Core Ultra 9 285K with up to DDR5-6400, and the very recently launched Core Ultra 7 270K PLUS with up to DDR5-7200.
Without the reported ECC on Infinity Fabric I would have had no idea that the errors were happening and affecting performance.
The main reason there aren’t more options here is that a) high-frequency memory usually isn’t sufficiently stable, unless binned, and b) there’s very, very, very little demand.
[1] https://v-color.net/products/ddr5-ocrdimm-amd-wrx90-workstat...
Aliexpress had $50 single socket motherboards for Xeon E5-26xx CPUs and $150 dual socket motherboards last time I checked that use DDR4 RDIMMs.
Before the prices went up due to AI foolishness DDR4 RDIMMs were $1.20AU/G means means something like $US0.80/G.
From that time, I have a few old computers with 64 GB or 128 GB of DDR4 ECC memory.
Unfortunately, after the passage to DDR5, DDR5 ECC UDIMMs have become hard to find, and even when you could find them, the price difference could be as high as 30% to 50%.
It is if you're doing something where corruption is detectable after the fact, and happens rarely enough that redoing affected work adds less than 3% overhead.
Without ECC, you need a custom operating system, which will compute some error detection code, e.g. a CRC, for each read-only memory page, i.e. including all code pages and constant pages, and all cached but not dirty file system pages, and which will check the CRC codes periodically in the background.
That would still leave the read-write data segments unprotected, though on those some errors may be benign, if the locations are overwritten later without being read again.