Comparing DDR5 Memory from Micron, Samsung, SK Hynix
eetimes.com
eetimes.com
All that said, there are still ECC (has an ECC memory) and no-ECC dimms for DDR5. So if the on-die ECC is concerning for anyone, they can still get a DIMM with a separate ECC memory. But the ECC happening at the interface between the DIMM and the CPU will still exist always and you will have to trust it.
* mandatory (an hypothetical DDR5 without could have error rates so high it would basically not work)
* an implementation detail (if the raw error rate was not that high, there would be no on-die ECC)
* not reported to the CPU
It's a complete different beast than real ECC. It's not that it is bad or concerning, it is that it does not provide RAS services and, like ECC-less DDR4, should be reserved for consumer electronics for basically only tasks like entertainment. Actually, in a better world most consumer electronic should have real ECC (instead of none at all or implementation detail on-die) -- but sadly for now vendors do not do that.
I doubt any manufacturer would make that public, however, but an estimate may be made if error rates actually start increasing in the real world due to ddr5 allowing this.
I agree that end-to-end ECC really should be the default for consumer products these days, but so long as the big players see it as a "Professional User" product differentiation point it'll always be more expensive than it should be.
Right. The more important Linus to speak up for ECC isn't Torvalds. It's Linus Sebastian, of Linus Tech Tips. He's made a few videos on ECC targeted towards gamers. Gamers drive the enthusiast PC market and when they start caring, more ECC gets made which will drive the cost down a bit. Last time I bought 32GB DDR4 UDIMM ECC there was literally one SKU. Not manufacturer. Not brand. SKU. One single item in production in the entire world. 16GB wasn't much better off, either.
It's a hard sell, though. Non-ECC will always be cheaper because it costs less to produce. Gamers don't really care that ECC prevents one crash in years because they are used to frequent crashes already. They are largely being fed dogshit from the AAA gaming industry and they have learned to just deal with it. Crashes are just part of being on the bleeding edge of gaming and Nvidia/Radeon drivers. One less crash in a sea of crashes isn't something gamers are lining up for. But a better model GPU or bigger SSD? It's an obvious choice.
I work on GPU drivers for one of those companies.
We regularly get reports and backtraces that cannot be reproduced, or "Cannot Happen" without some external factor (e.g. some other bit of code poking around our memory space). Often they're just silently dropped or ignored on the long tail of issues that nobody can get any traction on.
My understanding is the stats from hyperscalers is that ECC correction events happen a lot more than "Common Knowledge" may imply - I wonder just what proportion of things that are blamed on software may actually be due to hardware issues like this?
Again, without a significant change in the market (IE enough gamers start using ECC to actually be statistically relevant and comparing stability) this cannot really be tested, but I've wondered.
The real way to make ECC happen industry wide is for OS vendors like Microsoft to make it a platform requirement. A no ECC, no boot policy would change things overnight. Sadly, we can't even get DRAM manufacturers to fix row hammer properly, so the likelihood of this happening is pretty much nil.
So many people grumble, but I'm not really sure Intel should push ECC if desktops users aren't willing to pay a modest premium for it.
Many cheer AMD, which does not disable ECC on desktop chips, but neither do they promise ECC will actually work. It's a confusing mess between physical capacity (ram increases by 16GB when you add a 16GB dimm), and the actual correction of errors and telling the OS about the event. Only on the EPYC does AMD test and certify that ECC will work.
ECC is only officially supported on the Ryzen Threadripper PRO, and that's largely because it uses registered memory which will practically always have ECC capability.
Regular Ryzen PRO CPUs have the same ECC support as non-PRO Ryzen CPUs: functional, but not validated by AMD.
Ryzen PRO APUs have (non-validated) support for ECC, whereas non-PRO APUs don't support ECC at all.
It's not really a supported config, there's no guarantee it will work, no promise from AMD (other than it's not disabled), and you can't return a Ryzen CPU because the ECC doesn't work.
Various tests from various reddit posts show that some vendors "qualify" ECC dimms as working (an inserted Dimm increases ram available), some correct bits but don't correctly tell the kernel, and others actually do correct and report. So it's basically just a big mess. I'd buy an Epyc, but sadly unlike Intel the premium for a "server" chip is huge, if you can find them anywhere near MSRP. I looked for a Epyc 7313p near MRSP (announced in March 2021) without luck and finally gave up and bought a Ryzen.
With all that said, ASRock is from what I can tell the best AMD motherboard company to support ECC with AMD. I've got a ASRock X570D4U, other than it refusing to allow sharing a IPMI/BMC interface with the system, I'm pretty happy with it.
People don’t care, because people don’t know. RowHammer-vulnerable RAM is like canned food with botulinum and gasoline with lead. In 100 years, we’ll all be flabbergasted that everyone in the industry wasn’t arrested for letting it happen.
Track row activation counts for real, no estimates and no hash tables, and you're not vulnerable to this.
Do I want ECC, sure. Would bitflips ECC could prevent be in any in the top 10 reasons in year for file corruption, application crashes, and OS crashes etc... unlikely.
$50 more for the motherboard and $100 more for the RAM is not a "modest premium" for most of the desktop market. But aside from that, the bigger problem is that consumers (and especially gamers) will always prioritize the advantages that are more easily quantified. When forced to choose between ECC support and a few hundred MHz plus the option to overclock further, consumers will pick the faster processor.
If Intel offered a part at the high end of the product line that had both ECC and overclocking enabled, it would be guaranteed to sell quite a few units, because a lot of consumers will always want to own the top of the line part. And such a part would be an opportunity for a better experiment to determine how much premium users are willing to pay for ECC, if Intel also kept offering the current -K overclockable parts that don't have ECC.
What is the order of magnitude of ECC correction events?
I can actually give you a GAMING (or at least enthusiast) use case for ECC, too: memory overclock validation. Right now, it's kind of a dark art, but ECC failures would absolutely give you a very explicit early warning that you're pushing the CPU IMC too hard.
[0] Pronounced LIE-nus
[1] Pronounced lee-nus
I hope that ecc will give more warning that you're close to the edge (e.g. any ecc correction means it's unstable and you should back off), but fear people will just keep pushing until the ecc can no longer even correct the errors.
Professionally I work on gpu drivers for systems that do not allow overclocking, so shouldn't affect our stats, but it certainly is a mess for other teams. I've noticed that it's even become a common suggestion on fan forums for gpu manufacturers to disable any overclock and revalidate when changing vendors or even updating drivers, presumably as they're so close to the edge small changes in the access patterns will then cause it to explode.
A system with no error correction has to be so close to perfect. Even if you put in much worse cells, going from zero error correction to mild error correction should leave you significantly better off.
Although it does complicate things, much like happened when relatively stupid discs of spinning rust (which might lie about sync) transitioned to SSDs which are a blackbox with fancy algorithms (for cache, wear leveling, minimizing write amplication, etc), hidden storage, ram (for cache), even some with fast and slow flash. So now you have to be very careful to test long enough to exceed cache (flash or ram), and if randomly writing you have to sustain it till you consume the entire drive to ensure the housekeeping is included in your performance numbers.
So this raises the spectre of dimms with so many errors that a detectable number of lookups might require ECC corrections and additional latency. So we will need a SLA from the dimms, something like 99.99% of the time with less than X ns of latency.
Now that memory controllers are on the CPU (well, the I/O die for AMD), there's actually a whole lot more slack for the timing needed to do ECC. While the I/O transistors use big, slow transistors for driving the off die signals, the logic is all implemented using the same 3-5+GHz transistors used. This is in constrast to the transistors used in DRAM that are tuned to minimize leakage and are a heck of a lot slower. Putting complicated logic on the DRAM die is a fundamental mismatch with that goal, which encourages DRAM vendors to put the minimum possible logic on their dies. ECC gets better with more logic, and worse with less. Do you really want your DRAM vendor deciding how much ECC is Good Enough for you? I don't.
For similar reasons I wish that SSDs, MicroSDs and similar were just stupid devices with simple read/write logic. Then various filesystems, databases, or key/value stores could compete for popularity for things like load leveling, minimizing write amplification, error checking, write performance, IOPs, etc.
Seems trivial to extend that to writes per cell, so if you screw up you lose the warranty.
LPDDR on the other hand move the individual dimm chips as close as possible to the CPU and don't have any connector. This also makes it much easier to have wider memory. A 13" MBP can have a 512 bit wide memory system with at least 16 channels in a thin/light laptop that is quite power efficient. To get similar with DIMMs you'd have to buy a dual socket server motherboard with 8 channels per socket and would be lucky to fit that in an ATX size motherboard in a 1.75" thick chassis.
The only similarity between LPDDR5 and DDR5 is its name. Otherwise you could think of them as completely different technology just like HBM and DDR.
While HBM memory is fast, it is limited in size (eg: 48GB-96GB), whereas DDRx memory can be in the TBs easily and that too inexpensively. You can think of the HBM-DDRx hierarchy analogously to the L1-L2-L(n) cache hierarchy for CPU processors. IIRC, even the CPU HBM implementations will use this HBM-DDRx hierarchy.
[0] https://www.nextplatform.com/2021/04/12/nvidia-enters-the-ar...
hmm? 7GB/s is the performance that modern disks achieve
edit. nvm: When it comes to bandwidth, DIMMs provide 38.4 GB/s or 44.8GB/s
I do wonder when disks will be close to RAM sticks because gap seems to be closing.
It'd be nice to do not have to buy RAM sticks and just use small 1TB NVMe M2 disk as both - ram and disk
Usually not. Most SSDs that have DRAM only use it to hold the address mapping tables, not any user data. The DRAM cache for that internal metadata helps with random IO performance (saving an extra flash read before the real IO can be done), but has minimal impact on sequential read throughput, which is what has actually reached 7+ GB/s.
It is time software take a new look at memory usage again. RAM isn't cheap nor free. Now even 8GB memory is shared with GPU. Considering we now have 4x or 5x or pixel count we actually have less memory to use compared to the same 8GB 10 years ago.
Looks like a common issue from my research.