Why use ECC? (2015)
danluu.com
danluu.com
Probably not surprising, there's a naturally antagonistic relationship between Performance and Reliability here, and it's clear which way many of those "enthusiast" forums lean.
I haven't got actual numbers, but I feel that most [0] of the issues I start looking at just can never be reproduced, or even make sense from the backtrace or similar. I can't say it's 100% hardware issues for this, as many games are a little... loose... with reliability if it works "well enough", and is heavily interacting with code and data we work on so might also be a source for "impossible" issues. But even on straightforward code paths, no weird OS interaction, no allocation, nothing async etc. "Impossible" states happen pretty regularly.
I would love there to be enough ECC-using gamers out there to statistically see if it makes a difference.
[0] Most in terms of number of different issues, not total reports of the same issue. That's dominated by one or two things, normally around the latest game or update doing something dumb :P
Games have pretty terrible quality because games can remain playable in the face of bugs. GPUs have similar issues because numerical accuracy is an after thought but numerical accuracy issues can cause issues in code that assume stability of results.
Things have gotten better particular on the GPU vendor side of things, but still, memory safety issues and UB issues in c/C++ code far dominate as the root cause any issue you’re likely to see vs memory bit flips (unless people are running very overclocked and unstable machines).
And UB shouldn't be a scary thing 99% of the time, as you won't be hitting it anyway (or its actually the result of another real bug, like not handling overflows etc.), though as at some level you start trying on platform specific behavior and start /defining/ them. And again, there are tools and options around that, or highlighting areas where you might still be relying on them. They just might be toolchain and/or platform specific. It's never been the scary monster some online programming language fanatics seem to make it sound like, but is something to be aware of and managed.
And
> unless people are running very overclocked and unstable machines
Yes. Have you seen the Gamer Hardware Enthusiast community?
This is exactly how exploits are made. Accessing UB usually is the result of "another bug" because the UB isnt a bug itself.
> Yes. Have you seen the Gamer Hardware Enthusiast community?
Having seen it for the last 2 decades: A lot of them are interested in OC, but nobody is doing it. Maybe 5%, most likely <2%. Also, most modern Motherboards do a "RAM-Training" session before boot. On unstable machines, this will result in a "test failed" once in a while, with the motherboard showing a "Boot failed, returned every clock setting to default" message, which most likely is not being read by the user, who is either not at the PC the moment it happens or in "why is this taking so long? skip, skip, skip"- mode.
What does happen is that Motherboards have OC-like settings out-of-the-box/per default, which were stable most of the time and de facto-accepted by Intel, however is now hitting issues and diminishing returns:
https://www.radgametools.com/oodleintel.htm https://www.igorslab.de/en/intel-spielt-mit-dem-namen-und-de...
I had a PC that was having inexplicable game crashes, but only in stressful games - swapping the power supply from an existing ~10 year old one (which was still rated as enough capacity for my hardware) to a brand new ATX 3.0 unit resolved it.
I'd venture 95% of people run stock, 4% do undervolting and 1% do overclocking.
Which isn't that strange because GPU vendors have become relatively adept at pushing the silicon near its max. Its why you can run an aggressive undervolt with a few percentage point lower clocks and be rewarded with a 20-30% wattage drop with commensurate drop temperatures and fan noise, all whilst losing perhaps <5% performance.
But the GPU vendors want to put big number on box so they'll have users suffer loud fans :-)
As others are pointing out, there's an ongoing issue where it seems some products push it beyond the point of stability by default.
I think I'm more pointing out the trend that even "informed" purchasers just look at the benchmark charts. There's no other information required.
Perhaps things like the mentioned instability making the news will change things going forward? But not holding my breath. If there's still advantages to pushing it beyond the limit, then manufacturers will do so.
I’m talking about language-level UB - that’s not anything you can rely on on any toolchain / platform. And UB will also manifest as random unexplainable crashes in random spots just like memory corruption will.
As for tooling, unless you’re using something 100% memory safe with absolutely no call out to unsafe code (which definitely isn’t a game/GPU scenario), you’re going to have this risk & it almost doesn’t matter how much you test or use memory checkers because the long tail of issues are going to be the result of a really difficult to reproduce sequence of events. Additionally, if you have any multi-threaded code, all your testing goes out the window because I’ve seen many concurrency bugs that have hidden in plain sight in very public spots until someone figured out a repro (e.g. https://reviews.llvm.org/D114119, https://probablydance.com/2022/09/17/finding-the-second-bug-...). Race conditions are notoriously difficult to reproduce in controlled environments.
Oh and when I said you’re safe if you’re in a 100% memory safe language? I lied. There are compiler bugs that could be misgenerating your code + you’re likely running unsafe constructs somewhere in your code that has access to your memory space (whether the OS kernel, whatever is making syscalls to the kernel on your behalf, something that eventually calls the C library, something that is needed for performance, etc etc etc) not to mention bugs in the compiler/JIT that either allow unsafe/unsound constructs or just straight up miscompile correct code. And finally, there are sources of HW issues unrelated to memory bitflips (your CPU is super complex and has bugs too you know as does your memory controller).
> research has shown that the majority of one-off soft errors in DRAM chips occur as a result of background radiation, chiefly neutrons from cosmic ray secondaries,
> Recent studies[5] show that single-event upsets due to cosmic radiation have been dropping dramatically with process geometry and previous concerns over increasing bit cell error rates are unfounded.
So aside from the fact that DRAM susceptibility to cosmic rays has been decreasing (I’m not convinced with the Wikipedia explanation - an alternate explanation can be that the percentage of critically important data in DRAM has shrunk as a percentage as the overall capacity has increased), you’re argument would be that random cosmic radiation is going to randomly hit the DRAM cell containing your code / critical data. On the order of things that are likely, that’s the last thing.
Oh and all this relies on you correctly grouping related crashes correctly and I’ve generally seen that to be a significant challenge on any project I’ve participated on (e.g. Oculus for the longest time would group unrelated crash reports and not group the same ones correctly although I helped the team try to make progress on that).
Again, it’s not impossible that it’s a legit HW corruption issue. However, all the engineers I know frequently also blamed cosmic rays but at the end it’s all just shorthand for “not worth wasting time trying to track down because it’s in the tail of issues you’ll never get to“.
One I can relate to is if you have something with low code churn and extremely logically simple that does many gigabytes of I/O per machine over a short period. The maintainer of that code is going to see tons of bad disks. That was me once.
It's not a rocket or something dangerous, right?
Also, newer iterations of the car will be even more cheaper due to cut corners, so you may save some money in the process, too.
Same signs were easily detected by a Ford Puma I drove recently on a highway.
I assume training data for European roads are also is of inferior quality due to amount of miles traveled and because of sheer variety of the roads, road signs and norms around Europe.
There is an error correction layer in DDR5, could that be spotted statistically as a difference between Ryzen 5000 and 7000 processors?
We already shipped with debug assertions enabled in release builds, so it wouldn't have been unprecedented to do things like validate result bytes after doing big memcpy operations or things like that. Not sure if we did, though.
Lots of people’s gaming PCs (self built and otherwise) are very crashy - there’s a surprising number of people who get regular BSODs who just accept it as “normal”.
But because we had some equipment that would checksum itself we were able to see the bit flips happen, probably every 1-3 weeks or so I'd find a bit flip.
Intel was responsible for most of this. It is hard to be sad when seeing how they've lost the market lead.
I made the mistake of running a big PostgresSQL database on non-ECC memory once and I must say, it taught me some hard lessons.
anyone interested in the topic should absolutely start from reading up on https://en.wikipedia.org/wiki/Error_correction_code and only then start looking into the engineering side, starting with https://en.wikipedia.org/wiki/ECC_memory, notably some benchmarks reported 25% performance hits (not sure if it's real or sales propaganda.)
The number of ECC bits needed for any error correction level on n bits are proportional to log(n).
DDR actually operates in bursts of 4 or 8 64-bit words (+8 ECC), so you can make a better ECC using all 32 or 64 bits to cover the full burst rather than covering each 8 bytes separately.
[1] vs. the semi-religious flame war about DRAM ECC in which I won't engage. People get nuts over this, and IMHO the actual data is awfully inconclusive.
SECDED ECC is the academic example.
Only in the sense of "Intel is responsible for most of computation" ... No one uses ECC pervasively anywhere, that's sort of the point of the article.
Related post:
QNAP has a $600 1U short-depth (11") 4x3.5 2xM.2 2x10GbE 2x2.5GbE 4-32GB DDR4 SODIMM Arm NAS that would benefit from OSS community attention. Based on a Marvell/Armada CN9130 SoC which supports ECC, has mainline Linux support, and public-but-non-upstream code for uboot [2]. With local serial console and a bit of effort, the QNAP OS can be replaced by Arm Debian/Devuan with ZFS. Rare combo of low power, small size, fast network, ECC memory and upstream-friendly Linux. QNAP also sell a 10GbE router based on the same SoC.
Ryzen Pro (OEM) can support ECC [3].
[1] https://www.qnap.com/en-us/product/ts-435xeu
[2] https://solidrun.atlassian.net/wiki/spaces/developer/pages/3...
[3] https://www.tomshardware.com/pc-components/cpus/amd-confirms...
Dude, not even close. That repo is for an opaque blob:
https://github.com/SolidRun/cn913x_build/blob/master/binarie...
... which is mashed together-bits of other blobs, including the notorious will-never-be-open-source highly-radioactive "snps blob":
https://github.com/MarvellEmbeddedProcessors/mv-ddr-marvell/...
and its vile EULA which you have agreed to:
https://github.com/MarvellEmbeddedProcessors/mv-ddr-marvell/...
That shit contains an ARC core. Never heard of ARC? That's okay, most people haven't. It's an obscure niche architecture used for almost nothing except for ... drum roll ... the Intel Management Engine (up until 2019). Gee I wonder why that's in there.
Don't touch Marvell shit with a ten foot pole, it's blobs all the way down. They're just good at hiding it.
> That repo is for an opaque blob
Marvell patches have been going to upstream u-boot for a subset of CN-9130 functionality. Some were rejected as unwanted SoC-unique features, but that code is still public and could be consolidated into a public git repo for evaluation. Support for core features was merged, e.g. this thread from 2020 to 2023, https://lore.kernel.org/u-boot/ddd355c2-344e-4fbd-ace9-29d10...
> the notorious will-never-be-open-source highly-radioactive "snps blob" .. ARC core
Thanks for highlighting snps+eula. While we all want blob-free hardware like the expensive Talos OpenPOWER, all modern CPUs from Intel (ME) and AMD (PSP, MS Pluton) have management cores and firmware blobs (Intel FSP, AMD Agesa). AMD has a public roadmap for open firmware with OpenSIL, but PSP code is not public. One mitigation is to keep the device offline.
> Marvell.. good at hiding it
While this obfuscation is undesirable, many Arm vendors don't even bother with a fig leaf of upstream support. Would you recommend any Arm SoCs that have ongoing upstream Linux and u-boot coverage/fixes? Ideally, the SoC would also have OSS code for TrustZone (TF-A) and support for ECC memory. RockChip RK-3588 looks promising, https://www.collabora.com/news-and-blog/blog/2024/02/21/almo...
... which is 100% blobless. Or the Cavium MIPS64 chips (100% blobless). Or if you want Arm64, the Rockchip RK3399 (100% blobless). You have plenty of choices.
One mitigation is to keep the device offline.
Here's an even better mitigation: don't buy junk like this chip.
Stop blobwashing junk hardware like this with deceptive misrepresentations.
Does that include the Cavium Plus CN5020 in Ubquiti EdgeRouter Lite?
https://openwrt.org/toh/hwdata/ubiquiti/ubiquiti_edgerouter_...
https://www.insidegadgets.com/wp-content/uploads/2015/06/CN5...
I looked into running OpenBSD on ER3-Lite and was told that network routing performance was poor without a binary blob that ships in Ubiquiti router firmware. Cavium is owned by Marvell.
> if you want Arm64, the Rockchip RK3399 (100% blobless). You have plenty of choices.
The challenge is finding existing products that people can buy at retail. We're having this conversation in an HN thread about ECC. The only off-the-shelf Arm NAS (4xSATA, 2xM.2) that I've found with ECC + Debian is the QNAP above.
I will look for an RK3399 SBC that supports ECC memory, 4xSATA and at least one NVME slot, which could be installed into a 1U NAS chassis.
> Stop blobwashing junk hardware like this with deceptive misrepresentations.
ECC is a high priority requirement for an Arm NAS. Would be delighted to find a blobless Arm SBC with ECC support and enough I/O channels for use in a NAS. Otherise, ECC trumps blobs for an offline NAS use case in a 1U chassis, which can be physically secured in a rack against physical tampering.
The only downside in my view is the cost. Unbuffered ECC and the cost of using a workstation class chipset really pushes this into luxury territory. Plus, I'm never too sure what Intel's future plans are for successor processors and chipsets, which is why I settled on W680. I don't really want to go full blown Xeon.
The only thing I’m considering is the upgrade of my iPad Pro and probably MacBook at some point (not sure if I need it at all). My home computer is doing just fine with 4th generation Intel CPU and I don’t see any need for upgrade in upcoming years, if not a decade.
Dont be fooled (like me) by the DDR5-6000/6400/6800 ECC registered Modules, all Desktop Motherboards only support unbuffered Modules, and most dont even support DDR5-5600 ECC, only DDR5-4800/5200 ECC.
Some cheap AMD motherboards support ECC. But the future is unknown. Ryzen 8000 CPUs don't.
I like having 100% assurances and guarantees that it will work, making the W680 based platform I'm on a no-brainer for this application.
Interestingly, Jeff Atwood has changed his mind on ECC memory.
Extra weird is Nvidia singling out ray tracing as a use case which shouldn't use ECC... I suppose it's no biggie if a single ray goes the wrong way down the BVH, out of trillions.
Intel's marketing ploy on ECC is very desctructive and have costed many parties a lot of wasted time money & resources to handle problems caused by non-ECC memory ; and Linux Torvalds is absolutely right in roasting them for this.
Since ECC is seemingly not getting mandatory, I've been wishing CPUs would support "soft-ECC". That is, the OS could mark certain pages as needing "soft-ECC", and the CPU would then store (at least) three copies of that page in RAM. When reading such pages back from RAM the CPU would read all physical copies and compare. If the majority agrees it can use that, otherwise raise an error.
This could then be used for executable pages and important configuration data which occupies relatively few pages, and where integrity matters a lot more than speed.
There's probably some good reasons why this is non-trivial to implement, I've forgotten most of what I learned about the virtual memory implementation in CPUs. But a man can dream...
Of course you could trade implementation complexity for speed as always. My main point was to have effective ECC without any additional support from motherboard and memory modules.
I imagine soft-ECC would be more plausible if it could be applied/enabled:
1. Per-process (e.g. whole application)
2. Or fine-grained per-allocation within a process/application
I think both could be implemented through a (system) memory allocator by taking the advantage of page alignment LSBs and/or (non) addressable bits of the virtual memory. Those spare bits in memory can be used to store the encoding of ECC algorithm (Hamming, Reed-Solomon, or something more primitive but less robust).
And then, depending on the ECC algo, one could substantially minimize the performance impact of encode/decode by using SIMD.
Of course an application that marks GB of data this way could play havoc with other processes, so maybe some OS limits on how many non-executable pages a process can mark may be needed.
And there certainly might be some further dragons I'm not considering.
[1]: https://learn.microsoft.com/en-us/windows/win32/api/memoryap...
This allows the kernel and applications to protect important variables and data structures.
But sure, ECC all around would be best.
Btw, it was my understanding that CPU caches already use ECC?
The DDR5 standard allows either 40-bit channels or 36-bit channels in ECC DIMMs. (A DDR5 UDIMM has 2 channels, while the desktop CPUs use 4 channels provided by 2 sockets.)
The former choice corresponds with a 25% overhead, while the latter with a 12.5% overhead.
In the beginning there were only 80-bit ECC DDR5 UDIMMs, because the DRAM vendors preferred to make only x8 chips. More than a year ago, 72-bit ECC DDR5 UDIMMs have also apppeared, which use a mixture of x8 and x4 chips.
Nowadays no ECC memory vendor can justify a 25% overprice, because if they happen to use 25% more memory that is because they believe this reduces their production costs (by making a single kind of chip). Only an overprice slightly more than 12.5% is right.
Unfortunately, due to limited offer, the ECC DDR5 UDIMMs can still be up to 50% more expensive than non-ECC modules.
The FCC could just not allow computers to ship it without.
CPU makers like Intel and AMD could simply have their CPUs not work with non-ECC RAM.
Microsoft could e.g. require ECC RAM for Windows 12.
It is insanity that most computers shipping today do not use ECC and are thus unreliable.
With luck they'll crash, but most likely they will fail silently, while corrupting data.
Memory is different from all other resources in the system. We are conditioned as engineers, we know drives fail more frequently than other resources. When memory fails it is indistinguishable from a drive failure. There are some system behaviors that matter too, we tend to think that page allocation is random and on heavily loaded systems it appears to be, but on specialized systems it can be rather consistent so the verification can fail in nearly the same place, repeatedly. Riddle me this: what is more likely? A memory failure, a drive failure, or a postgresql bug that results in a corrupted row? Badblocks checks out on the server’s disks… if the data matters, it is extremely unpleasant going through that whole thing, it’s crystal clear after the fact but it’s a bloody nightmare in the heat of it all.
everyone must still be quoting numbers when we had 4mb of premium chips. now that all pcs have 8-128gb of the crumiest, cheapest silicon... i bet the failure rates are way more noticeable.
sadly, i got laptops for my company that have a PRO amd cpu and sodimm sockets... only to find out ecc sodimm ram is sold by one manufacturer gouging the NAS market with crazy insane prices.
Why Use ECC? (2015) - https://news.ycombinator.com/item?id=25167288 - Nov 2020 (98 comments)
Why Use ECC Memory? - https://news.ycombinator.com/item?id=23361577 - May 2020 (2 comments)
Should I buy ECC memory? (2015) - https://news.ycombinator.com/item?id=14206635 - April 2017 (224 comments)
Why use ECC? - https://news.ycombinator.com/item?id=10638324 - Nov 2015 (95 comments)
https://cr.yp.to/hardware/ecc.html (2001)
DEF CON 19 - Artem Dinaburg - Bit-squatting DNS Hijacking Without Exploitation (2011)
https://media.defcon.org/DEF%20CON%2019/DEF%20CON%2019%20vid...
My question is, how common are transmission errors over errors happening within RAM?
Adding protocol-level ECC on top only helps, although it is somewhat inefficient.
You have no idea if you have tons of errors and how many were corrected.
https://en.wikipedia.org/wiki/Chipkill
DRAM Errors in the Wild: A Large-Scale Field Study (2009)
https://static.googleusercontent.com/media/research.google.c...
So, let me ask again. I was to buy a NUC new or off Ebay, how can I be 100% sure it works with ECC RAM without having to spend half a hour researching CPU, mobo and BIOS specs for each single product I come across?
If I had a budget in the thousands, I would go with a Xeon server that comes with ECC pre-installed. I don't and have modest needs. I only want to splurge on ECC RAM to replace the original sticks.
(No "you don't need ECC for a NUC" reply please. That is not the point of my question, yet it is a far too common response)
There have been some NUC-like computers from ASRock industrial, Supermicro and others, with either Tiger Lake or Elkhart Lake CPUs, where you could enable in BIOS the so-called in-band ECC.
All these models are obsolete. Moreover, in-band ECC is an ugly and inefficient workaround. It can be used with soldered LPDDR memories, which do not have ECC variants, but it has worse performances than standard ECC. It is not cheaper, because it diminishes the memory capacity in the same ratio as any ECC and it requires a greater die area inside the CPU for its implementation (including a dedicated cache memory for the ECC bits).
There are many mini-ITX motherboards that support ECC (but you must check carefully the specifications, even if they are for AMD CPUs). For a smaller size than mini-ITX, there are 2 choices, either expensive industrial single-board computers, which usually have the 3.5" form factor of the PCB, or one of the so-called mobile workstation laptops from Dell, HP or Lenovo, e.g. a Dell Precision mobile workstation, which are also much more expensive than an equivalent NUC-like computer.
So, if a low price is desired and an up-to-date fast CPU, you cannot have ECC in form factors smaller than mini-ITX. If paying double or triple is not a problem, there are solutions.
If you want a preassembled small computer with ECC and a mini-ITX motherboard, there are some at companies like ASRock Rack or Supermicro, but they are much more expensive than if you get the best components and you assemble them yourself.
Sadly, my go-to Linux hardware manufacturers either don’t offer ECC RAM, or only offer it as an option on their absolute top-end machines. Yes, yes, the extra two thousand dollars for a machine with a six-year lifespan probably is worth it on a monthly basis, but man it still hurts.
or whatever dell 730 or something fits your budget
Old used enterprise server. None of them will be great at power/performance in typical (i.e. mostly idle) home use tho. Intel ones usually far better here
Not the cheapest, but I wanted to keep power consumption low for noise and reduced heating while still having good performance if needed.
I also considered a motherboard with IPMI on AM5(Asrock rack), but that was much more expensive.
Worked out quite nicely.
Nice. That got me curious, how many electrons are in today's DRAM capacitor? I tried searching but haven't found any recent info.
Binning is likely the problem.
Consumer GPUs, which do not have ECC, are notorious for frequent errors when they are used for general-purpose computation tasks instead of graphics or AI applications (where errrors happen, but they usually do not matter), so for any important computation it is recommended to repeat it and compare the results.
Only when non-correctable errors happen, which should be very seldom (e.g. less than one per year), the CPU generates a machine error exception, for which it can be identified that the cause was a (multiple-bit) error in one of the internal cache memories.
Therefore you could sense only extremely-high radiation levels, which would also require a customized OS kernel to recover from exceptions without crashes. Such great radiation levels could be much easier detected with some discrete reverse-biased diodes, which are also easier to place wherever the radiation exists, keeping the microcontroller or whatever device records the radiation levels either away or shielded from the radiation.
Nevertheless, I have never seen such a vendor giving a non-fuzzy description of those features. Moreover, after the sinking of Itanium, except for IBM the other server CPU makers use CPU cores that are also designed for consumer applications, so at most the server CPUs may have additional redundancy in some execution units (for validating the results) that is disabled in the consumer variants, to reduce the power consumption. However if that were true I do not see why the CPU vendors would not brag about it, to justify the premium price that is requested for the server CPUs.
For the applications that need very high reliability, it appears that in most cases it is preferable to use multiple CPUs operating in lock-step, which can be compared, ensuring that errors will be detected regardless in which part of the CPU they appear, instead of attempting to use a single CPU that is extremely fault-tolerant by using specially designed redundant execution units (because errors can appear anywhere, including on interconnections and buffers).