Why Use ECC? (2015)
danluu.com
danluu.com
Like others, I was super happy when I could by a Ryzen chip at retail prices and get ECC rather than buying a Xeon sku which was an extra $300.
[1] When I first built the system I was getting about 2 a week which I tracked down to a 'weak' ram stick. Replaced it (Corsair, and they replaced it for free) and now it's about one a month.
I ran this for a couple of years and never saw a change.
I believe that based on published memory error rates for memory at that time, I should have seen a few errors. I am not sure why I never did. Maybe consumer desktop memory was less dense than server or workstation memory, and the published error rates were for the latter?
From 2008 to 2017 I had a 2008 Mac Pro at work with ECC memory, and from 2009 to 2017 I had a 2009 Mac Pro at home with ECC memory. I'd occasional look in the memory section of the System Information report and never saw any ECC corrected errors reported.
With the Mac Pros, it is possible that I just never happened to check the report between the time of a corrected error and the time the counter was next reset (assuming it resets every boot...if it is cumulative than I have no idea how I never got an error).
Or put differently, your survey size was quite a bit different than mine has been. Also there is always a chance that the bit flip is in the ECC bits not the main memory bits. ECC DIMMS typically have 72 bits per 64 bit long word, vs 66 bits per 64 data bits for parity memory, and 64 bits for 64 bits on memory with no protection at all. So it's hard to make a 1:1 comparison without knowing the exact layout of the memory and whether or not error detection bits are present.
[1] Soft (and Hard) bit errors are generally spec'd as a probability per unit time per bit, so more bits means higher probability of a bit flip per unit time.
Of course, that's when I'm not paying the bills. I don't have ECC on my home systems, because of added expense, lower performance (no XMP ECC ram when I was shopping), and extra research needed. If a consumer oriented motherboard says it supports ECC, that may mean it lets you use ECC ram, but doesn't enable any of the reporting, which would be mostly useless; not completely useless, because used ECC ram from server retirements is sometimes very inexpensive, but it's often fully buffered, which is less likely to work in consumer platforms.
That is largely a feature of OS software, not the motherboard. Linux on AMD motherboards which support ECC has reporting. There is no reason why it wouldn't. If the ECC is there and active, Linux can get the information (EDAC).
On Intel though, they flat out fuse ECC off on consumer CPUs for market segmentation reasons. If you want a consumer motherboard with ECC, you go AMD.
Is there any special term which does describe this capability? How would I find out if that reporting exists? Only by reading the Manual and looking through the BIOS pages?
I think the point about chipkill is salient, why isn't it more widely used if the benefits there are obvious? I think ECC is completely not worth the price for home users if you can do regular backups, run memtest, and compare checksums for any data corruption. Even with errors near the hard/soft boundary, you'll likely catch those by just running memtest for a longer period of time. DRAM errors get progressively worse, so you'll catch any errors that way as well using memtest again.
Datacenters have more surface area which cosmic rays can attack and are also more likely to see weird hard errors which might not have been caught in QC/QA, which is very different from isolated home computers or phones which aren't on 24/7 or can restart easily. If you have a datacenter and have truly sensitive unreplicated data, get chipkill. If you are home user, do regular backups, and run memtest. It's that simple. Hard drives have moving parts, and more failure modes, so pick a file system with checksums. CPUs/GPUs and SSDs might or might not have ECC caches because they have other ways of reducing hard/soft errors in SRAM.
People may disagree, but the study I would like to see is one that takes into account the denser environment of a datacenter vs the isolated home user and all modes of data loss/corruption in each case. We know Google replicates data in GFS, and a home user can do the same with backups.
Don't assume that that means there are never any errors. DRAM errors often depend on access patterns: While dialing in the speed and timings for my RAM initially settled on a configuration that did not report any errors in memtest but later during heavy usage (e.g. compiling LLVM with make -j64) reported ECC errors.
With a median of 10,000 FIT or lower, it's zero errors every 5-10 years, and DDR4/5 might actually be below that if the rate of improvement between DDR1/2 was anywhere between 2x-10x continued. You'd have to overlook this important trend to see any benefit from ECC. Yes, there are DRAM errors, but if they can be found by memtest, bg scrubbing, self-tests and don't affect low-use scenarios, then it's meaningless to be talking about ECC for home use.
Setting speed and timings on ECC modules seems super flaky to me, do you know what speed and timings the separate ECC logic can handle? Can you turn off ECC and test DRAM independently? Maybe what you're seeing is the ECC logic throwing errors and not the DRAM.
The hard part when doing support is convincing the customer that the errors in your software are from their systems. I had to learn the failure patterns of a couple large filers so I could look for them.
Now we use git which is fine, and git uses much stronger checksums. My main concern is that since those SHA{1,256} checksums are expensive they are typically generated only once and then only used for lookups. If you intentionally corrupt git files, it is hard to have git notice without explicitly running a check. These checks should be done it cron nightly as part of backups.
Anyone know the git invocation to put in cron?
edit: probably this? https://git-scm.com/docs/git-fsck
For example, right now all the new Tiger Lake CPUs contain a so-called "In-Band ECC" device, which allows error detection and correction even when using LPDDR4x memories, which do not have an ECC variant, like the DIMMs or SODIMMs.
This works by storing the ECC codes in a reserved part of the memory, so increasing the reliability is paid by a slight reduction in the memory capacity and in the memory bandwidth.
Nonetheless, those who are risk-averse, like me, would prefer this trade-off and would enable the "In-Band ECC".
However, you cannot do that because Intel disables the "In-Band ECC" feature on all Tiger Lake SKUs, except on 3 SKUs intended for Embedded and Industrial Temperature Range, which are presumably more expensive.
"In-Band ECC" is even disabled on the other 4 Tiger Lake SKUs for the Embedded market, which leaves them without any visible advantage over the normal Tiger Lake SKUs.
When Intel competed with Intel, this kind of business decisions probably made money for Intel, but now, knowledgeable customers should better buy from competitors, e.g. the new Ryzen V2000 for embedded applications, which support standard ECC memories without problems.
I agree that this practice sucks, but to play devils advocate - is it possible that this is due to binning / yield maximization?
Now Intel 10nm still seems to be so bad that maybe it is interesting enough to bin on that, and maybe that's the cause Tiger Lake is even more limited on ECC capabilities than before in the various SKU. Although we already saw a restrictive move on Comet Lake.
Instead of continuing to multiply their SKU (I think they had enough even 5 or 10 years ago), Intel should go back to the drawing board and ship interesting microarch on better nodes...
Oh definitely not. They disable all kinds of arbitrary features that are totally unrelated to yield.
I put together a new SuperMicro server a few months ago and went with 256GB of ECC. Yeah it's very expensive, but it's absolutely worth it if you care about the integrity of your data.
See: https://media.blackhat.com/bh-us-11/Dinaburg/BH_US_11_Dinabu...
Where are you now Artem?
Clocked at 2400MHz, sure. I can't find a single non-sketchy listing for even 3200MHz unbuffered ECC memory, let alone the 3600MHz you want for current Ryzen chips. And those are just the speeds at reasonable prices, ignoring the 4000-5000 range.
At least DDR5 should improve things.
https://www.crucial.com/memory/server-ddr4/mta18adf2g72az-3g...
> let alone the 3600MHz you want for current Ryzen chips
Current DDR4 JEDEC-approved speeds top out at DDR4-3200. Anything beyond that is technically overclocking, even if it's overclocked from the factory and marketed to run at those speeds. Assuming that your platform supports overclocking DRAM, you can also run DDR4-3200 ECC memory at 3600 MT/s, if you feel that the performance uplift is worth the risk of stability issues.
I don't care about JEDEC. I care about the XMP spec that's promised to me by the manufacturer, so I can rely on it. I'm not against overclocking myself but if I'm trying to boost it by a significant amount then I have no idea what to set the timings to and there's a good chance it just won't work, with no recourse on my end.
Also that module costs way more than $65.
All of the ECC-capable platforms I've used (including my AMD-based desktop PCs) have supported patrol scrubbing. What have you used that doesn't?
Support for patrol scrubbing is a bit more variable. It's supported in my ASRock motherboard (B450M Pro4). It was also supported in two previous ASUS motherboards (AMD AM2 and AM3 platforms, respectively), although they called it something else that I can't remember. As well, a lack of a setting for patrol scrubbing doesn't necessarily mean that it's unavailable; it could just be enabled but not user-configurable.
The motherboard doesn't really get a say on whether the OS can correctly use ECC support on AMD as far as I know. As long as the memory is correctly initialized in ECC mode, the OS should always be able to use its EDAC driver to interact with it.
This isn't entirely true. The motherboard needs to physically support ECC (i.e., traces for a 72-bit memory bus need to be present) for it to work. ECC support can also be disabled within the BIOS firmware, either explicitly via a user-adjustable toggle, or by the motherboard vendor hard-coding the setting to disabled.
In any case, if the OS reports ECC as working, then ECC is fully functional.
It's a binary thing. Either the motherboard supports ECC, or it doesn't. If it supports ECC then you can always tweak the scrubbing from the OS.
And I wouldn't say faster memory is necessarily more error prone. But it doesn't really matter, because non-ECC products have to have extremely low error rates to be viable. Even if a speed is at the high end of "extremely low", once you add ECC you'll have an exceptionally solid component.
ECC isn't popular on desktops because Intel, which supplies the overwhelming majority of the processors for platforms that use DIMMs, doesn't support ECC at all. As such, the market for people who can actually use ECC in a desktop (much less would want to) is very small. The only reason unbuffered ECC DIMMs are even on the market at all is because low-end servers use them.
The real reason ECC is not common is because it is perceived as unnecessary for desktop workloads. Software crashes and corrupts data all the time because of bugs, not bitflips. Most consumers care so little about their data that a single hard drive crash will wipe it all out anyways. Why spend 12% extra on memory if consumers hardly see any value from it?
But since you don't know that you backup a corrupted file for years i'm pro ECC and pro FS-Checksumming.
I learned my lesson with some corrupted Sound-Files, the bad thing, i had to check every single file near that death block cluster, the good thing, all of them where ripped from my CD's.
You would have to have some way of detecting that memory is corrupted, which non-ECC platforms do by writing a known pattern to memory and accessing it to verify that the pattern matches. You obviously can't do that during runtime because that memory is being used by the kernel and applications.
The first generation of UltraSPARC CPUs with 8mb cache didn't use ECC on that cache. This caused some issues; so the "fix" cut that to 4mb mirrored and checked. These were the 250Mhz parts and everybody went to 400Mhz new process stuff instead anyway, but it was a bit of a thing at the time.
However, the fuss meant that generation of CPU hit ebay cheap and in volume, which I was quite grateful for.
I've heard of people collecting known-bad DIMMs for this purpose, or trying to blast their RAM with heat or radiation.
Why does this have to be so difficult?
In my experience getting the information isn't "difficult" so much as it isn't "common". On my FreeNAS device (which has the Xeon) it also has an IPMI controller which lets me ask it over the network how many ecc errors have been seen. (and it can act like a remote console as well but we'll leave that for another time :-))
So find out how the motherboard is tracking ECC, then find the tool that talks to that system to ask it for the logs of ECC errors.
The problem is how to determine whether your system is actually capable of detecting/logging errors end-to-end. There needs to be a standard way to inject an ECC error and see what happens.
"0 errors" is very different from "ECC not available," and that's a distinction I'd expect the operating system to be able to make.
For example, here's how Linux shows ECC and non-ECC memory reports:
# ECC
$ sudo edac-util -v
mc0: 0 Uncorrected Errors with no DIMM info
mc0: 0 Corrected Errors with no DIMM info
edac-util: No errors to report.
# Non-ECC
$ sudo edac-util -v
edac-util: Error: No memory controller data found.In case the OS/BIOS does not support overclocking or fails report memory error properly, certain data pins of the dimm slot could be physically shorted with metal wire to add a persistent stuck bit. Without ECC the system will not boot. I've seen it done on multiple occasions but web search is drawing a blank for the moment. IIRC it was being discussed a lot when first gen Ryzen cpu came out and people were arguing left and right whether ECC support is present on motherboard X.
Finally some memtest software claims to be able to inject errors on a software level. I can't vouch for them as I have very little experience but it would be very simple to find out.
Or play with rowhammer tests.
https://ark.intel.com/content/www/us/en/ark/products/77495/i...
When they are new, memory errors are very seldom, if at all.
Nevertheless, after several years of use, I had many cases when some DIMM modules began suddenly to have frequent errors and being notified by ECC allowed me to replace the modules (or the motherboards if they were very old) before causing unrecoverable data corruption.
But you can have far fewer redundant bits if you only want to detect error s. (Parity). In the worst case the computer reboots. But you must foresee this because hardware can die for many reasons (power failures, component failures etc)
Yeah or one bit in your checksum...imagine the fun :)
Quick search later - yes, you can: https://linux.die.net/man/1/edac-util
For all the discussion I've seen about this topic, I wish more people actually published their numbers.
I wanted to buy ECC but I have a hard time finding 2x32GB of DDR 3600, no XMP and I don't care about the data on that desktop as much as I care about the data on my nas.
While there's nothing technically preventing vendors from factory-overclocking ECC memory the same way, ECC's (perceived) main audience is reliability-critical systems, and so it doesn't make sense for vendors to market a feature that their target audience wouldn't appreciate.
What you can do is get two 32 GB DDR4-3200 memory modules (example: https://www.crucial.com/memory/server-ddr4/mta18asf4g72az-3g...), and then overclock them yourself to 3600 MT/s. As with any overclock, your results may vary, but that's also the case with factory overclocking as well.
2017 https://news.ycombinator.com/item?id=14206635
Discussed at the time: https://news.ycombinator.com/item?id=10638324
What about Apple? Will they support it now they're unshackled from Intel?
It's just Intel that doesn't.
I don't think there's a mass-market laptop with ECC support, but AMD's desktop CPUs have supported ECC for as long as I can remember. That being said, it's been an optional feature that is up to the motherboard maker to implement, and not all do.
I have a Dell Precision laptop from 2016, with a Skylake Xeon and ECC memory.
That being said, unbuffered ECC DRAM, all other things being equal, probably won't overclock as high since the extra memory chip places a slightly higher load on the memory controller than equivalent non-ECC memory. In that sense, you could consider it "slower."
Now that AMD CPUs are taking a big market share, hopefully we see a better market for ECC hardware.
>>When Google used servers without ECC back in 1999, they found a number of symptoms that were ultimately due to memory corruption, including a search index that returned effectively random results to queries.
Oh and not to mention debugging, can you trust your program or your hardware, who makes the wrong query, how to find out which node needs to be restarted?
And it outlined how a bit flip at google created a huge security risk, that an outside attacker can exploit.