ECC RAM should be a human right
dmitrybrant.com
dmitrybrant.com
FTFY
I never did figure out a way to verify that the ECC is working/if it is able to report errors (to the kernel?). It was also a bit hard finding the right ram, but there’s a brand called Nemix that I found. I was a bit sketched out by the brand, but the chips themselves were Micron.
Not all of them, afaik the am4 processors with graphics don't unless you have a pro branded version.
> I never did figure out a way to verify that the ECC is working/if it is able to report errors (to the kernel?
The easiest way is to induce an error. Some of the memtest tools try? But also, you can tweak memory voltage and timing to make things less reliable, and you should be able to get errors and error reporting as a result.
One thing that could change this quick is if Apple used ECC. Others would follow to not seem inferior. I don't hear Apple users complain much at all about lack of ECC options.
Comes up in almost every Apple Silicon thread.
ECC means not only that you know precisely when you've gone too far with overclocking, but potentially allows overclocking a bit further, relying on that some amount of trouble can now be tolerated.
It also means you're not going to break your OS by playing with this stuff. Memory corruption carries a huge risk of disk corruption, which can mean things like corrupt data, random crashes or an unbootable system that persists even after reverting everything to defaults.
Assuming stationary processes this is pretty nice. Maybe bump back 5% for margin.
Better than nothing, but not full ECC either.
https://web.archive.org/web/20170207162829/https://www.micro...
There are sites and services that do "binning" where they test the specific chips and you can buy ones that have been vetted to clock higher.
And the performance differences can be significant: 20% or more in certain workloads.
[1] https://www.pcgamer.com/what-are-xmp-profiles-and-how-do-i-u...
Good luck with that. https://chainsawsuit.krisstraub.com/20160223.shtml
Last time I was looking for a case, this was a real issue.
Some people prefer to trade off stability for a slight performance improvement.
In my PC, I've got a Intel Core i5-3570K (3.4ghz stock) overclocked to 4.5ghz or something silly like that. I used the motherboard manufacturer's "one touch overclocking" feature to determine that speed years ago and haven't touched it since.The performance improvement is more than slight. Stability is rock solid across thousands of hours of gaming. I'm not running advanced cooling. $60 Corsair closed loop cooler and three Noctua case fans. System runs cool and quiet, fans throttle up and down nicely. Near silent when idle and still rather quiet under load.
That is a pretty decent gain in performance for zero drawbacks and essentially zero effort.
Personally, I run a 4670K at 4Ghz, and it's been rock stable as my main machine for the last 8.5 years. It's finally beginning to show its age in 22-23, but the longevity and use I got out of a $200 processor is incredible.
But when CPUs auto-turbo to 5+ Ghz, I agree, overclocking sort of loses its luster.
I'm just saying that it has a potential appeal for gamers too, so it's not just a datacenter type of technology that some nerds want to play with.
At the very least it'd make overclocking safer and easier, so any manufacturer making gamer type boards with a lot of overclocking settings in the BIOS should like the idea of it.
Silent Data Corruptions at Scale https://arxiv.org/abs/2102.11245
Harish Dixit has two other papers available https://arxiv.org/search/cs?searchtype=author&query=Dixit%2C...
Revisiting Memory Errors in Large-Scale Production Data Centers: Analysis and Modeling of New Trends from the Field https://users.ece.cmu.edu/~omutlu/pub/memory-errors-at-faceb...
Justin Meza has a similar body of research. https://www.semanticscholar.org/author/Justin-Meza/144606145
https://kilthub.cmu.edu/articles/thesis/Large_Scale_Studies_...
If you're relying on your computer to be 100% error-free (eg: doing professional work), paying some extra for ECC RAM support isn't even a drop in the bucket. You either do or don't, money isn't an object.
If you're Joe Average browsing Hacker News or playing some games at home, who cares if you win the bit flip lottery?
Would it be nice for ECC RAM to be more mainstream? Sure. But the fact it is not isn't breaking anyones' backs either.
What a horrible failure mode! Kudos to the OP for even thinking to investigate this.
The mentioned memtest86 tool sounds quite useful - so useful, in fact, that it seems it should be part of the operating system. If my OS is perfectly willing to write corrupted memory back to disk, and it's designed to run on thousands of different OEM hardware configurations that may or may not have ECC RAM, then I would expect it to proactively monitor for faulty RAM and bitflips.
Is there a good reason why this isn't the default behavior of operating systems (maybe it is - I use Mac and don't know much about Windows)? It seems like a trivial diagnostic tool that could prevent a lot of headache. But perhaps the problem is there's no reliable test without some number of false positives, so popping a warning on the screen that your RAM appears unhealthy seems like a good way to confuse the average user. But then again, so does silently corrupting a video file.
Because any equivalent of memtestx86 will flag a LOT of hardware for the buggy piles of shit that they really are.
If Windows started flagging people's crappy cheap hardware, Microsoft would get a lot of grief and have to spend a chunk of money on customer support and PR.
The OS layer is probably the hardest place to do online application data integrity checking: it's too low level to know what data is important and which parts will change and why in a way which allows checksumming to work effectively and efficiently, but too high above the hardware to be able to check the integrity of memory without a massive performance penalty (especially when it comes to how memory is moving in and out of caches). Most solutions work at a higher level in the application itself or lower level with ECC RAM as mentioned in the article.
Apple Diagnostics check memory as well, though maybe not as thoroughly as the Microsoft or memtest86 do (multiple passes).
I've had memory errors that only appeared on two of eight passes through the whole memory test overnight. They certainly caused occasional issues with games though.
Windows has the Memory Diagnostics Tool, and your linux probably has a memory tester option in the boot manager. Of course you have to run them manually, which is a bit bothersome. Thee have been attempts at kernel patches for linux to test memory in the background, going at least back to 2005 [1], but there were probably some before that. The simpler versions just test whatever memory is free, the more sophisticated aproaches try to move stuff in physical memory to free each region periodically so it can be tested. But in the end it's spending CPU cycles on a problem most users never experience.
1:https://groups.google.com/g/comp.os.linux.development.system...
Does Apple hardware come with ECC RAM? If anyone could make it make sense as a business, it's them.
Printers and plotters from this era used ECC modules most of the time.
But by the end of the century, they were replaced by unbuffered, unregistered 16/32/64 bits modules.
Every mid range server still use ECC. Entry HPE Servers use ECC UREG (unregistered, 9 chips) modules, while mid range and more use ECC REG modules (9 chip + interface controller onboard). Ironically, UREG module are more expensive than ECC REG.
Also, most workstations used ECC modules. Less frequently since 4-5 years.
At work I have two working slot 1 PIII 800's each with 1GB ECC (4x 256MB DIMMS) on a regular Asus board (doing nothing but waiting to go home with me one day). The board reports the RAM is in fact ECC and that it is enabled.
Wikipedia says "By the mid-1990s, most DRAM had dropped parity checking as manufacturers felt confident that it was no longer necessary.". https://en.wikipedia.org/wiki/RAM_parity
I'd love to read a technical deep dive on RAM reliability over time. You'd think with increasing memory cell density and overall larger RAM the number of absolute errors on a desktop computer would be going up over time.
Thanks.
This was the reason I always used Mac Pro machines, and at the end, iMac Pro: those machines had ECC RAM, and incidentally always seemed more stable over months and years than the MacBook Pros I had, which didn’t.
(Of course, there are a lot of reasons why a laptop might crash or malfunction an order of magnitude more than a desktop machine. The benefits of ECC RAM are really hard to catch definitively in the wild. We know they exist though, so I’m in the “it’s insane not to use ECC RAM for any use case that persisting data and then reading it back later” camp.)
And god the horrific price of thoses sticks...
I've basically given up on maintaining large digital media collections for long term purposes. It puts my OCD in overdrive.
I’m running btrfs with integrity checks, in RAID 1 so it can automatically heal. Yet it’s non-ECC and therefore still has this gaping Achilles heel.
ECC RAM is not a prerequisite for using ZFS.
Matt Ahrens, co-creator of ZFS and still one of the main developers, said this [1]:
"There's nothing special about ZFS that requires/encourages the use of ECC RAM more so than any other filesystem. If you use UFS, EXT, NTFS, btrfs, etc without ECC RAM, you are just as much at risk as if you used ZFS without ECC RAM. Actually, ZFS can mitigate this risk to some degree if you enable the unsupported ZFS_DEBUG_MODIFY flag (zfs_flags=0x10). This will checksum the data while at rest in memory, and verify it before writing to disk, thus reducing the window of vulnerability from a memory error.
I would simply say: if you love your data, use ECC RAM. Additionally, use a filesystem that checksums your data, such as ZFS."
[1] https://arstechnica.com/civis/threads/ars-walkthrough-using-...
However, even if you're not using ECC RAM, you're much better off using ZFS or btrfs, because due to their frequent checksumming and checksum validations, these filesystems will usually detect memory corruption much sooner than if you didn't use them.
This could be immensely helpful in scenarios such as the great^4-parent poster.
Note that bad hardware is not limited to non-ECC RAM. ZFS and btrfs help just as much in detecting other kinds of bad hardware, such as bad SATA cables, bad disks, bad disk/SATA controllers, bad CPUs, bad power supply, etc.
But of course, once these checksum errors are flagged by ZFS/btrfs, it's a signal to test and fix/replace your hardware, not keep using a machine with bad hardware.
And yes, while ZFS and btrfs cannot fix errors that happen before the checksumming takes place (e.g. due to bad RAM, bad CPUs, etc), they can still detect these kinds of errors (in some cases, at least), especially when they happen after checksumming already took place. And they definitely can detect and fix errors in the rest of the data-to-storage-and-back path (e.g. bad disk cables, bad disks or disk controllers, etc).
AMD seems to support ECC on some consumer-level chipsets, but it's a nightmare to sort it all out and verify that all the bits are actually functioning properly.
It's about as reckless as running any other filesystem without ECC [1].
In fact, you're much more likely to discover memory errors earlier if you use a checksumming filesystem than if you use a non-checksumming one.
I was wrong twice huh!
ECC can only report >1 bit flip per word. It can't correct it. If your memory is going bad, the chance of multiple bit failure is going to be much higher than the stray bit flip from other causes (mind you, I have 32GB ECC running right now and have seen zero bit flips over 1.5 years).
I'm not arguing against ECC. I have it and run it myself. But this blog post is an argument for ZFS and regular backups. Failures can occur at many other places than RAM. I had a drive that had corruption and I thought went bad. Turns out, bad SATA cable. A week later I upgraded from cable internet to 1Gbit fiber. My speeds were not at all what I expected. Turns out... bad ethernet cable. Lightning can strike twice I guess.
> By the way, you’d better believe that your disk(s) have all kinds of error correction schemes built into them, which work automatically and transparently.
Well...
https://twitter.com/xenadu02/status/1495693475584557056
https://www.tomshardware.com/news/sk-hynix-sabrent-rocket-ss...
When a chip starts going bad, the thousands of error reports will more than make up for the errors that slip through.
Just make sure reporting is working.
> Well...
Well what?
Those drives were lying about finishing writes, which is barely related to losing data that was written, and especially data that was written more than a few seconds ago. And if drives didn't have ECC, data loss would be a million times more common.
It's completely possible for a single DRAM cell to go bad permanently.
> But this blog post is an argument for ZFS
ZFS won't save you if what you tell it to put into a file is wrong. And I think that corrupting ZFS' own in-RAM data structures voids your ZFS warranty.
> and regular backups.
Of course you should have regular backups. But do you really feel like spending hours restoring from a backup (and then maybe weeks finding things you didn't restore)? Do you feel like trying to guess how far back you have to go to get to a good backup of a corrupted file? Do you want to lose whatever new work got corrupted between the time your memory failed and the time you finally noticed it?
OP's machine was corrupting data for weeks. You can't just roll all your work back to several weeks ago. Most of us can't, anyway.
You need backups and ECC.
> I had a drive that had corruption and I thought went bad. Turns out, bad SATA cable.
If a bad cable did that, something wasn't doing SATA CRC checking the way it was supposed to (or wasn't reacting to detected CRC errors the way it should have). Similar in spirit to ECC checking for memory.
> My speeds were not at all what I expected. Turns out... bad ethernet cable.
It probably got slow because of retransmissions. If there hadn't been CRCs on both the Ethernet layer and higher layers, it could possibly have just silently corrupted your data.
You need to cover as much of a computer system as you can with error detection and correction, or you always lose. Even one big failure in a lifetime can pay for a lot of ECC RAM.
I would have thought that the additional cost would be minimal (additional wiring on the logic board in some cases), but maybe this is just more artificial market segmentation?
Thinking about it, when a DIMM doesn't have extra RAM chips for ECC, you could reuse the ECC pins to run a parity bit to each chip. It would cost nothing. And with DDR5 having 2x8 ECC pins, it would work with any chip layout: x16, x8, or x4.
> Unlike DDR4, all DDR5 chips have on-die ECC, where errors are detected and corrected before sending data to the CPU. This, however, is not the same as true ECC memory with an extra data correction chip on the memory module. DDR5's on-die error correction is to improve reliability and to allow denser RAM chips which lowers the per-chip defect rate. There still exist non-ECC and ECC DDR5 DIMM variants; the ECC variants have extra data lines to the CPU to send error-detection data, letting the CPU detect and correct errors that occurred in transit.
So in some ways it is better than previous generations, but it gives vendors another excuse not to implement full-coverage ECC. That's my guess of why GP said it complicates things.
I have had a scenario where the machine failed to boot at 7200 until I reseated the memory, so there's definitely a limit on the physical media being hit.
I very highly doubt so. I remember BSOD and Guru meditations (on the Amiga) before that. Once I switched to Linux suddenly no more BSOD and very little kernel panics. I've had at times my desktop reach six months of uptime.
I think bit-flips are a great excuse for unreliable software.
BTW I'm not saying ECC is unnecessary: I'm saying it's unlikely the lack of ECC was the reason for most of the BSOD you saw throughout the ages.
I’ll always buy ECC after multiple data loss events from ram suddenly going bad, worth every penny to have ECC ring alarm bells rather than put 2+2 together after random BSODs and silent corruption
“A large-scale study based on Google's very large number of servers was presented at the SIGMETRICS/Performance '09 conference.[6] The actual error rate found was several orders of magnitude higher than the previous small-scale or laboratory studies, with between 25,000 (2.5 × 10−11 error/bit·h) and 70,000 (7.0 × 10−11 error/bit·h, or 1 bit error per gigabyte of RAM per 1.8 hours) errors per billion device hours per megabit. More than 8% of DIMM memory modules were affected by errors per year.” https://en.wikipedia.org/wiki/ECC_memory
A random stick of non ECC memory might be far above average or have several errors per minute, but you just don’t know.
Source: https://www.anandtech.com/show/16900/samsung-teases-512-gb-d...
That just shows how useful ECC memory is not that these bit flips didn’t occur.
ODECC was added because they wanted to be able to use DDR5 chips which would have had unacceptably large error rates without it. In other words that improvement is before the binning process, so they are selling chips with a vastly higher innate error rate to the point where the average DDR5 stick could actually be worse than the average DDR4 chip, it’s hard to say without large scale testing from multiple manufacturers.
Mentioning W680 feels pointless. You've always been able to buy high-end workstation-class motherboards and stick ECC in them. The entire point of the article is that all computers should be using ECC RAM, not just the expensive, workstation class computers.
If you just /want/ ECC, then yes $450 bucks is expensive. You don't /need/ it, though, so this is neither here nor there.
The whole point of this article and discussion says otherwise. Don't come to me to complain about that. Your comment should be a top-level comment complaining at the author.
I fully agree with the author that ECC should not be reserved for expensive computers, of course, but I'm just here to point out that W680 is not a response to the author's concerns at all, period. W680 is a continuation of Intel's status quo.
Also, you can use asterisks to create italics.
Even so that's an extremely trivial task. It doesn't need a "more advanced" anything. It needs them to not deliberately disable the code or remove the tiny tiny amount of circuit.
Do game machines need ECC? I'd say no. They are optimized for cost-performance. The worse that can happen with a memory error is a lost game.
ECC memory isn't that much more expensive. I was fortunate to build my current machine a few months before memory prices doubled. Previous generation Xeon CPUs aren't very expensive. Same with motherboards. And another option is to just buy a used server. These are super cheap now as so many companies are moving to cloud computing.
> Memory manufacturers assure us that desktop RAM is so reliable that it doesn’t need ECC, that the probability of bit flip events is so low that it’s not worth the extra “cost” of ECC
Yea, and they are right.
Could you explain why you disagree?
Most people have very little need for ECC. The author didn't even know that they wanted ECC until they were unlucky enough to get a stick of RAM that failed, and failed in such a way that the OS booted, but a file silently corrupted (not that common, because if a chip fails it usually doesn't silently fail like that).
The author is basically asking for ECC for free.