Non-ECC memory corrupted my hard drive image [video]
youtube.com
youtube.com
Apparently this was conceived as a market segmentation scheme: people outfitting servers could get ECC when they pay a huge premium. They would thereby not be tempted to cheap out and buy consumer-grade equipment, otherwise wholly adequate to meet all their needs at a radically cheaper price.
That we cannot get laptops or even desk machines with ECC, and so have them crash frequently, is seen as a trivial side effect of the strategy. If you did not hate Intel enough before, you may increase your hatred accordingly. Intel doesn't hate you back; they simply care not even a little how you feel.
(Historically, just running Microsoft software was overwhelmingly more likely to be the cause of a crash than a memory bit-flip; and there were orders of magnitude fewer RAM bits at risk. Microsoft succeeded in getting customers to accept and even expect frequent crashes; before MS, a program crashing was grounds for a refund.)
But I did see a Lenovo model, IIRC, that had some kind of Xeon and ECC. Not sure what the noise and battery life situations on that thing were, though.
It's also my understanding that the DDR5 data-in-flight ECC is a mandatory feature because the link between the memory modules and everything else is so error-prone that the system would simply not function without it.
But in practice, it will probably be more reliable than DDR4 without ECC, since now you need 2 cosmic ray flips, or 1 plus a manufacturing defect flip, and the defect flips will probably be uncommon-ish.
It's too bad data in flight isn't protected without old fashioned ECC on top of that, but it will probably be a big step up, the same way that flash memory is now very reliable even though the actual uncorrected errors are probably worse under the hood.
But even without it, it would be no worse than DDR4 with no ECC at all, where the only indication of bad ram is just that it crashes or memtest fails.
It sounds like they intend to make the chips all slightly bad and then correct it, so a few errors here and there might not indicate a true bad stick that needs replacement anymore.
My particular laptop does have a "pro" CPU. However, I would be surprised to no end to learn that it supports ECC. This particular model sports an MBP-level price tag [0], but is absurdly cheaply built. Even for "customer facing components", that are easy to compare, such as the screen (terrible colors) and case (creaks if you look at it wrong). HP doesn't offer ECC RAM, not even as an upgrade, so I really don't think the additional lines are physically present.
---
[0] I don't remember the specific number, but it was within 100 € of a 14" M1 MBP with 32 GB RAM and 512 GB SSD. That's counting a RAM (8 -> 32) and SSD (256 -> 512) upgrade which were made with components bought separately (though they were rather high-end).
I’m not sure what you mean by “frequently”, but my non-ECC machines definitely do not crash “frequently”.
> before MS, a program crashing was grounds for a refund
Source?
The difference in significance is so great that by comparison a mere crash is no problem at all.
In fact you design systems with hair triggers to 'crash' on purpose as readily as possible. Trying desperately to crash at all times every millisecond all day every day.
IE, halting all operation of some subsystem, or the whole thing, the second a single bit wrong is detected. Better to kill the hd or the entire machine than let it keep running one second after getting any hint it might not be 100% trustworthy.
'crash' in quotes because really all you want to do is halt, and you're doing it on purpose, but that is still a crash, in the sense that your application does not want to halt and it isn't necessarily halted gracefully with any chance to conclude anything or save anything. Those are just more operations you can no longer trust to be correct, and so should not be allowed to do.
If it's crashing once per hour, it's probably unstable drivers/software or flaky hardware that needs to be RMAed, not random bitflips.
I get the claimed need for EEC, but yet I don't. I own a number of web properties and regularly update logos, software packages, html files, PHP files, ruby files, etc. No corruption.
I'm continuously writing metrics on a non-ECC system. There are 3 places that the bitflip can affect the data: pre-writing (checksum will be correct), when flushing to the drive (checksum mismatch), calculating the checksum (mismatch again). There may be some silent corruption (in 1 out of 3 cases), but I've been scrubbing the drive every day for ages and there's not a single error.
Non-ECC seems to be doing well. Other claims could really use actual numbers.
The bad habit that nailed me was probably this one: I would ball up a whole bunch of photos in a big .zip; (to enhance copying speed and maybe privacy) then next year unzip that, add that year's selected photos and zip it again. Year after year. That added up to a lot of trips through RAM in which one bit flip could take out an image. Bugs in compression software aren't impossible as the villan, but that's not my bet.
Solution: I now zip up a few years together and copy each chunk I don't ever rezip old large chunks. No new errors noticed.
Microsoft, uniquely exercising monopoly power, was able to institute a "no refunds, nohow" policy and make it stick: Windows95 machines commonly crashed several times a day, enough so that "crash" came to mean, instead, you have to reinstall the OS, which happened at least one per several months.
And here we are.
"Infrequently" would mean I had not heard of it, or that I heard only of isolated instances, well publicized.
The Xeon series of laptop processors does support ECC just at a quite large premium.
It’s added complexity and cost for something that rarely would benefit most consumers. Now, you can argue that the complexity and cost is a nonissue on modern setups and I would probably agree.
But Intel has long had desktop grade hardware with ECC support. The 440GX chipset supported ECC and I ran a Dell GX1 SFF with 768MB of ECC PC100 for yeeeears with a 450 MHz P3 and later upgraded to 1.4ghz tualatin-256 via a slotket adapter.
The 440HX /socket 7/ chipset supported ECC. And that’s a Pentium 1 chipset.
The 440BX/GX and 450NX supported ECC and that’s with desktop pentium 2 and 3 chips.
The 820/820E/840 supported ECC with desktop celeron and pentium2/3 chips
845/845e/850/850e/860 pentium4 chipsets support ECC
875/e7205/e7221/e7230 did with desktop pentium 4 and pentium d chips
925/925xe/955x/975x did with desktop pentium 4/pentium d/core 2
It’s more sparse now that they moved to the IMC, granted. But Intel has long had multiple chipsets per generation with ECC support for desktop grade hardware.
Let's say you are afraid to destroy the through holes of the multilayer PCB by de-soldering the broken parts, or heating bordering parts up too much, whatever.
In most cases you can pull the caps just off from their pins, clean the pins, (even with paper tissue only) and simply solder the new ones to the old pins still anchored in the board.
Looks weird, but works :-)
You slide the needle over the component leg and then melt the solder. The needle slips over the component leg. Wait for the solder to cool and pull the needle out. Hey presto the leg is separated from the PCB pad. Also works great for cleaning solder out of a hole after traditional desoldering.
BTW, AMD's Ryzen CPUs and "Pro" APUs (APU = cpu with integrated graphics) will all do ECC, but not "officially", and not all motherboards have BIOS that supports it. You need to google around or just try it.
And, by your own suggestion, this isn’t a fault of Intel and more of OEMs sucking.
I mean, I guess Intel could have tried mandating it or something but meh
How frequently would you say you encounter a crash that you can pin down to a lack of ECC memory in your laptop or desktop?
I have a Ryzen desktop with ECC, and it registers about one bit-flip per week. I don't know how many of those would become crashes, but I'm more worried about the ones that wouldn't.
This isn’t normal. You have a bad memory module.
A non-ECC machine should able to support memtest86 (a rigorous memory testing tool) for a week straight without a single bitflip.
Having a bit flip per week is so far away from normal that it’s definitely a bad memory part.
Memtest86 is not magic. It is designed to look for systematic failures, even if intermittent. It does that pretty well given it has only bus-level access.
Maybe OP's house is in the Van Allen Belt or something?
> This isn’t normal. You have a bad memory module.
> Having a bit flip per week is so far away from normal that it’s definitely a bad memory part.
It's hard to know what's normal. If you're somewhere in the world like northern Sweden the chance of cosmic rays goes up from being less protected from the suns rays (Hey, but northern lights!) and if you're at higher altitude; along with less interesting things such as how shielded the ram is (server chassis, if you removed the rack doors or not).
Memory density has a lot to go with it too, the chances of bitflips increases a lot.
Anyway; regardless of the theory behind it: I've run global operations with multiple-thousands of physical machines with high density ram (and 16-32 DIMMS in each server) and noticed at least 1 bit flip a week on average for each machine. But some locations were worse impacted than others.
Annoyingly, ECC is not created equal, there are some correctable bit-flips that are hidden from the iDRAC/iLO: this is the same type of ECC that the rPI is using (referred to as on-die): https://datasheets.raspberrypi.com/rpi4/raspberry-pi-4-produ...
As with all things like this YMMV, but ECC is doing a lot of heavy lifting.
yeah, you have bad ram. Or you overclocked it.
I have a Ryzen server with 32GB ECC RAM running 24/7 for a year. 0 bit flips. Not a single one.
Of course, overclocked or not, I don't get memory errors. Just the occasional EDAC event.
Deep in the archives of a well known tech company is a very well documented case of a bit flip that caused the wrong function to be executed in a C++ v-table. The big oof was that this function was the equivalent of an SQL "drop table", and just happened to be 32 bytes off of a very benign function that did something like stat(). Really funny stuff once the crisis is over :)
In limited cases, a checksum is good enough. If you checksum outgoing data, and verify it on reception, then it being corrupted in transit whether on the network card or the cable can be detected and transparently compensated for.
Really, we can do much better than to "get used to it".
It can (and should! whenever possible) be improved, not fixed. There's always that pesky gamma that can hit a specific transistor, even if it is deep underground. Gamma cannot be fully stopped. At certain scales data corruption becomes directly measurable. And yes, corruption levels vary between pieces of hardware.
The reason why we don't is laziness and market segmentation, mostly.
You can even just buy (well, chipaggedon aside) ARM cores that have 2 chips running in parallel and faulting when the result is different
See dual-core lock-step Arm chips (used for automotive).
You replied: it should be improved, not (it cannot be completely 100%) fixed
Your error: replying to a strawman
I thought L2/L3 is ECC (at least on Intel, though L1 I think is parity only)
Your statement is in practical terms false.
Error correction exists, and adding enough redundancy to make any corruption have such a tiny chance of being undetected that the whole history of the universe could go by without it happening is a problem that was solved over 60 years ago with the invention of the Reed-Solomon code,
Way back I had a Pentium 133 doing firewall duty in a closet. It did approximately nothing besides iptables, but of course any machine has logs, updates and so on going on.
After running fine for months one day it suddenly died. I rebooted it. A few days later it died again. Another reboot. Then it died for the last time and failed to boot at all. Examination showed the disk was corrupt and couldn't be mounted. Further examination showed that one of the memory modules was loose for some reason, could be that it was never firmly in and I just bumped the box when messing with something else.
Then came the wasted weekend of dealing with that my normal internet connection relied on the thing that was now completely broken.
And that was the luckiest case I can imagine, when the broken machine contains no data of actual value. Since then I'm very paranoid, always run memtest on any new RAM I buy overnight, and have ECC where it's possible to have it.
Yeah, I do the same, but I've learned that you have to do it regularly.
In one of my desktop machines, the RAM ran fine for like two years. Then, all of sudden, random Firefox segfaults, etc.
Whipped up a memtest ISO, and sure enough, one of the sticks was bad.
You normally have a scrub time that can be configured in the BIOS, which also adds a regular verification of the entire RAM at regular intervals, just in case something goes wrong in some rarely used part of the memory.
I'll go for #1 in most cases, as long as the system is to be relied upon for anything deemed important.
Few months later powersupply outright died (had ~8years at that point), replaced it with good one, no memory errors.
In this particular case, a bad PSU would be the end of the PC. It's an HP dekstop mini. Basically a laptop without a screen, powered by and external adaptor that puts out a single 12V line. All further conversions are done on the motherboard somehow.
Personally, I would've just checksummed the individual failing files rather than the disk image and only back up the bad files separately. There are all kinds of ways for a disk image to fail and I wouldn't spend a second longer on it than absolutely necessary. The whole memtest permutation setup also would've been too much work for. E, I would just declare the motherboard faulty when two sticks that otherwise pass the test fail in specific configurations. A new motherboard is cheaper than super specific RAM sticks.
Onto the discussion of ECC RAM. In a perfect world, all memory would be ECC... but try finding some high performance 16GB sticks of ECC DDR4 RAM like what you'll see on gaming computers. I don't even think they make anything comparable in terms of speed and definitely not costs. I guess you don't really know that you needed ECC until it's too late.
I spent many years on hardware consultation and was amazed at the all the times I had to explain it was just a what if insurance like any other things their business was mitigating against. Sometimes they'd even decided they needed to save costs in non-ecc ram when it was $4 a gb in difference, or (during the FB-DIMM era) there wasn't even an option to avoid it.
Never really understood the resistance towards it.
Maybe the lack of evidence before the Google study and people thinking RAM manufacturers were trying to rip them off or something.
The "never had a problem so why would I need" it attitude with no way to know if an issue was caused by a bit flip was most baffling.
Here ya go:
https://nemixram.com/16gb-ddr4-3200-pc4-25600-ecc-udimm-2rx8...
It doesn't have pretty lights on it, but it does seem to be in the same speed class that gets called "gaming RAM" by a _whole_ bunch of retailers.
Unless you have a particular reason to keep your order history in Newegg, just buy direct from NEMIX.
(IDK about NEMIX's relationship with Amazon, as I don't buy things from Amazon.)
Still the experience that synology btrfs provides is nowhere as good as ZFS (due to a lot of limitations).
I replaced it with non-ECC memory and never had an issue. In fact you can combine ECC and non-ECC memory without any problem.
Then the author goes through a process of randomly buying memory and hoping it works.
Anyone know what that "high density" memory problem is about? Maybe it's a misunderstanding of memory channels and ranks?
When I used to sling PC hardware, as memory densities increased, you would run into things like 'This Motherboard will only take a 256MB module if it has 8 chips on both sides (16 chips total), it will not take a 256MB module with 8 chips on one side (8 chips total)' Depending on the board, it might only register part of the capacity and be stable, might register part of the capacity and be unstable, or just not boot at all.
Sounds strange, as - without knowing any better - I'd expect the memory interface to be agnostic to the chips implementing it. eg "address 0x11111111" on 'high density' ram would be seen exactly the same as "address 0x11111111" on other ram
The mind boggles. :)
Your server is faulting and preventing itself from corrupting data.
An actual bad chip may fall into a different category, I think that's when technologies like Chipkill[0] might come into play.
Think it mysteriously crashed once or twice in the 4 years I had it and the HP diagnostic light came on.
https://community.fs.com/blog/ecc-vs-non-ecc-memory-which-on...
Unless you live over 4000 meters over the sea level, like to compile while flying or live close to an unshielded nuclear reactor, you don't need ECC.
And most memory problems you can fix by better cooling, and better shielding.
How would you know?
Yes though, when you have gobs of memory, you're unlikely to see the effects of bit flips. That's not the same at them not occuring.
Memory errors are heavily concentrated to bad modules. They’re not evenly distributed across all RAM. Always test the stability of your memory when doing a new build.
Except that it's not for many use cases. It's great for servers but for people on their personal and/or work computer, it's simply not that useful.
Seriously: which percentage of developers have ECC on their development machine(s)?
As developers we live in a world of SSH, cryptographic hashes, checksums everywhere, Git repositories (that is a big one), Merkle trees, digital signatures, reproducible builds (which are gaining traction), etc.
Heck, I'm torrenting the latest Debian or Devuan .iso image. My torrent client is using every known trick under the sun to make sure that should anything go wrong, the broken data shall be discarded and re-downloaded. Download is done, I dd the image to some installation medium. I can then verify its checksum matches the official one. A bit flip didn't slip by unnoticed.
All the music I carefully ripped from my audio CDs? They're all cross-checked with an online DB of known bit-perfect rips. There's an accompanying file containing each song's hash and I can verify at anytime that all my files are 100% correct.
But really most of all I live in a world of Git repositories. My entire Emacs config is versioned under Git (I know YMMV but I like it that way). Some people version under Git their entire user dir.
Tell me how my lack of ECC is going to really make life miserable here?
I have nothing against ECC... But if I want to upgrade my AMD 3700X to a 7700X, apparently I cannot get ECC.
And that's totally fine: I certainly won't discard the 7700X because I cannot get ECC for it.
And if anything looks suspicious, running Memtest is the first thing you should do.
I've had bad RAM at times. I'm still there.
The CPU, RAM and mobo manufacturers need to get together and make ECC RAM mandatory. It's absurd that we have machines with gigabytes of storage using microscopic (nanoscopic??) capacitors that doesn't have this basic protection. And honestly this should have happened years ago.
(Edit: And before anyone says DDR5 is ECC by default, that's not quite true although the difference is a bit subtle: https://en.wikipedia.org/wiki/DDR5_SDRAM#DIMMs_versus_memory...)
Devils advocate: if it is just some bits in character flipped it's entirely recoverable while flipping some bits in compressed stream for video would corrupt more.
Back in the 90s we had a database server which had a stuck bit in memory normally mapped to the page cache. This caused sectors to be written to the backing software RAID which couldn't be read back in (because I think some checksum was corrupted when written and then failed when read back). It took an absolute age to diagnose this. I think I only worked it out by eliminating everything else.
Every single one of my non-mini computers uses ECC ram except for my laptop. If someone would release a framework laptop motherboard that supports ECC ram (and preferably risc-v) I'd finally be able to close to reliability gap. It blows my mind that we say "Well sure, we COULD make infallible ram, but that would cost a tiny bit extra so instead lets just hope nothing bad happens." That's right up there with not wearing a seat-belt because I haven't needed one yet.
This. If I could, I would. The only reason my laptop doesn't have ecc is because the manufacturer doesn't offer the option in any macines I otherwise want.
That comment was very misguided in trying to suggest that there is any valid excuse to tolerate unreliable execution hardware. git and ssh and md5sums do not mean that it's ok if your very brain can't be trusted to deliver data from one part to another within itself, or spit back the same data that was put in a cell. Everything else is built upon that!
I can't speak for ECC at the moment but ZFS has definitely saved me from data corruption that would have been left to manifset otherwise.
That question won't help you evaluate the demand for ECC simply because the supply is strangled. Those who want it have to make compromises to get ECC: get a Xeon, or get one of few AMD motherboards with matching CPUs and overpriced RAM without a guarantee that it will end up working.
I mean, if you want a _guarantee_, then sure, get one of those certified-for-ECC motherboards.
But, like, as far as I know, going _at least_ far back as the Phenom II (released in 2008) AMD desktop processors have always supported ECC RAM. And -as far as I know- ASUS motherboards for said processors have always supported dropping in ECC RAM (and Linux and memtest and friends have always agreed that ECC was enabled and functioning in such a system).
Source: Personal experience with Phenom II, Threadripper, and Ryzen 5 CPUs and ASUS motherboards, and looking-from-a-distance at the rest of the AMD CPUs between the Phenom and the Ryzen 5.
Per 1gb of RAM, you can expect to see 266 bit errors per month[1-2] if you are using your PC 16h per day. Multiply that by 64GB or 128GB of RAM and it's crazy to think that you won't run into any of the stability issues.
[1-2] https://static.googleusercontent.com/media/research.google.c... [1-2] https://en.wikipedia.org/wiki/ECC_memory#Research
That seems like an overestimate. How can memtest ever pass on non-ECC RAM if errors are that frequent?
Memtest is looking for reliable failures, not evanescent one-off events.
With ECC in the refresh path, such an error could be corrected and the right value would be overwritten over top of the bad one. Then errors would not accumulate, but would instead be "scrubbed". Mainframe machines scrub their RAM. Disks too.
...so ? we're not talking about slow deterioration (which is why refresh is needed), we're talking about bit flip from cosmic rays where cell changes state completely.
But I didn't have time or desire to publish a study on our experience, so there's no hard numbers.
The Google study you linked says "25,000 to 70,000 errors per billion device hours per Mbit".
Assuming bits, {25000 to 70000}/1e9 * 3016 1000 = 12 to 33.6 bit errors per month. Assuming bytes, it's 96 to 268 bit errors per month.
Apparently you've meant bytes, not bits. (I'm a pedant, but was also just unsure and interested in the numbers.)
FWIW, a comment I ran across that confirms toast0's view[1]: "It's a bimodal distribution - you either have many errors (due to a defect somewhere) or basically zero. If you're on the good side of the distribution, with only extremely rare errors, then you probably don't need ECC. But without ECC, you don't know whether you need ECC!"
[1] https://www.realworldtech.com/forum/?threadid=198497&curpost...
> But really most of all I live in a world of Git repositories
That's great for you, but 99% don't even know what Git is, just like checksums, cryptographic hashes, Merkle trees, digital signatures and reproducible builds.
You just miss the one imortant thing: ECC isn't that helpful where you have the means to check and verify the data. ECC is the only way to at least know what something is happening with data when there is no way to check.
To give you a slight idea I would tell you an anecdote from my L1 life almost two decades ago:
I visited a client who claimed what the PC was working erratically and constantly threw weird error messages.
Welp, the usual deal, just some ugly software or a virus. In the first 3 minutes I got like 5 errors about failing to load a .dll from C:\WINDOWS\SYSTEM33\USER32.DLL, nothing unusual, just need to.. WAIT. Why system33? Stupid virus masking for a well-known folder? Doubt. So I go to C:\Windows and I see:
System32
System33
System34
If you are a smart fellow you probably already understood what that was a bit-flip error which somehow managed to be in the in-memory copy of MFT of the system drive. And this is the only reason the user noticed it - because sometimes programs wouldn't start and sometimes there was weird error messages. If that error was in the data area of some program - it wouldn't be discovered at all.> Tell me how my lack of ECC is going to really make life miserable here?
Git will break if RAM is bad just like anything else. That it checksums everything won't save you from checking in corrupt data, the filesystem itself being corrupt, or some internal git structure becoming corrupt. Losing your repo because something in it was written wrong is very much a possibility.
Having multiple machines involved helps, but it's not a complete fix, because the possibility exists of something damaged being transmitted from a broken machine to a good one, ensuring there's no good copy anywhere.
There's really nothing software can do to operate correctly with bad RAM all of the time. Instructions for the software are in RAM. The OS that the software expects to behave right is in RAM. Various buffers used for disk access and networking are in RAM. An application like git assumes all of that is performing correctly, and can't compensate for every possible malfunction that could happen.
For actual verification, unless there's a bug my understanding is that very trustworthy. But by that point it's already too late. Okay, you know something is broken, but that won't give your good data back.
But that only tells you that Git data is intact and that all the hashes match. If git got a corrupted file to start with, then correctly hashed it, everything will verify 100% and still be broken.
Recently I discovered that one of my SSDs was quietly failing without setting off any warnings; doing a chkdsk showed that some files had already gotten corrupted. One of them was my backblaze backup index!
Even though I have automated backups (backblaze + macrium backups to a NAS), recovering files from them is non-trivial. If I were to lose work to non-ECC ram who knows how long it would take me to reconstruct a known-good work environment and file set. Imagine if you're working on something huge and hard to validate like neural net weights where corruption can occur silently and be hard to detect after it's happened?
Apparently Ryzen 7000 cpus can use ECC. I've heard reports that AMD needs to release an AGESA update, though, and ECC DDR5 memory availability is terrible. I'm hopeful that the situation will improve, because I also want to update my desktop. I've been using ECC memory since losing a filesystem on a desktop when a DIMM went bad.