A DRAM Failure
complang.tuwien.ac.at
complang.tuwien.ac.at
The H, P and U series of Intel CPUs do not support ECC, so there is no chance of laptops using them with ECC.
As a notable change from the earlier mobile Ryzen CPUs, the Ryzen 6000 series, a.k.a. Rembrandt, support ECC memory (with DDR5 SODIMM; I mean real ECC, not the internal ECC of DDR5).
Unfortunately, I am not aware of any laptop with Ryzen 6000 that takes advantage of this specification change and offers ECC. I doubt that any such laptop will be introduced. Maybe in 2023, with the next generation of mobile CPUs we might finally see some competition for Intel in the mobile workstation market segment.
I am currently using under Linux a rather old Dell Precision with ECC memory (and with an NVIDIA GPU). I would also like an upgrade for it and the current integrated GPUs with 768 FMA units from Intel Alder Lake H or AMD Rembrandt would be good enough to avoid the high power consumption and cost of a discrete GPU, but no such luck yet.
I don't think its a good as memory with ECC, but its something.
edit (looking up question...)according to wikipedia, it has some ECC, but I'm failing to see the difference between the explicitly ECC and the regular memory:
"Unlike DDR4, all DDR5 chips have on-die ECC, where errors are detected and corrected before sending data to the CPU. This, however, is not the same as true ECC memory with an extra data correction chip on the memory module. DDR5's on-die error correction is to improve reliability and to allow denser RAM chips which lowers the per-chip defect rate. There still exist non-ECC and ECC DDR5 DIMM variants; the ECC variants have extra data lines to the CPU to send error-detection data, letting the CPU detect and correct errors that occurred in transit."
Real ECC must be end-to-end, from the memory controller inside the CPU to the memory chips and back to the memory controller.
This allows the CPU to be aware of any errors, which is essential e.g. for detecting the memory modules that will soon fail, so they must be replaced, like in this story, and like it also happened in my computers.
Moreover, not only the errors that happen in the memory cells are detected, but also the errors caused by electrical noise on the PCB traces or those caused by poor contacts in the DIMM sockets.
Solaris evicted the memory pages from the DIMM or bank, marked it bad, and reduced the amount of memory it 'saw' without crashing or rebooting, since the remedy worked before anything bad happened.
I only noticed it because of top showing less RAM than I knew I had installed...
On server-class machines, ECC errors often also show up in the system event log, so one can run "ipmitool sel list" and inspect the most recent messages, and they often point to the failing DIMM in a nomenclature that corresponds to how the slots are labelled on the mainboard or in its manual.
In this case, they are using a "gaming" mainboard, so this strategy probably doesn't work (no nice system event log).
Not sure if this would be seen in dmesg.
rasdaemon also attempts to report which physical DIMM / slot triggered the ECC error.
Memory breaks.
I had my share of broken memory too. In data centers you see this regularly.
Those are just two different stories. Nothing really connection them to make a scandal out of this
And all of them need to be so perfect that a billion single cells are refreshed, read and written every few ns.
It's very normal that sometimes a ram module breaks.
That's the reason why ECC exist.
ECC also exists btw for bit flip from space radiation.
"Good enough" is what he was told. It sure was, until it wasn't.
I guess people are too young to know this? But it was quite common in the old days to first check on faulty memory. And then it would be PSU, and Motherboard capacitor. Both are increasingly rare given how much improvement we have made over the past two decade. But faulty memory is still a thing.
The testing doesn't need much human interaction or babysitting, and if a problem is found it's usually the kind where you want to know ASAP so you can do an RMA/refund in a timely fashion.
Hardware checks including shorter memory tests were done, but this computer could not be on overnight. The prolonged test could be run when the user was out for a day.
This wasn't a corporate machine but a private one. As always several OS software configuration and driver issues were incrementally discovered and rectified, and as the crashes only happened very infrequently there was always the (wrong) thought of 'this might have been the problem', only to be refuted several days to a week later.
I'm not claiming I couldn't have done a better job on this one. Low frequency errors without a clear trace will always pose a challenge.
I wonder if the kernel couldn't perform ECC itself at a small performance cost by storing CRC bits in padding bits (that would require cooperation from the compiler), or other memory areas, checking them on load, and recomputing them before store operations. It may be easier to achieve on RISC architectures where load and stores are explicit.
I read a paper that seemed related to that some time ago, probably on an embedded microcontroller; the performance penalty was about 20% (you just need a few extra memory operations, performing ECC computation is pretty cheap, and is done 100% within the cache).
Alternatively, the kernel could checksum memory areas, and warn/panic if they change without having written to them. As a bonus, it could help protect against rowhammer attacks. Performance cost might be less.
Of course, ideally, ECC should be a standard feature, as it is on FLASH devices... CPUs should also support it even without explicit support from DRAM... Implementing the mechanisms I described above in hardware should be relatively easy.
Unfond, more recent memories of being asked by the Dell support tech to remove all the DIMMs in the server and start putting them back, one by one.
48 DIMMs. 15 minute BIOS boot time. Didn't even bother doing the math.
It was a better use of our time and money to mention "contractual four hour response time" and "We're about to send email to our corporate lawyer, do you want to reconsider your support response?"
Remember those problems where you are given n gold coins (among which one is fake) and a balance, and the task is to find the minimal number of weighting operations to identify the fake coin.
So, you go binary search looking for the borked DIMM, load 24 DIMMs see if it crashes, if no crash -> the borked DIMM is in the other pile, rinse and repeat...
As a simple for instance, out of what would be thousands of interacting and correlated issues, if some new line of RAM is put out that has more errors than before, your detector would have a hard time not seeing that as an increase in cosmic rays. And the real world deals out correlations that are very difficult to deal with. They could creep out slowly, or, AWS might order a whackload of these over the course of weeks and turn them all on at Tuesday at 9am. Trying to filter all those out is not necessarily mathematically possible.
Just aggregate the memory errors so you get a baseline for N machines and if that baseline goes to 10N then you likely had something interesting happen ... hopefully not WW3.
This is one of the great machine learning delusions.
My example was just to show that even MUCH noisier dataset (like phone motion) have been useful for detecting earthquakes that ECC should be MUCH easier.
Did they pull out the sticks of memory while the system was running? Can you remove and re-insert a stick and have any chance of the system continuing to operate?
I can easily imagine chips getting fried this way, unless they're specifically designed to handle such cases.
wondering is it common for people to specifically monitor their system log for correctable error related messages, do they consider the memory is faulty when there are correctable errors?
Is this true for servers? If I had a Ryzen based server, I’d use ECC RAM.
I think it is also not required for consumer Ryzen mainboards to support ECC but at least for the high end ones many do.
There are many ASUS and ASRock AM5 (and AM4) motherboards that support ECC, and for those it is typically writen in the memory section "supports ECC & Non-ECC unbuffered DIMMs".
When nothing like this is written, then ECC is not supported.
Moreover, all the motherboards with ECC support must have in the "Advanced" BIOS Setup an option for enabling ECC, which must be used, because the default is always to disable ECC.
With the Ryzen 7000 series there is an improvement over the previous Ryzen series, because in their specification it is written clearly that ECC is supported. Previously, the ECC support was not explicit, even if, unlike Intel they did not disable ECC, so you could hope that it works fine.
Now Intel no longer disables ECC in many Raptor Lake and Alder Lake desktop CPUs, but the motherboards with ECC support for Intel are much harder to find (because they must use a special workstation chipset, while for AMD it is enough to add the PCB traces for the ECC bits).
On both of my Ryzen ASUS motherboards (WS X570-ACE, and ROG STRIX X399-E GAMING) this is not true. I just slapped the DIMMs in there and powered the box up.
dmidecode thinks that the system has ECC enabled:
dmidecode --type memory | grep -e "Error Correction"
Error Correction Type: Multi-bit ECC
'amd64_edac' doesn't complain about being loaded on a non-ECC system.The closed-source version of memtest86 reports that it's running on an ECC-enabled system.
I also have the same ASUS Pro WS X570-ACE (bought in Q4 2019), which I use with ECC DIMMs, and I had to enable in BIOS the support for ECC.
In any case, one should always check for such an option in the BIOS, to avoid surprises.
It depends on the frequency. Occasional CEs are somewhat expected (on a large enough scale) and one can live with them, after all that's what ECC is for. When CEs start happening frequently on one machine, most likely a DIMM is going bad and will worsen over time, so one should replace it.
Random correctable errors are rare but they do happen - at least if you overclock your RAM ("gaming" RAM often is already pre-overclocked). Might just be confirmation bias but I noticed ECC errors and then later heard there was a solar flare around the time.
I also replaced a DIMM that was starting to get more frequent ECC errors once. As OP found the mapping for consumer boards requires to some trial and error - my motherboard documentation even had a table but the numbering was different from the one used in Linux :/
I don't think I'm ever going to use a non-ECC desktop again, the additional cost is not that high for the extra safety against silent corruption.
same here, but sadly you don't get to choose what you get when purchasing laptops, it is simply impossible to get ECC ram if you run mbp.
I can only remember one model of Lenovo that had an option to have a Xeon CPU with ECC RAM. I've never seen one with an AMD CPU.
When browsing their Web sites, these models are not obvious, because they are in the section for "enterprise" laptops, listed under "mobile workstations".
https://paste.debian.net/1257030/
It gets started via xdg autostart here, and will tell me about new "stuff" that happens. For it to work, your user will have to have permission to read the kernel event log/debug ringbuffer. I achieve that by setting the appropriate sysctl:
kernel.dmesg_restrict = 0However, you can also set a policy what the Linux kernel will/should do on its own when an ECC error condition has been detected: The `edac_core` module has options such as `edac_mc_panic_on_ue`, which, if set, will trigger a kernel panic upon detecting an Uncorrectable Error in system memory. Depending on your use case, this can be better or worse than just logging it.
Zfs based Nas with ECC, smart check for HDD, system check including ECC too.
However, you can cheaply extend Hamming codes with, e.g., a parity bit, so that errors in a slightly larger radius are detectable, though for obvious reasons you couldn't correct such an error.
No comment on what sort of algorithm is used for ECC, though it might be worth mentioning that the above is a pretty general feature of error correction, where it's possible to cheaply or even for free be able to detect errors in a larger radius than you're able to correct.
First, let's define a correctable error — that means an error where you have enough additional information (in your Hamming-code stream, in a parity bit, whatever) to repair the error, e.g. a one-bit error when you're using ECC RAM.
An uncorrectable error, then, is one where you do have enough information to detect the corruption, but not enough information to figure out the correct repair for the corruption. (The amount of information required to detect corruption is always less than the amount of information required to correct it.)
With ECC RAM, usually exactly two bit-flips will produce a detected, but uncorrectable, error on read-back.
With non-ECC but "with a parity-bit per word" RAM (which exists, and is a bit cheaper than ECC RAM), you can't correct anything, only detect. (Which is sometimes all you need, if you're willing to do the calculation over again.)
All that being said: completely separately from these hardware-level features, some operating systems (e.g. macOS) compress memory, and generate page-level checksums of memory pages as they're compressed. A checksum failure during memory-page decompression can also trigger the kernel to throw this kind of "uncorrectable error" itself.
There may also exist (i.e. it wouldn't be impossible for there to be) RAM modules that continuously calculate page checksums for each page on each write; and then check the contents of pages against these, perhaps asynchronously, in sort of the same way a ZFS scrub works. I've never heard of this being done, but it feels like the sort of thing you'd implement in hardware for an extremely "ruggedized" system like a Mars rover. If this approach were to be implemented, it would also emit "uncorrectable" errors.
>> HPE Fast Fault Tolerant (ADDDC)—Enables the system to correct memory errors and continue to operate in cases of multiple DRAM device failures on a DIMM. Provides protection against uncorrectable memory errors beyond what is available with Advanced ECC.
https://techlibrary.hpe.com/docs/iss/proliant-gen10-uefi/s_c...
https://www.hpe.com/us/en/collaterals/collateral.4aa4-3490.M...