ECC RAM on AMD Ryzen 7000 Desktop CPUs
sunshowers.io
sunshowers.io
[1] https://www.reddit.com/r/Amd/comments/lzxqod/list_of_am4_mot...
So this definitively settles it, that the AMD+ASRock combo is truly ECC RAM?
Kind of a shame they don't have a mini-itx/matx X670E board (only ASUS does).
https://www.supermicro.com/en/products/motherboard/h13sae-mf
0. Not supported at all (i.e., if you plug ECC RAM, your system won't boot).
1. ECC RAM can be plugged but the ECC functionality is not used (i.e., there is no relevant traces/circuitry/etc).
2. ECC functionality is present (i.e., the circuitry is there) but it was not validated by the motherboard manufacturer to be functioning correctly (i.e., detecting/correcting errors).
3. ECC functionality is present and was validated by the motherboard manufacturer. This is the level one would expect from the server-grade boards from a reputable manufacturer like Supermicro.
In case of the AMD processors, when you see "ECC supported", it's anyone's guess which level it is. This is in contrast with Intel, where if it says CPU/chipset supports ECC, then you know it really does. I bet Intel won't allow a motherboards manufacturer to sell a board with chipset like W680 without validated ECC support.
What separates Ryzen from its professional-grade counterparts is that ECC support is an optional part of the consumer Ryzen platform spec, which means that it's up to the motherboard vendor to enable support for it. Some motherboards don't have any support at all, some have ECC support as an explicitly-advertised feature, and many have ECC support but it's not explicitly advertised (simply listed as a footnote in the manual).
That's different from the way Intel does it, where Intel has explicit control over what the platform's feature are, and uses that control to aggressively segment their markets. Intel's approach makes easier to reason about ECC support as a buyer, but you pay for it with the lack of flexibility compared to AMD's platform (and you literally pay more for ECC).
You mean flexibility to claim ECC support but not doing any validation to make sure it actually works? Does any AMD motherboard manufacturer actualy states that "ECC is supported and has been validated"? I think the muddy waters that AMD has created would warrant such an explicit statement.
The actual ECC functionality is part of the memory controller, which is entirely AMD's domain. The ECC functionality of the memory controller is fully validated.
Whether or not ECC support is present on the rest of the platform is the responsibility of the system builder. However, if the platform isn't fully ECC-capable (e.g., you're using non-ECC DIMMs, the motherboard doesn't have the necessary electrical traces, ECC support is disabled in the UEFI, etc.), this will result in the memory controller disabling ECC support, which is visible to the operating system and is something that you -- the end user -- can verify.
> Does any AMD motherboard manufacturer actualy states that "ECC is supported and has been validated"?
Yes. ECC support will be listed in the motherboard's spec sheet and manual. There also motherboards marketed for professional use where ECC support is explicitly advertised in the vendor's marketing.
> I think the muddy waters that AMD has created would warrant such an explicit statement.
AMD's official statement regarding ECC support on Ryzen is that it's supported if the motherboard also supports it. I'm not sure how they can be any clearer without making ECC a mandatory feature of the platform.
Now all Zen 3 and Zen 4 CPUs, both desktop and laptop, have explicit ECC support, which means that ECC must be validated by AMD in all of them.
If any current AMD CPU happened to have defective ECC, that would be a completely defective CPU, which must be replaced by the vendor.
Despite the fact that all laptop Ryzen 6000 and Ryzen 7000 CPUs support ECC, I have not seen yet any AMD laptop or SFF computer with ECC support. On the other hand it is much easier to find AM5 MBs with ECC support than Intel W680 MBs.
(Edited to reflect the current state of the website, as it used to say ECC support was present[3].)
[1] e.g. https://www.amd.com/en/product/13186
[2] https://community.frame.work/t/responded-amd-batch-1-guild/2...
[3] https://web.archive.org/web/20230513075641/https://www.amd.c... (here “FP7r2” is the version for use with interchangeable DDR5 modules rather than with soldered LPDDR5)
I have saved the page from your link on the 17th of July. In the saved page it is written:
"ECC Support: Yes (FP7r2 only; Requires platform support)"
The same was written at all AMD Rembrandt and Phoenix models, and it has been written in all such mobile CPU specifications at least since the beginning of 2022, so at least during a year and a half, if not more.
The ECC support was available only with DDR5 SODIMM memory, not with LPDDR5 memory, and only in the FP7r2 package, and neither in the FP7 nor in the FP8 packages.
Perhaps AMD has discontinued the FP7r2 package, but this is not mentioned in the specifications. Either that, or they have decided right now that they may charge more for PRO models, or save on testing on non-PRO models.
Either way, it seems a stupid move for AMD to degrade right now their mobile CPU specifications, when Intel will launch Meteor Lake, which is likely to be better than AMD Phoenix. AMD mobile Zen 5 will be better than Meteor Lake, but until its launch it remains about a half of year, supposing that it will not be delayed.
This is not the case.
Ryzen, Threadripper, and EPYC use a unified memory controller that has fully validated/qualified/supported ECC capability. The only difference between the memory controllers in these CPUs is the number of them (Threadripper and EPYC will have multiple memory controller to support the extra memory channels).
When AMD claims that ECC on Ryzen isn't validated, they're talking about the platform as a whole, not the CPU specifically. ECC support on consumer CPUs depends on the motherboard supporting it. Unlike on Threadripper and EPYC (and also unlike Intel's approach), ECC is not a guaranteed feature of the Ryzen platform, so people who want that functionality need to explicitly looks for motherboards that have it (in the same way that you'd need to explicitly verify PCIe bifurcation support).
However, if the motherboard supports it, ECC on Ryzen is a fully supported, validated, you-can-RMA-if-it-doesn't-work feature.
(Edit: These are way more attractive for home server use because they consume a lot less power than the Ryzen CPUs when idle)
If anyone is looking for one of these standalone, outside of a full build, I bought a Ryzen 7 PRO 4750G about a year or 2 ago from AliExpress. It's been running a homeserver 24x7 since then and I never had any issues with it. YMMV of course.
It is expected that AMD will launch a desktop Zen 4 APU in the near future. If that happens, it remains to be seen whether ECC will remain enabled in it, like in the current laptop packages, but there seems to be no reason to disable it.
While during at least one year and a half all the laptop CPU specifications for Ryzen 6000 and Ryzen 7000 specifications have included a clear statement that ECC is supported, in the very recent past this statement has been removed from all AMD mobile CPU (non-PRO) specifications, for unknown reasons.
This is normally specified in the "Memory" section of the specifications, in something like "ECC & Non-ECC, Unbuffered Memory".
Beware of mentions of "On-die ECC", which is present in Non-ECC memories and which is irrelevant.
You also have to buy ECC DDR5 UDIMMs and you must be careful to not buy by accident ECC DDR5 RDIMMs, which are incompatible with AM5 motherboards. The ECC DDR5 UDIMMs have a width of either 80 bits or 72 bits. The width does not matter, as long as it is not the 64-bit width of Non-ECC DDR5 UDIMMs. (For some vendors it is cheaper to use identical x8 chips in ECC and Non-ECC modules, despite wasting some capacity, which results in an 80-bit width; there is a myth that DDR5 ECC DIMMs must have a width of 80 bits, the myth is wrong, because the standard includes 36-bit channels, which result in a 72-bit width for a dual-channel DIMM; for instance Micron makes 72-bit DDR5 ECC UDIMMs)
The last time when I have checked, ASUS had the most AM5 motherboards with ECC support. I like most the PRIME X670E-PRO WIFI because it has the best PCIe expandability beyond the slot occupied by the GPU.
However there are many other cheaper MBs, when less connectivity is enough.
I've seen allegations that for some vendors, text like this on their non-server/workstation boards turns out to mean that ECC modules will work in the board, but without the actual ECC function.
The only way to gain more confidence is if the downloadable MB manual has an exhaustive description of the BIOS options.
If in the BIOS options there is one for enabling ECC and perhaps additional related options, e.g. for configuring scrubbing, only then there is complete certainty about ECC support in the MB.
However most recent MB manuals no longer have a complete BIOS description, so they may be not helpful.
At least with the ASRock or ASUS AM4 or AM5 MBs that I have used, whenever "ECC & Non-ECC, Unbuffered Memory" was specified, the MB really had ECC support.
Hard to imagine asrock wouldn't run the lines given all the other components are there, but I guess it's conceivably possible. If the kernel says it has ECC, it's right. If not, return the board as defective and use a different vendor.
Some motherboards list ECC RAM support as a feature and have ECC RAM on their QVL for memory. Be warned ECC UDIMMS are expensive.
For checking if ECC RAM is working in linux use the dmidecode command.
For monitoring ECC errors use rasdaemon.
Asrock isn't the only brand that reliably supports ECC memory. See below.
https://www.reddit.com/r/homelab/comments/15zuj70/ecc_udimm_...
It's what I've been using for 15+ months as my main development desktop.
This is how they show up in the system (Linux) from dmidecode:
Handle 0x000F, DMI type 16, 23 bytes
Physical Memory Array
Location: System Board Or Motherboard
Use: System Memory
Error Correction Type: Multi-bit ECC
Maximum Capacity: 128 GB
Error Information Handle: 0x000E
Number Of Devices: 4
Handle 0x001A, DMI type 17, 92 bytes (this section is repeated 4 times though, once per DIMM stick)
Memory Device
Array Handle: 0x000F
Error Information Handle: 0x0019
Total Width: 72 bits
Data Width: 64 bits
Size: 16384 MB
Form Factor: DIMM
Set: None
Locator: DIMM 1
Bank Locator: P0 CHANNEL A
Type: DDR4
Type Detail: Synchronous Unbuffered (Unregistered)
Speed: 3200 MT/s
Manufacturer: Kingston
Serial Number: [snipped]
Asset Tag: Not Specified
Part Number: 9965745-028.A00G
Rank: 2
Configured Memory Speed: 3200 MT/s
Minimum Voltage: 1.2 V
Maximum Voltage: 1.2 V
Configured Voltage: 1.2 V
Memory Technology: DRAM
Memory Operating Mode Capability: Volatile memory
Firmware Version: Unknown
Module Manufacturer ID: Bank 2, Hex 0x98
Module Product ID: Unknown
Memory Subsystem Controller Manufacturer ID: Unknown
Memory Subsystem Controller Product ID: Unknown
Non-Volatile Size: None
Volatile Size: 16 GB
Cache Size: None
Logical Size: NoneYeah. The comment I'm replying to sounded like they're wanting to still use the AM4 platform. Maybe I misread it, or they've adjusted the post for clarity in the meantime. :)
$ sudo ras-mc-ctl --errors | tail -n5
14 2023-08-20 20:16:41 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x64e31c78, cpuid=0x00a50f00, bank=0x00000011
15 2023-08-23 17:17:49 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x64ea5188, cpuid=0x00a50f00, bank=0x00000011
16 2023-09-03 16:52:15 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x64f4d227, cpuid=0x00a50f00, bank=0x00000011
17 2023-09-15 21:37:59 +0200 error: Corrected error, no action required., CPU 2, bank Unified Memory Controller (bank=17), mcg mcgstatus=0, mci CECC, memory_channel=0,csrow=0, mcgcap=0x0000011c, status=0x9c2040000000011b, addr=0x36e701dc0, misc=0xd01a000101000000, walltime=0x65071ed7, cpuid=0x00a50f00, bank=0x00000011
This is with an ASRock B550M-ITX/ac and a AMD Ryzen 5 PRO 5650G. It used to work the same with a Ryzen 5 3600 (using a dedicated GPU for video output) before I upgraded the CPU.To detect and log ECC activity on modern GNU/Linux, you will want to have the "rasdaemon" service active. I will decode MCE (and other hardware-related) errors and persist them to the database that is shown being queried above.
Edit: thinking about this a bit longer: that frequency is actually so high that you may well have a broken module in there. Note how it is the same module and the same address every time.
Along the lines of "computer gets toasty doing work, ecc errors start happening"?
Stuff like that could just mean the memory sticks need pushing in a bit more.
No Memory errors.
No PCIe AER errors.
No Extlog errors.
No MCE errors.
Do I need to enable something so these errors get logged or have I been misled by dmidecode and dmesg?
Some/most(?) AM4 boards can enable "PCIe AER" (Advanced Error Reporting) in their firmware, which will tell you about stuff going awry while components are communicating over said bus (but every instance of PCIe error I have ever seen, even on rather faulty hardware, was recoverable/correctable), and rasdaemon will also persist those.
I do not know what "Memory errors" are supposed to be, since ECC-related problems will be dropped into the "MCE errors" bucket. Neither do I know what "Extlog errors" are.
For correctable errors, we wouldn't swap unless the error counts got pretty big; some systems went for a long time with one or two errors a day, which is fine. Others went zero errors for a long time, then a couple days at a small number, and then big numbers. There was one system that managed to get thousands per hour and the system was unusable because handling machine check exceptions was too expensive; unfortunately the reporting interval was one hour, so we didn't realize why the system was unusable until the next report.
The correctables were trickier to judge, because system halt on UCE is easy to recover from and easy to diagnose (system is down, check console, see UCE message); system slow as heck because of constant machine check exceptions is hard to diagnose and a slow system can disrupt a distributed system a lot more than a dead system.
Tracking down almost working networking without access to the switch metrics was kind of fun, ish. :P
I'm considering upgrading end of the year. I briefly considered thread ripper (for it's pcie lanes) until I found there is still no zen4 thread ripper and the price is likely yo be eye-watering.
So the choice is, upgrade to Ryzen 5900x and keep the rest of the system, or spend a lot more and upgrade to a AM5 ryzen plus a new mobo(I considered Intel too, until I found it tops at 20 pcie lanes).
I've always been buying mostly gigabyte and asus mobos, but I got burned more than once by them so I might go with another manufacturer this time.
What do you think? Is it worth upgrading to AM5 for just single core and disk performance? What about Intel? Considering I'd love more pcie lanes for a 10gb adapter perhaps(so I can move some of these spinning disks out) Intel is probably not a good choice.
I'm happy enough with the multi-core performance of ryzen 3700x. If I got a new MB I'd definitely want at least two nvmes in a mirror(for speed) - perhaps more and sata ports for my 6 spinning disks.
I kind of want an OLED but they don't come cheap, vertical resolution is the same, and they are not amazing for coding. Maybe in time for a Zen 5 based 3d cache chip.
Potentially because of resizable BAR. 3** CPUs didn't support it, 5** did.
For me it was in the end much faster to get a new board with pcie 5.0 support and a large enough single SSD. Total IOPS and throughput both turned out higher than my previous RAID 0 array.
> If I got a new MB I'd definitely want at least two nvmes in a mirror(for speed)
raid1 is a mirror, and can have speed benefits over just a single drive (able to read from multiple drives in parallel).
But yeah, you might be right. When talking casually, things like "mirror", "raid" etc can be a bit fuzzily applied. :)
https://www.hetzner.com/dedicated-rootserver/matrix-ax
The AX52 has optional ECC ram as an upgrade, while the AX102 comes with ECC by default.
It'd be surprising if they'd been offering ECC that didn't really work.
It is unnerving most computing is done in fragile non-ecc systems.
A very messed up form of artificial market segmentation.
Something that works absolutely fine 99.9999% of the time is not fragile.
ECC is a really good idea. It’s only expensive right now because it’s a “premium feature”. If it were a standard part of all ram sticks, it’d be cheap and we’d all benefit.
For example, I download software from the internet then hash it. The hash matches. Before the bytes are written to disk locally, a bit flips in RAM. The corrupted data is written to disk and used.
Likewise, dnssec doesn't protect you against DNS bitsquatting attacks[1] because the domain name can be changed before the DNS request is made. So the DNS response your computer makes for a-azon.com might be totally valid and signed. It can come through DoH or whatever. The problem is that your browser thought it was the response for amazon.com and chrome send a bitsquatter your amazon cookies. (Oops).
This was fine when consumer computers were for games and porn. Now they store birth certificates, submit information to the tax man, documents to court, sometimes deal with matters of life and death
Haven't you seen those congressional hearings with Zuck and others where congresspeople ask about how to use their phone ? How do you expect them to legislate about ECC ?
Probably silently wrong values happening in memory somewhere. If it's not in something where that value is not actively running code code, you probably won't have a crash.
If it's in (say) Excel calculation formula it'll probably just screw up the calculation. Which may or may not be obvious. Similar thing for 3D cad, it could be completely non-obvious.
How about hitting the RAM with a warm stream of air from a hair dryer? I've seen that technique used in the past to generate errors.
Link training can be disabled in the BIOS, which would make searching for borderline bandwidth settings a quick operation. The results would not be very repeatable, but that's not important.
If you're trying to get chips up to something less than 100C, but instead get stuff underneath it with significant thermal mass up to above 200C, you're really messing up.
Hair dryers tend to emit like 65C-70C air on their "high heat" setting (as opposed to hot air guns used for heat-shrink and reflow soldering).
https://hackaday.com/2022/01/29/blast-chips-with-this-bbq-li...
The mode of action here is you're producing a very wide band pulse but it's most strongly couple to the PCB traces long before anything within the chip itself. Guessing again but when people are using this to hack chips, they're probably just causing voltage swings on the power traces that are effectively voltage glitching the chip. The problem is you're over volting the part rather than under volting it which may cause permanent damage.
Incidentally the RAM on a Raspberry Pi 4B (and CM4) is feported to use ECC RAM, but this is not the same as discussed here. It's on die ECC and the purpose is to improve yield of the chips. ECC errors are corrected and not reported via H/W. ECC errors that cannot be corrected (I suppose) are just read out as incorrect data. I wonder if modern RAM sticks use chips with on-die ECC.
DDR5 does, AFAIK, but it is not a replacement for running it 72 bits wide like previous off-die ECC systems.
Agreed, but is it better than no ECC at all?
ECC reporting for these processors appeared in Linux 6.5, so Debian Stable users will have to either wait for it to appear in Backports or stray from the beaten path.
I really dislike Auto settings in the BIOS. It wouldn't be so bad if there was a way to see the effective value, but 9 times out of 10 - it's no obvious.
https://www.reddit.com/r/truenas/comments/10lqofy/
[1] Part of the firmware on AMD systems, that brings up core system components. https://en.wikipedia.org/wiki/AGESA
Had no idea that 7000s series doesn’t official state support of ECC.
ASRock's Rack line supports it, for example: https://www.asrockrack.com/general/productdetail.asp?Model=1...
This ASUS motherboard also claims to support ECC: https://www.asus.com/us/motherboards-components/motherboards...
Neither of those were out at the time I originally got my AM5 motherboard (which was right as they came out -- the performance numbers were so good I couldn't help but spring on one early.)
I would be interested in hearing someones experience with getting this board running as a type 1 ESXI hypervisor.
You've misinterpreted the author.
All currently-available Ryzen 7000 series CPUs have official ECC support, but it requires motherboard support as well.[1]
This conditional ECC support has been the case for all past AMD consumer CPUs going back to the original Athlon 64, but the Ryzen 7000 series is the first time I can recall that this support has been explicitly listed on AMD's marketing materials.
What the author is saying is that mention of ECC support had disappeared from ASRock's motherboard documentation. This was a notable change for ASRock, as they had explicitly mentioned ECC support on their previous Ryzen motherboards.
[1] Example from the Ryzen 5 7600 spec page: https://www.amd.com/en/product/12756#:~:text=ECC%20Support,R...)
I've got two daily driver machines - a 3970x threadripper with 96GB ram, and an M1 Macbook Pro - neither of which have ECC ram.
I've been using them both for over 2 years, and not once in that period (that I'm aware of) have I found myself with a problem due to faulty RAM, but I do regularly find myself wishing both were faster.
What practical benefit would I get in exchange for the performance hit of ECC memory?
Some industries require to be able to reproduce exact bit by bit data.
Anything that requires absolute bit certainty, for example digital signatures, encryption, or financial information, will require ECC.
If you don't care if your files eventually don't match a sha256 checksum, you'll be fine. for example source code, it doesn't matter because in the worst case, your code won't compile because variable names got flipped.
And even then, if you're using Git for source code control, you're already using hash signatures to detect data corruption, as well as a very large redundancy and replication. Git is, in a philosophical way, ECC by software.
Totally, and I get that. We've got people in this thread who are saying that non-ECC ram shouldn't be trusted [0], or that it should be standard [1]. I get the use cases, but why do I want that on my workstation.
> your code won't compile because variable names got flipped.
I have _never_ seen or even heard of this or anything vaguely resembling this happening. Are there any writeups of this anywhere at all?
ECC is expensive just because Intel had a monopoly over the server segment, and unilaterally forced desktop not to use ECC.
Otherwise people would use desktop computers as servers, and that would reduce profits.
At work, my workstation Desktop computer has ECC. The colleague's sitting next to me has on their screen sometimes kernel warnings of ECC errors.
And cases of bit flips happen every day, that's why some problems are solved by just restarting your computer or the software. I'd argue that many times we blame software bugs on pure hardware bit flips.
Here's a story of a variable name bit flip: https://alexbakker.me/post/did-cosmic-rays-break-my-linux-bu...
Bit flips usually affect operating system stuff, here's another story: https://blogs.oracle.com/linux/post/attack-of-the-cosmic-ray...
The ability to detect (and possibly correct) physical memory errors.
> I've been using them both for over 2 years, and not once in that period (that I'm aware of) have I found myself with a problem due to faulty RAM
The "that I'm aware of" part is the key issue. ECC provides visibility into the health of the physical memory that you otherwise wouldn't have.
> What practical benefit would I get in exchange for the performance hit of ECC memory?
Side-band ECC (which is what you'd use on desktops/workstations/servers) doesn't have a performance hit, as the overhead of ECC is fully neutralized by the additional bandwidth and capacity of an ECC DIMM.
This is a great example of theoretical benefits (note I'm not saying that those theoretical benefits are real, I'm asking how they practically benefit me).
> The "that I'm aware of" part is the key issue.
Ok so tell me what it looks like. _That's_ what I want to know.
> into the health of the physical memory that you otherwise wouldn't have.
And does what, on my windows workstation or my linux workstation?
> Side-band ECC (which is what you'd use on desktops/workstations/servers) doesn't have a performance hit, as the overhead of ECC is fully neutralized by the additional bandwidth and capacity of an ECC DIMM.
That statement doesn't hold water. If it takes extra bandwidth and capacity to provide ECC, I can use the extra bandwidth as memory instead of error correction, no? (Or I could if the non-ecc DIMM utilised that range). Either way, it's an extra x bits for ECC that _could_ be used for storage.
Suppose an application running on your machine suffers some type of malfunction -- a segmentation fault or a seemingly-random kernel panic that you're unable to reproduce.[1] And suppose it happens often enough that you want to fix it, and understand the root cause.
You research the issue, and are pointed to potentially faulty hardware. That then raises a question: how do you know your physical memory is working properly?
You could run a diagnostic application like Memtest86. However, that has the following issues:
- It only tests a particular point in time. If the test "passes", that doesn't say whether your memory encountered a fault in the past, nor does it guarantee that it won't fault in the future.
- What does it even mean for a test to "pass?" Just a single run through a test suite? Running it for some period of time, like a few hours?
- A diagnostic like Memtest86 is invasive. While it's running, you can't use your machine for other things.
Troubleshooting these types of hardware errors without ECC is tricky, because there's not really a way to conclusively link a software fault to a hardware error.[2][3]
However, on a machine equipped with ECC, this type of troubleshooting is a lot more straight-forward. If bits are flipped in memory, the memory controller can probably detect that (assuming only only one or two bits are flipped), and can raise a machine exception that the OS can catch and do something with (e.g., logging the error, terminating the process impacted if the error is correctable, etc). That saves a lot of time and headache that you'd otherwise be spending on guesswork.
You may still be asking, "how big of an issue is this, really?" I suppose whether or not you care about having hardware that can detect memory faults is up to you.[4] However, during the portion of my career where I was managing large machine fleets, memory failures were the second most common type of hardware failure, behind only mechanical disk drives.
> That statement doesn't hold water. If it takes extra bandwidth and capacity to provide ECC, I can use the extra bandwidth as memory instead of error correction, no?
No. The memory controller is unable to use the additional capacity for anything other than error checking (hence why it's called "side-band").
What you're describing is "in-band ECC", which is something that you sometimes see on GPUs or small form factor systems aimed at professional markets.
---
[1] An application crash is actually one of the better outcomes of a memory error. The worst-case scenario is silent corruption of data stored in memory that your application assumes to be valid.
[2] Errors as a result of faulty memory have been particularly frustrating to OS developers, as it can result in nonsense error reports. Linus Torvalds has bemoaned the lack of ECC options on consumer hardware, claiming that it was the industry cheaping out. During the lead-up to the launch of Windows Vista, Microsoft officially encouraged the use of ECC. I don't follow development of the BSDs, but it wouldn't surprise me if they similarly wished their users would use ECC across the board.
[3] When Google first started building out the infrastructure for their server farms, they opted not to use ECC memory for their servers as a cost-savings measure. They subsequently had a memory fault on one their machines that resulted on corruption of their search index. While they worked around the problem by adding logic to verify the contents of the index, all subsequent generations of machines that Google has deployed have ECC support. For specifics, see the second footnote in https://danluu.com/why-ecc/
[4] Note that many of your other components that store data in some way (e.g., caches, persistent storage, etc.) are likely to have ECC capability. The only notable exceptions I can think of are GPU memory, the main system memory on consumer hardware, and the processor's registers.
> Side-band ECC (which is what you'd use on desktops/workstations/servers) doesn't have a performance hit, as the overhead of ECC is fully neutralized by the additional bandwidth and capacity of an ECC DIMM.
This is true as far as I know but only for comparing ECC ram vs non ECC ram that is otherwise equivalent. But if you want the fastest ram available, it doesn't support ECC, so you're still taking a perf hit by buying slower ram with ECC support than you could get without ECC
Now I'm on an M1 MBP with 32GB RAM, and I get unexplained crashes roughly every 3-4 months, but I've been chalking it up to MacOS. Behavior also seems to degrade after a month or two of uptime. Hmm, maybe it's time for Apple to go ECC ...
Thanks. This is a great anecdote/example that is much more what I'm interested in. I don't see the same level of system instability on my machine, but it's a good example. Appreciate it.
> Behavior also seems to degrade after a month or two of uptime.
Ah, here's another thing. I do tend to reboot daily/weekly which may help
> Unfortunately, when the AMD Ryzen 7000 “Raphael” CPUs were launched along with the brand new Socket AM5, all mention of ECC support was gone.
1. This was effectively a big point of concern for me. Previous ASRock Taichi motherboards officially advertised full ECC support, but not the latest X670E for AM5 model. Older version of the X670E Taichi page mentioned ECC support ("Supports DDR5 ECC/non-ECC, un-buffered memory up to 6600+(OC)") [0], and this Level1Tech review [1] had a segment on ECC support (they fully tested it). The segment was around the 10min mark, but it looks like the video was swapped to remove it during the last 30 days (the original video was 17min41sec long, the current one is only 16min32sec). So I assumed that ECC was supported, but it's a bit concerning that there's no official mention. The article reassures me.
2. Finding ECC memory for desktop computers is hard. For the X670E Taichi, you need DDR5 unregistered DIMM. The best solution I have to find ECC memory is to check QVL lists for similar server motherboards, for example this ASRock Rack [2] or this Gigabyte Motherboard [3]. I decided to go with two 32GiB Kingston sticks [4] for my own build.
[0]: https://web.archive.org/web/20221002103925/https://www.asroc...
[1]: https://www.youtube.com/watch?v=PhrqEV-VAjE
[2]: https://www.asrockrack.com/general/productdetail.asp?Model=1...
[3]: https://download.gigabyte.com/FileList/QVL/server_mb_qvl_MC1...
[4]: https://www.kingston.com/en/memory/server-premier/ddr5-4800m...
Unfortunately AMD has been going the wrong way with this. Oddly enough, Intel has quietly added a few cheaper models with ECC support.
Unlike the author I don't have time for building a new PC and a bunch of maybes. I found this machine that's pretty cheap:
https://www.dell.com/en-us/shop/desktop-computers/precision-...
This one has a CPU option with a TDP of 65 W. What do y'all think?
I suspect the group that wants to physically maintain their laptop but doesn’t care about data integrity isn’t too large.
The Dell may support ECC, but as any workstation it's pricier than consumer grade desktop with equivalent performance. If you need it, you need it but usually individuals don't pay for that, their employers do.
And never been given the option.
I believe that there might be issues with it not being enabled properly if you buy the machine with a minimum of non-ecc memory intending to switch it out, because lenovo does something fucky if you don't get ECC from the get go. I could be misremembering that last bit, it's been a couple of years since I looked at ECC capable workstation laptops.
But yeah, when framework comes out with an ECC workstation laptop is when I'll finally consider getting a new laptop.
That Dell has the Intel W680 chipset that is responsible for giving any Intel CPU ECC, just more pointless market segmentation by Intel. I don't know if I would give Intel a glowing review for that, you can also buy an OEM system with an Ryzen Pro processor if you want that kind of ECC guarantee.
You're not going to find Intel offerings with ECC solutions for light laptops either right now.
Or just buy one?
> buy an OEM system
Who sells certified ECC systems with AMD at similar prices? I did look but didn’t find much already assembled.
I don't think you can blame AMD for the OEMs choices. Lenovo does compete in a way. You have to buy a GPU on their Threadripper workstations because of no iGPU for the computer to fall back on like the Dell's Intel iGPU. Also your Dell comes with 1 non-ECC 8GB DIMM standard, you have to pay $247 to upgrade to an ECC DIMM. Lenovos Threadrippers come with the ECC standard and can even run registered DIMMs.
But yeah it is a $2500k Threadripper workstation with no GPU compared to a $1300 I5 workstation with no GPU.
If you want to build an ECC AMD system some manufacturers want your money: https://www.supermicro.com/en/products/motherboard/h13sae-mf
Frankly "shopping" OEM sites makes me crazy and it's why I build my own.
Epic! Some people, really...
> I don’t quite have the courage to physically short pins, nor the patience to slowly overclock my RAM, waiting multiple minutes for DDR5 link training each time
It's not clear to me: when is the DDR5 re-trained? I assembled my Ryzen 7000 desktop PC myself and I don't remember ever having to wait minutes for it to boot, not even the first time. It's a bit slower at boot I'd say than non-DDR5 PCs I had in the past but it's mere seconds, not minutes.
And last but not least: DDR5 specs for ECC vs non-ECC, which price difference and speed difference are to be expected?
ECC RAM uses 1/8th more memory chips, so as a rule of thumb it costs 1/8th more.
They can have very different dynamics on the used market though, because servers tend to be upgraded to larger modules over time, but there aren't many buyers for smaller modules. Not a big factor for DDR5 because it's too new, but for DDR4 you can get some great deals for 8GB and 16GB modules.
...But are there any ECC sticks with "known" good Hynix ICs? I got lucky with a pair of sticks that can do DDR5 6000 CL30 at 1.3V (and maybe better), and I wouldn't want to sacrifice a ton of RAM performance for ECC.
I haven't tried overclocking it.
You should do it and post it. Ryzen 7000 likes OC'd RAM, and overclocked ECC ram is kind of a desktop hardware myth/legend because its so rare and oxymoronic. Not sure if you've done it before, but DDR5 overclocking is quite a rabbit hole, and there are some "easy" tutorials out there like this: https://www.youtube.com/watch?v=dlYxmRcdLVw
Test the stock primary timings with the modded subtimings at 6000 first, then jump straight to cl34 for the primary timings and reduce the voltage/primary timings from there.
I would grab these myself, but the kit is too rich for my blood at the moment.
To be honest the only RAM overclocking I've done is by going into the XMP settings and turning on the shipped overclock configuration. For CPUs I tend to undervolt them a bit to reduce power consumption.
This doesn't seem to be the proof that ECC is working. Error detection on things like this is different than bit flip correction.
Ryzen has the additional disadvantage vs TR of only being able to run RAM fast if you limit yourself to two sticks. My TR build used 4x16GiB but for my 7950x build I had to get 2x32GiB to be able to run them at 6000MHz.
I'm curious what your use case is. I mostly use mine for software development (including VMs) and some family media processing. I like the peace of mind that ECC brings, but I have no way to really know if RAM speed is performance factor.
TR has it similar, just doubled; you are able to run 4 sticks faster than 8. On my TR2920X+X399 (probably the same config as yours, Asrock Taichi) I'm able to run any four of them at 3066 MHz and all of them only at 2933 MHz.
(My 2920x was on an X399 Phantom Gaming 6 and my 7950x is on a B650 PG Lightning. I find the Taichis of all generations to be overkill.)
I'll reboot it occasionally. Every once in awhile it will exhibit odd behavior that seems to resolve on reboot.
Is it worth getting something that has ECC? I think I'm running like a Sandy Bridge i7 for reference.
On the more immediately practical side.. if this is happening frequently enough then running memtest86 or GSAT google stressful application test and see if it can pick up any errors.
You also might be able to improve things with a bios upgrade, or by lowering the clock speed of the ram.
That, and most memory problems in PCs tend to be from dodgy overclocking or just bad sticks rather than cosmic rays. ECC won't really save you from that.
It's just neurotic to worry about issues when you couldn't tell they existed without being informed of them.
Error detection and correction is actually ideal when overclocking. It'll not only save you, it'll let you know you need to back off a bit, because the handful of errors that may happen once a month despite passing days of tests will show up in the system logs instead of staying silent and causing problems. I speak from experience as I overclocked ECC memory in a first gen threadripper system, playing around with overclocking as a hobby for a bit. ECC memory was fantastic to work with.
However the dodgy or outright lack of CPU/Motherboard support results in trying to market that feature to gamers/overclockers having too much friction. Much like what went on with intel's confusing proprietary optane/flash combo M.2 drives that needed certain motherboards in order to fully access both parts of the drive.
And with bad sticks, instead of hard to explain crashes, you'll end up with system logs filled with corrected/detected errors and maby crashes if it's really bad. Then you just RMA the sticks like normal. So honestly that's a win too, because that's the whole point of detecting and correcting errors. It's in every other part of the system already, continuing to leave this capability out of memory is an oversight that needs to be.. corrected.
I just don't really get the strange resistance people have to ECC, it's not some sacred cow. It's just good sense.
No, it is not "the same reason"
Disks have ECC, even consumer hard drives have checksummed blocks, that's how you get media errors detected. RAM does not. Well, technically with DDR5 it has internal one but it does not give any feedback to the machine so you might not know your RAM has any problems.
> It's why you don't need an automatic fire suppression system in or homelab
But you want smoke sensors. ECC is the smoke sensor too
> That, and most memory problems in PCs tend to be from dodgy overclocking or just bad sticks rather than cosmic rays. ECC won't really save you from that.
But it will tell you the problem
Very ignorant take
Sure, and disks have random errors even with error correction. Adding hamming codes just makes errors less likely and easier to detect, but they aren't fool proof, in ram or on disk or anywhere else. The protection offered is a bit questionable because they rely on the very sketchy assumption that errors are statistically independent events, which is rarely how storage errors work.
> But it will tell you the problem > Very ignorant take
Hardware isn't perfect. No matter how much error correction you pile onto it, you will still have errors and some cases will still be undetectable. It's a matter of pushing the rate of errors into an acceptable range, as a pure cost-benefit tradeoff, statistics through and through.
It's a bit pricy, but you also get IPMI.
Why?..
No.
> One friend felt like he was going crazy
Tell him about memtest86.
Set it back down to a supported frequency, ran a full memtest suite again with no errors.
Never had any issues since.
It's not a matter of overclocking. Bit flips are a fact of life running with 32+ GB RAM. Leaving your machine on 24/7 (even if in sleep) stacks the odds against you.
Cool. You tested your memory at some point in the past.
How do you know it's still working properly and hasn't flipped any bits?
You don't. Because you have no practical way of testing the integrity of the data without running an intrusive tool like memtest86 that basically monopolizes the use of the computer.
Being able to detect these types of memory errors at a hardware level while the processor is doing other things is the fundamental capability that ECC gives you that you otherwise wouldn't have, no matter how thoroughly you run memtest86.
(Of course, I don't run ECC on my personal systems, but at least I'm wandering knowingly into the abyss)
https://forum.level1techs.com/t/ecc-capable-verified-motherb...
>"I have ECC working on ASUS ProArt X670E-Creator with AMD Ryzen 9 7950X. But, you have to explicitly turn ECC on in the bios. If left on ‘Auto’, it will be off. I use four sticks of Supermicro (Hynix) 32GB 288-Pin DDR5 4800 (PC5-38400) Server Memory (MEM-DR532MD-EU48)."
https://www.reddit.com/r/truenas/comments/10lqofy/ecc_suppor...
<Read yourself>