I *detest* the crazy industry politics that made ECC memory so “special”
lkml.iu.edu
lkml.iu.edu
Since desktop ECC gets around this by having physically more RAM ICs (usually 9 instead of 8, for example), what is the impediment from having a similar solution to Nvidia? I'd readily take a hit to memory capacity* and performance in exchange for ECC.
Why can't the memory controller already do this?
I should note, I'm mostly thinking of my NAS. I know ZFS can be run without ECC and some consumer solutions do. However, it seems ZFS should be run with ECC. I've already experienced observable bitrot with older images and video files, I'd rather not let it progress.
[*] in this case, 12.5% if we follow typical desktop ECC allocations
[0] https://www.nvidia.com/content/Control-Panel-Help/vLatest/en...
[1] https://docs.nvidia.com/cuda/cuda-c-best-practices-guide/ind...
Market segmentation on what platforms/controllers support ECC: fine, whatever. But market segmentation of what is an “ECC DIMM” vs. a “regular DIMM”? It makes no sense that the commodity memory manufacturers have any leverage to enforce that segmentation.
Is it just laziness on the part of the platform vendors (who do have leverage) not simply allow ECC with any DIMMs by giving over 1/k {bits, lines, pages, chips, whatever-granularity-they-reason-about} to parity?
On the AMD side, no.
However Intel guarantees ECC will work on their "workstation" chipsets. AMD doesn't guarantee ECC will work on their desktop/workstation chipsets. You have to go up to Epyc to find a guaranteed/tested ECC.
(And nvidia Tegra does in-band ECC too)
By the way, RTX 4090 doesn't have ECC disabled: https://techgage.com/article/nvidia-geforce-rtx-4090-the-new...
[0] - https://jrs-s.net/2015/02/03/will-zfs-and-non-ecc-ram-kill-y...
"...We further demonstrate that ZFS is less resilient to memory corruption, which can lead to corrupt data being returned to applications or system crashes..." - https://research.cs.wisc.edu/wind/Publications/zfs-corruptio...
"Please Use ZFS With ECC Memory" - https://louwrentius.com/please-use-zfs-with-ecc-memory.html
...ie, ZFS is 'less resilient' in comparison to its robust disk fault handling, not that it's less resilient to memory corruption in comparison to other filesystems. The parent quotation above implies that ZFS is more sensitive to memory corruption than other fs but that is not claimed in the referenced paper.
In-band ECC means significant sacrifice of performance on a system not designed for it. Random read throughput doesn't go down by 6.25%, it goes down by half.
But adjusting DDR for that could be pretty easy. Instead of a burst of 16 transfers, do 18. It's already set up to stream longer transfers when desired.
There will be more overhead than making the sizes properly match, but it shouldn't be anywhere near cutting throughput in half.
That's not really how DDR5 works. The granularity of column addresses is (iirc) 32 bytes, and you cannot do transfers that are of any length other than 64 or 32 bytes (and 32 bytes only with burst chop, which means that the bank is busy for the remaining 8 cycles). Bursts longer than 16 are really just multiple adjancent requests, with an optimized command.
You could change this, by completely changing how the memory modules themselves work, and by widening the column address for more granularity. Can't do it well by just tweaking the memory controllers.
I wasn't trying to suggest you could do it by changing only the memory controllers and not the DIMMs.
Especially since the ECC overhead on DDR5 is so high.
From my understanding, the only risk to your data from non-ECC is a bit flip in RAM, pre-checksum calculation. In that unlikely scenario, you commit bad data to disk as good data(valid checksum). Bitrot isn't a factor, at all.
If your data is important enough to warrant ECC RAM, you should get ECC RAM whether you use ZFS or not.
If you want to use ZFS (for its volume management, compression, mirroring, healthchecks, whathaveyou), you should do so whether or not you have ECC RAM.
If you’re using ZFS for other reasons, then you be you I guess.
Wouldn't an option to do it twice in different memory regions be nice? I'm pretty sure in many use cases scarifying performance for greater reliability wouldn't be an issue. Given how many cores we have available nowadays it could potentially even not have that much impact on performance.
Also are there any software solutions (like a kernel patch) which would do "software ECC"? I imagine in this case performance hit would be quite devastating but it still could be acceptable trade-off for NAS-like systems where you want to have lots of RAM for dedup and cache but it's not a busy system.
Otherwise a bit flip that early during read shouldn't matter because you're checking it against the disk checksum.
If you don't have disk checksums then ECC memory is not where you should be putting effort to keep things safe.
It’s so uncommon that the PAR3 format was never really finished and no one has created a replacement that handles subfolders.
Why I’m surprised: Not only does it solve the problem of bit-rot, but the parity files can be moved to USB sticks, NAS drives, Mobile devices, etc and the original files can be verified/repaired by any device that understand the parity file format. PAR2 is still great for photos/audio/video, as well as any flat-folder assets.
It often makes more sense for the file system to deal with ECC in my opinion. PAR probably makes more sense for archived files that aren't expected to change, but may be moved across file systems.
PAR2 handles subfolders by the way, just not empty folders.
Files with on-disk ECC can be moved from cloud to cloud, cloud to desktop, filesystem to filesystem, desktop to stick, then stick to NAS all without losing ECC protection. No single filesystem can do that.
So specifically: photography archives, videos (including b-roll for content producers/videographers), project backups, personal files, important documents, etc. Up to and including anything that could be posted to r/datahoarders
1. “Lots of tiny files”
Some folders unavoidably have tiny files in bulk (Document backups can be like this. One other example that jumps to mind: macOS applications with translation files)
In these cases, PAR/PAR2 have issues with the block size (can only have one file per block which leads to a lot of wasted space)
2. Tracking changes across filenames
This is counterintuitive, but I’ve run in to this enough to mention it: if the item to archive is a folder where the contents might change over time, any single file might get renamed and it’s contents slightly modified. A parity file tool could look at the blocks that have not changed, recognize the rename, and “correct” the reference before doing more processing. If it’s a valid change to the file: saving the work required to recalculate the whole file and if it’s damage to the renamed file: being able to repair it simply.
3. Being able to update in-place
Sometimes the ideal is to create parity files for a folder, even if that folder is actively used (say for example b-roll that changes by 10% maybe once a month). A parity tool could update that 10% without having to recalculate the whole thing (Ideally this would be adding files similar to ‘git add’ so that someone does not accidentally add file damage to the parity set)
But updates are problematic. You could 'delete' parts of a recovery file if the data is present. However a file being updated typically means that the original data (before the update) is no longer present, meaning you can't really 'delete' that part of the recovery to replace it with the new data. You could try and retrieve the old data by recovering it, but at this point, it may just be easier to recreate the entire PAR set again.
If it's just adding to the recovery set, PAR3 does have provision for that.
After our discussion I went and found the current work on PAR2/PAR3 (which, for those in the know: the current PAR3 format is not the old on the was never finished, but a spec that’s been re-written and is close to completion with real-world tools close behind) and I have a lot more hope for the future of parity files.
I still think they are wildly under-utilized (BackBlaze has always used them, but they are the only business I know of), but we might be having a different conversation in 2-3 years.
-
Mushkin Essentials DIMM 32GB, DDR4-3200, CL22-22-22-52 - 85.79€
Mushkin Proline DIMM 32GB, DDR4-3200, CL22-22-22-52, ECC - 143.00€ (= 1.67×)
-
Kingston ValueRAM DIMM 32GB, DDR4-3200, CL22-22-22 - 117.39€
Kingston Server Premier DIMM 32GB, DDR4-3200, CL22-22-22, ECC - 147.90€ (= 1.26×)
-
Samsung DIMM 32GB, DDR4-3200, CL22-22-22 - 117.39€
Samsung DIMM 32GB, DDR4-3200, CL22, ECC - 152.89€ (= 1.30×)
-
It should be noted however that a bunch of cheap brands do not even offer ECC variants, and those may dominate the lower end of the price spectrum. So getting ECC memory may also involve choosing a more pricey brand.
-
References:
https://geizhals.de/mushkin-essentials-dimm-32gb-mes4u320nf3...
https://geizhals.de/mushkin-proline-dimm-32gb-mpl4e320nf32g2...
https://geizhals.de/kingston-valueram-dimm-32gb-kvr32n22d8-3...
https://geizhals.de/kingston-server-premier-dimm-32gb-ksm32e...
https://geizhals.de/samsung-dimm-32gb-m378a4g43ab2-cwe-a2328...
https://geizhals.de/samsung-dimm-32gb-m391a4g43bb1-cwe-a2755...
… which does not mean other factors are "irrelevant"…
… and also the price differential in ECC memory is not only a factor but also a metric as it shows that ECC memory has a "non-mainstay" surcharge created by the vast majority of users going for non-ECC. If demand were equal, the price would be 9/8 = 1.125×.
The current situation demonstrates that the customer has no ability to negotiate product quality.
It's debatable if that's due to ignorance or the game is just so syacked against consumers
Political economy as an area of inquiry took off in the late 18th century--in no small part due to Adam Smith--and was already very well established before Marx was born [1]. Absent evidence to the contrary it seems questionable to assert that Marx's use of the term was anything other than the common use of the time.
[1] https://books.google.com/ngrams/graph?content=political+econ...
Anyway, try to take the technical specification of any consumer oriented motherboard and discover if it allows ECC.
Having to pay a 20% premium on RAM for the stability of ECC is one thing, having to pay double, triple, or quadruple the system price to disable arbitrary locks is another.
Well, yes, but it's not the only tool. In practice you'd use some validation software suite to find a reasonable stable configuration and afterwards prey. Overclockers, who rather than spending extra money on a CPU which assuredly delivers requested performance stress their hardware beyond specifications, are least likely to pay for the extra bits.
That doesn't match up with my experience, even though I've seen it mentioned several times by people as if it's the case.
For example, when I went looking for ECC UDIMM sticks for my Ryzen 5600X build a year or so ago, I went looking at the website of my local computer supplier.
Many potential options available:
https://www.scorptec.com.au/product/memory/ecc-&-registered
Now, there aren't as many different options as for non-ECC stuff. eg:
https://www.scorptec.com.au/product/memory/ddr4 (many pages of sticks available)
But it's not like theres any kind of availability problem. And the Kingston ram I bought happily overclocked to 3200MHz without any effort on my part. Using an ASRock B550M Pro4 motherboard for reference, if that's useful.
So sure you can run dimms every so slightly faster and fix the occasional single bit flip, but even a single double bit flip and an process or your kernel crashes.
Seems much saner to go for a safe, robust, and reliable system at standard clocks in ECC instead of trying to get slightly more performance which increases the chances or errors, corruptions, crashes, and shorter service life.
Some people can’t afford to buy anything but the cheapest but in most cases it comes down to not thinking that they use it enough to matter (e.g. the common homeowner advice to buy a cheap tool the first time & replace it better if you break it), not having a good way to tell whether something is actually better, or the market being such that there is a huge gap between the product bands. ECC falls into the latter two categories: the average buyer isn’t familiar with the issue and probably thinks the outcome would be a crash rather than silent data corruption, and Intel’s market segmentation means that you don’t have a choice in the consumer space and have to move into far more expensive and limited categories. It’s not reasonable to say price-sensitive consumers are the problem when that’s also saying “stop buying laptops”.
My 4 years old laptop has ecc. To see if anything has changed I just checked the recent models. The manufacturers website didn't provide any possibility to filter for ecc or mobile Xeon, but an explcit text search returned some results. Ecc still exist in laptops, its just a bit difficult to find.
On the plus side, 19:10 screens are back.
Just yesterday, I bought an extra 64GB for my home Linux PC. I absolutely couldn’t care less about it crashing or calculating the wrong result every blue moon (in practice: never), but I did choose the RAM sticks that were $10 cheaper.
When I look at my daily home computer usage, it’s remarkable how little I calculate on my local computer that’s actually with protecting.
The computing world is full of more and less featured products.
It's not like it wasn't available, Linus just forgot to put a reminder to buy some once available (backordering them would have been even more prudent).
Bitflips aren't uncommon, it's just compression makes them much more noticeable. Without compression users might just see an occasional application crash, and just relaunch the proces.s Seems common for people collecting photos, music, etc for years go back and find some corrupted. It's hard to say exactly what happened, but with compression a bit flip it turned into a seriously corrupted file.
Keep in mind that, radiation caused bit flips are rare, some failure in the chip, pin, dimm, dimm connector, motherboard, CPU pin, CPU are much more common. It's MUCH nicer to see "error on dim X, row Y, column Z" then a randomly crashing machine. Even Linus sounds like he spent hours tracking it down, and thought he was hitting a kernel bug.
Isn't it worth 1/9th more memory chips to make the system robust in the face of wide variety of errors that can corrupt memory, which can lead to corrupted storage?
Memory issues can be caused by the PSU, motherboard, the CPU, or the memory itself. Personally, I always run memtest86 and memtest86+ for 2 days on any new components.
Out-of-date but pertinent: https://cr.yp.to/hardware/ecc.html (c. 1999-2006)
Other risks: https://defcon.org/images/defcon-21/dc-21-presentations/Schu...
The mitigation of bitsquatting requires ECC in network gear also. Furthermore, ECC isn't just about the memory type but having integrity in data buses, caches, storage, and network protocols also.
Home system is 96 thread, 512 GiB (16 x 32GB ECC Registered DDR4-3200), numerous SSDs and HDDs running RAID.
[0] - https://ark.intel.com/content/www/us/en/ark/products/96149/i...
[1] - https://ark.intel.com/content/www/us/en/ark/products/134591/...
[2] - https://www.anandtech.com/show/17308/the-intel-w680-chipset-...
Intel's marketing is really in a league of their own. As soon as their branding starts making sense they will change it.
I am hoping for a good successor (ECC, enough bays, not too fast but cheap enough) to come out.
It would be nice to have a nice little server which consumes 15-20 Watts.
The regular G10 one is pretty useless obviously. Soldered AMD CPU, cheap design etc. No iLO. It's more of a home entertainment thingy, nothing enterprise-class at all.
I agree I'm looking for something lower power too. I have three G8's but they are each doing 50W idle so I can't run them 24/7. And this is with the lowest-TDP processor that is available for them! (E3-1220Lv2, 17W TDP). The iLO alone is consuming 5W even when it's off.
I'm not sure how much power the G10 Plus draws but I haven't really considered it. It's too expensive still (I bought all my G8's for 175-200 euro each and 2 of them even had a 60 euro cashback on top of that!! So I barely paid more than 100 for them brand new. Crazy cheap pricing for a well-built 4-bay server. Each of the drives inside them cost more than the server itself :)
But 15-20W I think is very ambitious with 4 3.5" drive slots. 30W would be doable I think.
Joking aside, power supplies are probably the next largest source of random hardware issues in PCs today
Also with some bad memory causing really really weird issues all across the system.
Power supplies? And some of the cheap ones I've used have just done their job and not caused any problems.
Do you have any examples of what's going on with this?
If you get a stock grey metal 500W power supply and try to run a modern graphics card you’ll have a far worse time than some Corsair 400W PSU.
I haven’t build a PC in a decade so maybe this is old news.
Not everyone is an electronic engineer, and if the overall wattage doesn't give the information people need, the sales/marketing should be forced to provide the necessary details right there on the box in "plain language"; RAM speeds are still a pet-peeve of mine.
That said, I've been building PC's for ~25 years and I'm yet to have a PSU fail on me. Anecdotal I know, and with some skew as I've always tended to build near or at the top end.
That said, the Ryzen hard lockups were a real disappointment for me, after spending north of £3500 on my desktop in the OG Ryzen era with an R7 1800X, I was left with a machine that frequently hard-locked when compiling code on all cores and AGESA is such a mess that if it boots at all, it takes several minutes at times to make it past POST.
I really hope the next gen build I make, whatever it is is more stable because I'm not planning to go back to Intel any time soon.
https://www.phoronix.com/review/new-ryzen-fixed
I was an early adopter, and I accept that but it still left a bitter taste none the less, and ultimately it just meant that my pretty expensive computer barely got used because I hated dealing with it (still have it but barely boot it these days).
The stock grey metal models are going to be out of date, but they're not that out of date. At this point I'd expect them to have a reasonable rail balance for a modern computer.
Cheap cases used to come with crap power supplies even 15 years ago. It's great that's no longer the norm. Everyone wised up a bit on this one
Maybe I've just been lucky or I've just been buying the quality stuff, two of the ones I'm running now are 1,200 Watts, the other two are in computers that don't use much power because they are basically idle servers all the time. They are also semi reasonable quality.
But I have been running a computer on a 450 w power supply that was very much having its limit pushed.
I also did run on a bottom of the barrel super cheap 500 w supply that I bought in the middle of 2020 when all of the supply crisis was happening.
That is why I was curious. My experiences didn't line up with what the OP was talking about, so I was interested to see examples of what was going on.
edit: I misrembered some aspects of this, The 3000 series was actually spiking much higher than rated TDP and tripping overcurrent protection versus just browning out. NVidia did recommend 850w PSUs for that first release of cards for that reason but iirc some 850w still had issues.
The power supply is just doing its job, when you far exceed its rating for even a short amount of time it's going to try to shut down to protect the computer because it thinks something inside of your computer is burning up.
Took me 3 SSDs, 2 motherboards and 2 PSUs to figure out what the cause of the problem was.
It was a ThermalTake TR2 500W PSU. No graphics card (integrated graphics) so 500W should have been fine.
ThermalTake is just a reseller, Season is an example of a company that makes power supplies. Seasonic sells to others for relabeling, but also sells direct to consumers.
My guess was that bad power was a contributor to many other issues, but the nature of the SLO was such that more complex issues resulted in a device swap. Also, any kind of dirty environment drives higher AFRs.
I reused a 10 year old "gold" 850W Corsair PSU in my zen3 build. It had some noticable inductor whine when the system was idle and I moved the mouse. After a year I got a 6800xt GPU, and it kept trucking along for another few months. I did some GPU overclocking and stress testing one Saturday, and my PC wouldn't start the next day. It had a good run for PSUs I suppose; it outlasted its 7 year warranty. I bought another 850W Corsair PSU to replace it, with "titanium" efficiency rating and a 10 (or was it 12?) year warranty.
In prebuilt PCs you don't see such issues because the manufacturer controls everything. They might not make the PSU itself but they will certainly get it made with the right specs.
In a home-built PC you need to consider not just the total power (wattage) but also the number of rails and current (amperage) per rail. Exceed that and you will run into stability problems. Also, cheap means you get what you pay for. Get something made by Delta and Seasonic and you will have far fewer problems. Some of the cheap no-name brands don't even specify what kind of rails are in there.
However failures in the memory chip -> chip pin -> dimm -> dimm connector -> motherboard -> CPU socket -> CPU pin -> CPU are pretty common. Sure ECC helps with random bitflips, but it's also very useful to diagnose something is broken in the CPU <-> memory chip pipeline. It's very frustrating to debug something when the main sign of the problem is a reboot.
I miss the old, less politically-correct Linus who didn't pull his punches.
How much was ECC ram at that point? Linus is incredibly wealthy and was building a machine that would be used to gate Linux releases, so I'm very curious how much was not "sane" for his use case.
This feels very wrong when the only difference is one additional chip that in terms of material only should increase the BOM price about 12.5% (going from 8->9 memory chips).
Also, onr of the primary markets for ECC UDIMMs are server/pro-workstation, thia alone.adds at least 10%
I don't think Linus likes wasting money for the sake of wasting it.
How common are RAM bit flips?
Research has shown that a computer with 4GB of memory has a 96% chance of having a random “bit flip” every three days. That's a crazy high chance of data corruption occurring on your computer.18 Mar 2021
https://www.macobserver.com/columns-opinions/devils-advocate...
How often do ECC errors occur?
These can all not be corrected, but are extremely rare. A 1 Gigabit ECC DRAM contains 16 Million blocks of 64 bit datawords. Per each of these 64 bit words, one error is correctable. In other words: Statistically one out of 16 million hits might be a double-bit error.
https://www.intelligentmemory.com/support/faq/ecc-dram/how-o...
such second hand server parts are cheaper than normal consumer grade parts you get extra protection & performance
256GB (32GB x 8) Samsung DDR4-2933 ECC Reg RAM for USD $550, it is pretty hard to say no to that.
Sure enough on memtest64 those issues were clearly diagnosed.
Haven't changed to ECC hardware, disabling xmp to ddr4-2333 helped with stability.
ZFS was reporting all sorts of errors, yet drive tests were showing no issues. I bought about 3 new drives before I realised what the root cause was. A real PITA. Next time I do a hardware refresh, ECC is definitely on the menu.
Haven’t bought RAM for awhile, what’s he talking about? ECC RAM should be at least 1/8 more expensive (plus something for the handler)
Just try and buy a reasonable non-server class machine that has ECC.
It's not unofficial, AMD lists it in their specifications.
And that, until Ryzen 7000.
PS: and being a pre-ryzen AMD user, I don't know much about new CPUs too :))
All Alder Lake chips (the ones shipping in volume today) support ECC, check ark.intel.com to be sure. For example the popular i7 model (https://ark.intel.com/content/www/us/en/ark/products/134591/...) As do all Ryzen desktop chips (5000 and 7000) do as well.
Granted popular consumer products don't have ECC, but it's pretty straight forward to built it yourself. Just make sure the motherboard you buy (Intel or AMD) supports ECC.
I suspect that the final phrase above is sarcastic, as well as instructive and even possibly bragging, and Linus is perfectly aware that his machine gets an exceptional amount of use compared to the average Peanuts character. RAM wears out, "And some of the degradation is noticiable if you use it intensively (as servers do)."[1]
[1] https://superuser.com/questions/1568933/does-ram-degrade-ove...
Which is what? Double the 30% duty cycle? Would it be so beyond the realm of belief that Linus' box has a 80%-90% duty cycle? I'm not exactly sure where your position falls, that the RAM was defective, or that it lasted longer than expected given its increased use.
And that's just about the law. Don't get me started on resource consumption and on my perfectly fine 10 years old computers I'll be forced to stop using soon because of litteral planned obsolecence (of MS, that's another story but a major actor of the computer industry too, and that will yield even more environmental destruction)
Expected by you.
From reading this, I guess one has to do some special setup to let a system use ECC?
Been thinking about ECC myself. What would I need to do, apart from buying the DIMMs and putting them in? Some BIOS settings? Jumper settings?
You need to have a compatible CPU/motherboard/chipset. For normal CPUs: AMD Ryzen non-Pro APUs don't have support for it, the rest of AMD's CPUs and chipsets have unofficial support for it. You'll have to check the motherboard vendor's support page if a certain board also has support for ECC. Then you need ECC memory modules and you should stick near the qualified vendors list (QVL) here since systems are kind of pickier with ECC memory. For Intel, you're out of luck except for the W680 chipset, but motherboards seem to be scarce.
For high-end desktop (HEDT) and workstations CPUs: AMD's Threadripper lineup have official ECC support, but still check with the motherboard vendor first. For Intel, most Xeons should do it, but check before you buy. The same caveat about motherboards applies here, too: Check if there's ECC support first and stick to the QVL to be safe.
Here's the link: https://www.asus.com/support/FAQ/1045186/
Otherwise, a good starting point would be to boot linux with `mce=0` so it panics ASAP upon uncorrectable errors.
In Ubuntu, there are the rasdaemon, mcelog, and edac-utils packages.
Linus uses a Threadripper machine which supports ECC, and non-ecc.
Most Mobos support non-ecc and maybe ECC if it's AMD and the supplier wired it up. (ASUS and someone else I can't recall seem to do so, Gigabyte does not appear to)
https://rog.asus.com/forum/showthread.php?112750-List-Asus-M...
I just built a desktop machine with such a board, a matching Ryzen and 128GB Kingston ECC memory. Works like a charm, the only problem is the on-board Intel Ethernet chip which ignores Wake-on-LAN (although it's supposed to handle it) so I had to add a PCIe ethernet card to get WoL running. Asus and Intel seem to discuss whose fault it is since two years, sigh.
AMD CBS -> DDR4 Common Options -> Common RAS -> ECC Configuration
memtest= [KNL,X86,ARM,M68K,PPC,RISCV] Enable memtest Format: <integer> default : 0 <disable> Specifies the number of memtest passes to be performed. Each pass selects another test pattern from a given set of patterns. Memtest fills the memory with this pattern, validates memory contents and reserves bad memory regions that are detected.
ECC also doesn’t fully protect against Rowhammer, in particular if errors remain unreported: https://news.ycombinator.com/item?id=18508692
So it's not just the error reporting, it's protecting against errors on the whole pipeline, not just inside the chip.
it's a clear part of the original email.
The most damaging Intel decision was about the ECC memory, but there were also many others that are less impactful, e.g. the various ugly workarounds for the laziness of Microsoft of adding various necessary features to Windows, e.g. the System Management Mode or the Management Engine.
The problem with the ECC memory has been created by Intel from 1994 to 1995, when Intel has split their top line of CPUs into two branches, Pentium (the second generation of Pentium, @ 90 or 100 MHz) and Pentium Pro.
For more than a decade, since the introduction of the IBM PC, all compatible personal computers had implemented memory error detection, even if it was possible to use memory modules without error detection, if one did not care about the reliability of the computer.
With Pentium and Pentium Pro, Intel has decided to introduce a market segmentation feature and they have reserved the use of ECC memory for the "professional" Pentium Pro, while removing the support for memory error detection from the Triton chipsets made for the Pentium CPUs (at that time, before AMD integrated the memory controller, the memory controller was still a part of the external northbridge chip).
The successors of Pentium Pro have been rebranded as "Xeon" and they have continued for a long time to be the only Intel CPUs with support for ECC.
The so called "market segmentation", even if it is practiced by a large number of companies, is just a combination of fraud with blackmail, which should have been forbidden by law in most cases.
To introduce market segmentation, a company takes advantage of the fact that the majority of its customers are naive and they are not able to evaluate correctly the quality of a product that they purchase.
The company then uses this fact to extract much more money from the fewer customers who actually know how to evaluate the quality. For this, the company convinces the naive customers that a lower quality is good enough for them and then the company lowers by various means the quality of the products sold at a decent price, in order to able to request an overprice from the quality-aware customers, who are forced to pay, because they do not have an alternative, since the products at the right price do not have the right quality.
This scheme would not work in a competitive market, but when there are few competitors they usually follow the example of the first company which did that and they introduce the same market segmentation policy, because this will increase the profits for all.
Now, because AMD did not disable ECC in Ryzens, even if AMD has provided much worse software support for this feature than Intel (in the EDAC device drivers), at least until recently, Intel has been forced eventually to enable ECC in many models of Alder Lake and Raptor Lake.
Nevertheless, after many decades of lack of support in consumer CPUs, there is a lot of inertia to overcome in the availability of ECC.
Even if now it is easy to find Intel desktop CPUs with ECC support, the ECC support on motherboards requires the special workstation chipsets, so the socket LGA 1700 motherboards with ECC support are hard to find and they are either expensive or with underwhelming features.
All Intel CPUs for mobile applications (the U, P and H series) continue to lack ECC support. Only the HX series for laptops, which are desktop chips packaged in BGAs, have ECC support.
Previously, even AMD had implemented a market segmentation by disabling ECC in their laptop CPUs. That has changed in the Ryzen 6000 Rembrandt series, which have ECC support. However the ECC support remains theoretical, because until now no laptop manufacturer has introduced any laptop with an AMD mobile CPU and with ECC memory, so there is no competition yet for the Intel mobile workstations.
Currently, it is difficult to buy a NAS (where bit flips are arguably more important to avoid) with ECC memory, unless custom-building and carefully selecting parts.
- The amount of memory that leaves the factory that has errors is a low percentage, not sure what it is but we can all agree it's low.
- You can run memtest when you first install memory for several hours or so to be 99.99% sure your memory is good
- There are outside influences and EXTREMEY RARE cases (cosmic radiation etc) where you may still get an error. If you had ECC it would protect you.
- If you start getting errors (as in more than the one off cosmic flare) later you can test the memory again and determine it's damaged
- Most memory you buy for builds has a lifetime warranty
- ECC costs more and reduces performance by some amount
So the only danger of not having ECC is exceedingly rare memory errors. If those rare instances mattered that much then you should get ECC but for everyone else how is it worth it?
The argument in this post is that he wouldn't have had to go through this since the ECC would have covered the error. I just don't see the value of adding this safety system.
TL;DR sometimes memory starts off good and goes bad later.
- Most memory you buy for builds has a lifetime warranty"
running a memtest overnight is hardly a good choice.
The only useful ECC for the user is the one computed in the memory controller inside the CPU, stored in the DRAM and verified after returning to the memory controller.
This allows the CPU to be aware of any error and it also corrects or detects the errors caused by electrical noise on the PCB traces or by bad memory sockets, not only those caused by bit flips inside the memory cells.
https://en.wikipedia.org/wiki/XFS
(and I don't use XFS any more.)
It feels like arguing in bad faith. How did you arrive to the point XFS is non-journaling to begin with?
And if it doesn't help, then why you decided to bring it as an argument?
NB: even if you use ZFS you still need backups.
https://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres...
And the whole points made above are pretty much the exact political points Linus is referencing. It’s clear in his opinion that ECC should just be normal, or at least not so elite.
For example, 12th gen intel CPUs that already exist, can now support ECC because a new chipset enables it. It was "policy" to not release ECC support for a period.
I then play TFA [2] and make the claim that I like it when computers work for at least 3 years.
I end my turn.
[1]: https://en.wikipedia.org/wiki/Hitchens%27s_razor [2]: https://news.ycombinator.com/item?id=33224680
I have always used only ECC memory in any computer larger than an Intel NUC.
When the memories were new, they always had very low rates of correctable errors, e.g. one error after 3 to 6 months of continuous operation.
Nevertheless, I had several cases when a certain memory module started to have very frequent errors after several years of working fine. Due to ECC, I was able to identify it and replace it, before causing irreparable data corruption in files.
Moreover, in one case I had a laptop with ECC memory which seems to have used some poor quality SODIMM sockets. After being not used for several months (which made it more sensitive to air humidity, by not being hot as during use) it seems that the contacts of the sockets had oxidized so when using the laptop again I have seen very frequent memory errors.
Eventually, after some time wasted with investigation, I have scrubbed the contacts of the SODIMM sockets and I have reseated the memory modules, and the errors have disappeared.
ECC is somewhat less necessary in those laptops and small computers that have soldered DRAM chips, both because the total amount of RAM is small (the error frequency is proportional with the total amount of RAM) and because there are no sockets and long PCB traces (which are susceptible to electrical noise) between CPU and RAM.
At least for all computers that have socketed memory, there should have been a customer protection law forbidding the sale of such computers without ECC memory, because it is not acceptable to use a computer that may produce at any time undetectable errors.
- While the error rate of memory is "low" (however you define that to mean), it is not zero, so the risk of memory errors persists.
- A machine without ECC memory has no reliable way to detect memory errors without some type of external diagnostic.
- While a memory test can (hopefully) detect faulty memory, it takes the computer out of operation for however long the test is run, and even then, it's simply a point-in-time test. It cannot detect memory errors that happened in the past, or that will happen in the future once the test has ended.
- ECC provides a mechanism to reliable detect memory errors as they occur, continuously, while the machine is running and performing useful work.
- While many memory manufacturers offer lifetime warranties on their memory modules, they cannot possibly warrant against data corruption and malfunctions caused by memory errors, which can have a much higher cost to the user than the modules themselves (and would almost certainly be more than the BOM cost difference between ECC and non-ECC modules).
- ECC has been cited by Microsoft and Linus Torvalds as desirable and something that should be broadly adopted, and ECC is commonly found in a wide variety of memory products (e.g., caches and solid state storage), with the glaring exception of main memory on consumer PC hardware.
- While ECC does cost more (all else being equal), the side-band ECC being discussed is the same effective speed as non-ECC memory. The overhead of ECC is canceled out by the ECC DIMM's extra capacity and bandwidth relative to the non-ECC DIMM.