I still think it's important, but it makes sense that people don't want to pay more for something to detect errors that are quite unlikely.
Are you talking about ECC ram being more expensive? That's probably true but why not give people the option whether to use ECC? ECC being optional is a good idea anyway because ECC memory is generally slower.
I don't think the actual ECC implementation in the processor is very expensive to implement.
Consumer CPUs not having ECC is pure product segmentation, without technical justification, IMO.
But supporting ECC ram also adds more expense. To properly support it, BIOS engineers have to support and test it; motherboard makers have to support and test it.
I appreciate that AMD offers it for their consumer oriented processors, but because it's a best effort feature, and it's hard for an end user (or reviewer) to test, you never really know if you're going to get full support, or if the support is really just that you can use ECC ram, if you want to spend a little more for your ram, and get none of the benefits of ECC.
It certainly adds something to the cost (die space) of the memory controller, but I agree it's probably not much.
Realistically the mainstream CPU manufacturers have ECC solutions, so including them in the mainstream processors shouldn't be a huge issue, it's just a market segmentation ploy to exclude the feature from the consumer processor designs.
I think the cost pressure is too high. Early IBM PCs required parity ram (a 9th chip to store if the sum of bits was even or odd), and would fault if the value was incorrect on reads. Ram module manufactures made innovative fake parity modules that calculated the parity value on access, replacing the 9th ram chip with a very simple circuit and saving money.
It would be hard to convince the whole industry not to make fake ECC ram, if ECC was mandatory.
We will have to see with DDR5 (because it supports "internal" ECC) if its worth it to the memory industry to build RAM that is internally denser, but more error prone (as is the case with modern flash) or continue to attempt to build 100% reliable ram (and failing).
I'm betting some clever person figures that out. Which leaves only the memory bus itself unprotected. Which IMHO, is foolish and serves only to create product segmentation. So, for a DDR5 dimm with internal ECC, generating bus ECC should be a trivial addition.
> It would be hard to convince the whole industry not to make fake ECC ram, if ECC was mandatory.
Assuming you don't use memory-mapped IO, that's easy to fix. On startup, generate 4 random bits a,b,c,d. Parity bit is data line 4a+2b+c, with d?even:odd parity, data bit 4a+2b+c is on data line D8. ECC on 64/72 uses more random bits, but is otherwise similar, although for modern chipsets it would probably have to be scrambled in the northbridge or southbridge (or equivalent) rather than the CPU, to allow for DMA and such. Note that there's no gate delay involved here; the multiplexing can be done with pass transistors.
I also disagree that the benefits are meager. I have suffered from several computers with flaky memory (fine when you build the computer, flaky several years later; then you do an overnight memory test and find that indeed the memory has gone bad), and a strong software signal that "hey your memory is bad" is very actionable. You also have to think about it from a programming standpoint -- what happens if this variable isn't what I set it to? What if "for 1..10" is actually "for 1..2147483658"? Do you have time to debug that? How much data do you lose when you persist that to disk? To me it is insane to not get this nearly-free consistency check if you ever plan to persist any bytes in RAM to long-term storage. Even consumer GPUs have ECC memory these days. It's a no-brainer.
What surprises me is that the industry segments RAM sticks by ECC/not-ECC, when it's just a software function performed by the memory controller. I think everyone with a HEDT setup would be happy to enable ECC and get better reliability at the cost of 12% less RAM. I know I would be. (I built a Threadripper workstation recently and just couldn't get any reasonable prices on QVL'd ECC memory. So I skipped it and paid like $600 for 128GB of 3600MT/16-CAS memory... would be happy to flip a switch and have that be only 112GB.)
The reason why ECC isn't widespread is it's because it's used to segment the market. Home users are cost sensitive, so hardware vendors have to think of ways to get server users to not buy home user equipment. ECC is one of these levers. That's all there is to it. Everyone would have ECC if it were a mere 12.5% more expensive.
Standard memory modules become (a multiple of) 9 bits wide, with the additional bit stored near the other 8 bits.
The problem is _INTEL_ deciding its a premium feature, and the memory manufactures charging 50%+ more for 12.5% more hardware.
So instead of being a couple percent of the system cost, it ends up being a noticeable double digit percentage.
Of course none of this explains why apple/etc haven't done it in their phones.
ECC memory physically requires additional wires, necessitating different parts that don't sell at the same volumes.
And as far as "wires", traces are basically free as long as they don't force additional pcb layers, which shouldn't be the case given the careful pin/ddr chip/processor designs focused on exactly that. The extra pins are there regardless, and in the past so were the "wires" given that CBx/DMx pins were muxed. Further, packetized ram interfaces have been known to bury the error correction in the protocol same as PCIe/etc. Meaning its a slight efficiency loss.
Look at it this way, every part of the system _EXCEPT_ the ram has some kind of error correction on it at this point. Its a cheap way to not only increase robustness but its also a security mechanism.
If I were guessing, I would say that DDR5 is the last version that isn't ECC end to end, the idea that the manufactures can shrink the dies at the cost of some BER and make it up with a bit of embedded ECC will just be to tempting. Particularly with AMD on the rise, intel will have a much harder time playing product segmentation games if AMD doesn't do the same.
So I save $50 on the CPU and that just about covers the ram/motherboard premium.
(I think there are probably issues with this in terms of how ddr itself actually works with its interleaving and adding a third channel to access from but I don't know enough about how it works to know for sure if that's the case)
At the end of the day, electrical/computer engineers do actually generally know what they're doing.
For a memory controller to add extra bits, they would have to come from somewhere. For every 64 bits (8 bytes read), you now need to read 72 bits (9 bytes).
However “DIMMs are printed circuit boards that carry multiple packaged DRAMs and support 64bit or 72bit databus widths, the latter to enable eight error-checking and correction (ECC) bits to protect against single-bit errors.”.
So now every read from 64 bit (non-ECC) DRAM now needs two reads, one for the 64 bits you want, and another to read the 8 bits ECC.
If your access pattern is random, then you will slow down memory access by 100%. For long run sequential reads/writes the slow down will be 12.5% assuming you can lock the memory accesses for 9 sequential reads/writes (avoiding invalid memory states is essential!).
The “cost” of ECC implemented by the memory controller is:
* you lose 1/9th of your memory (as you point out)
* the speed of your computer drops by up to 100% (speed of many operations is limited by random memory access speed, not CPU)
* CPU atomic instructions for memory access are complicated https://en.wikipedia.org/wiki/Compare-and-swap and would also incur a further speed penalty in critical sections - which can very significant.
The idea seems useful (clearly it would be a fantastic feature to be able to switch on if you suspect faulty memory) but there are likely technical reasons for why your idea is not implemented (not just price discrimination).
> dumb
That appears condescending to me. Generally you should assume people are smart. ECC RAM is designed by smart people to be the way it is and start with the assumption that if an idea isn’t implemented that perhaps there are good reasons why not?
(Edited to make reply flow better. Disclaimer: I only have a very shallow knowledge of the design constraints, and I expect there are other more serious problems with the idea).
By the way, for detection you don't need the full error correcting code, but you can use an error detecting code which can use fewer bits.
Does it, though? I don't remember all the prices from when I was researching memory close to a year ago for a new PC, but one big notable difference between ECC and non-ECC I've seen is that almost all non-ECC sticks are overclocked. You buy something rated at 2667MHz, and you actually get is e.g. 1833MHz chips overclocked at that speed. Which makes them cheaper. But I'd expect non-overclocked non-ECC sticks to be that far in pricing to ECC sticks (well, apart from the obvious difference they'd have more chips). And because the market is focused on cheap non-ECC sticks, they don't do cheaper ECC overclocked sticks, but there is no reason they couldn't make them. There aren't that many ECC unbuffered sticks already, it's easier to find registered ones. Essentially, the market is skewed by Intel not supporting ECC on non-Xeons.
Been there. If your servers don’t have ECC memory, you’re eventually going to get bit.
I built a ML workstation for my work team and was disappointed that my options were: expensive, low clock speed server CPU + ECC, or inexpensive, very fast, desktop CPU without ECC. Even if I were willing to pay more money, it was really hard to get the same performance if I needed ECC.
[1]: https://hardwarecanucks.com/cpu-motherboard/ecc-memory-amds-...
On the other hand, AMD doesn't do that.