When JEDEC was standardizing DDR5 they mandated on-die ECC which checks for errors that happen while the data rests on RAM. This was done to increase reliability because of high memory density. This is good and every DDR5 module has it.
However there’s another ECC: the one where errors are checked each time data is transferred between CPU and RAM. This requires the support of the CPU, the RAM, and the mainboard. This is important for data integrity and every computer that runs for multi-day periods must have it. I’ve found by reading papers about it that 1 bit error per 4 GB RAM per 3 days is guaranteed to happen.
If RAM is marked DDR5 it has on-die ECC. If RAM is marked DDR5 ECC it has on-die and transfer ECC. That is what I want and it is not optional at all in my opinion.
So for my 64GB of RAM I get 2 bytes worth of errors / day. And a few hundred for 2 or 3 months since I didn't restart my PC.
Memory can be assigned to all sorts of things so the impact of a single bit change will depend entirely on what that memory is being used for at that moment.
For example, I load a program into memory, but I'm not running all of the code in that program all the time. Let's say the program includes code for exporting a list to excel, but that's a feature I never use.
Or it has a procedure which uses a local variable as a loop counter. Outside the confines of that loop the variable is meaningless.
Plus its unlikely you are even nominally using all 64g at the same time. The errors might all be happening in unused ram.
Even if it flips actual data, it may end up being no more than a typo in a document.
So yes, a single bit in 4gb might flip from time to time. But the probability of the flip being "meaningful" is less than 1.
There are even worse potential outcomes:
* A single bit error in a stream of compressed or encrypted data stream that renders the stream unreadable.
* A permissions check being yes instead of no.
* Backups being silently corrupted
Saying that say 1 bit per 4Gb will flip every day is a long way from saying that a _significant_ (or even detectable) bit flip will happen every day (per 4Gb). From a simple feel for "what memory is used for", coupled with "how much of the ram is actually used" makes me suspect that it's a very low probability of it being an issue.
Which makes the parent comment, about 1600 bits of data being flipped since the machine last rebooted [1] being both "true" and likely "meaningless".
[1] I'm not sure what rebooting has to do with anything though.
Like isn't this the kind of thing that Turing and Zuse had to deal with?
And if it isn't, when is it a problem that we will finally and absolutely excise?
Is it more of a 2030 thing? 2040?
That is very, very dubious indeed. I can't believe there's no ECC between ram and cpu, it's not credible for anything but a cheap home machine. Any sources?
Most paper characterizing errors are very old (so they used memory chips with large feature sizes), and they scaled their error rates by memory amount (because they didn't know better). This results in massively overinflated error bit rates. It's possible to prove that this is wrong very easily by just writing a program that allocates a large pool of ram, and constantly reads it and writes back to it over a period of time, and then check results after a while. I once did this for an argument on the internet about this -- on my old Intel Sandy bridge system with 16GB non-ecc ram, I allocated a pool of 8GB, wrote ascending integers from 0 to 2^31 to it, and then iterated over it, just reading it in, checking if it's still correct and terminating if it's not, and writing it back out to same address. This program ran for a week with no errors before I had to end it because I needed the machine. 1 bit error per 1.333GB*days is completely not credible.
> . I can't believe there's no ECC between ram and cpu, it's not credible for anything but a cheap home machine.
it is absolutely not wrong: there is no ECC between CPU and RAM for most consumer and desktop computers, and even tons of embedded systems, and that's basically the only bus lacking even basic integrity checks, so there really should be some ECC also there, nearly everywhere.
As for embedded, I'd be surprised if there's ECC at all, never mind on the bus (with the exception of high reliability stuff like spacecraft and factory controllers. Clocks, cookers, microwaves, etc don't count).
Kind of, that's why I was I was saying it is not wrong. IMO though, there should be ECC on everything but the really cheapest computers and if the application even allows it, meaning ECC for business desktop computers, and home computers marketed as using state of the art tech, like Core processors. Maybe you can avoid it in game consoles, or computers marketed for 99% gaming (except they risk being expensive and using good quality processors, so what's the point), but that should be it.
IIRC, the error rate is more like 1 bit error per 4 GB RAM per 30 days (in the relatively recent studies made by Google for their servers).
Nevertheless, that is still not negligible. For workstations and servers with 64 GB DRAM or more, the errors can become frequent. Also, at high altitude the error rate is higher.
Moreover, besides the errors due to radiation, there are errors caused by electrical noise, which can become more frequent in old computers with cheap DIMM sockets, whose contacts have oxidized.
(I actually had this problem in some HP laptops, which had been stored for a long time without being used; being cold made the SODIMM sockets more vulnerable to humidity; after identifying the cause and scrubbing the sockets, the frequent memory errors disappeared.)
This estimate is several orders of magnitude too high. We’d be seeing weird bit flips and file corruption all over the place if this was even close to being true.
I have ECC systems that will report the number of errors detected and corrected. Still waiting to see any errors on my main machine with 64GB of RAM.
At server farm scale, the majority of memory errors come from a very small number of faulty memory sticks. It’s not an even distribution of errors across all memory.
I don't know if internal ECC errors ever get reported.