Normal good DIMMs, at least when new, should not have errors more frequently than one error per many months.
Frequent errors may appear with memories that are not seated well in their sockets, or which are old, at least several years old.
Frequent errors may also be caused by more general computer problems, like a bad power supply unit.
(Correlated bit-flips are likely from a manufacturing error... or, at least in one case, from excessively radioactive ceramic packaging.)
The more atmosphere you have to randomly absorb high-energy photons, the better.
Higher density memory drops fewer electrons in each well, which means a lower-energy photon can change the state.
Corolary: submarine datacenters are much better than orbital datacenters. Cheaper to cool, construct, shield and access.
This gets said a lot but with little evidence. The mario speedrun bit flip has been replicated with marginal connector insertion.
Whether it is the primary source of errors, depends on the design of the DRAM die.
In a badly designed DRAM there could be many other error causes.
The designers of a DRAM certainly attempt to minimize all the error causes that they can control. If they succeed, the cosmic radiation would remain the primary error source.
Without access to internal documents of the DRAM vendors, we cannot know if they have indeed minimized all other error causes.
In modern DRAM chips, a major error cause can also be the disturbance of some memory cells when some other cells are accessed in their neighborhood, which is the basis of the Rowhammer attacks, but which can also happen during the normal use of the memory, when certain access patterns happen by chance.
The DRAM vendors do not provide details about the behavior of their devices, with the hope that this makes harder the task for someone who wants to design a variant of the Rowhammer attacks, but this also makes impossible for the owner of the memory to predict whether an application program will not perform by chance an access pattern that will unintentionally flip some memory bits, which without ECC will not be detected and corrected.
Avoiding returns for ECC DIMMs is a harder challenge when they are used in an environment where the server will have a management system that tracks errors over all time and when there might be a network management system tracking all errors over all systems.
Avoiding returns for non-ECC DIMMs is easy when a few errors per month won't be noticed by most users who assume that Windows always crashes anyway.
Has any DRAM company tested chips and made non-ECC DIMMs from the ones that have lower quality?
As an aside are the 2^N*3 sizes of DDR5 DIMMs (like 24G and 48G) made from chips that had errors in one section and got reconfigured to have 3/4 the capacity to not use the bad parts?
Are you disputing that the Earth's atmosphere absorbs and reduces the amount of ambient cosmic radiation? Or where is the skepticism coming from?
Muons are nasty little buggers.
Yea, except datacenters also need connectivity. And submarine connectivity, while not particularly expensive, is quite vulnerable - even more so than on land.
If you only need 100M deep water then you don't need to go far offshore, there are places where it's only about 1KM out. 12M of water is used for storing spent nuclear fuel assemblies and less than half of that is needed for protection so 5M should be adequate for stopping cosmic rays and that depth is common in harbours for recreational boating, such harbours shouldn't have issues of cable breakage.
But the real issue for comparing submarine and orbital computers is cooling. Space is cold but doesn't allow easy transmission of heat.
I would think most servers and workstations would be RDIMM (Registered DIMM) by now and consumer stuff uses soldered down memory. Memory failing because it's old is definitely a thing, and very possible in this scenario, but I feel like I haven't seen errors due to physical insertion, that were not caught immediately by POST, in years.
But maybe it's just me. Happy to be lucky. :)
The computers in plastic cases can sometimes be affected by appliances with electrical motors that are used nearby them.
The cosmic radiation is attenuated only inside a big building, e.g. one with many stories above you. A really great attenuation is obtained only underground.
But overall, pc cases are probably too thin to appreciably affect the radiation levels. Cosmic rays and their secondary radiation is fine with going through buildings so a 1mm thick case isn't going to do much.
Google and Facebook report a few unreliable CPU cores out of thousands of systems. Meaning maybe a 1/10,000 rate of repeatable CPU errors.
https://www.cs.toronto.edu/~bianca/papers/sigmetrics09.pdf
Google reports 8% of DIMMs being affected by errors. While the differences in metrics prevent comparing the two things directly it seems that DRAM errors are more common than CPU errors.