It's nearly impossible to find a desktop CPU that supports ECC ram now even though 5 years ago it was commonplace.
Trying to run a NAS with some sensitive data is now impossible unless you buy their server chips
That's not correct. All 6th Gen Core i3 have ECC support, and the 7th Gen Core i3 that have 'E' in the product name support ECC:
http://ark.intel.com/products/97130/Intel-Core-i3-7101E-Proc...
Quad (and more) core i5 and i7 that do have near-equivalent Xeon parts do have ECC disabled.
Getting an Intel Xeon E3 for socket 1151 isn't particularly hard or expensive either.
Btw, is there any updated paper/source with some stats on bit-flip odds in modern computers?
Last I've read was this one [1] but it's been debated to death. [2] [3]
[1] http://lambda-diode.com/opinion/ecc-memory
[2] https://news.ycombinator.com/item?id=1109401
[3] https://www.reddit.com/r/programming/comments/ayleb/got_4gb_...
Smartphones and laptops should also come with ECC-RAM.
It is that important for reliability.
It's important but it's not that important outside of specific applications/use cases.
They do fail. Lots and lots of times. And it gets worse over time, as chips degrade due to hot-electron effects or electromigration.
You'll have to use your best guess or intuition on this.
On some systems, ECC is flat out broken or silently ignored. On many others ECC errors aren't reported to the OS in granular enough batches to do anything about them.
Edit: relevant portions:
> Before we begin to discuss our results, we must first discuss ECC protected systems and the fact that there are no standards on what constitutes ECC protection or ECC event reporting. At its base level, ECC protection simply means that a server can handle or correct single bit errors, although some systems advertise the ability to correct multiple bit errors. Generally, it is our belief that ECC events should be reported to the operating system so that savvy users can gauge the health of their infrastructure.
> Unfortunately, server vendors routinely use a technique called ECC threshold or the 'leaky bucket' algorithm where they count ECC errors for a period of time and report them only if they reach certain levels of failure. From what we understand, this threshold is commonly above 100 per hour, but this remains a trade secret and varies based on the server vendor. So, to see ECC errors (MCE in Linux or WHEA in Windows), there generally needs to be 100 bit flips per hour or greater. This makes “seeing” Rowhammer on server error logs more difficult.
> In addition, we have observed some server vendors will NEVER report ECC events back to the OS, although they might get logged into IPMI. Typically, users expect to see correctable ECC errors logged directly to the OS or that halt the system when they cannot be corrected. During our investigation into this phenomenon, we even encountered one server that neither reported ECC events to the OS nor halted when bit flips were not correctable. The end result was data corruption at the application level. This is something, in our opinion, that should never happen on an ECC protected server system.
> ...
> Using these advanced techniques, we were able to observe ECC events within the first 3 minutes of test time. Generally, this system would lockup or reboot within 30 minutes. Once again, this was on a Rowhammer mitigated system using both ECC and double refresh. So it follows that dual mitigations, on some systems, appear to be flawed and can be exploited as a denial of service.
On the fact that it's broken on some systems. Isn't this like complaining about air bags that have been disabled ? It however very good you report that they are broken on some systems.
That said if you want hardware level attacks then attacks against the cachelines of the CPU are considerably more dangerous and reliable and there is very little one can do to mitigate against those.
How does ECC compare to the fiat money ponzi scheme?