I was always thinking that ECC can prevent blue screens/kernel faults, but I haven't seen those in years on my laptop without ECC.
I was always thinking that ECC can prevent blue screens/kernel faults, but I haven't seen those in years on my laptop without ECC.
Anyone using ZFS will (or should!) care about ECC support. [0]
Lots of people build their own NAS/SAN boxes, so ECC support on a desktop CPU at a reasonable price point would be very appreciated. Currently you need to buy specific model CPUs (Celeron or Xeon, IIRC) to get ECC support from Intel. [1]
[0] https://serverfault.com/questions/454736/non-ecc-memory-with...
[1] https://ark.intel.com/search/advanced?ECCMemory=true&MarketS...
You are getting the same benefits using any filesystem while running with ECC.
Also the "myth" that wrong checksum calculations due to a bitflip will degrade a ZFS filesystem even faster is not valid.
There are some articles/newsgroups around explaining this in much more detail.
So: Using ZFS on non ECC is not inherently unsafer than other filesystems on non ECC.
https://arstechnica.com/civis/viewtopic.php?f=2&t=1235679&p=...
However, the checksum and the data is also only as good as the memory in which it is stored during computation. Since ZFS takes care of everything else other than memory, you'd use ECC if the data being stored is important enough that you want to ensure that ZFS gets healthy checksums.
Once ZFS writes corrupted metadata with corrupted checksums to the disk, it's very hard to recover that data. Yes, ZFS is not backup, but the implications of serving clients messed up data is also worrying in certain business cases.
Given that ECC RAM isn't (much?) more expensive, why not just use ECC? (For my microserver I remember it being the same price as non-ECC.) Well, because CPUs with ECC support are rare, etc.
Without ECC, ZFS loses one of the guarantees that it otherwise provides. But it only degrades down to how bad every filesystem is in the face of memory corruption - not worse.
If Zen client chips have full ECC support and can handle 64GB of memory I'll be sold easily on an upgrade for my TrueNAS box, then I'll anxiously await some lower-cost (4/8c) dual-socket server CPU's and swap out my TD340 with some supermicro barebones build.
The E3-1220 is usually idle and frequency scaled back unless something like a ZFS scrub is running, I've got ~45W of PCIe cards (SAS HBA, 10GBe NIC and 4x1GBe NIC), 4x8GB sticks of DDR3 UDIMM's probably uses 12W.
I'd say without drives it probably draws no more than 150W on average, unless I'm doing something CPU intensive. My drives all sit in an external SAS enclosure, since I wanted more than 4 drive bays that the LFF expansion bracket provided.
Anyway, the ML10 is pretty decent for light-medium work. It's got 4 really fast cores and 32GB of RAM is adequate for most home server use (this was why I got the TD340 though, I've got 72GB in it) - just beware that the Gen1 units don't include the drive bracket so if you want more than a single HDD you have to purchase one or get an external SAS enclosure.
EDIT: the ML10 is basically silent too, even under load - I can't hear the fans unless I try (though my SAS enclosure makes up for this by being the loudest bit of kit in my lab).
For way less, if you're willing to five up on ECC, you can get an i5-7500T and a pretty good motherboard, bringing your TDP to 35W but having a passmark of 7055, and most importantly, a single thread rating of 1924.
This might not matter much if you're use is exclusively NAS, but I will probably end up running some virtualized or containerized server, or streaming video possibly transcoding on the fly, and I'm afraid the Atom might become a bottleneck.
Is there a solution that is somewhat competitive with the i3/5/7 on price and power, and has ECC? And, ideally, that comes in a Mini-ITX form factor?
[0] Obviously the PassMark score is only a ballpark estimate to get an idea of how fast is a chip, but still.
The 80W TDP worries me about though, cooling, noise and money/pollution wise (the extra 45W make around 400kWh/year).
Considering the CPU will be idle most of the time, do you have figures about the actual idle consumption of an E3 based machine?
[0] https://www.theregister.co.uk/2017/02/06/cisco_intel_decline...
Coding Horror To ECC or Not To ECC [1], What Every Programmer Should Know About Memory [2], Memory Errors in Modern Systems [3], and an analysis of memory errors in the entire fleet of servers at Facebook over the course of fourteen months [4].
[1] https://blog.codinghorror.com/to-ecc-or-not-to-ecc/ [2] https://people.freebsd.org/~lstewart/articles/cpumemory.pdf [3] https://www.cs.virginia.edu/~gurumurthi/papers/asplos15.pdf [4] https://users.ece.cmu.edu/~omutlu/pub/memory-errors-at-faceb...
RAM bitflips randomly, period. It's just how it works. A cosmic ray can hit the memory chip just right and flip it, there's no way to predict or control that no matter how "stable" your machine is. ECC still does the same, it just has a parity bit on each line to confirm against and flip it back, as needed.
Cosmic radiation bitflips are BS. No cosmic rays reach ground level. The chance of a, say, Al-28 nucleus successfully penetrating the entire atmosphere is as close to zero as it could be possible to get. Basic physics. Something with such a high charge density won't penetrate ~100km of atmosphere and magnetic field. Even a basic muon wouldn't get through a sheet of aluminum foil and those are still capable of actually getting (barely) through the atmosphere.
The chances of cosmic radiation causing a bitflip are pretty much in the range of "Elvis coming into town on Nessie." Radiation originating from inside the system itself is much more likely a cause.
Your ground-level bit flips are most likely caused by terrestrial radiation sources, not extra-terrestrial ones. This is just basic physics.
Only a small minority of main memory data corruptions lead to OS crashes, mostly the in-memory application or filesystem data just silently gets corrupted.
Not exactly. The prices between comparable i7 to Xeon are nearly the same.
I did check and it is still true that with laptops you always pay a premium for Xeon v Core on an equal performance basis (even within the same model).
Because I care about my data. Data corruption may kill your main storage and the first backup, too.
Companies I worked for used those only on the servers where we run critical stuff, not a single developer laptop had ECC.
Not sure I agree. Most science datasets (both from simulations and experiments) are sufficiently noisy that if your scientific end results and conclusions change as a result of even thousands of bitflips in your 16 GB of data, you're Doing It Wrong and your article isn't worth the paper it's printed on. (There are probably exceptions, as always, but those working in those few specific subfields should be aware of it.)
There is way too much overreach in this statement. Imagine doing Finite Element or Computational Fluid Dynamics analyses; bitflips of the floating-point values in the field solutions, which could easily make those values completely unphysical, are not the kinds of errors the solvers are written to guard against. In order to do so, you'd need to sanity-check every value, and if you had to use a "guardrail" value, it could easily take a significant number of iterations to recover to the more correct value. Solvers can be easily crashed by corruption of numbers. Sure, if you're lucky enough to have bitflips in low-order mantissa bits, no real harm done. Just don't expect the bitflips to cooperate in this way.
Maybe the "big data" and machine learning crowd don't care about some corrupt values, but most numerical/scientific computing is not so sanguine about corruption.
It is a little frightening how we are moving into significantly larger computational solutions, but are simultaneously increasing our exposure to the fragility and lack of guarantees regarding enormous quantities of perfect bits at all times.
a) bit flip happens so high in the mantissa it makes your code crash. detectable, you run simulation again
b) bit flip happens somewhat lower in the mantissa, enough that it shows up in your analysis. detectable as unphysical result, you run simulation again
c) bit flip happens even lower, you don't catch it as unphysical in your postprocessing and analysis. Here's where I'm saying: if a bit flip happens like this and affects your simulation in such a way that you don't see it's an error, but it still changes your end result and conclusion, and you're not running replications of your simulation to test robustness etc., you're Doing It Wrong.
d) bit flip happens even lower, same order of magnitude as numerical errors. nothing bad happens
There's a reason why HPC systems universally use ECC.
I was under the impression that HPC systems universally use ECC because at that scale, the probability of memory errors in the OS are large enough to cause constant instability of some of the nodes?
More generally, it's not unreasonable to say that our entire computing paradigm rests on accurate RAM. Above 16GB, the risks just become too big for anybody doing serious work, and not just messing around with prototypes.
Exactly this. The size of your code is positively infinitesimal compared to your data. And unless you're writing your code and then running it exactly once, which is a) even more unlikely and b) bad practice, you'll catch any of those bit-flip errors in your code or data structures.
This has been discussed a lot in the literature, especially for GPUs where ECC carries a performance penalty both on speed and available memory, e.g. in this paper where they've tested it on a GPU cluster:
http://www.rosswalker.co.uk/papers/2014_03_ECC_AMBER_Paper_1...