https://forums.freenas.org/index.php?threads/ecc-vs-non-ecc-...
The reality is that the worst case consequences for a bit flip is the complete loss of all data, regardless of the filesystem used. While the automated repair routined bundled in fsck utilities are often able to fix problems on other file systems, the problems that they do fix simply do not occur on ZFS. The kinds of problems that kill ZFS are not among those that an automated repair tool can fix. e.g. overwrite all ext4 superblocks with random data and then see if fsck.ext4 can fix it.
That said, I am usually able to resuscitate a pool that another person would have considered to have been killed by a bit flip. I not only find such failure modes to be incredibly rare, but I find those that I cannot fix to be a rarity among the cases where things did in fact go wrong.
I have had a VERY hard time finding a situation where there was corrupted data in ZFS, and where there are ZFS pools that absolutely will not import back into full operation.
In other words, it is pretty difficult to "accidentally" corrupt a ZFS pool to the point of not being able to recover your data.
ECC RAM greatly minimizes the potential for a corrupted file in ZFS, but ECC RAM also greatly minimizes the potential for a corrupted file in ext4 or XFS.
So long story short, I have experienced the same thing as ryao, and come to the same conclusions.
ZFS warned me of that while I reading its installation guide. Which is why I never installed it.
https://news.ycombinator.com/item?id=8438416
Would you provide a link to that guide? If there is a guide out there that says that you should use something other than ZFS when a system lacks ECC, I would like to know so that I can try to get it corrected.
Not using ZFS because it cannot provide full protection without ECC is like not getting a flu shot because you can still get sick anyway. It is true, but opting to be even less safe in the name of safety is counterproductive.
You are wrong. Don't spread potentially dangerous maladvice on the internet when you have no idea what you're talking about.
The official ZFS documentation[1] tells you to use ECC Ram and why. The first google hit for "zfs ecc ram"[2] further elaborates on the risks of using ZFS without ECC memory.
[1] https://pthree.org/2013/12/10/zfs-administration-appendix-c-...
[2] http://louwrentius.com/please-use-zfs-with-ecc-memory.html
If your house burns down ZFS will not save your data. If too many disks fail zfs will not save your data. If your disks' unrecoverable read error rate is too high and your array is rebuilding zfs will not save your data.
If your computer is fundamentally broken (which is what a machine with a persistent memory error is) zfs will not save your data.
Take backups.
If you had a 'stuck bit' anywhere in your memory space you'd be way deep into nasal demon unspecified behavior the first time you tried to dereference a pointer that crossed that bit. When your hardware is that broken you can't count on the OS to not stab your dog much less ZFS to do anything sane.
Note that this is just as true for all storage systems. Regular filesystems might corrupt themselves or do other insane things. Hardware RAID or kernel software raid will happily propogate the error. How many bits separate the kernel record for "this disk is totally cool" from "this is a new disk and should be zeroed"?
As for the articles that you link, they are correct to say that you want to use ECC RAM. However, there is nothing specific to ZFS that makes it require ECC any more than any other filesystem. It should also be noted that I wrote the official documentation on this subject. It can be found at the Open ZFS wiki, rather than the pages you linked:
there is nothing specific to ZFS that makes it require ECC any more than any other filesystem
This statement still troubles me. I was under the assumption that ZFS scrubbing may indeed lead to compounding corruption as described in the blog post I linked. Hence the ongoing(?) debate in the ZFS community.
However, I consider myself disqualified from this discussion and sorry again for my tone. I'll read up on the current state of the argument.
That said, the only thing in the code remotely similar to "compound corruption" is logic for preventing such a scenario, which has the following comment:
/*
* Don't rewrite known good children.
* Not only is it unnecessary, it could
* actually be harmful: if the system lost
* power while rewriting the only good copy,
* there would be no good copies left!
*/
https://github.com/zfsonlinux/zfs/blob/master/module/zfs/vde...That touches more than obvious at a glance because the mirror code is used for processing both mirrors and ditto blocks.
I'm not familiar with ZFS at all, so I don't know what the terminology 'ditto blocks' and 'mirrors' mean, so there's a good chance I'm not talking sense here at all.
The main argument for ZFS on non-ECC using systems being worse than other less sophisticated file systems seems to be that just reading a file (or during periodic 'scrubbing') while you have bad ram can cause good files to be corrupted when ZFS notices checksum errors and tries to correct them. (There are other arguments about not being able to mount pools due to bad ram, but that seems no different to any other FS).
People seem to be quite insistent that this is possible, and at first glance the code you linked above doesn't seem to be relevant since it refers to 'known good' data, but the problem scenario is when data has (erroneously) been determined to be bad due to faulty ram. Will ZFS overwrite good data in that situation in an attempt to correct what it sees as a disk error?
A few links where people are making this claim:
https://forums.freenas.org/index.php?threads/ecc-vs-non-ecc-...
https://forums.freenas.org/index.php?threads/ecc-vs-non-ecc-...
However, thanks for linking to my blog! :)