I'd much rather manage a simple local disk with offset backups.
[0]: https://www.jodybruchon.com/2017/03/07/zfs-wont-save-you-fan...
I'd much rather manage a simple local disk with offset backups.
[0]: https://www.jodybruchon.com/2017/03/07/zfs-wont-save-you-fan...
The article is interesting in his contrarian view, however, when it comes to bitrot, it counters anecdata with other anecdata:
> One bit flip will easily be detected and corrected, so we’re talking about a scenario where multiple bit flips happen in close proximity and in such a manner that it is still mathematically valid. While it is a possible scenario, it is also very unlikely. A drive that has this many bit errors in close proximity is likely to be failing
I detected bitrot once or twice, and in neither case the drive was failing. This is anecdata though - is it valid? Who knows.
I'm personally skeptical about blanket statements (which the author makes) without seriously backing data.
I have a ZFS setup, and it's arguable whether it's a hassle in itself. At least for RAID-1 setups (I have two), once installed, it's not inherently harder to maintain than other FSs. Installation is manual, and that's definitely a hassle, but users are definitely intended to be advanced ones.
Regarding SMART: it's not as easy at the article author states. I have a laptop that periodically pops up with new instances of a certain error, but the SMART guides says that this is not an error one needs to consider, so I'm confused. Additionally, the smart-notifier of Ubuntu (at least up to 18.04) is broken. I agree that SMART is important to consider, but it's not straightforward as it seems.
It is perfectly possible to read corrupted data from a disk. I know this because I've seen it happen several times over the years. If your system is making decisions (ie generating new data) based on read information, this can actually be quite harmful. Like it or not, transitional errors on to-be failing disks may cause data corruption. It is easy to say "hey, restore from backups!", but it may happen weeks go by before an actual failure happens. By then you don't know if your backups are tainted, and how long was your storage misbehaving. ZFS actually helps with this, because it can tell you explicitly that your file/block is tainted, even if readable. This provides a level of confidence on the system, based on observability. And snapshots can actually refer to different blocks, so it is often possible to recover a previous version of a given file without firing up the backup system.
Also, the idea you need RAID for healing is nonsense - ZFS can keep multiple copies of your block even on a single-disk system.
To finalize on the "bitrot" topic, keep in mind network communications, as most serial protocols, have varying degrees of CRC and checksum checks at different levels of the stack, following the end-to-end principle. We even use compression and encryption on top of that, that also provides multiple verification methods. Yet most relevant files such as iso images have a checksum file to verify your download - and sometimes it doesn't match. ZFS provides you the same functionality, but for your storage.
I can say that at both work and home, I have only ever seen groups of mirrors or RAIDZ in use - I've never seen just striped pools or single disk ZFS.
I know it's anecdotal, but I have seen ZFS recover data flawlessly with drives returning incorrect data for some sectors with no I/O errors, or from total and sudden drive failure with no SMART warning. I personally think drive hardware is rather more fallible than this article assumes.
With that said, of course ZFS is not a magic bullet, and there's no substitute for backups - but ZFS does make that easier too, because snapshotting is trivial, and zfs send | zfs receive is very useful for transferring the snapshots to another pool for backup. And it does require an amount of reading and understanding before you set it up.
And SMART tells you that a disk is dying, it doesn’t tell you the a disk is not dying.
Furthermore the disk health tools can’t catch e.g. a dying cable.
Some SMART implementation can report E2E errors.
However, you could also use MDRaid and LVM2 Thin Volumes with any file system you like to get almost the same, without additional kernel modules.