I have seen a few servers with ext4/mdraid over the last five years have serious corruption but have had to reset a ZoL server maybe twice.
I transitioned an md RAID1 from spinning disks to SSDs last week. After I removed the last spinning disk, one of the SSDs started returning garbage.
1/3 reads are returning garbage and ext4 freaks out, of course. It's too late and the array is shot. I restore from backup.
This would have been a non-event with ZFS. I've got a few production ZoL arrays running and the only problems I've had have been around memory consumption and responsiveness under load. Data integrity has been perfect.
ZFS-on-Linux devs say it's ready for production[1].
Lawrence Livermore laboratory stores petabytes of data using ZoL[2].
If we're sharing anecdotes, ZoL has served me fantastically for several years.
[1] https://clusterhq.com/2014/09/11/state-zfs-on-linux/ [2] http://computation.llnl.gov/newsroom/livermores-zfs-linux-po...
1) Seen users complaining about data loss on issues on github. 2) Had the init script fail on upgrade and had to fix it by hand when upgrading Ubuntu. Probably a one time issue.
Need a bit more reliability from a file system.
https://github.com/zfsonlinux/zfs/issues/5535
We're strongly considering using something else until this gets addressed. The problem is, we don't know what, because every other CoW implementation also has issues.
* dm-thinp: Slow, wastes disk space
* OverlayFS: No SELinux support
* aufs: Not in mainline or CentOS kernel; rename(2) not implemented correctly; slow writes
If that were me, I'd see how quickly it was fixed before strongly considering something else.
In our opinion, when we made the switch, it was much more important to trust the integrity of the data, than any possible kernel panic.
The setup process was a bit painful given some interesting delays when using some HW storage controllers that caused udev to not make some HDD devices available under /dev before the ZFS scripts kicked in and we have been bitten a couple times by changes (or bugs) in the boot scripts, however the gains provided by ZFS in terms of data integrity, backup, and virtual machine provisioning workflow were definitely worth it.
But it's at a point where it safely stores your data correctly. Perhaps some init scripts fail on boot to import your pool/etc. but the data is there.
We do run it production, but we also have in-house tooling built around it.
i'm perfectly willing to believe there may be some rare situations where zfs on linux will cause you a problem. but i bet they're rare enough it'll have saved you a few times before it bites you.