In all honesty I can't understand why would anyone use anything other than zfs nowadays for important data.
In all honesty I can't understand why would anyone use anything other than zfs nowadays for important data.
At a personal level, I opted out of it based on simplicity and familiarity. If I have a data drive where I need to move and mount in another box, I do not want to mess with the complexities of supporting ZFS to get at that data or the risk that the only OS I have available can't read it.
Again, this is a non-commercial environment, but I consider my data important as well. I've decided to go the route of keeping multiple backups of my data spread across multiple drives. I've already experienced bitrot several times over the years, but for me, this approach is more practical than relying on ZFS.
(And somebody's probably going to correct my ignorance and point out that beyond some release 'x', the ZFS stuff is already built-in.)
For me ZFS has been rock solid, even with power cuts during running vms(linux and windows) and a scrub.
I even managed to steal away enough ram from a server so that ZFS produced a stacktrace in dmesg, no data corruption, even on running vms.
I haven't done this to recover ZFS as written by Linux or Solaris/illumos, not sure how well that works but wouldn't be surprised if it does.
Why people aren't just using ZFS (or BTRFS) ?
Well...
Maybe because XFS is well proven in certain critical environments, especially ones which might involve frequent manipulations of large quantities of small files (think mailserver or database). Indeed, some developers will tell you in the docs that they only support XFS or ext4 and that you're on your own if you choose something else.
Maybe because ZFS and BTRFS still have their bugs, small and big (hello RAID5 in BTRFS ;-) ) and can be heavily dependent on particular implementations.
Maybe because in a virtualised environment, you loose many of the benefits of ZFS or BTRFS (they like to see the raw disk, not some abstract notation).
Maybe because the combination of XFS/ext4 plus LVM is more than good enough for most people as that route still provides snapshots etc.
Maybe because BTRFS needs babysitting[1] (don't know about ZFS, I suspect it might).
Maybe because not everyone feels the need to "keep up with the Jonses".
If I had to pick one place to put precious data that I needed to access for 10-20 years, ZFS would be it. Of course, I wouldn't put it in one place, so if I picked places to put that data, ZFS would definitely be one of them.
I've been using ZFS for over a decade, both for my personal storage systems and for work (largely backup servers), and despite heavy use and abuse, I've never once had data loss while using it. I've had some white-knuckle times, but in the end ZFS was able to recover all the data.
I say "abuse" because I was able to write a stress test program based on my usage that uncovered several ZFS+FUSE issues early in that port.
Just a data point.
The one time I lost data from a ZFS system (which was a faulty controller silently writing garbage to multiple disks in the array), ZFS told me exactly which few files were uncorrectable, and I was able to restore them from backup.
Not one single second of doubt as to which storage system I want my data on.
A) ZFS-on-Linux was always behind BSD ZFS until recently.
B) The license issue, so you have to compile it as a separate module using DKMS, which means having a compiler, the linux headers, etc etc. No where near as low friction as just choosing ext4 or btrfs.
Sure, but the sort of people who need ZFS are the same sort of people who are doing their storage on a SAN, and so the sort of people for whom choice of OS (for the SAN itself) can be contingent on choice of storage layer, rather than the other way around.
This is half the reason ZFS-on-Linux didn't have much momentum: most ZFS deployments were happy to just run BSD on their SAN, then serve SAN volumes out over the network (or into a hypervisor cluster) for Linux systems to consume.
FreeBSD was never really working on ZFS in a way ZOL was. Whole reason for the switch is because main contributor upstream used by FreeBSD stopped using ZFS.
ZOL had more features for a while now as well.
With btrfs you're still waiting for stuff to be properly implemented.
Redhat gave up on waiting, and is removing btrfs as a supported option in RHEL. Mostly, I assume, because the btrfs team has yet to ship something anyone in enterprise would be comfortable betting their customers' data on.
It's just... Snapshot thing and replication (zfs send) are so damn easy in zfs...
BTRFS is a science experiment by comparison. They tried to copy zfs features because Oracle won't release it under GPL and did a fairly poor job of it. The fact we're this many years later and their RAID5 code still has total data loss bugs pretty much sums up btrfs.
I managed to extract most of my data to another disk on another system, but it was pretty shocking. I haven't seen unfixable filesystem corruption like that since ext2 in 2001. I'd completely forgotten it was a thing, but there I was in 2019 trying to use a second linux box to salvage my data.
Of course, whenever I mention this the btrfs zealots always come out of the woodwork saying "you must not know what you're doing if you broke btrfs" or "you obviously didn't really try if you couldn't get it fixed" or a half-dozen other excuses, but in the end I've just completely given up on btrfs. It's just not worth the hassle.
> Maybe because in a virtualised environment, you loose many of the benefits of ZFS or BTRFS (they like to see the raw disk, not some abstract notation).
ZVOL exist...You can snapshot, close, send/recv it just like dataset. You can fine-tune it like any other dataset. ZVOL works just fine for VMs.
> Maybe because ZFS and BTRFS still have their bugs, small and big (hello RAID5 in BTRFS ;-) ) and can be heavily dependent on particular implementations.
Everything has bugs. Again, you put BTRFS here just to drive your point like if it's try BTRFS then it's true for ZFS?
> Maybe because BTRFS needs babysitting[1] (don't know about ZFS, I suspect it might).
Same as above.
> Maybe because the combination of XFS/ext4 plus LVM is more than good enough for most people as that route still provides snapshots etc.
Very different kind of snapshots...
I've been using ZFS on my desktop and laptops for over a decade now. You gotta make really kick-ass FS for me to switch off ZFS even for desktop.
> Maybe because XFS is well proven in certain critical environments, especially ones which might involve frequent manipulations of large quantities of small files (think mailserver or database). Indeed, some developers will tell you in the docs that they only support XFS or ext4 and that you're on your own if you choose something else.
XFS is great, except it doesn't do any of the things that ZFS does well. For example: snapshots (to the fantastic extent that ZFS does), multi-volume support, shared storage pooling, etc. The only thing XFS offers that ZFS doesn't is support for reflink copies on Linux (i.e. "cp -R --reflink=always src dst" for near-instantaneous copies of large directories with no extra disk usage).
> Maybe because ZFS and BTRFS still have their bugs, small and big (hello RAID5 in BTRFS ;-) ) and can be heavily dependent on particular implementations.
ZFS has been, in my experience, amazingly bug-free over the last however long I've been using it (5 years in production, more personally). ZFS has working and reliable RAID support; only BTRFS is lacking it.
BTRFS, on a stock Ubuntu install with default settings, corrupted after a power outage, and could only recover most of my data, but not all. I haven't had filesystem-related data loss since ext2 in 2001.
Also, for everyone except people running Solaris, there is (now) effectively one implementation, and it works fine.
> Maybe because in a virtualised environment, you loose many of the benefits of ZFS or BTRFS (they like to see the raw disk, not some abstract notation).
You still get point-in-time snapshots on a copy-on-write filesystem, which can support different block sizes per subvolume, on-the-fly compression, and incremental or full exports, all from a shared, extendable storage pool. Those alone are huge enough to justify it for me.
ZFS wants to see the underlying disks so that it can make sane judgements about block allocation, alignment, etc. On a VM, that shouldn't matter, because your host filesystem/SAN/etc. should be doing that anyway.
> Maybe because the combination of XFS/ext4 plus LVM is more than good enough for most people as that route still provides snapshots etc.
LVM snapshots are awful from a performance standpoint.
In ZFS, when you create a snapshot, any new data is written to new blocks (read the old block, make the change, write it to the new block), so your snapshot points to the old blocks on disk.
From what I can tell[1], LVM snapshots mean that when you write data to a block, LVM reads that block, writes it somewhere else, and then updates the original block in-place. This makes every write a synchronous write, because you have to write the old data to its new block, sync to make sure it actually gets stored, and then do your new write.
BTRFS snapshots cannot be recursive (i.e. you cannot snapshot /data/ and /data/gitlab and /data/svn and /data/backups atomically). They argue that this is a feature, but I would consider it a massive bug. LVM snapshots by their nature are not recursive because LVM has no concept that I can find of nesting.
> Maybe because BTRFS needs babysitting[1] (don't know about ZFS, I suspect it might).
ZFS, as far as I can tell, does not need babysitting. I used it for storing backups at my last job; we had a QNAP NAS exporting over iSCSI, ZFS was using the iSCSI storage, and we wrote data to it. Nothing ever went wrong, and I literally never checked on it unless the power went out or something of the sort. It was three years before I ever got around to setting up monitoring for it, because literally nothing ever went wrong so I completely forgot that there was anything to monitor. Eventually I did set it up, but it never went off. Literally set it and forget it.
In summary: ZFS is fantastic, and basically never needs to be taken care of. It does its thing and it works great, and that's it. BTRFS has corrupted itself arbitrarily after a power outage (which I thought we'd solved with ReiserFS back in 2001), LVM snapshots are slow and awful, and the tooling for LVM and BTRFS is a usability disaster while also losing a lot of features that ZFS has (like on-the-fly compression, block deduplication, etc.).
In datacenters many probably do, however most home/soho NAS machines are severely limited wrt memory and CPU power, which could be a limit. My self assembled NAS uses a Atom board to keep power requirements low since it stays always on, and I can't complain about its performance, however that CPU doesn't support more than 4GB RAM which is considered the minimum to properly use ZFS (NAS4free). Many cheap commercial NAS boxes used in small offices are even more limited. I wonder if there are any technical reasons preventing ZFS to operate (or be adapted to) in low memory environments.
This is mostly a myth with origins in the high-memory requirements of ZFS deduplication (which few should use, anyway). Of course, more memory allows for more caching, but that’s true of any filesystem on a “modern” OS.
I might upgrade it to something with a better processor because I do want to run de-duplication on my 14TB drive at some point.
My work backup server has "only" 32GB of RAM which is nowhere near enough to do dedup, even though there's a ton of stuff that could be deduplicated. Looks like it needs at least 64GB RAM to dedup, and that machine is maxed out.
But at home I'm in a similar situation to the parent. I'd like to set up a NAS, but I use it mostly for backups long term storage and don't want a big machine sitting idle all the time. A tiny Atom with Optane for cache and DDT might be ideal, in my thinking.
But for most of the storage use Backblaze is probably the right answer.
You could definitely put the L2ARC on Optane, but Optane is probably better suited to the ZIL (ZFS Intent Log, basically the journal so writes can be quickly acked) if you aren't using literal NVRAM.
Breaking the DDT out so it could exist on a Optane isn't currently available (again, AFAIK, but I recall hearing about some changes in that department, I don't recall specifics and may be misremembering).
I've basically never had a machine big enough to comfortably use dedup except on trivial loads. My backup boxes, where I'd really like to use it, have all fallen over when enabling dedup because of the RAM requirements.
https://forums.servethehome.com/index.php?threads/zfs-alloca...
Generally speaking, I will go with lower specs machines until I run into some actual problem. Haven't so far.
If you're talking about GFS, that's been gone for a decade, and was not a filesystem as far as the kernel was concerned.
But maybe there is also an automatic cleanup tool that deletes old snapshots after some time?
The snapshot refers to the storage blocks/records on disk, not files as such, an important distinction since ZFS can expose block storage (zvol) as well as a "regular" file system.
Since ZFS is copy-on-write, the only storage you pay for with a snapshot are the blocks that have changed since the snapshot was taken (plus a little overhead). Thus for data that does not change much, a snapshot is almost "free".
The blocks are reference counted. Once a snapshot is deleted, ZFS decreases the reference count of the blocks referenced by the snapshot. Any block with a refcount of zero is considered free and thus that space is reclaimed. This happens when the block has changed since the snapshot was taken and there were no other snapshots referencing that block.
ZFS itself has no automatic deletion of old snapshots AFAIK, but there are tools built around ZFS that allow for periodic snapshotting and cleanup.
Turns out a DAG, as is used in both ZFS and git, is a good data structure in many use cases.
Furthermore, you only "pay" their storage cost for data which actually changes. On a 2TB volume, last snapshotted 1 week ago, if only 5GB have changed since then, that's the only storage overhead (ignoring for simplicity the folder structure itself).
Also, if a single data block changes 20 times since its last snapshot, you still only "pay" storage costs for its current version + the one sitting in the last snapshot.
We need to encourage more of this :)