ZFS has been quite feature-full and production ready for more than a decade... why do you think it is not finished?
Yes. When it does what it advertises, properly.
In what way has your experience differed?
ext4 had at least two critical data corruption bugs in the past 5 years in stable kernels (personal anecdote: one ruined the root filesystem on a server that I used, after which I stopped using ext4).
No it hadn't. Those bugs were lower in the stack in block layer.
The one you're referring to is https://lwn.net/Articles/774440/ which is a different bug.
I should add that as a user, what matters to me is the reliability of my data. If a bug exists outside of fs/ext4 in the kernel but affects only ext4 and not other filesystems (such as https://www.phoronix.com/scan.php?page=news_item&px=MTIxNDQ which was caused by an ext4-related commit, and similar others in prior years), this makes ext4 unreliable for me.
> which was caused by an ext4-related commit
Now that's blatantly wrong. It wasn't caused by an ext4-related commit. The commit to blame was scheduler code in the block layer. Nothing filesystem specific.
The point is, ext4 isn't a monument of reliability like you delude yourself with, and it had, and will continue having data corruption bugs.
> Now that's blatantly wrong. It wasn't caused by an ext4-related commit.
Yes it was. This is very clear from the context: https://patchwork.ozlabs.org/project/linux-ext4/patch/500F1C...
You can try and weasel your way out with the semantics, say that it didn't change any files under ext4/ (which does happen with xfs/btrfs/... related patches too BTW, simply because fs/ contains common fs code), but the reality is, it was an ext4 related fix, it appeared in "linux-ext4" mailing list, and Ted Tso, the ext4 maintainer, signed off the patch.
Lastly, idolizing a piece of code is nonsensical. Yes, btrfs had its data corruption bugs, but you can't pretend that ext4 and other filesystems didn't.
I'm currently on Bcachefs, which has erasure coding with similar promises and works better for me.
ZFS is more rigid, I have to have all matched disks for best performance, plus RAM unless I like bad performance (I don't). This could be solved by striping disks into 1TB partitions, but ZFS doesn't like living on a partitioned disk nearly as much.
Mixed disk has a bunch of cool properties for homelab users. Bcachefs also solves another problem; caching and speed. I could designate my 1TB NVMe as fast and my SATA 2TB SSDs as slow. I could also define that /tmp requiers no Erasure Coding and everything else requires 2 replicas. Then I would be able to combine the performance of my NVMe with the capacity of my SATA SSDs while gaining a simple redundance solution. My NAS has a dozen of mixed disks that I could combine more easily than ever.
The issue with ZFS is not that it's not ready to use. The issue is that it's born from enterprise and requires a costly enterprise setup to run, rather than a cheap homelab setup.
Don't know how true is this, since it's not even possible to create a zpool on the whole unpartitioned device on linux. It automatically creates GPT label with zfs and a small efi partition.
Every ZFS pool I have is composed of GPT partitions, and this is also recommended in ZFS books as good practice to make it easier to identify a failed drive. Since you see the GPT partition name in the "zpool status" and other tools' output, it's handy.
You can't unfortunately tell it that different disks have different speeds.
It also works just fine w/ a small (1GB or so) ARC if you want, just at an obvious performance penalty compared to having a larger cache.
For instance, a common threshold for "ready to use" is "included in a released version of the upstream Linux kernel, and not under CONFIG_BROKEN or CONFIG_STAGING". Under that definition, neither ZFS nor bcachefs are "ready to use".
ZFS on FUSE should be deprecated by now, and ZoL I believe merged with OpenZFS for the 2.0 release.
FreeBSD ZFS is fine, and just check your versions if using it on Linux. Ie- don't use anything pre-2.0.
(Nb, avoid SMR hard drives - the rebuild took more than a week!)
It does however require a bit of care to maintain performance, and you need to know your expected workload going in. Otherwise you can find yourself in a situation with a very poorly performing pool where the only realistic route to recovery is a send/receive to a fresh pool and back.
If you do not have a lot of sync workload (VMs, DBs) and don't have super-high performance needs, say a home NAS, then mainly you just need to think about not filling up the pool too much.
A nice way to do this is to create a root dataset where you set a quota to say 75% of capacity, and then create all other datasets below this one. You should not go above 85% space usage, as ZFS switches allocation strategy then to one which can significantly increase fragmentation.
If you do have a lot of sync writes, a SLOG device is basically mandatory. The SLOG device does not have to be large, it only stores about 5-10 seconds worth of writes, so 10-20GB can be plenty. I've partitioned up my SSDs and created a mirror out of two small partitions, using the remaining SSD space for other things.
One thing to keep in mind is that while an L2ARC device sounds like a great thing, depending on your configuration you can actually slow things down with one. A bunch of disks has a lot more bandwidth than a single SATA SSD. An L2ARC device also requires some memory overhead, so reduces your primary ARC. Again depending on load this can be detrimental.
And finally, don't ever think about using deduplication, unless you've read about the consequences, measured the performance benefits and ensured the memory overhead is acceptable. It sounds great on paper but has a lot of associated downsides that can ruin pool performance, and disabling it does not make it go away.
At least that's what I've picked up so far.
I think recreating the entire thing is about my only option at the moment.
My pool, 2 vdevs each a 4-way RAID-Z1, has been used and abused for almost 7 years now. I've gone over the 85% mark but it still works fine, performance too.
Some VM stuff but mostly media and similar.
Do you mean slow IOPS or also sequential performance?
But yeah sadly that's the one area where ZFS is less stellar. Once it's fragmented it's hard to fix. Easiest is to send/receive to another pool, but as you note that is not always feasible.
Assuming your free space fragmentation is not too bad (check output of zpool list, "frag" column is free space fragmentation level), you could just move files back and forth between the two. Assuming you don't have snapshots holding on to the files, this can help reduce the fragmentation.
Otherwise you're stuck with the send/receive.
Though I'd try the mailing list[1] to see if any of the gurus can help identify what's going wrong before attempting random ailments.