The only case where I suspect btrfs might significantly outperform ZFS involves getdents() when the caches is cold. This is because ZoL does not presently do readahead on directory lookups. This has a noticeable impact on the performance of `ls` on the first access of a directory on mechanical storage, but it is not a problem for subsequent reads or on solid state storage. Prefetch support in getdents() will likely be implemented in either the next release or the one that follows.
This is especially important because on use cases where performance differences are significant, that is, not on general desktop usage, the maturity of the FS is funamental, and btrfs is discouraged right now.
https://github.com/ClusterHQ/flocker/blob/zfs-on-coreos-tuto...
CoreOS uses btrfs as its rootfs. I imagine that the CoreOS developers managed to avoid issues like ENOSPC by virtue of not writing to their rootfs very much. I did not have that luxury since I compiled a Gentoo GNU userland on top of it during the course of development. I encountered numerous ENOSPC errors on btrfs when developing the ZFS port to CoreOS and even hit ENOSPC errors when trying to correct the btrfs ENOSPC errors with `btrfs balance /`. I would not consider btrfs ready for production use, but your mileage will vary.
$ df -h .
Filesystem Size Used Avail Use% Mounted on
/dev/dm-1 42G 41G 280M 100% /home
$ btrfs fi df .
Data, single: total=40.48GiB, used=40.21GiB
System, single: total=4.00MiB, used=12.00KiB
Metadata, single: total=1.01GiB, used=649.18MiB
$ uname -r
3.15.2-hardened
As you can see, I've been working with a pretty much full btrfs volume. It used to be TERRIBLE to deal with btrfs in such a situation, but I haven't had an ENOSPC issue in ages.https://www.gentoo.org/proj/en/gentoo-alt/prefix/bootstrap.x...
CoreOS uses Linux 3.15.y.
2: http://www.phoronix.com/scan.php?page=article&item=linux_313...
Edit: Nevermind about the advertising, it's still intrusive, I'm just running adblock now.
As for Phoronix's ZFS benchmarks, the test hardware used drives that misreported themselves as having 512-byte sectors, which handicapped ZFS performance. Phoronix rejected all suggestions that it correct for this as end-users had been doing. Phoronix refused to meet half way by posting two results (one with proper configuration and one without), and also refused the suggestion that it to mention the existence of that problem in its test hardware. I eventually wrote code to identify drives known to misreport their sector sizes so that ZFS will automatically use the correct settings on them. That lead to the Phoronix August 2013 benchmarks showing a remarkable improvement in ZFS performance in FIO. It was so great that it sparked a discussion among the btrfs developers:
http://comments.gmane.org/gmane.comp.file-systems.btrfs/2754...
Later that month, I publicly criticized Phoronix for posting misleading benchmarks:
http://phoronix.com/forums/showthread.php?83731-ZFSOnLinux-0...
Phoronix has not posted ZFS benchmarks since that time.
https://blogs.oracle.com/brendan/entry/a_quarter_million_nfs...
More recent derivatives of OpenSolaris were reported to do 1.6 million IOPS last year:
http://www.high-availability.com/high-availability-zfs-partn...
Your mileage will vary, but ZFS has always had strong performance.
Here's a benchmark from 2013 with most results showing that it has worse performance on Linux than XFS and EXT4:
http://www.phoronix.com/scan.php?page=article&item=zfs_linux...
As for the benchmarks you cite, there are several key problems:
1. You say that they show ZFS having worse performance than XFS and ext4, but XFS is not in those benchmarks and the FIO tester shows ZFS as outperforming its competition by a significant margin.
2. They use a single disk. No server does this and while desktops and laptops do this, it is not clear how the benchmarks are of any relevance there. Additionally, LZ4 compression is not in use, when practically everyone deploying ZFSOnLinux would configure it to use LZ4.
3. ext4 manages to perform better than the theoretical limit of the SATA II interface and there is no discussion as to why.
4. ZFSOnLinux 0.6.2 is an old release. The most recent 0.6.3 release includes a new IO elevator and other improvements that enable ZFSOnLinux 0.6.3 to outperform its precedessor by a significant margin in many workloads.
Would you post something constructive that you actually did yourself? I am beginning to think that you have never even used ZFS.
As with most things, this isn't a binary decision, but typically ZFS is a good (if not the best) solution for most storage arrays.
As for XFS, it is a single block device filesystem that relies on external shims to scale to multiple disks. ZFS can outscale it when various shims are put into place to allow multiple disks to be used. One user on freenode had difficulty getting good sequential performance from XFS + LVM + MD RAID 5 on 4 disks. He reported that he could not get better than 44MB/sec writes while ZFS managed 210MB/sec. I had a similar problem in 2011 with ext4 + LVM + MD RAID 6 on 6 disks. In that case, I could only manage 20MB/sec. It is why I am a contributor to the ZFSOnLinux project today. To make these anecdotes constructive, it would be nice if we had documentation on how to configure XFS + LVM + MD RAID 5/6 in a way that sequential performance does not suffer. In my case, my performance issue involved KVM/Xen guests. I never confirmed whether that user ran his tests on bare metal, but I suspect that he did.
XFS' inability to scale past a single block device is not its only issue. Until a disk format change occurs that will add checksums, it has none to protect itself against corruption. When it gains them, it will not be able to do anything about that corruption when it detects it (aside from keeping the kernel from panicing) and it does nothing to protect your data.
That said, ZFS has always focused on obtaining strong performance. This is why it has innovations like ZIL, ARC and L2ARC. The only instances where ZFS purposefully sacrifices performance is when getting a few extra percentage points means jeopardizing data integrity. There would have been no sense in developing ZFS as a replacement for existing filesystems if it did not keep data safe.
Usually I hear people complaining about slow file creates/deletes/renames with ZFS or slowness navigating or dealing with large directories. A quick Google confirms that there are lots of people that have seen these behaviours. But I don't use it myself so I don't have any specific data. I was just trying to say that ZFS has a reputation as being a better filesystem, but not a faster one, at least with the people I talk to about storage.
ARC is a technology from IBM. Intent logs go way back. Sure ZFS is doing some interesting things, but it's a bit ironic to talk about innovation given the fact that it was more or less based on Netapp's WAFL filesystem, and all of the lawsuits that followed on from that.
I am beginning to think that not only do you not use ZFS, but that you have never used ZFS. I suggest that you try running your own benchmarks and workloads on it. It is a joy to use. I think you would agree if you were to use it in a blind fold comparison test.