ZFS v0.8.2
github.com
github.com
It's incredible how far they've come. We're using ZFS on Linux on about 120 servers at work and it's rock solid. Snapshots are a life saver in our day-to-day ops.
There's one command for handling storage pools (zpool, https://manpages.debian.org/unstable/zfsutils-linux/zpool.8....), e.g. adding/replacing disks, monitoring I/O utilization, etc.
And then there's another command for dealing with ZFS datasets (zfs, https://manpages.debian.org/unstable/zfsutils-linux/zfs.8.en...), e.g. setting properties for datasets (quotas, compression, delegation of privileges to non-root users), managing snapshots.
Both CLI commands are scriptable (e.g. -H can be used to suppress human-readable headers and -p turns off "friendly" formatting for numbers) and there are libraries (libzfs, libzpool) which can be used to access their functionality (e.g. for managing snapshots) in your own programs.
I don't think I can do ZFS on Linux (or btrfs for that matter) justice in a (short) comment on HN. There's just so much to it (e.g. SSDs as cache/log devices, sending/receiving snapshots over the network, compression, etc.)
There are some rough edges, of course. Until a few years ago ZFS on Linux had massive problems with resource management, i.e. it would regularly panic in low-memory situations. These appear to be fixed though - at least I haven't seen any panics in the past two or so years.
Due to its license ZFS on Linux will probably never be part of the upstream kernel. We use it on Debian which means we get to compile the ZFS kernel modules on each box when there's a kernel update (with DKMS). Plus there are some pitfalls when using ZFS as your root filesystem (on Debian stretch without systemd services would start before /var/log was mounted).
We've had plenty of disk failures which ZFS detected well before SMART did because we run monthly scrubs for our ZFS pools (i.e. basically ZFS verifies its checksums for all the data that's stored in a pool). Recovery is as easy as popping in a new disk and running "zpool replace".
Ok, this ended up much longer than I had planned. This has to suffice for now.
btrfs uses a more traditional command, subcommand, args setup, which by itself is not a problem, but the actual command and subcommand structure is a dumpster fire to be charitable.
Try managing large groups of snapshots without custom tooling. The whole thing makes me sad really.
btrfs has a lot going for it at the architecture and disk format level (dedup is actually something theoretically useful, while the zfs design was flawed from the start), but the implementation has just never made it.
So the only place you can do dedup is inline, as the data is being written the first time, not after-the-fact.
In addition, this requires you keep a huge indirection table (the DDT, or dedup table) that needs to be read for all writes, so either that has to be kept in memory or on fast storage, or you've just turned every write into one or more random reads, plus writes.
This also means that even if you turn off dedup after turning it on, the performance implications remain until the DDT no longer contains any blocks (e.g. you rewrote all the data after turning dedup off).
There are feature proposals to make the performance of dedup less pathological, but nobody's taken up implementing them so far. (Someone even did a proof of concept implementation of one of them, and it still hasn't been finished and integrated.)
With the exception of accidentally enabling deduplication on a 16 TB array on a system with 8 GB of RAM, ZFS has been fantastic. Everything works, and works much as you'd expect. The tools and concepts are clear and concise and it's obvious that a lot of work was put into everything from the design to the documentation.
btrfs, on the other hand, has been a nightmare. I set it up on a new desktop machine earlier this week, and I've spent most of this afternoon trying to recover data from it. Due to a power outage, the filesystem is corrupted. I didn't know that right away because I enabled zstd compression in the filesystem but GRUB couldn't boot/read zstd files/filesystems, etc.
A lot of btrfs is counterintuitive, or downright worrying. Most of the threads I've seen regarding filesystem errors end with "I reinstalled and everything is fine". The documentation suggests not trying to `btrfs check --repair`, because apparently that's the wrong way to do things? You should try to mount with the 'recovery' mount flag instead, which is not intuitive.
In my case, the kernel is throwing an error about the checksum map, but it can't rebuild it or repair the filesystem. In a rescue image it errors because it tries to call pthread_cancel but can't load libgcc for whatever reason, so I can't rescue my system from a rescue system.
Even when it did work it was confusing. Unlike ZFS, btrfs subvolumes seem... counterintuitive? On ZFS I can take a recursive snapshot of a volume or subvolume and any subvolumes it has, which allows me to divide a subtree up into multiple subvolumes but still treat it the same (e.g. if I want multiple entries in /var/lib/mysql/<mysql_instance> or something similar). On btrfs, you cannot do recursive snapshots, and I've even seen some people describe this as one of the (few) reasons you'd use subvolumes: to exclude something from snapshots.
After ten years, it feels as though btrfs is 90% done; that is to say, it has 90% of the functionality it should, that those features are 90% done, and that what is done works about 90% of the time.
Honestly, with the state that btrfs is in, I don't understand why it's in the kernel at all, and why distros support it (but I do understand why RHEL pulled it). Working with it directly makes me worry for the data I have stored on my synology, and now I'm wondering if I should have just built a FreeNAS box instead.
TL;DR btrfs is a giant mess and you shouldn't use it. ZFS has more up-front overhead (getting the package versions configured for Ubuntu) and is more difficult to go all-in on (e.g. ZFS root) but at least you won't lose your data out of nowhere.
https://rkeene.org/projects/info/wiki/BtrFS
Overall similar. I ran into more critical bugs with BtrFS, but ZFS (on Solaris) was not perfect (most of our Solaris kernel panics were ZFS-related on Solaris).
I found the same as the comment above, btrfs was a dream to set up and a nightmare when something went wrong, which was several times a year it seemed.
ZFS on the other hand, I had to write scripts to handle my snapshots and snapshot expirations but the filesystem itself has been rock solid.
I haven't put ZFS through it's paces. However with btrfs I had so many issues on a dead simple workload it just didn't seem stable to me. This sucks and I wish I could contribute because I loved their goals and thought they were designing everything right, but stability wins the long game, as it were.
I first came in contact with it through solaris, and looking back at that stack with zfs, zones & software defined networking it just seems like such a contemporary fit. Recently I’ve looked a bit at smartos and joyents offering that kind of takes it the way I thought it would take, back in the day, pre-oracle.
Been using ZoL for years now, both for private filers but also professionally as of late for things like physical docker/container hosts.
Just rock solid!
Just a shame that Solaris wasn't Open Sourced sooner, as if it'd grown more traction we'd be in a different world now.
I think Linux kind of creeped up on them, and in the meantime Oracle got its grubby its on Sun, then embraced Linux itself.
Zones are /amazing/.
DTrace for example, that stuff is proper. Great to see whats happening with BPF on Linux these days.
I guess it speaks to the engineering efforts behind it all. Not an easy feat to integrate or bolt on. And not to mention the inertia of eco-systems of tooling on a well established platform such as Linux.
Interesting times ahead and great work, my favorite community! :)
FreeBSD forked ZFS on Linux GitHub repo and now its called ZFS on FreeBSD. You can try it already by installing these two ports:
- sysutils/zol
- sysutils/openzfs-kmod
Regards, vermaden
Correct me if I am wrong, So we now have
Oracle ZFS, FreeBSD ZFS, ZFS on Linux and ZFS on FreeBSD?
This is correct - and good news - but there should be a reality check involved with regard to timelines ...
ZoL code is most likely not going to be part of any 12-RELEASE distributions, which means that:
IF you only run -RELEASE in production (a good policy) AND if you do not put a x.0 release into production (historically, a good policy) then you are 2-3 years away from running this ZoL code:
http://freebsd.1045724.x6.nabble.com/How-many-quot-productio...
As it stands now, our customers want encryption and raw send which we provide to them using ZoL, on 12, in ports form, in a VM.
This is how people zfs-send encrypted snapshots to rsync.net.
Our base ZFS platform, however, will most likely not run a ZFS version that doesn't come with a -RELEASE version of FreeBSD with a x.0 version number. Not a big deal when everyone is using borg and restic anyway, but I digress...
The license for this is CDDL? Wasn't that the license for the original ZFS, which was the whole reason people had to port it? (apart from kernel sys calls, etc)
FreeBSD moved to this new codebase because it is getting a lot of attention and better support.
Maybe v0.8.2 is close to the point where 0.8 is stable enough to upgrade?
0.8 also has trim support :-)
The incident where updating your Ubuntu kernel resulted in data loss probably could not have occurred under the BSD model due to close coupling between the kernel and the filesystem.
https://news.ycombinator.com/item?id=16797919
The whole downstream/upstream model with different parts of the OS kernel / filesystem / userland moving out of sync with each other through multiple levels of backporting is not a positive one for reliability. FreeBSD is FreeBSD and the buck stops there. Responsibility is too diffuse in the Linux model to get the kind of reliability that FreeBSD has enjoyed.
I realize that I'm being a bit of a fuddy-duddy but it really didn't take too long at all for the ZoL team to start making dumb mistakes that cause data loss. Hopefully it's a learning experience and didn't happen again.
I myself have a dataset that I can't send from a FreeBSD system to my Ubuntu 16.04/ZoL backup server. It worked at one point, as of about a year ago it no longer does. When the send is finished, the client machine just spins forever. I've tried everything short of formatting and reinstalling the backup server. I tried incremental sends, I tried re-sending the whole thing, I tried killing the pool, updating everything and re-creating the pool, I tried messing with microcode in case it was a wayward regression from Spectre, etc etc. ZoL just won't receive that dataset anymore.
Best of luck to the ZoL guys but I guess FreeBSD is working well enough as a data storage layer for me. Too much weird shit going on with ZoL.
Have you tried reporting the bug you're having with send/recv? I imagine people would care about that kind of reproducible failure.
Re: FBSD versus ZoL, go with what works for you. If something is working well enough, I'm not going to advocate changing it. (This tends to be an unpopular opinion among other people though.)
BTRFS is developed since 2007 ...
... and Red Hat decides to reimplement this kind of pooled storage filesystem again with XFS on LVM now called Stratis ... not very bright.
I suppose the benefit of using XFS and LVM is that they're both very well understood and widely used technologies. Sometimes incremental evolution is all you need.
I'm following the development with interest, because I think the approach has merits. Dave Chinner's talk about teaching XFS how to do snapshots and subvolumes was particularly interesting. The gist of it, as far as I understand, is that XFS can pretty much do those things already, but it needs better integration for performance and usability.
Obviously it is going somewhere, upstream is very active, and SUSE uses it by default everywhere both enterprise and openSUSE. And they have the developers to support it. Facebook uses Btrfs quite a lot in production both servers and desktops.
Given how many kernel developers and filesystem developers they have, it would not be hard for them to hire btrfs developers or ask some of their kernel developers with a background in file systems to specialize in btrfs. I think the reasons are different: either it's because they have a lot of XFS developers on staff, fighting a transition to btrfs nail and tooth; or they simply believe that btrfs cannot be made good enough to support their enterprise developers within a reasonable time frame.
I have a data store (nightly build archive) that gets almost 3x reduction in space used with dedupe, but the performance hit is just too big for us to use it.