LLNL openzfs project: https://computing.llnl.gov/projects/openzfs Old presentation from intel with info on what was one of our bigger deployments in 2016 (~50pb): https://www.intel.com/content/dam/www/public/us/en/documents...
I believe it was called ZFS On Linux or something like that.
Nice how things have evolved: from FreeBSD to linux and back. In my mind this has always been a very inspiring example of a public institution working for the public good.
ZoL, if my ancient memory serves, was at LLNL, not based on the FreeBSD port (if you go _very_ far back in the commit history you can see Brian rebasing against OpenSolaris revisions), but like 2 or 3 different orgs originally announced Linux ports at the same time and then all pooled together, since originally only one of the three was going to have a POSIX layer (the other two didn't need a working POSIX filesystem layer). (I'm not actually sure how much came of this collaboration, I just remember being very amused when within the span of a week or two, three different orgs announced ports, looked at each other, and went "...wait.")
Then for a while people developed on either the FreeBSD port, the illumos fork called OpenZFS, or the Linux port, but because (among other reasons) a bunch of development kept happening on the Linux port, it became the defacto upstream and got renamed "OpenZFS", and then FreeBSD more or less got a fresh port from the OpenZFS codebase that is now what it's based on.
The macOS port got a fresh sync against that codebase recently and is slowly trying to merge in, and then from there, ???
Note that we do not use ZFS for content, since it is incompatible with efficient use of sendfile (both because there is no async handler for ZFS, so no async sendfile, and because the ARC is not integrated with the page cache, so content would require an extra memory copy to be served).
One thing that makes the interfaces in Linux much messier than the FreeBSD ones is that a lot of the core functionality you might like to leverage in Linux (workqueues, basically anything more complicated than just calling kmalloc, and of course any SIMD save/restore state, to name three examples I've stumbled over recently) are marked EXPORT_SYMBOL_GPL or just entirely not exported in newer releases, so you get to reimplement the wheel for those, whereas on FreeBSD it's trivial to just use their implementations of such things and shim them to the Solaris-ish interfaces the non-platform-specific code expects.
So that makes the Linux-specific code a lot heavier, because upstream is actively hostile.
I wish someone would come in and convince the kernel devs that "hey, if you want EXPORT_SYMBOL_GPL to have legal weight in a copyleft sense then you can't just slap it onto interfaces for political reasons"
IMO Linus should stop being half-and-half about it and either mark everything SYMBOL_GPL and see how well that goes or stop this nonsense.
Linus himself also made remarks about ZFS at one point that were pretty...hostile. [1] [2]
> The fact is, the whole point of the GPL is that you're being "paid" in terms of tit-for-tat: we give source code to you for free, but we want source code improvements back. If you don't do that but instead say "I think this is _legal_, but I'm not going to help you" you certainly don't get any help from us.
> So things that are outside the kernel tree simply do not matter to us. They get absolutely zero attention. We simply don't care. It's that simple.
> And things that don't do that "give back" have no business talking about us being assholes when we don't care about them.
> See?
Note that there's at least one unfixed Linux kernel bug that was found by OpenZFS users, reproducible without using OpenZFS in any way, reported with a patch, and ignored. [3]
So "not giving back" is a dubious claim.
[1] - https://arstechnica.com/gadgets/2020/01/linus-torvalds-zfs-s...
[2] - https://www.realworldtech.com/forum/?threadid=189711&curpost...
(/s)
Also, ZFS goes with the same licensing interpretation as Linus used with AFS.
Which three?
As far as know the point of EXPORT_SYMBOL_GPL was to push back on companies like Nvidia who wanted to exploit loopholes in the GPL. That seems to me like a reasonable objective.
Relevant Torvalds quote: https://yarchive.net/comp/linux/export_symbol_gpl.html
But if you're marking interfaces as GPL-only, or implementing taint detection that means if you use a non-SYMBOL_GPL kernel symbol which calls a GPL-only function it treats the non-SYMBOL_GPL symbol as GPL-only and blocks your linking, it gets a bit out of hand.
Building the kernel with certain kernel options makes modules like OpenZFS or OpenAFS not link because of that taint propagation - because things like the lockdep checker turn uninfringing calls into infringing ones.
Or a little while ago, there was a change which broke building on PPC because a change made a non-SYMBOL_GPL call on POWER into a SYMBOL_GPL one indirectly, and when the original author was contacted, he sent a patch reverting the changed symbol, and GregKH refused to pull it into stable, suggesting distros could carry it if they wanted to. (Of course, he had happily merged a change into -stable earlier that just implemented more aggressive GPL tainting and thereby broke things like the aforementioned...)
All ZFS needs to do is just have one of Oracles many lawyers say "CDDL is compatible with GPL". Yet, they Oracle don't.
It's explicitly not compatible with GPL, though. It has clauses that are more restrictive than GPL, and IIRC some people who contributed to the OpenZFS project did so explicitly without allowing later CDDL license revisions, which removes Oracle's ability to say CDDL-2 or whatever is GPL-compatible.
So even if someone rolled up dumptrucks of cash and convinced Oracle that everything was great, they don't have all the control needed to do that.
But "call an opaque function that saves SIMD state" is obviously not derivative of the kernel code in any way. The more exports that get badly marked this way, the more EXPORT_SYMBOL_GPL becomes indistinguishable from EXPORT_SYMBOL.
The main "legal effect" I see is that you are not willing to take that risk, just like Oracle isn't.
If enough symbols get restricted without valid copyright-based reasons, modules that legitimately have non-GPL licenses will have to lie to the kernel to get loaded. And that's a stupid situation all around.
The ZoL project lead said at one point there were a variety of reasons this wasn't initially done for the Linux integration [1], but that it was worth taking another look at since that was a decade ago now. Having looked at the Linux memory subsystems recently for various reasons, I would suspect the limiting factor is that almost all the Linux memory management functions that involve details beyond "give me X pages" are SYMBOL_GPL, so I suspect we couldn't access whatever functionality would be needed to do this.
I could be wrong, though, as I wasn't looking at the code for that specific purpose, so I might have missed functionality that would provide this.
[1] - https://github.com/openzfs/zfs/issues/10255#issuecomment-620...
But sure, it's certainly an older issue, and given that the ABD rework happened, I wouldn't put anything past being "feasible" if the benefits were great enough.
(Look at the O_DIRECT zvol rework stuff that's pending (I believe not merged) for how a more cut-through memory model could be done, though that has all the tradeoffs you might expect of skipping the abstractions ZFS uses to minimize the ability of applications to poke holes in the abstraction model and violate consistency, I believe...)
[0] https://www.kernel.org/doc/Documentation/filesystems/dax.txt
Back to the topic at hand, it’s actually scary how few software expose control over whether or not sendfile is used, assuming support is only a matter of OS and kernel version but not taking into account filesystem limitations. I ran into a terrible Samba on FreeBSD bug (shares remotely disconnected and connections reset with moderate levels of concurrent ro access from even a single client) that I ultimately tracked down to sendfile being enabled in the (default?) config - so it wasn’t just the expected “performance requirements not being met” with sendfile on ZFS but even other reliability issues (almost certainly exposing a different underlying bug, tbh). Imagine if Samba didn’t have a tubeable to set/override sendfile support, though.
No gotchas / issues, works well, easy to setup.
I am looking forward to the Direct IO speed improvements for NVMe drives with https://github.com/openzfs/zfs/pull/10018
edit: one thing I forgot to mention is, when creating your pool make sure to import your drives by ID (zpool import -d /dev/disk/by-id/ <poolname>) instead of name in case name assignments change somehow [1]
[1] https://superuser.com/questions/1732532/zfs-disk-drive-lette...
See also "Scaling ZFS for NVMe" by Allan Jude at EuroBSDcon 2022:
I’ve been out of this particular game for a long time.
We typically use Proxmox. It’s a convenient node host setup and usually has a very up to date zfs and it’s stable
I just wouldn’t use the Proxmox web ui for zfs configuration. It doesn’t have up to date options. Always configure zfs on the cli
What options are missing for you?
Would be great if you could open an enhancement request over at https://bugzilla.proxmox.com/ for tracking this (no promises on (immediate) implementation though).
The fact how I can replicate live data of arbitrary size (many TB sized filesystems) in small chunks to other hosts every minute greatly increased my deep sleep quality. Of course databases with multi-host write etc are nice, but in our use case, all customers are rather small with just lots and lots of files (medical and otherwise), the database itself is rather small and doesn't need replication.
Best thing, on the receiver side of the backup ZFS ensures due to its architecture that the diff is directly applied on top of the existing filesystem, while in normal differential backups one might find out months or years later that one diff snapshot was damaged in transfer or is not accessible.
zfs scrubbing with S.M.A.R.T monitoring also helps a lot to ensure drive quality over time.
# Gotchas
ZFS:
- There is no undo etc, this is unix, so beware of wrong commands. - ZFS Scrubbing can be stopped (it does sometimes affect io speed), but ZFS resilvering cannot. This can lead to performance issues. - There must be enough RAM for the caching to work well and synchronous workloads do well with good write cache drives (ZIL) - Data usage patterns should fit well with the Append Log schema of ZFS. E.g databases such as LevelDB worked really well. Others are not slow, but need a good ZIL more then when the pattern fits.
SmartOS: Some minor gotchas with how vmadm deletes zfs filesystems or in general with SmartOS, e.g when having too many snapshots, but everything quite predictable.
A slightly annoying property of snapshots and clones is the inability to fully re-root a tree of snapshots, e.g. permanently split a clone from its original source and allow first-class send/receive from that clone. The snapshot which originated the clone needs to stick around forever[2]. This prevents a typical virtual machine imagine process of keeping a base image up to date over time that VMs can be cloned from when desired and eventually removing the storage used by the original base image after e.g. several OS upgrades.
I don't have any big performance requirements and most file storage is throughput based on spinning disks which can easily saturate the gigabit network.
I also use ZFS on my laptop's SSD under Ubuntu with about 1GB/s performance and no shortage of IOPS and the ability to send snapshots off to the backup system which is pretty nice. Ubuntu is going backwards on support for ZFS and native encryption uses a hacky intermediate key under LUKS, but it works.
[0] https://github.com/openzfs/zfs/issues/12014 [1] https://github.com/openzfs/zfs/issues/12594 [2]https://serverfault.com/questions/265779/split-a-zfs-clone
Some discussion: https://www.reddit.com/r/NixOS/comments/ops0n0/big_shoutout_...
They have both FreeBSD and Linux based stuff, targeting different use cases.
It can be made somewhat[0] stable with using no (or a simple compression like ZLE), avoiding log devices and caches it can work, but it's way simpler and guaranteed stable to just use a separate partition.
[0] even setting the myriad of options there still exist reports about hangs, like https://github.com/openzfs/zfs/issues/7734#issuecomment-4167...
On a VM host you have a bit of leeway for memory overcommit, but ideally if you have a box that frequently has to rely on swap, something is bad wrong. I don't think it's really fair to criticize a worst-choice configuration for being poorly supported here. There are simply too many opportunities for catch-22/deadlock scenarios in the interaction of the fs and allocator during OOM...
We've successfully survived numerous disk failures (a broken batch of HDD's giving all kinds of small read errors, an SSD that completely failed and disappeared, etc), and were in most cases able to replace them without a second of downtime (would have been all cases if not for disks placed in hard-to-reach places, now only a few minutes downtime to physically swap the disk).
Snapshots work perfectly as well. Systems are set up to automatically make snapshots using [1], on boot, on a timer, and right before potentially dangerous operations such as package manager commands as well. I've rolled back after botched OS updates without problems; after a reboot the machine was back in it's old state. Also rolled back a live system a few times after a broken package update, restoring the filesystem state without any issues. Easily accessing old versions of a file is an added bonus which has been helpful a few times.
Send/receive is ideal for backups. We are able to send snapshots between machines, even across different OSes, without issues. We've also moved entire pools from one OS to another without problems.
Knowing we have automatic snapshots and external backups configured also allows me to be very liberal with giving root access to inexperienced people to various (non-critical) machines, knowing that if anything breaks it will always be easy to roll back, and encouraging them to learn by experimenting a bit, to the point where we can even diff between snapshots to inspect what changed and learn from that.
Biggest gotchas so far have been on my personal Arch Linux setup, where the out-of-tree nature of ZFS has caused some issues like a incompatible kernel being installed, the ZFS module failing to compile, and my workstation subsequently being unable to boot. But even that was solved by my entire system running on ZFS: a single rollback from my bootloader [2] and all was back the way it was before.
Having good tooling set up definitely helped a lot. My monkey brain has the tendency to think "surely I got it right this time, so no need to make a snapshot before trying out X!", especially when experimenting on my own workstation. Automating snapshots using a systemd timer and hooks added to my package manager saved me a number of times.
[1]: https://github.com/psy0rz/zfs_autobackup [2]: https://zfsbootmenu.org/
I do that with sqlite to keep a selection of snapshots from the last hours, days etc.
We also use ZFS for testing of our own block device. My biggest problems with it are related to this activity. Hot-plugging or removing block devices from the pool often leads to unrepeatable pool, and the tests have to be scrambled because the whole system needs to be rebooted.
I don't understand what you mean by this comment. Are you creating a zvol on your zpool and loopback mounting it with a different filesystem? It sounds like you are very much using ZFS as a filesystem, even if you are not booting from it or using it for any kind of high availability purpose.
I do agree with you on the matter that there should be a way to clear out an (UNAVAIL vdev / SUSPENDED pool) without a reboot. This has been a thorn in my side for as long as ZFS has existed. The deadlock-by-design line is really quite nonsense.
Is nice to see it advancing in useful ways. Hopefully this leads to offline dedup without the runtime memory costs.
The PersistentVolume API is a nice way to divvy up a shared resource across different teams, and using ZFS for that gives us the snapshotting, deduplication, and compression for free. For our workloads, it benchmarked faster than XFS so it was a no-brainer.