OpenZFS – add disks to existing RAIDZ
github.com
github.com
How cool would it be if we had a great GUI for ZFS (snapshots, volume management, etc.). I could buy a new external disk, add it to a pool, have seamless storage expansion.
It would be great. Ah, what could have been.
As a workaround, you can create a sparsevolume to store your parallels volume. Sparsevolumes are stores in bands, and only bands that change get backed up. It might be slightly more efficient.
How cool would it be if we had a great TUI for ZFS...
Live in the now: https://github.com/kimono-koans/httm
As a Rust project, it's something of a duty to hype and self-promote, and to troll C++, of course. So let it be known: If you're not using httm, you're not using modern ZFS. Or -- httm is the cppfront of ZFS utilities? Or -- httm is the 'auto' keyword for filesystems, but less confusing (to me)?
I'll work on it.
See QNAP HERO 5:
https://www.qnap.com/static/landing/2021/quts-hero-5.0/en/in...
NAS Options:
https://www.qnap.com/en-us/product/?conditions=4-3
// This was about MacOS, and yes, that would be cool. But QNAP is remarkably MacOS-friendly, including Thunderbolt and Time Machine support.
And HERO OS is only available on the Enterprise range of NAS, not even Prosumer class NAS.
I have TVS-h1288X and I am really happy with it, maybe it's a bit overkill for most.
That said, I could not use zfs send from my prev fbsd server, there are some compat checks in QNAP's version. Used nfs to migrate from my decade++ old pool.
Using Tailscale for access, they have QNAP builds on GH.
So it's not just a check for check's sake, I would suspect in practice their send/recv is just completely incompatible at this point.
If you're happily using their devices, this may not matter to you, but since the post is about OpenZFS...
https://macoverdrive.blogspot.com/2008/10/using-zfs-to-manag...
I would certainly not put it past Jobs. While he had some genius qualities, his personality was pretty deeply flawed. But on the other hand I wonder if his sense for business would not have prevailed.
If only Apple had absorbed sun instead of oracle, their legacy would have been better off. Of course it couldn't have been worse.
Which also made me wonder, of course Jobs and Ellison were big friends so perhaps Jobs just left it for him to take? Oracle always wanted to get their hands on java obviously.
Ps the whole ZFS on Mac adventure came a little after time machine was already in production.
ZFS is indeed an amazing desktop filesystem. I'm using it now on FreeBSD, I moved away from Mac because it became ever more closed down. I wanted more control, not less.
But when all this happened they still did, yes.
Apple would have killed Java in half a second. They don't need it, it was built to do the opposite of what they want their tech to do. Same for Solaris, burdensome duplicate. As for all the desktop-oriented FOSS projects (OpenOffice, VirtualBox etc), they would not have just been killed - Apple would have viciously bullied anyone to preclude them from forking them. The stuff they'd keep (MySql, probably), they would have made osx-only, and then strangled them quietly after they got out of the server game. They would have ripped anything technologically interesting from the hardware divisions and then shut them down - because they were no match, in terms of supply efficiency and value, for the Apple equivalents; and anyway they were never seriously interested in the server space.
Oracle did what Oracle does, and realistically the culture clash was never going to result in a smooth transition; but from a commercial perspective they kept around the best of what Sun was making (Java, servers) and largely let the community get on with the stuff they weren't interested in (if rebranded/forked). The only real crime they committed, IMHO, was trying too hard to make money from VirtualBox, effectively spooking the market; there is a different timeline where VBox ends up being what Docker became. But everything else was par for the course.
You still can't do that with ZFS though, that's what this PR is for.
So if over the years I got, say, five disks of varying quality - I shouldn't add them to a pool as five individual vdevs :)
Yes, that's why you don't have single-disk vdevs.
ZFS's separate cache meant that the better choice was APFS. Once they got around to writing it.
And on drives at least, TM backups ARE APFS snapshots.
I believe that was something built into the file manager in OpenSolaris, and then into illumos OS's. OpenIndiana has it in their file manager within their themed MATE desktop environment.
Not in workstation OS, maybe not "great" but tools like FreeNAS had ZFS GUI sorted out[1] pretty decently and user-friendly for some time. I've seen storage laymen set these up, learn on the go and not screw up.
[1] https://windows-cdn.softpedia.com/screenshots/FreeNAS_1.jpg
[1] https://www.ixsystems.com/documentation/freenas/9.3/freenas_...
https://github.com/openzfs/zfs/pull/12225#issuecomment-16101...
What worries me the most is them doubling down on it after it was shown to be a mess and totally insecure. I imagine to protect the corporate image. But such corporate interests don't belong in FreeBSD. Just admit you hired the wrong guy (or the right guy at the wrong time) and take the hit. Because it wasn't even the company's fault really.
The lack of serious review on the freebsd side was also a big one but I find the 'strings attached' the biggest issue myself because it made mitigation of serious issues such a problem.
The large amount of corporate kernel work in Linux is why I went for FreeBSD and to see whatever little go so wrong was worrying.
One can trace its legacy to BSD while it was still at Berkley and has worked with many of the most respected names in the BSD space.
The other bought a domain name of a fork and used it to post disparaging messages and Hitler "Downfall" memes slandering them. Source: https://www.wipo.int/amc/en/domains/search/text.jsp?case=D20...
While I have some disagreements with iXsystems' pivot to ZFS-on-Linux offerings and the like, they're not the clowns that Netgate / pfSense are. 10cm to my right are TrueNAS and opnSense boxes - you won't catch me dead using pfSense in my network after the crap they've pulled.
Does anyone know why this is the case? When expanding an array which is getting full this will result in a far smaller capacity gain than desired.
Let's assume we are using 5x 10TB disks which are 90% full. Before the process, each disk will contain 5.4TB of data, 3.6TB of parity, and 1TB of free space. After the process and converting it to 6x 10TB, each disk will contain 4.5TB of data, 3TB of parity, and 2.5TB of free space. We can fill this free space with 1.66TB of data and 0.83TB of parity per disk - after which our entire array will contain 36.96TB of data.
If we made a new 6-wide Z2 array, it would be able to contain 40TB of data - so adding a disk this way made us lose over 3TB in capacity! Considering the process is already reading and rewriting basically the entire array, why not recalculate the parity as well?
IANA expert but my guess is -- because, here, you don't have to modify block pointers, etc.
ZFS RAIDZ is not like traditional RAID, as it's not just a sequence of arbitrary bits, data plus parity. RAIDZ stripe width is variable/dynamic, written in blocks (imagine a 128K block, compressed to ~88K), and there is no way to quickly tell where the parity data is within a written block, where the end of any written block is, etc.
If you had to instead, modify the block pointers, I'd assume you have to also change each block in the live tree and all dependent (including snapshot) blocks at the same time? That sounds extraordinarily complicated (and this is the data integrity FS!), and much slower, than just blasting through the data, in order.
To do what you want, you can do what one could always do -- zfs send/recv between a filesystem between and old and new filesystem.
Sure, but that involves having enough spare disks, enough places to put them, and enough places to connect them.
This way, while the initial expansion is not ideal, it works. If you really need the space gains from a wider distribution, you can do this expansion and then do the make a new copy, then replace old copy with a new copy dance... although that's counter productive if you have snapshots.
So long as we are pointing out things: this presumes that the user doesn't have enough space to make the copies on their own disks? If one has a 3GB dataset and 9GB of free space, one can easily zfs send/recv that dataset and destroy the old copy.
> This way, while the initial expansion is not ideal, it works.
Yep.
> If you really need the space gains from a wider distribution, you can do this expansion and then do the make a new copy, then replace old copy with a new copy dance... although that's counter productive if you have snapshots.
That's why you do a zfs send/recv, instead of a copy. ZFS will copy your snapshots for you!
Because snapshots might refer to the old blocks. Sure you could recompute, but then any snapshots would mean those old blocks would have to stay around so now you've taken up ~twice the space.
AFAIK in ZFS snapshots are just a pointer to (an old) merkle tree[1], if you go about changing the blocks then you need to update that tree, but you can't due to copy-on-write (without paying for the copy, like I mentioned).
[1]: https://openzfs.readthedocs.io/en/latest/introduction.html#d...
You can do this yourself though when convenient to get those lost TB back.
The RAID VDEV expansion feature actually was quite stale and wasn’t being worked on afaik until this sponsorship.
The stuff about rewriting data I just wrote to clarify
The reason the parity ratio stays the same, is that all of the references to the data are by DVA (Data Virtual Address, effectively the LBA within the RAID-Z vdev).
So the data will occupy the same amount of space and parity as it did before.
All stripes in RAID-Z are dynamic, so if your stripe is 5 wide and your array is 6 wide, the 2nd stripe will start on the last disk and wrap around.
So if your 5x10 TB disks are 90% full, after the expansion they will contain the same 5.4 TB of data and 3.6 TB of parity, and the pool will now be 10 TB bigger.
New writes, will be 4+2 instead, but the old data won't change (they is how this feature is able to work without needing block-pointer rewrite).
See this presentation: https://www.youtube.com/watch?v=yF2KgQGmUic
So you lose data capacity compared to "dumb" RAID6 on mdadm.
If you expand RAID6 from 4+2 to 5+2, you go from using 33.3% data for parity to 28.5% on parity
If you expand RAIDZ from 4+2 to 5+2, your new data will use 28.5%, but your old (which is majority, because if it wasn't you wouldn't be expanding) would still use 33.3% on parity.
Edit: I suppose I could cover this with a shell script needing only the spare space of the largest file. Nice!
On btrfs that's a rebalance, and part of how one expands an array (btrfs add + btrs balance)
(Not sure if ZFS has a similar operation, but from my understanding resilvering would not be it)
Not that it matters much though as RAID5 and RAID6 aren't dependable upon, and the array failure modes are weird in practice, so in context of expanding storage it really only matters for RAID0 and RAID10.
https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
https://www.unixsheikh.com/articles/battle-testing-zfs-btrfs...
Regardless, my entire point is that you still lose a significant amount of capacity due to the old data remaining as 3+2 rather than being rewritten to 4+2, which heavily disincentives the expansion of arrays reaching capacity - but that is the only time people would want to expand their array.
It just seems to me like they are spending a lot of effort on a feature which you frankly should not ever want to use.
I don't use raidz for my personal pools because it has the wrong set of tradeoffs for my usage, but if I did, I'd absolutely use this.
Yes, your data has the old data:parity ratio for older data, but you now have more total storage available, which is the entire goal. Sure, it'd be more space-efficient to go rewrite your data, piecemeal or entirely, afterward, but you now have more storage to work with, rather than having to remake the pool or replace every disk in the vdev with a larger one.
ZFS really deeply assumes you're not gonna be rewriting history for a bunch of features, and you'd have written a good chunk of a new filesystem to reimplement everything without those assumptions.
I cannot stress how expensive and invasive that would be enough.
Not hard, but it does require sufficient free space. Once it's done you can destroy the original dataset and reclaim the space.
A feature I am excited to see is being added!
Coincidentally one of the drives of my 1st mirror died few days ago after rebooting the host machine for updates and I replaced it today, it’s been resilvering for a while.
So for example, a pool with 2 RAID-Z vdevs each containing 5 disks effectively has only 2x the IOPS of a single disk, while a pool with 5 mirror vdevs (i.e. RAID-10) has 10x the IOPS of a single disk.
It's not a big problem if you mostly do sequential I/O but it's a huge difference if/when you do small random reads (e.g. traversing a non-cached directory tree).
In the context of resilvers, currently RAID-Z pools can only be done by traversing the block tree which, due to fragmentation of CoW filesystems, usually leads to a lot of random reads/writes, while a RAID-10 pool can basically resilver a pool while doing almost fully sequential I/O which can be much, much faster.
https://jrs-s.net/2015/02/06/zfs-you-should-use-mirror-vdevs...
You still need 6/Z2 if you want to have reasonable fault tolerance, unless you want to waste ungodly amount of space.
We did had a big case so we just ran 2xRAID6 setup and that one time where it was needed we just replaced to bigger drives one by one, while using the removed ones as spares for other machines. But that's benefit of scale bigger than "a NAS server".
I got 2 so I’ll also replace the sibling of the dead 4TB one.
Listed survival probability of N-disk failure in 8-drive/4-vdev mirror.
1: 1; 2: 0.857; 3: 0.667; 4: 0.400; 5: 0; 6: N/A; 7: N/A; 8: N/A
Proper survival probability:
1: 1; 2: 0.857; 3: 0.571; 4: 0.229; 5: 0; 6: 0; 7: 0; 8: 0
Comparatively, survival probability for 8-drive RAIDz4 with equal amount of usable space:
1: 1; 2: 1; 3: 1; 4: 1; 5: 0; 6: 0; 7: 0; 8: 0
Personally, I'd use multi-vdev mirror pools only for data that is either backed-up or data I can afford to lose completely.
To replace drive in RAID1 you need to read the entirety of one drive.
Instead of reading one whole drive worth of stress, you're reading N drives worth of stress
But it is funny that the people making arguments about "stressing drives" don't even fucking know how RAID works...
Drive stress (for me) is mostly a concern about data loss.
... except now you have more drives that can fail.
With RAID6 yes, you can still fail one more time and still keep your data, which is why it is recommended.
Data safety wise I'd go RAID6 -> RAID1/10 (particularly linux implementation can do raid10 on odd number of drives which is nice) -> RAID5
Imagine a gambler telling you "you can either throw one k-faced die and you lose if it comes up on 1, or throw n k-faced dice and you lose if two come up on 1". Depending on k and n, the second really can be the better choice
Often people cite the faster rebuild of mirrors as a safety advantage, but the same amount of rebuild IO will occur regardless of how long it takes. Yes, there will be more non-rebuild IO in a larger time window, but unless that routine load was causing disks to fail weekly then I doubt it will change the numbers non-negligibly. It will of course affect array performance though, so mirrors for performance is a good argument
Realistically though, RAIDz recovery is longer and more stressful, so more of your drives can fail in the critical period, and, assuming you have backups, your storage is there for for usability - mirroring gives you a performant usable system during a fast recovery for the price of a small chance of complete data loss (but you have backups?) vs RAIDz that gives you long recovery pains on a degraded system, but I expect a smaller chance of data loss on a lightly loaded system.
Granted, that was a bit of uncommon case where:
* someone forgot to order spares after taking last one
* we still had our consumable buying pipeline going thru helpdesk
* helpdesk didn't had any importance communicated about it and because of some accounting bullshit the purchase got delayed long enough
* the drives in question were all from some segate's fuckup of a model with much higher failure rates.
One disk failed with some media errors, remaining 2 got kicked out of array for same reason during resilvering
We ddrescue'd the 2 on the pair of fresh ones and the bad blocks didn't land in the same place on both drives so it made full recovery. But we did learn many lessons from that..
"Production" systems should not even be considered production unless you have a backup of them,
Some of us think people should be hosting stuff from home, accessible from their mobile devices. But the first and to me one of the biggest hurdles is managing storage. And that requires a storage appliance that is simpler than using a laptop, not requiring the skills of an IT professional.
Drobo tried to make a storage appliance, but once you got to the fine print it had the same set of problems that ZFS still does.
All professional storage solutions are built on an assumption of symmetry of hardware. I have n identical (except not the same batch?) drives which I will smear files out across.
Consumers will never have drive symmetry. That’s a huge expenditure that few can justify, or much afford. My Synology didn’t like most of my old drives so by the time I had a working array I’d spent practically a laptop on it. For a weirdly shaped computer I couldn’t actually use directly. I’m a developer, I can afford it. None of my friends can. Mom definitely can’t.
A consumer solution needs to assume drive asymmetry. That day it is first plugged in, it will contain a couple new drives, and every hard drive the consumer can scrounge up from junk drawers - save two: their current backup drive and an extra copy. Once the array (with one open slot) is built and verified, then one of the backups can go into the array for additional space and speed.
From then on, the owner will likely buy one or two new drives every year, at whatever price point they’re willing to pay, and swap out the smallest or slowest drive in the array. Meaning the array will always contain 2-3 different generation of hard drives. Never the same speed and never the same capacity. And they expect that if a rebuild fails, some of their data will still be retrievable. Without a professional data recovery company.
which rules out all RAID levels except 0, which is nuts. An algorithm that can handle this scenario is consistent hashing. Weighted consistent hashing can handle disparate resources, by assigning more buckets to faster or larger machines. And it can grow and shrink (in a drive array, the two are sequential or simultaneous).
Small and old businesses begin to resemble consumer purchasing patterns. They can’t afford a shiny new array all at once. It’s scrounging and piecemeal. So this isn’t strictly about chasing consumers.
I thought ZFS was on a similar path, but the delays in sprouting these features make me wonder.
I would love to have more options for expandable redundancy.
The only thing that can work for your mom and your friend is, in my opinion, a pair of disks in mirror. When the space finished, buy another box with two other disk in mirror. Anything more than this is not only too complex for the average user but also too expensive.
Both schemes are vulnerable to the (admittedly rarer) errors where both drives fail simultaneously (e.g. mobo fried them) or are just ... destroyed by a fire or whatever.
A periodic sync (while harder to set up) will occasionally save you from the deleting the wrong files which mirroring doesn't.
Either way: Any truly important data (family photos/videos, etc.) needs to be saved periodically to remote storage. There's no getting around that if you really care about the data.
Which ever solution for disks you end up with, you should definitely always be using something that keeps around periodic snapshots.
If they want network attached storage, I’d just use a single disk NAS, and remotely back it up.
I started using snapraid [1] several years ago, after finding zfs couldn't expand. Often when I went to add space the "sweet spot" disk size (best $/TB) was 2-3x the size of the previous biggest disk I ran. This was very economical compared to replacing the whole array every couple years.
It works by having "data" and "parity" drives. Data drives are totally normal filesystems, and joined with unionfs. In fact you can mount them independently and access whatever files are on it. Parity drives are just a big file that snapraid updates nightly.
The big downside is it's not realtime redundant: you can lose a day's worth of data from a (data) drive failure. For my use case this is acceptable.
A huge upside is rebuilds are fairly painless. Rebuilding a parity drive has zero downtime, just degraded performance. Rebuilding a data drive leaves it offline, but the rest work fine (I think the individual files are actually accessible as they're restored though). In the worst case you can mount each data drive independently on any system and recover its contents.
I've been running the "same" array for a decade, but at this point every disk has been swapped out at least once (for a larger one), and it's been in at least two different host systems.
I didn't want to point applications at 100 T of "free" space only for attires to start blocking after 8.
Am I mistaken about that?
I use "existing path, least free space". Once a path is created, it keeps using it for new files in that path. If it runs out of space, it creates that same path on another drive. If the path exists on both drives for some reason, my rationale is this keeps most of the related files (same path) together on the same drive.
I see there's some newer "most shared path" options I don't remember that might even make more sense for me, so maybe that's something I'll change next time I need to touch it.
[1] https://github.com/trapexit/mergerfs
[2] https://github.com/trapexit/mergerfs#policy-descriptions
That's not how it works.
The policy picks what branch to use and then once selected mergerfs will clone the relative path as needed. With "ep" policies it will never select a branch that doesn't have the full relative path. "msp" will always rerun the check one level up in the hierarchy if nothing is found at the current level.
Will MergerFS mount points behave the same as on the host inside of a docker container if passed as a bind mount?
* Rebuilds are semi-offline, as you said. Almost every other solution, even Unraid, will immediately emulate the data from a failed drive. On Snapraid you have to wait for each file to be restored, and depending on your union setup this may mean you have directories with half the files missing
* Whenever you modify or delete a file between syncs, your parity is now out of sync. This means some other files may fail to restore if you have a failure now. However using more than single parity will make this far less likely to happen. Another way to fix this completely the "snapraid-btrfs" tool, which runs Snapraid on snapshots of independent btrfs disks (somewhat like Synology?), meaning the old data is still available
* It saves almost no file metadata, not even owner and mode. Restored files just use the umask of the user running the restore command. A minor one, but surprisingly annoying
However one big advantage over Unraid it has is how transactional and rigorous it is. A power failure can cause Unraid parity to desync, with no clear way to know which disk is right. Snapraid OTOH is designed to survive interruptions gracefully, and checksums all files so should never accidentally restore corrupt data. And can detect silent drive failure ("bitrot") as a bonus
And for a typical home NAS storing movies and family photos (mostly append-only), those downsides are probably no big deal anyway
RAID 1, you mean? Because that way you still have a complete copy of your data if one drive fails.
BTRFS is excellent for the use case of a wide variety of mismatched drives, because it supports adding and removing drives and rebalancing the array. But for the moment only the RAID 1 modes are really trustworthy. I have a NAS consisting of drives whose advertised capacities in GB are: 1000, 1000, 1920, 2000, 2048, 2050, 3840, 3840, 4000. But it started as a pile of infamously unreliable 3TB hard drives.
Then glue it together with LVM and hope for best.
The fact that total data loss happened regularly on full volumes many years after a “stable” release means BTRFS is either fundamentally flawed in design or incompetently implemented. I have no idea how it was accepted into mainline Linux.
The fact that a certain large NAS vendor used BTRFS by default (without documentation at the time of purchase) has cost my org and others lots of downtime and money.
I didn't say it was. There are plenty of disk full scenarios that can be handled by BTRFS without trouble, especially when adding more disks (even temporarily) is an option. There are ways to get stuck with a full fs, but it's not a guarantee you'll end up in a "wipe and restore from backup" situation. (And to be fair, ZFS also gets very problematic when completely full; this is a general problem for CoW filesystems.)
And with plain RAID6 I can still add hard drives if needed. Yeah the case of buying bigger drives is still a bit iffy but btrfs RAID1 gonna lose me more space than RAID6+LVM+xfs anyway...
Were you assuming that BTRFS RAID 1 meant every block of data gets mirrored across every device? I've seen that assumption before, from people who don't realize there are two ways to generalize RAID1 to more than two devices.
However you get no striping, and data is only read from one drive, so performance is limited to that of one drive for reads and writes. Plus with mismatched drives, smaller drives go unused unless you write enough data.
You're going to have to support that claim a bit better. The core idea of RAID 1 is mirroring data, which BTRFS RAID 1 mode definitely does. Striping is not an essential part of RAID 1 (hence RAID 10), and reading data from two disks in parallel is an optional performance optimization that is not performed by all RAID 1 implementations (but could be implemented for BTRFS RAID 1: https://stackoverflow.com/questions/55408256/btrfs-raid-1-wh... ).
> Plus with mismatched drives, smaller drives go unused unless you write enough data.
Yes, the allocation is suboptimal from a performance perspective, as I've already said. But it is simple and straightforward and reasonably good at avoiding putting you into a situation where manually issuing a rebalance command is necessary. If you do need better performance, there's a RAID 10 mode. But since my NAS is currently a motley pile of SSDs, I don't need to to anything extra to have decent performance.
It'd be way easier to talk about if it had a unique name, and you could say "It's like RAID1".
Despite all that I do like the mode, and use it in a few places.
You have raid1c<drive count> mode in BTRFS which does this.
Ceph actually works quite well for that, althought obviously far more complex than anything for home use. You just tell it "x chunks with N parity" and it will spread it over available drives. Just need more drives than x + n and not too egregious size differences.
> Weighted consistent hashing can handle disparate resources, by assigning more buckets to faster or larger machines. And it can grow and shrink (in a drive array, the two are sequential or simultaneous).
It would need to be more complex than that. Putting chunks on say 2, 6, 8, 12, 22 TB (which might be what you'd get if you just buy "cheapest GB/$" for your NAS over last 10 years!) is more complex than that.
Do any of you people actually do UX or DX for a living? This is starting to worry me.
You will take your Btrfs and you will like it.
It’s promising, but there’s… bugs. Right now only performance oriented ones, that I’ve noticed, but I’d wait a bit longer.
By all accounts mainline is at best not interested, if not actively against ZFS on Linux. The last few kerfuffles around symbols used by the out-of-tree module laid out the position rather unambiguously.
Which one?
Source, for someone who isn't following kernel mailing lists?
https://lore.kernel.org/all/20190110182413.GA6932@kroah.com/
Similarly, Linux probably has too many contributors, some of which are no longer living, to come to an agreement on a fixed license either.
Although, license changes do happen; OpenSSL did one recently, but I think there are fewer contributors.
VLC did, too, although with probably even less contributors.
Doing a clean room rewrite would be a huge job.
ZFS is already a very awkward citizen on Linux (e.g. can't use the page cache so it has to duplicate all that in the ARC; won't be able to take advantage of multi-page folios because of the lowest-common-denominator SPL; ...) and it will get worse, not better.
It would be nice to have a precedent deciding on this bullshit argument once and for all so distros can freely ship prebuilt binary modules and bring Linux to the modern times when it comes to filesystems.
The whole situation is ridiculous. I'd understand if this was about money, but who exactly gets hurt by users getting a prebuilt module from somewhere, vs building exactly the same thing locally from freely-available source?
Linus Torvalds doesn't feel like being the guinea pig by risking ZFS in the mainline kernel. A totally reasonable position while the CDDL+GPL resolution is still ultimately unknown. (And honestly, with OpenZFS supporting a very wide range of Linux versions and FreeBSD at the same time, I have the feeling that mainline inclusion in Linux might not be the best outcome anyway.)
Here I do not see this argument applying since the source is freely available to use and extend; the license explicitly allows someone to compile it and use it. In this case providing prebuilt binaries is more akin to providing a "cache" for something you can (and are allowed to) build locally (using ZFS-DKMS for example) using source you are once again allowed to acquire and use.
What prejudice does it cause to Oracle that the “make” command is ran on Ubuntu’s build servers as opposed to users’ individual machines? Have similar cases been litigated before where the argument was about who runs the make command, with source that either party has otherwise a right to download & use?
I don’t think there’s any issue with canonical shipping a kmod, but similar to 3ᴿᴰ party binary drivers, it would need to be treated as a different “work”.
it's definitely a GPL issue more than a CDDL issue
ZFS has generally been easier to go from zero to working well once you accept that life doesn’t need to be so complicated as legacy md, lvm, and fs layers made it. It also has a pretty big footgun that causes me to carefully read the man page before using “zpool add” or “zpool attach”. Vdev removal is a big help for this.
I can just slap in any drives of any size and it'll just adapt. As long as the new drives go in an empty slot or are bigger than the previous one, I get more space. Zero CLI commands needed.
Now get this feature in unraid or TrueNAS and I'm switching.