Bcachefs Merged into the Linux 6.7 Kernel
phoronix.com
phoronix.com
This has been a solitary and largely self-financed effort by Kent over many years. He must feel pretty great to finally see this happen!
Edit: I currently see 279 supporters for a total of $2,328 per month, so $8.34 average per month per supporter.
The pre-populated options are $20 and $100 a month, which is a lot. I think he'd get a lot more supporters if he dropped those asks to more like $1 and $5.
Either way, $5 is probably both the median and the mode for the payment distribution, whether you include non-paying members or not.
From memory I stopped because I was adding other creators and I got the impression he was doing okay.
I've played around with it a few times, since it's been easily available in NixOS for a while. Didn't run into any issues with a few disks and a few hundred GB of data.
Some very interesting properties, including (actually efficient) snapshots, spreading data over multiple disks, using a fast SSD as a cache layer for HDDs, built-in encryption (not audited yet though!), automatic deduplication, compression, ...
A lot of that is already available through other file systems (btfs, zfs) and/or by layering different solutions (LVM, dm-crypt, ...) , but getting all of it out of the box with a single FS that's in the mainline kernel is quite appealing.
Snapshots and compression OTOH are better done on the FS level.
Simpler yes, but since generally no extra metadata is allocated it is vulnerable to some cryptanalysis notably comparing different snapshots of the drive. Doing this properly requires storing unique keys for different versions of data. Doing this with typical blockdev-level encryption is very expensive (you either need to reduce the effective block size which disrupts lots of software that assumes things about block size or store the data out-of-line (typically at the end of the disk) which requires up to 2x writes. Doing this in the filesystem allows strong encryption with minimal performance impact (as the IV write is co-located with data that is changing anyways).
and because writes to separate sectors aren't atomic, you probably want to add journaling or some kind of CoW for crash safety, and oh look now you're actually just writing a filesystem and it's not simpler anymore.
Most importantly, you're not duplicating the effort for every filesystem that you want to support encryption for, and the code can largely remain fixed once mature.
Yes there have been issues in the past but since quite a a while its stable and feature-rich. Easily the most advanced free file system
> Snapshots are writeable and may be snapshotted again, creating a tree of snapshots.
* https://bcachefs.org/bcachefs-principles-of-operation.pdf
> bcachefs provides btrfs style writeable snapshots, at subvolume granularity.
* https://bcachefs.org/Snapshots/
Every other implementation of the concept has the implicit idea that snapshots are read-only:
* https://en.wikipedia.org/wiki/Snapshot_(computer_storage)
The word "clone" seems to have been settled on for a read-write copy of things.
-s|--snapshot
Create a snapshot. Snapshots provide a "frozen image" of an ori‐
gin LV. The snapshot LV can be used, e.g. for backups, while
the origin LV continues to be used. This option can create a
[…]
* Also: https://manpages.ubuntu.com/manpages/lunar/en/man8/lvcreate....The word 'frozen' to me means unmoving / fixed.
As mentioned in the Wikipedia article, the analogy comes from photography where a picture / snap(shot) is a moment frozen in time.
I've admined NetApps in the past, used Veritas VxFS back in the day, and currently run a lot of ZFS (first using it on Solaris 10), and "snapshot" has meant read-only for the past few decades whenever I've run across it.
So I think LVM2 snapshots are indeed read/write. Perhaps that manpage sentence was not updated since LVM1 read-only snapshots ?
(I agree with you that 'snapshot' to me strongly suggests read/write; I'm just saying that you can't actually rely on that assumption because it's not just bcachefs that doesn't use that meaning.)
> This option can create a COW (copy on write) snapshot, or a thin snapshot (in a thin pool.) [...] COW snapshots are created when a size is specified. The size is allocated from space in the VG, and is the amount of space that can be used for saving COW blocks as writes occur to the origin or snapshot.
Likely snapshots _were_ originally read-only, and the description of creating thin and COW snapshots was added later, but the man page text was not re-written completely; rather the description of thin and COW snapshots were added to the end of the existing text.
Bcachefs is safer to use than btrfs and is also shown to outperform zfs in terms of speed and reliability
So what makes it more reliable? I can't find a simple overview of the design / reasoning behind the whole thing and what makes it 'better' than the rest.
It's still a very exciting filesystem so I'm sure we'll be seeing third parties test it rigorously very soon.
Citation needed. With a sample size of 1 it ate my data and BTRFS has been running perfectly fine on that system (after I bailed off of bcachefs) and other systems. I think it is great that they consider data safety very important but it will take lots of testing and real-world experience to validate that claim.
That is a pretty bold claim given that Facebook runs btrfs in prod across the majority of their fleet and almost nobody uses bcachefs.
Btrfs was terrible in the early days, took ages to "git gud", and given a filesystem is supposed to be among the most stable code in the OS, that burned a lot of bridges. It wasn't until fairly recently that btrfs could tolerate being completely filled.
I have no idea how valid the claims are, but bcache's developer claims that btrfs suffers from a lot of terrible early design decisions that can't be undone.
The show-stopper for me is that bcachefs lacks a scrub:
> We're still missing scrub support. Scrub's job will be to walk all data in the filesystem and verify checksums, recovery bad data from a good copy if it exists or notifying the user if data is unrecoverable.
The only argument I see for btrfs is that it supports throwing random drives into a pool and btrfs magically handles redundancy across them.
...but it doesn't support tiered storage like bcachefs does. We're well past "I want to have a redundant filesystem I can randomly add a drive to and magically my shit is mirrored." These days people want to have a large pile of spinning rust with some SSD in front of it, and doing that in all but ZFS is kind of a pain.
find . -xdev -print0|xargs -0 cat > /dev/null
(ideally replacing cat with something that continues after errors)The claim for reliability comes from the idea that bcache has been in heavy production use for a decade, and considered rock solid with plenty of testing of corner cases, and that bcachefs builds a filesystem over the bcache block store so that most of the hard things (like locking and such) are managed by the underlying block store, not the bcachefs layer. This way, the filesystem code itself is simple and easy to understand.
Do you guys know why someone should get excited by bcachefs?
I use btrfs in preference over ext4 for Linux filesystems and turn on zstd compression for performance and a bit of space saving. It seems simple enough for my use case, though I'm not doing any snapshots etc.
What are some of the potential pitfalls?
- RAID5/6 are still not stable
- it will not mount a RAID in degraded mode automatically, failing the high availability promise that might tempt you towards RAID.
- swapfile support exists, but it breaks snapshots (and I don't want to snapshot the swapfile)
- Just an Ubuntu/Debian thing, but snapshots are not integrated into the update process unless you install `apt-btrfs-snapshot` (and know that package exists)
Thanks - I did not know about that.
I agree with the stance of not mounting a degraded RAID automatically as then the danger is that someone might not notice it and be subjected to total data loss later on. The best option would be to allow over-riding that choice if the RAID is otherwise monitored.
Why does everyone insist on RAID5? I am completely fine with btrfs RAID1, survived few disk crashes as advertised.
- it will not mount a RAID in degraded mode automatically
it will if using -o degraded, it's the tooling which generally sucks and won't support this. I even had problems booting at all with / on multi-device btrfs using recommended tools (dracut, grub-mkconfig), the bugs are known and unfixed, ended up rolling out own initramfs.
- swapfile support exists, but it breaks snapshots
you are supposed to have all your stuff in a subvolume, and snapshot that, not whole toplevel root filesystem...and yes it should be documented better that it interferes with snapshots
- but snapshots are not integrated..
yep tooling sucks
But generally, RAID1 gets you 50% of storage space, whereas RAID5 gets you 66% (and any odd combination of disks).
But… you have to weigh it against losing all your data. With raid5 if any 2 of your disks die at the same time then you've lost your volume and most likely 100% of your data (no matter how many disks you have).
With raid1 you have to lose _all_ your disks before that happens. Typically that's also just 2 but you can mirror 3 or more drives if you need some data to be _really_ resistant to disk failures and you don't care about "wasting" n-1 times the space.
So you end up trading off efficiency of storage space and resiliency to data loss.
Myself, disks got cheap enough that I always just buy 2 disks and mirror them. I find it easier to reason about overall, especially in the face of a degraded array.
I remember when btrfs was very young and they announced the ability to create mirrors. I tested this out and was pleased, and then I tested the failure scenario: pull a drive, try to boot.
It wouldn't boot!
I jump into IRC and ask if it's expected that you can't boot from a degraded mirror and the answer was "not supported" which means the mirror is pointless.
Obviously it has improved since then as there's a way to force it to work, but I returned to ZFS on FreeBSD and never looked back
With btrfs, you can easily add a new drive so that new writes can be accepted and mirrored immediately, then start rebuilding the old data with lost redundancy (possibly after running another incremental backup, which will be less stressful to the drive than the full rebuild). But the "add a new drive" step happens outside the kernel, in userspace and possibly in meatspace, so it can't be a default action for the kernel to take.
Also, if you're happy running off a single disk in a faulted mirror, then I'd question why you've got the mirror setup at all.
RAID-5/6:
* https://btrfs.readthedocs.io/en/latest/btrfs-man5.html#raid5...
I remember when Btrfs was announced in 2007 as I was already running Solaris 10 with ZFS in production (and ZFS had non-"experimental" RAID-5-like RAID-Z from day one). Here we are 15+ years later and Btrfs still doesn't have it?
There are abuse patterns that are toxic for ZFS pools (and all other filesystems). Btrfs appears to be able to repair this damage.
https://www.usenix.org/system/files/login/articles/login_sum...
It should instead be giving the user error messages written to their terminal, in logs, etc instead of breaking the entire system until the user finds the manual
If the system doesn't have spare capacity ready, the only sane response is to not boot/mount normally. "giving the user error messages written to their terminal, in logs, etc" isn't a real solution for something like a NAS with no terminal connected and nobody looking at the logs as long as they can still establish a SMB connection; it's too likely to be a silent failure in practice. Mounting the filesystem degraded but read-only makes sense if it's necessary to boot the system so that the user (or their pre-configured userspace tooling) can decide how to deal with the problem, but a lot of Linux distros aren't happy with the root filesystem being read-only.
In summary: there's no single right answer to the problem of a failed drive, and btrfs defaults to what is the safest behavior based on the information available to the filesystem itself. Userspace tooling with more information can make other, less universal choices. A distro that tries to simply adopt btrfs as a drop-in replacement for ext4 probably doesn't have all the tooling necessary to make good use of the unique features of btrfs.
It doesn't need the spare to "boot normally" and the system can turn on a scary LED, ring bells, call you, text you, hit you up on WhatsApp, DM you on Instagram, or whatever method you want your NAS to use to notify you there's a degradation. (You're monitoring it right??)
This explanation of "it's dangerous to boot off a degraded array" is lunacy. I will not take this terrible advice from armchair experts when I've been doing this for over 25 years
There's nothing terrible about advice against responding to a drive failure by putting the system into an even more precarious state without user interaction.
Having things fail to boot would just mean you haven't configured your system appropriately for your environment. If you are using a btrfs RAID filesystem for your root filesystem, and you need that fs to be writeable in order to boot, and you want it to boot even if it's missing a drive, then you need to add an extra mount option and a few lines to your init scripts to persist new downgraded RAID settings in the event a degraded mount was necessary.
But that's hardly the only valid use case for btrfs; plenty of users want strong guarantees about the redundancy of their data rather than silent downgrading.
Also, do you really expect me to believe that any of the large shops still running enough spinning rust to have daily drive failures are still booting off those arrays instead of having separate SSDs as their boot drives? Separate storage of the OS from storage of the important data is such a common and long-ingrained practice that it is embodied in the physical layout of typical server systems, and the primary reason for it is the need for different tradeoffs between performance, redundancy, capacity and cost.
What seems particularly interesting about Bcachefs is how much it seems to be using database concepts to implement a general filesystem. Ultimately, it seems inevitable that filesystems and databases will converge as they're both supposed to manage data.
[1] by heavily, I mean that I use it for my home directory
• On Btrfs the `df` command lies. You can't get an accurate count of free space.
• There is no working `fsck` and the existing repair tools come with dire warnings. Take these very very seriously. I have tested them. They do not work and will destroy data.
• The main point of Btrfs is snapshots. [open]SUSE, Spiral Linux, Garuda Linux and siduction all use these heavily for transactional updates.
But the snapshot tool cannot test that there's enough free space for the snapshot, because `df` lies. So, it will fill up your disk.
Writing to a full Btrfs volume will corrupt it. In my testing it destroyed my root partition roughly once per year. It was the most unstable fs I have tried since the era of ext2 in the mid-1990s. (Yes I am that old.)
Thankfully, all of those incidents were with some non-critical, throw-away VMs where the data loss wasn't really an issue.
I've also used ext4 under the same circumstances for years, and I can't think of a single time that I've lost data, nor have I experienced corruption that fsck couldn't easily deal with.
I, too, would have to go back to the 1990s to think of a filesystem I used that was that unreliable.
After what I experienced, I don't trust Btrfs at all, and I have no plans to ever use it again.
I worked at SUSE for 4Y and used it every day. The company is in deep denial about its problems, or that there are any problems, and when I pointed at ZFS as a more mature tool, this was actually mocked.
df is not lying because the fs layer reports corrupt data but because of dedup and fs layering.
> df is not lying
To me, that reads as "df isn't lying because $EXCUSES."
I disagree. I don't care about excuses. I want a 100% accurate accounting of free space at all times via the standard xNix free-disk-space reporting command, and the same from the APIs that command uses so that applications can also get an accurate report of free space.
If a filesystem cannot report free space reliably and accurately, then that filesystem is IMHO broken. Excuses do not exonerate the FS, and having other FS-specific commands that can report free space do not exonerate it. The `df` command must work, or the FS is broken.
The primary point of Btrfs is that it is the only GPL snapshot-capable FS. The other stuff is gravy: it's a bonus. There are distros that use Btrfs that don't use snapshots, such as Fedora.
Some Btrfs advocates use this to claim that the problems are not problematic. If the filesystem is of interest on the basis of feature $FOO, then "product $BAR does not exhibit this problem" is not an endorsement or a refutation if $BAR does not use feature $FOO.
Btrfs RAID is broken in important ways, but that is not a deal-breaker because there are other perfectly good ways of obtaining that functionality using other parts of the Linux stack. If no feature or functionality is lost considering the OS and stack as a whole, then that isn't a problem. However, this remains serious and an issue.
Additional problems include:
• Poor integration into the overall industry-wide OS stack.
Examples:
- Existing commands do not work or give inconsistent results.
- Duplication of functionality (e.g. overlap with `mdraid`)
• Poor integration into specific vendors' OS stacks.
Examples:
- SUSE uses Btrfs heavily.
But SUSE's `zypper` package manager is not integrated with its `snapper` tool. Zypper doesn't include snapshot space used by Snapper in its space estimation.
Snapper is integrated with Btrfs; licence restrictions notwithstanding, I would be much reassured if Snapper supported other COW filesystems.
(This has been attempted but I don't think anything shipped -- https://github.com/openSUSE/snapper/issues/145 . I welcome correction on this!)
The transactional features of SUSE's MicroOS family of distros rely heavily on Btrfs. As such, this lack of awareness of snapshot space utilization deeply worries me. I have raised this with SUSE management, but my concerns were dismissed. That worries me.
What I want to see, for clarity, is for Zypper to look at what packages will be replaced, then ask the FS how big the consequent snapshot will be, and include that snapshot in its space estimation checks before beginning the operation so that at least the packaging operation can be safely aborted before starting.
A better implementation would be to integrate package management with snapshot management so that older snapshots could be automatically pruned to ensure necessary space is made available, while also ensuring that a pre-operation snapshot is retained for rollback. That's harder but would work better.
As it is, currently neither is attempted, and Zypper will start actions that result in filling the disk and thus trashing the FS, and there are no working repair tools to recover.
- Red Hat removed Btrfs support from RHEL. As a result it has had to bodge transactional package management together by grafting Git-like functionality into OStree, then building two entirely new packaging systems around OStree, one for the OS itself and a different one for GUI-level packages. The latter is Flatpak, of course.
This strikes me as prime evidence that:
1. Btrfs isn't ready.
2. Linux needs an in-kernel COW filesystem -- because much of the complexity of OStree, Flatpak, Nix/NixOS, Guix, SUSE's `transactional-update` commands and so on would be rendered unnecessary if it were in there.
In that case, df also lies on ZFS, bcachefs, and LVM+snapshots. It's in the nature of thin allocation and CoW; if you ask two things sharing storage how much space they have available, that doesn't mean that space is available to both of them at the same time.
It is possible. Got any examples?
And more saliently do you have reports of disk corruption occurring as a result of falsely reported free space which was not in fact available? As has corrupted multiple Btrfs volumes for me, on multiple machines, in production, and with disastrous consequences and resultant need for complete system reinstallation?
> I want a 100% accurate accounting of free space at all times via the standard xNix free-disk-space reporting command
You don't get that on many modern storage systems, because of thin provisioning.
You don't get that even on old school unix if you ask separately in two different directories, and fail to account for whether they're on the same filesystem or not.
[[citation needed]]
All xNix owners who can use the shell competently know that different directories may have different free-space counts. The problem is when the filesystem won't tell you a real free-space count at all.
This stems from a misunderstanding. Fsck and fsck-adjacent tools have three purposes:
1. Replay journal entries in a journalled filesystem, so that the filesystem is repaired to a good state for mounting
2. Scrub through checksums and recover any data/metadata that has a redundant copy
3. In rare cases, a fsck tool encountering invalid data can make guesses as to how the filesystems should be structured -- basically, shot in the dark attempts at recovery.
Btrfs does not need #1 because it is not journalled. Assuming write barriers are working, any partially written copy of the filesystem is valid and will simply appear as if the pending writes had been rolled back. This aspect alone greatly diminishes the need for a fsck tool that filesystems like ext4 have.
As for #2, Btrfs already has scrub support. No issues there.
As for #3, it's questionable whether you should ever rely on such functionality, and fsck tools that do implement such functionality tend to have little maneuverability in the first place.
I disagree.
You seem to be attempting to justify a profound failing by quibbling about the meaning of words or commands.
The real problem is: Btrfs corrupts readily, and it lacks tools to fix the corruption.
What the tools are called, what their functional role is theoretically meant to be, and whether this is justified in a tool of a given name is tangential and relatively speaking unimportant.
Here's a recent example of corruption unearthed by users after Fedora started defaulting to btrfs: https://bugzilla.redhat.com/show_bug.cgi?id=2169947.
Unfortunately, btrfs is not that smart, and you must trigger a rebalance event to rewrite every block in the filesystem to return to full redundancy.
This rebalance behavior is a deal-killer for many uses.
It has a more limited feature set and said to have simpler codebase than zfs/btrfs. It has a single outstanding non-stable feature.
It seems to be in active development, while btrfs seems to have become stagnant/abandonware before it was finished/stabilised completely. I have read several horror stories about data loss, so I have avoided it so far.
On the other hand it is not widely deployed yet, there is less accumulated knowledge than in case of zfs.
I'm looking forward to trying it in my NAS when buying new disks next year. The COW snapshots would fit my needs (automatic daily snapshots, weekly backups).
(Now using LUKS+LVM+ext4, this would give a better, more integrated, deduplicated solution, I have lots of duplicated data right now)
Why would you think so? I can't remember the last time a kernel was released without something at least a bit exciting about btrfs
https://kernelnewbies.org/LinuxChanges#Linux_6.5.File_system...
I found using mirrored vdevs in ZFS much easier to manage and much more stable.
That's not exactly a fair comparison. If you restricted your usage of btrfs to a similarly narrow range of features, you would probably have had a much better experience.
I can't say for sure that this never happens, but that's certainly not been the failure mode for any of the drive failures my btrfs RAID1 has experienced. I don't think I've ever needed to reboot or even remount my filesystem, just replace the failed drive (physically, then in software). But I always have more than two drives in the filesystem, so a single drive failure only puts a fraction of my data at risk, not everything.
That's not correct, you can rebuild online. The readonly mode is only relevant when you reboot during the failure and don't have the right options set on the volume.
But you totally can replace a live drive without affecting the availability.
They are fixing the fixable issues, but the on-disk format still makes some gotchas inevitable. It sounds like there's never going to be a great solution to live rebuilding of redundancy.
The last benchmarks from Phoenix are a few years old but look promising: https://www.phoronix.com/review/bcachefs-linux-2019
> The design features of this file-system are similar to ZFS/Btrfs and include native encryption, snapshots, compression, caching, multi-device/RAID support, and more. But even with all of its features, it aims to offer XFS/EXT4-like performance, which is something that can't generally be said for Btrfs.
I was surprised at that as I believed that btrfs is generally faster than ext4. Looking ahead to the last page, the geometric mean or the benchmarks supports that view too.
Btrfs goes very fast at first but slows down when it has to start pruning/compacting its on-disk btree structures, and then the performance suffers bad. Thus, btrfs works best when you have a spiky workload that lets it "catch up", and you never fill the disk. Concrete example: historically, removing a snapshot while under load was a disaster, with IO waits over 2 minutes. So, both sides are correct: btrfs is very fast and btrfs is very slow.
XFS was never crazy fast, and has the smallest feature set of bunch, but it just kept chugging at the same pace with almost no change, regardless of what the workload did. In more complex use, you had to avoid triggering bad behavior; e.g. there was a fixed number of write streams open, something like 8, and if you had more than that many concurrent writes going your blocks got fragmented. It was very much a freight train; not particularly fast but very predictable performance, and no serious degradation ever.
ext4 was sort of in between those; mostly very fast, with some hiccups in performance. Great as long as your storage is 100% reliable -- we had scrubbing in-product on top of the filesystem.
We ended up recommending xfs to most customers, at the time. Predictability trumped minor gain in performance, for most uses.
It does something that ZFS can't- be merged into the kernel.
bachefs has a really flexible design here where you basically add all of your disks to the storage pool and then you can pick redundancy and performance settings per folder (arbitrary subtrees, not just datasets decided at setup time) or even file. For example you can configure a default of 2 replicas for all data, but for your cache directory set it to 1 replica. If you have an important documents folder you can set that to 3 replicas, or 4.2 erasure coding.
Similarly you can tell it to put your cache folder on devices labeled "ssd" but your documents folder should write to "ssd" but then be migrated to "hdd" when they are cold.
And again, all of this can be set at any time on any subtree. Not just when you initially set up your disks or create the directories.
It's literally the second item listed on the main web site
Full data and metadata checksumming
Maybe I'm missing something, but the pool abstraction makes this very clean and clear.
That sounds backwards. Don't you have to manually define the layout of the vdevs first in order to establish the redundancy, and then allocate the volumes you use for filesystems or iSCSI? If you just do a `zpool create` and give it a dozen disks and ask for raidz2, you're just creating a single vdev that's a RAID6 over all the drives. There's an extra step compared to the btrfs workflow, but if you're not using that opportunity to micromanage your array layout I don't see why you'd prefer that extra step to exist.
> and I can't have multiple file systems (well, you can have sub-volumes but you can't avoid a big file system in that collection of devices).
Isn't this a purely cosmetic complaint? With at least btrfs, you don't even have to mount the root volume, you can simply mount the subvolumes directly wherever you want them and pretty much ignore the existence of the root volume except when provisioning more subvolumes from the root. You can pretend that you do have a ZFS-style pool abstraction, but one that's navigable like a filesystem in the Unix tradition instead of requiring non-standard tooling to inspect.
> you are unavoidably constrained by the capacity of the smaller device
Sure, so what does bcachefs actually do about it? ENOSPC?
Answers to the question on synchronous write behavior also welcome.
It looks more like you're questioning whether the filesystem that Linus just merged does obviously wrong things for the simplest test cases of its headline features. Has something given you cause to suspect that this filesystem is so deeply and thoroughly flawed? Because this doesn't quite look like trying to learn, it looks like trying to find an excuse to dismiss bcachefs before even reading any of the docs. Asking if everything about the filesystem is a lie seems really odd.
This person was asking pertinent questions, no need to bash it.
I lost some data with XFS after it was declared "stable" in the linux kernel. Every day at around 00:00 the power will fail for around 2 seconds. Other filesystems (ext2, jfs, reiser) will do a fsck, but xfs was smarter. After 2 or 3 crashes the xfs volume will not be usable anymore (no fsck possible).
So yes, some of us do need more than a "trust me, it's ok" when we are talking about our data.
In the common case that you mentioned, data present on the full SSD would be overwritten "in standard LRU fashion"; meaning the "Least Recently Used" data would no longer be cached. New data would be written to the SSD while a background "rebalance thread" would copy that data to the HDD. I assume that the "sync" command would wait for the "rebalance thread" to finish, though I will admit my own ignorance on that front.
(LVM + LUKS + BTRFS does it for me right now)
I stand corrected, last time I checked they didn't support freeze/unthaw on ZoL, it's always worked on FreeBSD though.
I too am not a big fan of the GPL, especially if the text of the license is longer than the program itself. But any filesystem (let alone a modern one) is very much a non-trivial feat of engineering; the author should have the full right to protect their (and their users') interests.
Different OS's prefer different filesystems, because filesystems tend to be both complicated and heavily opinionated in design and implementation - just like different OS's. Linux is the odd one by supporting several dozen, all the other OS's stick to 1 or 2 (usually "old" and "new" like HFS+/APFS, FAT/NTFS, etc), plus UDF&FAT as the lowest common denominator for data interchange. There is very little precedent / use cases for sharing volumes like you suggest: non-removable disks tend to stay in one machine for their lifetime; dual-booting is extremely niche (where Linux/BSD themselves are all niche) and mostly a domain of enthusiasts.
Maybe in 10-20 years if someone white-rooms a bcachefs or ZFS implementation, I guess.
I recommend this talk by Bryan Cantrill, which provides more context on the whole licensing story for Solaris: <https://www.youtube.com/watch?v=-zRN7XLCRhc&t=1375s> TL;DW: lots of effort and good will was put into making this code as free as possible, with the intention of making it broadly usable.
There are other reasons why e.g. OpenBSD won't adopt ZFS, the main one being the sheer complexity: <https://flak.tedunangst.com/post/ZFS-on-OpenBSD>. Again, different projects, different goals.
Apple heavily considered ZFS (even advertised it as an upcoming feature), but then gave in to NIH. Probably because it didn't fit their plan for mobile, and seeing the runaway success in that dept you can't really blame them.
But we digress! Is ZFS even a good system for removable media? Heck no. What do you want to use it with? Digital cameras? Portable music players? 2007 called and wants its toys back. Backups? Yeah, that can work, but there's little value in being able to recover backups on a foreign OS. Giving someone a random file on a thumb drive? Use ExFAT (or even plain old FAT32), just keep it simple!
A bigger question would be "why bcachefs when we have a stable ZFS in base?"
Yes, you can bundle GPL software with BSD software, much the same as I can bundle EULA'd software with BSD software: under the stricter terms. Obviously that's not the intent of the question, but sure, it's theoretically possible in some future fictional world where the BSDs are GPL'd.
The problem is this sort of code is strongly tied to the operating system, and porting it would require significant effort, if even feasible at all.
How stable it to use for day to day desktop tasks?