Bringing Bcachefs to Linux Mainline
lwn.net
lwn.net
How did they get burned? (I'm thinking of switching to Btrfs on my personal laptop (running Arch), mainly for the snapshots, subvolumes, checksums+deduplication, etc.)
> We prioritize robustness and reliability over features and hype: we make every effort to ensure you won't lose data.
This is exactly what btrfs failed at, this is not just an issue with the code itself, but also with the developers. Why should i trust feature-oriented developers with writing my infrastructure code, where reliability and robustness are top priority?
There's a list of features and readiness at https://btrfs.wiki.kernel.org/index.php/Status - there were people who didn't check it and used unstable features and experienced issues. This is partially on the project for not making it more explicit.
The second is that there are people who read that, see the unstable raid5/6, and are vocal about btrfs being overall broken / not ready rather than accepting that feature just not being available yet.
Here's a list of failures from zfs, including many kernel panics for example https://github.com/openzfs/zfs/labels/Type%3A%20Defect - is it broken and shouldn't have been released?
Maybe to put it differently: new file systems should offer improvements wrt reliability for stable features of the file systems of the previous generation; regressions are unacceptable.
How long are we expected to wait? It's been over a decade.
btrfs was touted as a spiritual successor to ZFS. raidz5 is a core feature of ZFS, not an incidental extra feature.
https://www.usenix.org/system/files/login/articles/login_sum...
In the early days after Fedora made it available, I lost my root twice after a kernel panic and later after a hard power reset.
EDIT: in comparison, I've been using xfs for almost a decade now with numerous hard power offs and kernel panics. I've only ever lost select files that were open for write at the time of restart, and not the entire filesystem.
Never used btrfs but I see corruption reports in here everytime I read about it, and it's been in development for over a decade and this thing needs to be put to rest before it eats more people's data.
Hopefully bcachefs becomes feature parity and stability against zfs to provide a modern filesystem to wider Linux audience at last.
Funny to see we're still using ext4 as the default filesystem which is nothing fancy forever.
Because the Fedora community wanted to, basically. They're not obligated to go along with whatever Red Hat wants to do, although there is naturally some alignment there. Plus having a different default filesystem really isn't that big of a deal.
disclosure: I work for Red Hat.
The fact that it turned out that the scrub function spent years corrupting data by mistake was only one of the many flaws with it.
There are very few filesystems that can shrink like this:
btrfs filesystem resize -1G /mount
If the "btrfs fi df" command shows any of the data/system/metadata elements close to limits, it is time to take action. They can all push you into read-only mode.the filesystem was always cleanly unmounted, I didn't do any crazy experiments with it, it had plenty (>40%, hundreds of GiBs) of free disk space; on one day I noticed that accessing some files was slower than usual, the noise of the HDDs had a weird pattern and from there it went downhill (at that time there was no fsck utility so I didn't even know what to do - if I'm not mistaken, BUT I MIGHT, the general opinion was that a fsck-util was not even needed? but again, I might have just dreamed this) and finally stopped working with a huge amount of error messages in the kernel logs => I had to reformat everything (I went back to mdadm-raid5) and restore from the secondary NAS' backup.
In my opinion there were 2 main things that contributed to the bad reputation of the FS:
1)
In those days there was a looot of hype about btrfs (as well for me it was a dream come true) which is the reason why I decided to use it.
As far as I remember I didn't have to enable any special "experimental" feature in the kernel to use it => that gave me the impression that "it works under normal conditions, the DEVs are probably being a bit conservative for corner cases".
2)
All functionalities (e.g. raid5 with rebalancing in my case) were all available from day 1 (at least when I started using it) so if something's available I tend to think that it works (I mean, from a user's point of view why make something available that does not work?) - but no, apparently many people had many problems with many different configurations/usage patterns.
-
As others have written I hope that Bcachefs will primarily focus on reliability and only secondarily on performance & features, but I admit that that's difficult to balance :). And I hope that the fsck-utility will be available from day1 (at least with a reduced functionality). And I hope that "experimental" features will need some kind of very explicit command line option (e.g. "-might_break_everything") to use them through userland utils.
Go Kent!!! :) I'm since a while one of your Patreon follower - If you manage to create something usable I'll print the pic shown on LWM and I'll hang it up in my PC-room, maybe I'll even take my twitter account out of cryosleep to tell Musk that he should use Bcachefs as FS for his Teslas :))
"The RAID56 feature provides striping and parity over several devices, same as the traditional RAID5/6. There are some implementation and design deficiencies that make it unreliable for some corner cases and the feature should not be used in production, only for evaluation or testing. The power failure safety for metadata with RAID56 is not 100%."
https://btrfs.readthedocs.io/en/latest/btrfs-man5.html#raid5...
RAID56 Stability: Unstable (do not use for other then testing purposes, known severe problems, missing implementation of some core parts)
There's also the issue with BTRFS performance degrading hugely with any significant fragmentation, to the point that "autodefrag" has to be a feature, which of course also adds SSD wear...
ZFS has never burned me.
My Linux laptop, running OpenSuse Tumbleweed, got its root fs into a state where it would hang when trying to mount RW. I was able to boot off of a rescue image, mount it RO, and copy my data to another machine. None of the repair suggestions were able to get the fs into a mountable state, I had to reinstall it and copy my data back.
Every VM at work that uses btrfs gets into a state where some maintenance process (rebalancing, I think - I'm not one of the Linux admins) will consume all CPU in the VM and prevent it from doing any work. The only thing that varies is how long it takes - it seems to be related to how much write activity there is on the fs, VMs with heavier write activity get into this state faster. This has led to outages and the Linux vendor was unable to resolve this issue. We've since removed all traces of btrfs from our data center.
[1]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
We didn't use Docker at work, at least not on the VMs that I was affected by. This one was a gradual degradation over time - the maintenance process would run fine, but after a while it would start hanging the machine. It really felt like the fs was only good for some amount of writes before the maintenance process got overwhelmed and would hang the VM whenever it ran.
I'm happy to see that things are getting better, but the reputational damage will take a long time to repair.
All that being said, I still use btrfs on my Linux laptop, but I've also become more conscientious about backing it up.
I've been running full Btrfs on my workstation since December 2021 with 0 issues (kernel 5.10). It survived multiple hard reboots and power offs, Ryzen 5000 CPUs seems to lose power every now and then on Linux (it's a known issue for years now that no one has figured out yet).
Having compression is great. compsize reports I've saved 33GB on my /home subvolume so far (https://github.com/kilobyte/compsize). I don't use snapshots, I just backup the whole thing.
Obviously this doesn't mean that in the future Btrfs won't eat my data, but so far it hasn't and I'm happy.
I have a comparison of features between BtrFS, ZFS, and EXT4+LVM at [0], and as I note at the top of that post I was looking to move to BtrFS. I did so and lost data.
Attempting to create a btrfs filesystem returned a prompt with this warning (until 2013):
Btrfs is a new filesystem with extents, writable snapshotting,
support for multiple devices and many more features.
Btrfs is highly experimental, and THE DISK FORMAT IS NOT YET
FINALIZED. You should say N here unless you are interested in
testing Btrfs with non-critical data.
That prompt was removed when things stabilized, and Fedora's use of btrfs as the default root has not been accompanied by complaints of excessive filesystem corruption (although there are still many rough edges and unfinished features).https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
* ZFS will forever be problematic because of licensing reasons * BTRFS stability is widely untrusted after their problems many years ago * ext4 and xfs don't have data checksumming * To get the same benefits of ZFS you currently need to stack a lot of device mapper plugins
https://news.ycombinator.com/item?id=31420769
Apparently the incompatibility wasn't even intentional when Sun wrote CDDLv1.
But if you're on Ubuntu or any other distro that require less efforts, you're missing much by using legacy filesystems.
If you're on Ubuntu, even if you have ext4 on root, all you do is install a single package and allocate an empty partition for zfs and you're good to go and it's stupid to miss out on what zfs offers by sticking to ext4 forever.
Flexible space allocation between volumes, instant snapshotting, transparent compression and backing up the entire filesystem to a remote server is already available if you choose to use zfs.
Far better than using btrfs with endless report of corruption.
Writing it as BCacheFS would have helped, but CamelCase is not very popular these days it seems
On a sidenote, I integrated bcachefs as a vmx driver into VMWare ESXi for a CI/CD build server a few years back. The build system ran in VMs and on containers, but the non-essential target directories sat on bcachefs volumes with the caching layer directed first at RAM, then at SSD, then finally at the HDD. Managing all the Unity3D and Unreal caches was amazingly fast across dozens of different SKUs of what the same project.
From my project write-up on my LinkedIn:
Bcache-like caching layer for VMWare ESXi
Reduce latency and read/write delay even if using SSD as your storage.
Written in C and inserting itself as a storage tier into VMWare ESXi to handle read & write storage requests this caching system accelerates all accesses to the underlying backing store.
Can work in both write-back and write-through caching modes. (Native C kernel device driver)It's a long list. It's better to compare with btrfs
https://lackofimagination.org/2022/04/our-experience-with-po...
https://www.percona.com/blog/2017/12/07/hands-look-zfs-with-...
Unfortunately, this has been just around the corner on Linux for over a decade now, and the two filesystems that promised to deliver it are unlikely to reach mainstream support on Linux. ZFS has licensing issues, and uses too much memory for a desktop system. BTRFS tried to do too many things, and has had too many reliability issues for me to trust it.
How is zfs using too much memory? zfs can run on a 2GB server (with some swap). Any laptop would have enough memory to run it. You might want to change "zfs_arc_max" as zfs tries to use half the memory available on a system if you don't set it and don't use deduplication.
Lopks like it can replace/complement ZFS, Btrfs
Not being at the mercy of a company like Oracle is a huge plus in many ways. A huge plus for future development and risk-free adoption.
I use ZFS on a server of mine, but I am one of those paranoid people that would switch just to get away from a project that could be hamstringed at any moment if Oracle has one of its episodes again.
ZFS (hopefully) never finding its way into the mainline kernel is kind of a meta-disadvantage.
They can't just relicense OpenZFS. The only attack vector is suing Ubuntu for distributing zfs because Canonical's lawyers think that's fine but even then, you're still free to use it by adding it to the system yourself like any other distribution is doing.
Running a new filesystem is far more dangerous than your theoretical legal concerns.
The thread you linked is someone who was using software with a proprietary license against its terms and is being asked to pay for its usage -- obviously I think Oracle is being scummy but it's not a comparable situation at all. This would be like saying that you won't use VS Code because Microsoft once demanded that someone who was using a cracked copy of Windows pay them -- it's a complete non-sequitur.
Would you refuse to use ZFS on FreeBSD as well?
Empty threats are void, IMHO. Might take a bit of bravery to step over the fence, but if no one comes after you after a certain amount of time has passed, the fence is effectively not there. Ubuntu already did that work, though, so...
Also "Empty threats are void" is a silly stance to take; betting that nobody who will own copyright to code in the CDDL is going to sue anytime in the next 80ish years would be dumb even ignoring the fact that Oracle owns copyright to at least some of that code right now, and they are quite litigious.
> betting that nobody who will own copyright to code in the CDDL is going to sue anytime in the next 80ish years would be dumb even ignoring the fact that Oracle owns copyright to at least some of that code right now, and they are quite litigious.
When I wrote this, I was under the impression that US law more or less required copyrights to be regularly defended in order to maintain any "legal weight", but on Googling, it seems that while the awards in a case where a copyright was undefended until some significant time had passed, were lesser, it was still quite defensible. So I guess, fuck, that kind of sucks in a way, because it incentivizes patent trolling/hoarding/squatting, but it also makes it possible for someone without significant financial backing to benefit from a patent without having to afford good and regular legal defense.
That's one hell of a damn tradeoff.
A possibly better system would perhaps be one similar to what happens in the music biz in which a company takes a cut of any wins or profits by providing someone substantial legal representation in the defense of their copyrighted works.
Some of us will NEVER accept ZFS's licensing. The ability to remain absolutely vigilant in regards to licensing immensely helped IBM during the SCO v. IBM debacle in regards to the JFS file system. We have learned this E-X-A-C-T lesson before.
Additionally, since the code originates from Oracle, many of us progress from being vigilant to outright paranoid. Oracle has never been forgiving, warm, soft, nor cuddly. Hell, even Torvalds himself wants Larry Ellison to personally sign off on ever including ZFS into the Linux kernel [0]
[0] https://www.phoronix.com/scan.php?page=news_item&px=Linus-Sa...
I'm less sympathetic to sour grapes about ZFS. Kernel developers including Torvalds have made false negative remarks about ZFS vs Btrfs; the SFC deployed implausible legal theories made veiled threats against Ubuntu for maintaining a patch to Linux at their own risk allowing OpenZFS to be used.
OpenSSL recently went through a such a relicening and rewriting project. It took years, alas.
> 1998: Many major companies such as IBM, Compaq and Oracle announce their support for Linux.
https://en.wikipedia.org/wiki/History_of_Linux
Maybe they could stop giving that 2.6% as well, https://lwn.net/Articles/839772/
And yet the kernel community has been remarkably relaxed about Intel's golden screwdriver code and NVidia's binary blobs. Nothing if not inconsistent.
The opposite is true for every linux circle I am familiar with.
> It's just that I'm so damn tired of this whole thing. I'm tired of people thinking they have a right to violate my copyright all the time. I'm tired of people and companies somehow treating our license in ways that are blatantly wrong and feeling fine about it. Because we are a loose band of a lot of individuals, and not a company or legal entity, it seems to give companies the chutzpah to feel that they can get away with violating our license.
See also https://lwn.net/Articles/603131/ among others
The only concern is the distribution and you would mean you don't accept the fact that it can't go in the kernel or be released as part of the distro installer?
It's not hard to download zfs separately and use it like any distro is doing without having any licensing problems.