How did they get burned? (I'm thinking of switching to Btrfs on my personal laptop (running Arch), mainly for the snapshots, subvolumes, checksums+deduplication, etc.)
How did they get burned? (I'm thinking of switching to Btrfs on my personal laptop (running Arch), mainly for the snapshots, subvolumes, checksums+deduplication, etc.)
There are very few filesystems that can shrink like this:
btrfs filesystem resize -1G /mount
If the "btrfs fi df" command shows any of the data/system/metadata elements close to limits, it is time to take action. They can all push you into read-only mode.the filesystem was always cleanly unmounted, I didn't do any crazy experiments with it, it had plenty (>40%, hundreds of GiBs) of free disk space; on one day I noticed that accessing some files was slower than usual, the noise of the HDDs had a weird pattern and from there it went downhill (at that time there was no fsck utility so I didn't even know what to do - if I'm not mistaken, BUT I MIGHT, the general opinion was that a fsck-util was not even needed? but again, I might have just dreamed this) and finally stopped working with a huge amount of error messages in the kernel logs => I had to reformat everything (I went back to mdadm-raid5) and restore from the secondary NAS' backup.
In my opinion there were 2 main things that contributed to the bad reputation of the FS:
1)
In those days there was a looot of hype about btrfs (as well for me it was a dream come true) which is the reason why I decided to use it.
As far as I remember I didn't have to enable any special "experimental" feature in the kernel to use it => that gave me the impression that "it works under normal conditions, the DEVs are probably being a bit conservative for corner cases".
2)
All functionalities (e.g. raid5 with rebalancing in my case) were all available from day 1 (at least when I started using it) so if something's available I tend to think that it works (I mean, from a user's point of view why make something available that does not work?) - but no, apparently many people had many problems with many different configurations/usage patterns.
-
As others have written I hope that Bcachefs will primarily focus on reliability and only secondarily on performance & features, but I admit that that's difficult to balance :). And I hope that the fsck-utility will be available from day1 (at least with a reduced functionality). And I hope that "experimental" features will need some kind of very explicit command line option (e.g. "-might_break_everything") to use them through userland utils.
Go Kent!!! :) I'm since a while one of your Patreon follower - If you manage to create something usable I'll print the pic shown on LWM and I'll hang it up in my PC-room, maybe I'll even take my twitter account out of cryosleep to tell Musk that he should use Bcachefs as FS for his Teslas :))
"The RAID56 feature provides striping and parity over several devices, same as the traditional RAID5/6. There are some implementation and design deficiencies that make it unreliable for some corner cases and the feature should not be used in production, only for evaluation or testing. The power failure safety for metadata with RAID56 is not 100%."
https://btrfs.readthedocs.io/en/latest/btrfs-man5.html#raid5...
RAID56 Stability: Unstable (do not use for other then testing purposes, known severe problems, missing implementation of some core parts)
Attempting to create a btrfs filesystem returned a prompt with this warning (until 2013):
Btrfs is a new filesystem with extents, writable snapshotting,
support for multiple devices and many more features.
Btrfs is highly experimental, and THE DISK FORMAT IS NOT YET
FINALIZED. You should say N here unless you are interested in
testing Btrfs with non-critical data.
That prompt was removed when things stabilized, and Fedora's use of btrfs as the default root has not been accompanied by complaints of excessive filesystem corruption (although there are still many rough edges and unfinished features).https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
> We prioritize robustness and reliability over features and hype: we make every effort to ensure you won't lose data.
This is exactly what btrfs failed at, this is not just an issue with the code itself, but also with the developers. Why should i trust feature-oriented developers with writing my infrastructure code, where reliability and robustness are top priority?
There's a list of features and readiness at https://btrfs.wiki.kernel.org/index.php/Status - there were people who didn't check it and used unstable features and experienced issues. This is partially on the project for not making it more explicit.
The second is that there are people who read that, see the unstable raid5/6, and are vocal about btrfs being overall broken / not ready rather than accepting that feature just not being available yet.
Here's a list of failures from zfs, including many kernel panics for example https://github.com/openzfs/zfs/labels/Type%3A%20Defect - is it broken and shouldn't have been released?
Maybe to put it differently: new file systems should offer improvements wrt reliability for stable features of the file systems of the previous generation; regressions are unacceptable.
How long are we expected to wait? It's been over a decade.
btrfs was touted as a spiritual successor to ZFS. raidz5 is a core feature of ZFS, not an incidental extra feature.
https://www.usenix.org/system/files/login/articles/login_sum...
I've been running full Btrfs on my workstation since December 2021 with 0 issues (kernel 5.10). It survived multiple hard reboots and power offs, Ryzen 5000 CPUs seems to lose power every now and then on Linux (it's a known issue for years now that no one has figured out yet).
Having compression is great. compsize reports I've saved 33GB on my /home subvolume so far (https://github.com/kilobyte/compsize). I don't use snapshots, I just backup the whole thing.
Obviously this doesn't mean that in the future Btrfs won't eat my data, but so far it hasn't and I'm happy.
There's also the issue with BTRFS performance degrading hugely with any significant fragmentation, to the point that "autodefrag" has to be a feature, which of course also adds SSD wear...
ZFS has never burned me.
In the early days after Fedora made it available, I lost my root twice after a kernel panic and later after a hard power reset.
EDIT: in comparison, I've been using xfs for almost a decade now with numerous hard power offs and kernel panics. I've only ever lost select files that were open for write at the time of restart, and not the entire filesystem.
Never used btrfs but I see corruption reports in here everytime I read about it, and it's been in development for over a decade and this thing needs to be put to rest before it eats more people's data.
Hopefully bcachefs becomes feature parity and stability against zfs to provide a modern filesystem to wider Linux audience at last.
Funny to see we're still using ext4 as the default filesystem which is nothing fancy forever.
Because the Fedora community wanted to, basically. They're not obligated to go along with whatever Red Hat wants to do, although there is naturally some alignment there. Plus having a different default filesystem really isn't that big of a deal.
disclosure: I work for Red Hat.
My Linux laptop, running OpenSuse Tumbleweed, got its root fs into a state where it would hang when trying to mount RW. I was able to boot off of a rescue image, mount it RO, and copy my data to another machine. None of the repair suggestions were able to get the fs into a mountable state, I had to reinstall it and copy my data back.
Every VM at work that uses btrfs gets into a state where some maintenance process (rebalancing, I think - I'm not one of the Linux admins) will consume all CPU in the VM and prevent it from doing any work. The only thing that varies is how long it takes - it seems to be related to how much write activity there is on the fs, VMs with heavier write activity get into this state faster. This has led to outages and the Linux vendor was unable to resolve this issue. We've since removed all traces of btrfs from our data center.
[1]: https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
We didn't use Docker at work, at least not on the VMs that I was affected by. This one was a gradual degradation over time - the maintenance process would run fine, but after a while it would start hanging the machine. It really felt like the fs was only good for some amount of writes before the maintenance process got overwhelmed and would hang the VM whenever it ran.
I'm happy to see that things are getting better, but the reputational damage will take a long time to repair.
All that being said, I still use btrfs on my Linux laptop, but I've also become more conscientious about backing it up.
I have a comparison of features between BtrFS, ZFS, and EXT4+LVM at [0], and as I note at the top of that post I was looking to move to BtrFS. I did so and lost data.
The fact that it turned out that the scrub function spent years corrupting data by mistake was only one of the many flaws with it.