Merging bcachefs
lwn.net
lwn.net
As someone not familiar with the filesystem related parts of the kernel, it's quite surprising to hear this. It sounds like filesystems are a lot more integrated (or at least less modular) than other kernel subsystems.
Anyone else who found this surprising, I recommend reading this other LWN article that is referenced as well: https://lwn.net/Articles/886708/
However, that mostly discusses issues caused by having 32-bit data structures, and how they will cause issues in 2038 when it's no longer able to handle the timestamps required. Specifically for ext3, NTFS, and ReiserFS.
But other than that issue, I don't really understand why it's difficult to simply rip out support for a specific filesystem. Compared to something like a driver for a PCIe or USB device, what makes filesystems so much more integrated and difficult to remove?
AFAIK, on Linux filesystems are closely coupled with the memory management subsystem and the directory and inode caches. It's part of the reason why filesystem access on Linux is so fast.
> But other than that issue, I don't really understand why it's difficult to simply rip out support for a specific filesystem. Compared to something like a driver for a PCIe or USB device, what makes filesystems so much more integrated and difficult to remove?
It's not that unusual for users to have a partition containing a filesystem (sometimes, but not always, on an external drive) surviving unchanged through several migrations to newer Linux distributions and/or hardware. Compared to something like a PCIe or USB device, filesystems have a longer life.
Ripping out a filesystem is, technically, near-trivial. The downsides are fully human: the Linux kernel-userspace ABI is normally considered a golden promise that shall not be broken. The developers don't want users to have a bad morning on which their old filesystems no longer work.
"including the ext4 filesystem, but also F2FS, FAT, GFS2, HFS, ISO9660 (CDROM), JFS, NTFS, NTFS3, and the device-mapper layer"
Side note: I feel almost gaslit, it's really hard to find any mention of guestmount running a VM outside this single page https://libguestfs.org/guestfs-internals.1.html
It's too bad linux can't already run any in-kernel filesystem as a user process via FUSE, when you prefer greater isolation in exchange for worse performance, at the flip of a mount option.
There's no technical reason for this to not be possible IMO... it's just a product of the tight coupling of everything in-kernel, as implemented today.
I'm short on time to confirm at the moment, but I believe it was this [0] talk that left me with the impression kernel devs were exploring general solutions of this nature.
Hard to remove because someone, somewhere uses it and Linux doesn't like to break things for the user.
This is exciting because db workloads are btrfs's kryptonite. The only way to avoid crippling fragmentation on btrfs is to disable copy-on-write, which also disables checksumming hence nullifying one of btrfs's main selling points. ZFS seems to handle such loads much better, and it would be interesting to see how bcachefs deals with them.
First mention of bcachefs here on HN (13years ago) and now it's gets merged into kernel.
What an awesome achievement!
Does anyone know how this actually works, is it going well for that patchset, is it getting closer to being merged?
It is extremely unlikely that bcachefs will be merged as-is. This is true for anything of it's size. (These days... There have been large subsystems in the past that were merged in unacceptable state with the promise that eventually they will be fixed. They usually weren't. Which is why the bar is so high now.) But this just means that there will need to be debate and work to hammer it into a shape that is acceptable for the kernel. This can be a long process, but I don't think there is anything fundamentally wrong in bcachefs that would exclude it.
bcachefs was discussed recently on HN – https://news.ycombinator.com/item?id=35899527 – and is a file system with COW, a GPL compatible license (licensing is an area where ZFS can be tricky, to put it lightly), and ext4 levels of performance.
McKusick did this for FreeBSD UFS/FFS:
* https://wiki.freebsd.org/ExampleUfsSnapshots
* https://people.freebsd.org/~rse/snapshot/
* https://man.freebsd.org/cgi/man.cgi?query=mksnap_ffs
Various papers at:
I like ext4 for being simple and fast and having options to turn off the unnecessary stuff. It's great for gigabytes of caches, tempfiles, build artifacts, scratch files etc. The bookkeeping and indirection needed for those features would be a waste of CPU cycles for shortlived data.
I know how these two features have some conflicts, but it meant we can't have database workload with decent performance on btrfs. -- brtfs snapshot just can't scale, the cow have too high performance hit.
ZFS handles snapshots with database workload just fine.
If you're willing to switch OSes, FreeBSD UFS is a more traditional filesystem, with optional snapshots (modifications to files in a snapshot have to be cow, of course)
Btrfs' implementation of cow, of course.
They have no intention to make it work for database workload. When you ask why it's slow, they just ask you to disable cow ( which requires recreating the file -- something you would die to avoid with multi TB database)
> btrfs don't allow nodatacow with snapshots
What's the use case for a snapshot at that point? Isn't it the same as making a copy of the files?
I'm not even sure what would this look like... CoW is what enables snapshots. Without CoW you can't get a consistent copy anymore.
As you say, it's a consistent copy. cp(1) won't give you that.
Under this paradigm, rather than trying to make a filesystem that understands applications well enough to snapshot them, you instead make applications that have filesystem snapshots as part of their conceptual model.
In the lower-effort version of this approach, you have software like Postgres, where the application layer can be told "I'm going to use filesystem tooling to take a consistent snapshot, so make your on-disk state consistent for a while." In PG, you'd call pg_start_backup(), which will flush all pending writes to the table files to disk, and then spool all future writes purely in the WAL journal until the backup completes (i.e. until you tell it pg_stop_backup()) — at which point all those pending changes get replayed out to the table files.
In the higher-effort version of this approach, you have software that has one or more CoW filesystems it natively understands and integrates with — where you never directly address the filesystem at all, but rather, you tell the application to take an application level snapshot; and then the application uses the filesystem CoW features, together with its own consistency primitives, to efficiently achieve that (which might not necessarily result in something that's a "filesystem snapshot" from the FS's perspective, but rather just a bunch of individual CoW-cloned files in a directory.) I believe that Oracle DBMS does this, though I might be wrong.
You need to copy the file over to enable/disable cow
Uhh. "You can't do overwrite-in-place if you want to keep a snapshot copy of the old data" simply makes sense, and you'll find every overwrite-in-place design that supports snapshots will take a write performance hit around the time the snapshot; either the snapshot has to atomically copy the whole data, or the first write after a snapshot can't be overwrite-in-place (or it has to make a separate copy of the original data for the snapshot; similar but worse).
And "move the snapshotted data out of the way upon writing" is going to give you better performance in a lot of cases.
> btrfs don't allow nodatacow with snapshots.
You can absolutely mix chattr +C with snapshots. It's CoW-when-needed in the face of snapshots, just like everything has to be in order to support snapshots.
And yes I can find many sources saying it's supported.
Btrfs was merged in 2009, but didn't gain wide acceptance until quite recently. Among the biggest distros, only Fedora uses it as default, and RHEL has actually dropped support. Even today, you can still find people claiming that they lost data because of it, though whether they are telling the truth I cannot say.
I'm not saying you should switch all your critical systems to bcachefs on the day it gets merged as you can never be sure about the absence of bugs (even the relatively simple ext4 had some data-eating bug a few years after introduction), but the path to "recommended filesystem" will be a lot shorter than btrfs.
At this point, I would already feel comfterable running bcachefs on my laptop – the only reason I don't is that I just can't be bothered running a custom kernel for it.
I haven't lost data on it, but my 8 drive Btrfs RAID6 filesystem locked up read-only on me, which wasn't fun. Switched to ZFS after that.
It is totally understandable that if the underlying driver starts returning errors, btrfs (or any other filesystem) may be unable to provide access to my data; but I was not at all happy about having to reboot the entire machine.
(I must admit, it is possible that the "freezing" was happening in underlying block device driver code and not in btrfs. I don't remember if I ever checked wchan to see which it was. My impression from reading dmesg output was the issue seemed to be with btrfs.)
I did once have a similar experience with ZFS as well. Sigh.
Only raid5/6 has known issues.
With mdadm, I don't get auto-repair on top of the checksumming.
A nice (and lone) data point regarding btrfs, though: My dentist told me that he was storing patient data using btrfs last year. :)
ISTR hours-long repair operations and a level of desperation that a lot of people wouldn't have had time for. This is assuming you didn't use its RAID or other features that definitely would eat your data.
(This is experience from 10 years ago when I was looking for the latest and greatest features to support a hosting platform)
Main problem with Bcachefs is its in unknown state. There was no major testing done on that. When some company (like Oracle or Suse) puts it through automated stress test matrix, it may do exceptionally well or completely fail.
For example Suse participated over 25 years on Ext3,4, ReiserFS/4, XFS, BTRFS.... This know how is not public!
Still a pretty good endorsement in my mind.