Bcachefs Pull Request Submitted for Linux 6.7
phoronix.com
phoronix.com
https://www.patreon.com/join/bcachefs
You can do a smaller pledge via the custom pledge button if you want.
Just to be clear: This is apparently unrelated to the (not that unusual) not-yet-accepted PRs and their lkml review discussions.
From Overstreet's Patreon post:
"The company that's been funding bcachefs for the past 6 years has, unfortunately, been hit by a business downturn - they've been affected by the strikes in the media production industry. As such, I'm now having to look for new funding.
This came pretty unexpectedly, and puts us in a bit of a bind; with the addition of Hunter (who's already doing great work), I'm stretched thin, without much of a buffer."
(https://www.patreon.com/posts/your-irregular-89670830)
I was concerned that without this context some might assume this is a "sponsors pulling out" type situation.
My understanding is that at some point during development of bcache, Kent realized "hmm... if I just got rid of cache eviction and cleaned this up a bit, it would make a nice light-weight CoW SSD-optimized filesystem".
My understanding is that there are very few untried ideas in bcachefs, and it's a relatively low-risk project. Having been burned by over-enthusiastic early adoption of btrfs, I'm hoping bcachefs becomes a conservative alternative for distros that don't currently use a CoW filesystem as the default filesystem. (Okay, I did use rsync to weekly mirror to ext4, so maybe "burned" is too strong a word, but I did wedge btrfs a few times by over-filling it.)
https://lore.kernel.org/all/169870980914.17163.3315949638870...
https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
As an outsider, there seems to be too many egos to placate when it comes to contributing to Linux.
Theoretically there is also a CoC.
The problem's more that you're going to have strong conflicts of opinions that you can't hash out face to face between people who are already overworked and under stress
It's hard for devs, but you also have people quitting because the maintener position is not just thankless, but fucking thankless.
It is not as simple as reviewers abusing developpers because of a lack of structure. A large part of the kernel is under resourced, lacking mainteners, and that is stressful for everyone
I have had the linux kernel ntfs driver corrupt data more than once, so I use it read only mode.
BTRFS is slow and unofficial on Windows, and its not that fast in linux to begin with.
F2FS would be perfect, and Samsung claimed they were building a Windows driver, but it never materialized.
Back when I needed something like this (over decade ago?), the Linux NTFS driver had many warnings to only use it read only. For kernel tied FSs, this is probably a good thought.
But isn’t there a Windows driver for ext4 (or ext2 at least, which is compatible with ext4)?
Honestly, I’m not sure I would try to write mount a filesystem from a different OS. Instead, I’d probably try to run the foreign OS in a VM (attached to the drive/image) and then export it over NFS/CIFS. This should work both directions and lets each OS deal with their own FS.
Of course, now you can mount ext[234] on Windows using WSL these days, so it's kinuva moot point.[4]
[1] https://sourceforge.net/projects/ext2fsd/ [2] e.g. https://www.diskgenius.com/ [3] e.g. https://www.paragon-drivers.com/en/lfswin/ [4] https://devblogs.microsoft.com/commandline/access-linux-file...
I guess we have different dreams.
That said, BTRFS is also not the rampant layering violation that ZFS is, for the good and the bad of that statement in so many ways. The way ARC and L2ARC is handled... is just crazy from a modern prospective. They may have cleaned it up since I last looked, but I tend to doubt it. It was pretty baked in.
BTRFS / ZFS really share 0 with bittorrent other than checksumming for integrity. That's it. Otherwise they are so different they might as well come from different planets.
ceph has a ways to go. It'll get there, but distributed filesystems are hard.
As a wise man said to me: Normal file systems take ~10-15 years to mature. Distributed ones take even longer. So ceph... may be getting usable, but I'm not sure I'd bet the farm on cephfs. (Their S3 stuff maybe. But that's much simpler.)
minio is a great example of a product not trying to be everything to everyone... and succeeding. Kudos to them and their team.
I've always much preferred to keep things as "down the middle" as I can for all of those reasons. Never had a CEPH cluster fail, haven't had any issues with ZFS (recently), nor any issues with BTRFS (recently), and that's not magic, it's all just because all of that is by the book or not done at all. It's a constraint at times, to be sure, but, hard to scoff at the results.
Also you're correct in assuming ARC/L2ARC is still.... weird. Sometimes you get wirespeed cache hits, sometimes you just don't, with no rhyme or reason outside of "well it just decided that file didn't belong in cache".
If you layer caches over each other and they don't pass hit rates to each other bad things happen, especially for designs like ARC. The ZFS caches explicitly take hit rate into account... but with a page cache over them... they don't know the true hit rates. Heck they are a secondary cache ;). Then L2ARC has the same issue one layer down. And there's the issue.
It isn't magical. It is just annoying.
Honestly on linux: btrfs is probably better for your sanity than ZFS.
As far as ceph goes, if you understand it, and can do the systems engineering.. yeah it is great stuff. My only complaint with ceph is just too complex for most mortals. (Despite that once you get to any real large scale... you can't avoid that complexity.)
ceph you can do some crazy shit if you know what you are doing. Want to do a seamless cluster migration? We got ya. ;) Not as good as AFS probably. But it is pretty scary.
But anecdotally I think I'm in the same spot many others are. Btrfs let me down a couple times (once lots of truncated files and another time I ran out of inode space on a mostly empty disk iirc) and I'm in no hurry to give it another chance.
zfs has just happily worked for me. I ran it as a root filesystem for years. All the snapshot/send/receive stuff just worked and stayed out of my way.
Note that this was for personal use and with SSDs and later nvme so I wasn't comparing it head to head with anything. Which might be why I didn't notice issues like page caching behavior.
I'm sure there's parts of btrfs that still suck. Don't get me wrong. And you can do the running out of space, so you can't remove files trick with ZFS too ;).
Personally, I only put btrfs into use on my laptop about 1.5 years ago. So... Take that for what you will. I'm pretty conservative.
I would posit however, if you need to guarantee cache hits... you don't really need cache, you need to roll your own flash storage and find a way to intelligently manage what's present on it at a given time. I've done this, namely for an environment that had very, very clear, scheduled needs for hot data, warm data, and archive data. Fairly trivial to take files and just scoot them to a different storage array when they're X days old on a nightly basis. That system is completely agnostic to, well, everything, I've done it on windows hosts and, while it's more of a pain in the ass, also worked fine there. That was back in the hardware raid or nothing days, when having a terabyte of flash available was fairly exotic and expensive, but, as a model, programmatically structured storage isn't old by any means and probably fits many environments better than ZFS and "as much memory as accounting would approve".
Also, I don't really 100% agree that ceph is too complex for mere mortals, it's more that, it very, very quickly becomes too complex for your average linux sysadmin to handle if you're not careful.
I really should dip my toes deeper into BTRFS for production usage though. I've played with it in the lab, I run it at home almost exclusively, but, never have done any big boy britches production stuff with it and it very well might be a good tool to have in my pocket.
I think this is mostly incorrect. In my understanding, the Linux page cache is mostly skipped. The problem, such as there is a problem, is occasionally duplicated data. My understanding is this data is only mmap files and perhaps the dentry and inode cache?
https://www.theregister.com/2021/08/02/paragon_ntfs_linux_ke...
I don't recall what, if any, relation the Paragon ntfs support has to ntfs-3g. I do recall when ntfs-3g was the new kid on the block, it was the only alternative for read-write ntfs support vs. paying Paragon software for their proprietary (at the time) kernel driver.
Or, well, as production ready as NTFS is...
Quoting an example from https://bcachefs.org/GettingStarted/
bcachefs format /dev/sd[ab]1 \
--foreground_target /dev/sda1 \
--promote_target /dev/sda1 \
--background_target /dev/sdb1
mount -t bcachefs /dev/sda1:/dev/sdb1 /mnt
This will configure the filesystem so that writes will be buffered to /dev/sda1 before being written back to /dev/sdb1 in the background, and that hot data will be promoted to /dev/sda1 for faster access. bcachefs format --help
--promote_target=(target)
Device or label to promote data to on readBut perhaps I can use bcache instead of cachefilesd to have some more controllable NFS cache.
Bcache seems to use block devices and I don't think it can stand in front of FUSE.
there is also [1]rclone with its own caching layer and own support for various backends directly. I don't remember anymore why i did prefer s3ql though, but i usually have some reasoning with things like this..