I do have a lot more pull requests to merge than I did before. I don't know if you want to count "Kent isn't reviewing PRs fast enough" as drama :)
also tricks to make it easy to convert a root FS to ZFS now that Ubuntu Server 24.04 added native root-on-zfs support: https://github.com/pirate/zfsify
In fact they delivered the erasure coding for parity raid back in march this year.
The thing is that as soon as you seriously give a chance to Bcachefs you see how good it is. I can only tell you that mixing different device tiers and having a per-file/directory replication setting is a god send specially in these times where storage costs more than gold.
I'm pretty sure Bcachefs is amazing and better than Btrfs. I also think Zfs is amazing and better than Btrfs. Even so, I still use Btrfs because I know it is guaranteed to always be present on any Linux without any effort on my part.
> That might be the best thing happened to the project since now development can happen at its own pace without the clicky bait influencers.
A better approach might have been to just paused mainline merging instead of forcing being kicked out?
Eg "Hey Linus, Bcachefs is still in early development and I need to merge changes in a pace that is not compatible with Linux development process. So I'm going to pause for a while now and once it reaches maintenance status I will focus on submitting patches in a healthy pace that you can digest".
As for BTRFS I think its also pretty good. Its just that I have the impression its development is guided by the needs of its sponsors and sadly for us META doesn't need RAID5.
A lot of things were tried, people did try to mediate.
The particularly galling thing though was when I finally started looking - post split - comparing bcachefs PRs to other subsystems and especially XFS - I was being more conservative with what I considered a critical bugfix.
There was never a clear statement on what the issue was. What you guys got in public was about as much as I got.
All I can say is - going fast when you're stabilizing and getting bugfixes out the door is what you can and should be doing when you've invested in test coverage, test automation, keeping the codebase clean and asserted, and building up a community that works well together on testing and shaking things out.
I genuinely do not know what they were thinking.
If you are referring to why bcachefs was removed from the Linux kernel, here's a discussion on bcachefs being removed from the Linux kernel.
They already know what was discussed.
(Good? Bad? Indifferent? I don't know and I don't have a dog in this race. I'm just here connecting the dots.)
Yes, that's why those remarks on how it's a mystery how bcachefs was pulled from the kernel are perplexing. To me they sound like gaslighting.
100%. My system is rock solid and the last thing I need is rolling the dice after every update on whether my system will boot. https://www.reddit.com/r/archlinux/comments/eywcp7/linux_551...
I'm impressed with bcachefs's accomplishments though, and if they ever reconcile with the kernel I'll surely give it a fair shake.
Actual distro support, and doing it right with people actually communicating with each other, has always been a priority for the project.
How large is large? I've deleted files with sizes of tens to hundreds of GBs and not seen that, and can probably whip up a test with a single-digit TB file if motivated.
Do you perhaps have 'discard=sync' in your mount options, or are using a kernel earlier than 6.2, which is the version -according to the docs- where async discard became the default?
for 1TB compressed (probably 5tb uncompressed) it is reproducable 100% reliably for me.
Here is some discussion: https://www.reddit.com/r/btrfs/comments/1mok440/filesystem_l...
I made a ~3TB btrfs FS and mounted it with 'force-compress', put 5TB of zeros on it (which compressed down to like 160GB), and did a delete along with some concurrent operations on that same FS. Based on what I saw, btrfs doesn't "block access for minutes" while a large delete is in progress, but access to the volume that has the delete in progress is dreadfully slow. I used vim to create a new file in the mountpoint and saw that write delays were between ten and twenty seconds. Really bad, but still functional. Not at all blocked.
Someone in that Reddit discussion that you linked to says that all btrfs filesystems hang during an extremely large delete. This is not what happens for me. The only btrfs FS made slow was the one that had the ongoing delete. I have four other btrfs filesystems mounted and they're all just fine, whether or not they're on the same physical disk that has the ongoing delete. space_cache is v2 on all of my btrfs filesystems.
On my system, it looks like an events_unbound kworker was eating 100% of a single CPU while the big delete was in progress. No other kernel threads seemed to be consistently occupied.
For fun, I re-ran the thing I document below when mounted without compression, and then with an uncompressable file when mounted with non-forced compression. I had to reduce the size of the file to 2TB for both scenarios, but omitting compression writes out like 10x the data to disk, so it still seems like a fair test.
I'm not going to the trouble to provide a transcript for those two runs, but both when mounted without compression enabled and when an uncompressable file was written to a compression-not-forced volume, I saw absolutely no delays in filesystem operations while I was deleting that 2TB file. FWIW, putting 2TB of /dev/zero on that compression-not-forced volume and deleting it behaved the same as it did on a 'force-compress' mount.
Whatever is causing the dreadful slowness is directly linked to transparent compression, rather than being something you get when you run btrfs in all configurations. "Why run btrfs if not for transparent compression?" you might ask. I would answer: "Snapshots and reflinks, and yes, I know that XFS has reflinks too.".
A lightly-edited terminal transcript follows for if you want to double-check my work up to the end of the 'force-compress' run.
# lvcreate --size 3T --name testlv --stripes=2 testvg
Using default stripesize 64.00 KiB.
Logical volume "testlv" created.
# mkfs.btrfs /dev/mapper/testvg-testlv
btrfs-progs v7.1
See https://btrfs.readthedocs.io for more information.
<extra crap removed>
# mount -o compress-force /dev/mapper/testvg-testlv /mnt/test/
# btrfs fi usage
Device size: 3.00TiB
Device allocated: 2.02GiB
Device unallocated: 3.00TiB
Device missing: 0.00B
Device slack: 0.00B
Used: 320.00KiB
Free (estimated): 3.00TiB (min: 1.50TiB)
<extra crap removed>
# dd if=/dev/zero of=/mnt/test/5TBFile bs=4MiB count=5TiB
1310720+0 records in
1310720+0 records out
5497558138880 bytes (5.5 TB, 5.0 TiB) copied, 2510.22 s, 2.2 GB/s
# /usr/bin/time --format='** fi sync %e' btrfs fi sync /mnt/test
** fi sync 0.00
# btrfs fi df /mnt/test/ | grep Data
Data, single: total=160.00GiB, used=160.00GiB
# /usr/bin/time --format='** totalTime %e' bash -c "
/usr/bin/time --format='** 20gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/20GBFile bs=4MiB count=20GiB status=none; du -h /mnt/test/20GBFile; rm /mnt/test/20GBFile'&
/usr/bin/time --format='** 10gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/10GBFile bs=4MiB count=10GiB status=none; du -h /mnt/test/10GBFile; rm /mnt/test/10GBFile'&
wait"
10G /mnt/test/10GBFile
** 10gbTime 2.42
20G /mnt/test/20GBFile
** 20gbTime 4.95
** totalTime 4.95
# date
Sun Sep 20 10:19:46 PM PDT 2026
# /usr/bin/time --format='** totalTime %e' bash -c "
/usr/bin/time --format='** 05tbTime %e' rm /mnt/test/5TBFile &
sleep 5 # It takes a few seconds for the delete to start making things slow when the file has been compressed. The operations on the 10GB file will complete in a normal amount of time if this sleep isn't present.
/usr/bin/time --format='** 20gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/20GBFile bs=4MiB count=20GiB status=none; du -h /mnt/test/20GBFile; rm /mnt/test/20GBFile'&
/usr/bin/time --format='** 10gbTime %e' bash -c 'dd if=/dev/zero of=/mnt/test/10GBFile bs=4MiB count=10GiB status=none; du -h /mnt/test/10GBFile; rm /mnt/test/10GBFile'&
wait"
10G /mnt/test/10GBFile
** 10gbTime 183.64
20G /mnt/test/20GBFile
** 20gbTime 259.47
** 05tbTime 1026.21
** totalTime 1026.22
# date ; /usr/bin/time --format='** 20gbTime2 %e' bash -c 'dd if=/dev/zero of=/mnt/test/20GBFile bs=4MiB count=20GiB status=none; du -h /mnt/test/20GBFile; rm /mnt/test/20GBFile' ; date
Sun Sep 20 10:36:52 PM PDT 2026
20G /mnt/test/20GBFile
** 20gbTime2 23.04
Sun Sep 20 10:37:15 PM PDT 2026I have 'discard=async'
I used to think that about ReiserFS, too. It was in the mainline kernel, development was snappy, and it solved some performance problems. I used it all over the place.
Things then subsequently... changed. :-/
And I did move away from it.
But I had once expected ReiserFS to be permanent, especially since this Hans Raiser dude who was driving the ship seemed to be sharp AF, and this seemed doubly-true when his filesystem got mainlined.
But it was not permanent. It did not last forever.
This idea of permanence, or rather the lack of it, was the whole of the point that I was responding to and also trying to impress upon.
At the end of the day: We do not know the future. Things can change.
I've outlived this one high-performance Linux filesystem in my life that I was using. This does not in any way mean that I will be the last to outlive other Linux filesystems.
Permanence is not guaranteed.
Except RHEL. They don’t include it in their kernels.
Alma Linux started including it again though.
It can never be easy.
Instead it got kicked out because Kent constantly ignored the kernel's contribution rules and is unlikely it will ever be accepted back into the kernel.
And it went in when it did because Redhat was pushing for it and claiming to be supportive - but that never materialized. They wanted to get something for free without investing, or putting in the absolute bare minimum.
A _lot_ of people were saying publicly and privately "dear god yes we need something better than btrfs" - but no one from the existing kernel community was interested in stepping up.
Community's still growing, though. A lot of people have gotten active in making sure bcachefs actually works well for people end to end, and there's a hell of a lot more to shipping a filesystem than just writing kernel code.
The FS was marked experimental, so there is no urgency in fixing bugs or providing features in a certain cycle. Everyone using it knows what they got themselves into. You can still provide the DKMS module for faster fixes and features for anyone who wants to use BCacheFS more seriously for the time that the upstreaming process takes, but eventually it would have all been on mainline.
Asahi is taking a similar approach where they have their downstream kernel and push things upstream once they are mature.
That means the upstream kernel is not useful for running on that hardware now, but things are moving there eventually.
All this has been discussed to death, we don't need people armchair quarterbacking a year later. It's over, it's time to move on.