Bcachefs: “the COW filesystem for Linux that won't eat your data”
bcachefs.org
bcachefs.org
Apparently they were not aware of the fact that bcache was a) GPLd code, or b) developed before the company existed, first as a hobby project and then at Google. After a couple of years, they noticed that Kent was in fact posting the bcache source code on his personal web site. At this point they fired him and threatened to sue. I quit the company then (along with a number of other people, for mostly unrelated reasons, such as the fact that the CTO was a notorious brogrammer). Kent got a litigator and when it was made very clear to them that they had no case, they backed down, but not before wasting a ton of money.
As far as I know, they're still actively violating the GPL by shipping a product containing modified kernel code in it without releasing the source, nor do they acknowledge that they did not develop the key component of their product.
The "commercial" version had a rather broken and messy snapshots implementation and had diverged a bit from the open source bcachefs at that point, mostly because snapshots were poorly implemented. It's also kind of funny because after we left the company we still knew of some tricky data corruption bugs, and it's likely they're still there in the "commercial" version, because backporting the latest fixes would be non-trivial and I don't think their testing or development methodology would have caught them.
Anyway, I gave up on startups and enterprise storage after this, but Kent is still developing bcachefs on his own time and money, so if you use it please consider donating some money to support its development.
Sad given how much a modern filesystem would help Linux : (
Maybe he can get extra money from a lawsuit?
From 2012-2014 it was mostly breakage every other month. From 2014-2016, it was semi-annual issues.
For the last ~18 months I have had ~30 machines running btrfs with no issues, some servers, some personal computers. The release notes are boring, the bugs are boring, and to me its definitely in a state I would strongly consider trusting it with any workload.
I worry that btrfs is just going to remain doomed. It wasn't stable half a decade ago, so it - for some reason - cannot be more stable now. But it has seen so much work put into it to make it as mature as it is now, and in my experience it is pretty damn mature now. All I want to see is another year and a half of perfect stability before I would start arguing to drop zfs entirely.
bcachefs seems to have a much more coherent design.
[1] "Oh, yeah, I don't know how to handle this code path yet, let's stick a BUG_ON in there! I'm sure we'll figure something out later."
Status page says so: https://btrfs.wiki.kernel.org/index.php/Status
- RAID 1 with more than 2 disks is not what you think it is: the data will be mirrored but only once, no matter how many disks you have (meaning if you have a mirror with 3 disks, you only have 2 copies of your data). Because in BTRFS lingo, RAID 1 means '2 copies of the data' https://btrfs.wiki.kernel.org/index.php/FAQ#What_are_the_dif... which is not what people expect from RAID 1 with more than 2 disks
- RAID 1 always needs 2 disks to be working, if not you can't mount the thing... Well you can, but only once... https://btrfs.wiki.kernel.org/index.php/Gotchas#raid1_volume...
- RAID 10 inherits these special RAID 1 cases as a result
Most of that stuff is described here: https://btrfs.wiki.kernel.org/index.php/Status
So as they say on the status page: "mostly ok"
My experience has been mixed, but haven't had any data loss. There was a bug for a while regarding free space so occasionally the system would seem to be full but wasn't ....and it was a real pain to correct it.
I now have a cron job that does a monthly btrfs balance along with a mount -oremount,clear_cache. I also run the latest kernels from http://elrepo.org instead of the CentOS7 kernels so that I get the latest patches.
I remembered thinking DKMS might solve the problem, but I ended up having to use recovery media just to get an environment to reinstall an older kernel and let DKMS do its thing after a botched update started provoking panics. I suspect a version mismatch based on the errors but never investigated it beyond fixing the problem and moving to the prebuilt modules. Things may have changed, but the Arch ZFS+DKMS packages were a bit flaky and required some manual modification just to boot (should've taken this as a warning!).
Granted, it was my fault entirely for being a bit too enthusiastic with ZFS on Arch. To be honest, if I were to use it again, it would be on FreeBSD. Not Linux. I recognize it's fine for other people, but in my use case it wasn't.
It is more like Arch Linux is a disaster. Upgrading the kernel package replaces the current one! Come on, any distribution worth it's salt just installs new versions alongside and you can select any of them in the boot screen. This is a ridiculous packaging policy regardless of ZFS or any other DKMS modules.
I will, however, agree that having no fallback to the prior kernel version is a problem. In practice, it's never caused me much trouble except when I do something stupid like using ZFS from the AUR. initrd generation has historically seemed to be more problematic under Arch, but I'd argue that's mostly fixed with install hooks.
In all honesty, it was probably more the fault of the zfs-dkms packages than it was either the kernel packaging policy or ZoL+DKMS itself (for reasons I elaborated on in my original post).
But, that's also what you get when you use packages from the AUR or using a distro like Arch for something that really only benefits from a wider installation base (like Ubuntu does, for instance).
I do agree. There are circumstances where Arch's packaging is brain dead (they only recently, within the last 2 years or so, started validating packages against signatures!). I use it for a number of applications, and as my desktop OS among others. However, I'll freely admit at least part of my choice is perhaps the fault of masochistic tendencies. After all, I migrated to Arch from Gentoo, and I used Gentoo for years! :)
In all honesty, I've been bit more by the initrd and mkinitcpio's failings than the lack of a fallback kernel. That's mostly fixed with packaging hooks that essentially guarantee it will run, but it's still a problem with the ZFS packages and may require running it manually (which is annoying). However, that wasn't always the case, and sometimes the generated initrd would be missing something important. You can imagine what happened next.
So far it's been running absolutely great for me on several Ubuntu installs.
I'm not sure I'd be brave enough to run ZoL again, but given Ubuntu's FAR wider install base and availability of binary packages, it's the better option if you have to choose.
My personal preference would be to stick with ZFS on FreeBSD. Performance is probably better.
The other problem is that at the time, the ZFS packages for LTS were pinned at a version that had a known issue with arc_reclaim encountering a deadlock essentially causing the file system to become unresponsive after a substantial transfer (think rsync).
Now, obviously, it wouldn't be that difficult to modify the PKGBUILD to pull a newer version of ZFS, but there's a point in time where the maintenance required to update starts to outweigh whatever benefit you can glean from the LTS kernel.
That's not the case now since the LTS packages appear to be at v0.6.5.9, which has the fixes, but I don't remember this being true about a year ago.
ZFS does seem to work better overall, but I wouldn't call either filesystem great at this point in time.
I believe there was a bug in the clear space cache. This could cause the system to think it didn't have free blocks... you'd have to mount another device to create more space in order to rebalance and fix.
Eventually I saw a bug fix report about a corruption in the cache... I never investigated to see if my current kernel has the fix.
Also been using it as / since 2012 with no issues.
Makes me wonder if anyone's tried something like Jepsen for filesystems.
I use it pretty much every day.
It is ok for a single disk FS. It is no ZFS though which is its largest problem, people keep marketing it as "linux awnser to ZFS"
No, it is not.
Maybe some day, but today it is no ZFS. I love zfs..
Further the ZFS utilities are far easier to use and understand. zfs and zpool commands well documented, and intuitive. btrfs utilities are are not, IMO
I am fine with using btrfs as a replacment for ext on my OS drive, but for my large data arrays of multiple disks it s ZoL all the way
I moved back to EXT4 and it never happened again since then.
Snapshots are the #1 feature of COW filesystems. I've been using them for a bit in btrfs and this feature is game-changing (and no, it hasn't eaten my data yet).
https://btrfs.wiki.kernel.org/index.php/Status
The problem areas are mostly RAID and exotic features. RAID can be handled by a different layer and most users don't really need the exotic features.
Judging from the media silence in the last months I'd say either people stopped using btrfs or it just about works good enough for everbody.
I'd say that for the OpenSuSE folks, btrfs falls squarely into the "good enough" category.
* https://blog.pcbsd.org/2012/07/9-1-feature-multiple-boot-env...
* https://blog.pcbsd.org/2013/06/pc-bsd-status-update/
* https://www.ixsystems.com/blog/a-closer-look-at-the-changes-...
* https://www.ixsystems.com/blog/the-revamped-life-preserver/
3rd option: there aren't really much people who ever used btrfs.
from the feedback I could gather when I enquired about it, it's 33/33/33:
33% who tried and say it's fine
33% who tried, hit a few bugs and stopped
33% who say it's known for not really being finished, not a good idea to go for it.
I wouldn't want to be the poor sucker supporting a large BTRFS array.
None of that virtual block device or driver-level software junk, let alone hardware RAID controller solutions.
That isn't to say that using high performance block storage isn't still a win even when the redundancy is multiplied at a higher level. The higher level redundancy is also about colocating more data with the code - i.e. it's not just redundant for integrity, but to increase the probability it's close to the code.
Even virtual memory for that matter. Now ancient concept:
I thought defrag (or things like defrag) was the most complicated thing to implement ?
ZFS has always had snapshots but I am told defrag is a long, long ways away ...
Apparently copygc is off right now because reasons, though (I'm going to assume it's almost certainly the related extent/compression issue that's holding this up from being enabled, which you can see referenced on the home page, at the bottom).
In parallel I see XFS as the long term evolution for Linux file systems. It will continue to scale slightly up from where it sits today and address fail in place, flash, metadata checksums, snapshots etc where total storage management is done by overlays like HDFS, object stores, etc.
- For non-business users who want a RAID, ZFS is too inflexible. You can't add or remove disks to a RAIDZ vdev. If you want the space efficiency of RAIDZ, you have to expand your array in units of entire vdevs. If you want replicas, you have to expand in at least pairs of disks. BTRFS and bcachefs both allow you to replicate more flexibly and reshape your array.
- ZFS doesn't work particularly well with SSDs as caches. ZIL and L2ARC are nice but they're not as nice as a full bcache-style tiering setup. bcachefs tiers let you do crazy things like a 4-tier storage setup with Nearline HDD -> 15k SAS HDD -> SATA SSD -> NVMe SSD.
- ZFS is pretty complex to manage in general and major features like ZIL and L2ARC are arcanely documented. So far, bcachefs is pretty straightforward to use.
This isn't a simple filesystem project, but plays in the next-gen space ZFS opened up. There will be a lot to do, especially IO scheduling, RAID safety with shitty drive firmwares, consistency guarantees with fsync/partial flushes etc.
I'm pessimistic about it being mainlined in the near future, the core team will be weary of a second btrfs.
What I would like to see is a APFS/exFAT crossover with COW and data checksums without all the volume mgmt with ports for all possible operating systems so everyone can use it for their SDcards, usb-sticks and external drives without making tradeoffs and using fuse.
The fact that the raidz volume is not an opaque block device allows ZFS to be aware of data corruption when comparing checksums and self heal if the data can be re-constructed from the array.
I'm not saying any attempt at a new filesystem should have to bundle the two layers together, but they should allow for communication between the abstractions.
+1. Filesystems without bit-rot protection on flash drives are going to become at least as big a problem as optical disc rot.
What's the problem with fuse? It allows sharing code between Linux, OS X, (Free)BSD and even Windows (via dokan).
Yes, it will not offer you the same performance as an in-kernel driver (due to context switches), but given that CPU power always increases, no big problem there.
This might be the case if you're running something incredibly easy on I/O like large sequential read/writes, but if you do anything at all challenging on I/O like opening desktop applications (Photoshop, lots of random reads), editing or viewing high bitrate video (very high throughput) or god forbid running a database, this is a huge problem.
1. Only available on Android when rooted.
2. Support varies between OSes. For example OpenBSD's FUSE does not have the default_permissions/allow_other flags, which makes for example encfs (and any other virtual filesystems that are backed by multiple files) a pain to use since OpenBSD 6.0 removed user mounting.
2. most non-fuse filesystems won't be ported to your BSD of choice anyway
The point is a lot of the things you bring up are already covered by bcache. Bcachefs "just" adds a filesystem layer on the bcache tree structure.
Or that someone sits down and does the reverse engineering work.
If you think we need an alternate effort and/or competition to build an advanced, native filesystem for Linux (I do), please consider a subscription on Patreon (https://www.patreon.com/bcachefs). Kent has a long history of shipping sophisticated, high-quality code.
I suspect many have lost patience with the promise of COW and unfortunately for bcachefs this history will cast a shadow on its development and potential.
Database performance remains problematic on COW and while things like snapshots and adhoc disk and volume management are interesting even exciting one soon realises unless one has a pressing need they are just nice to have. Eventually boring ext4 ticks all the boxes and one may as well forget about the fs and focus elsewhere.
The fact that some COW filesystem perform poorly does not mean all COW filesystems do.
Funny, but every so often I wonder what it might be like in a parallel world where Apple bought Sun instead of Oracle.
[1] http://macoverdrive.blogspot.com.au/2008/10/using-zfs-to-man...
For example, if you have a dirty page in a cgroup, and the cgroup OOMs, the kernel will trigger writes. If any of these writes require memory allocations, they'll probably fail since the current cgroup is OOM. ZFS subsequently gets stuck in an infinite loop, and locks up. See: https://github.com/zfsonlinux/zfs/issues/5535
I understand that a lot of ZFS works comes from LLNL & government funding. I'm not blaming them, as it works for their use case of machines that are running dedicated, controlled workloads.
We're experimenting with Btrfs, and we'll see how it goes.
There are reasons to still want it, despite its newness; for example, the latest updates bring huge improvements in metadata efficiency (low metadata overhead -> more metadata in the cache -> larger working set). Someone on the IRC channel reported it's somewhere around 20x faster than most filesystems when it comes to "iterate millions of files recursively", blowing everything else out of the water. (This seems somewhat synthetic, and I'd say it mostly is -- but OTOH, "tons of files in a directory" being really slow is life, and has bitten me multiple times in a prior job). In general, improved metadata efficiency helps everywhere, though. For example, if you're doing backups on a really big filesystem recursively, you'll have to traverse the metadata inodes a lot to get e.g. last modified time. bcachefs will likely do awesome here in terms of performance.
Another unique feature I recall is the fact it has very very good tail latency -- bcachefs almost never blocks on I/O unncessarily, so you do not get random 'lag spikes' when things like the page cache get flushed out (which may halt some other I/O ops). This makes the system feel much more consistent, in general.
There's lots of good info in the architecture document and Patreon posts from Kent:
https://gist.github.com/liloman/d525131fab9b9a440140905921e9...
I'll give it a try on bcachefs. :)
The script needs a 512MB spare disk partition and some basic changes but the fundamental work is there.
I do however see some big red flags in the linked page:
> Starting from there, bcachefs development has prioritized incremental development, and keeping things stable, and aggressively fixing design issues as they are found
Which is it? Big design changes or stable FS?
I wonder why Hans Reiser doesn't take up filesystem work again ? He has plenty of time on his hands ...
Guess that's probably the answer.
I imagine ultimately TRIM will be supported, though (I don't see a reason why it wouldn't be, and considering Kent is focused on hammering out the design I imagine it'll inevitably fit in well).
from site:
> Bcachefs is not yet upstream - you'll have to build a kernel to use it.
> Snapshot implementation has been started, but snapshots are by far the most complex of the remaining features to implement -
Yes. Very mature.