Had there not been uncertainties about licensing which prevented ZFS from being in mainline there wouldn't be any Btrfs. How can you ever really trust a filesystem developed by adding features first, trying to add stability and usability later anyway?
I did think it was odd that installing the ZFS package took such a long time, but that was a one-time wait.
- RAID5/6 mode is only experimental in btrfs, while it is really stable in ZFS.
- I don't have concrete data for that, but in my experience, BTRFS has high latency (>1 second) even for small file operations when under load.
- While that should not be a problem for production systems, I have some crappy hardware where BTRFS oopses or corrupts data, while other filesystems (Ext4, ZFS) work fine.
I even had to add eatmydata support to the schroot tool as a command-prefix option to allow every command to be run via eatmydata when using btrfs snapshots. (I've since dropped btrfs support entirely; it was too unreliable with large snapshot turnover rapidly unbalancing the filesystem. Unusable in production.)
When there's only a single filesystem, and that filesystem is btrfs (or ZFS), it should however be possible to optimise this away and delegate everything to the filesystem. But even here, maintainer scripts may issue their own fsyncs as they update their own databases, kernel images or whatever.
Not if file-change notifications were supported robustly by dpkg and the kernel (to a lesser extent). Getting to that would, however, require massively restricting the compatible-kernel-versions set of dpkg, and would also probably require undoing some of the more . . . misguided pieces of history with regard to file-change notification systems in Linux.
The problem is that the system state needs checkpointing for every package state change. It must allow for recovery on failure, termination, abortion or power loss, amongst other scenarios. And the package database must remain in sync with the filesystem state.
When every managed file is on one snapshot-able filesystem, this could be rolled back atomically, and the fsyncs skipped. But as soon as you have a non-snapshot-able filesystem or multiple filesystems in use, the fsyncs can't be skipped.
While working in that specific industry it seems like there is a motivation to specifically go out of their way to make things unreliable, extra complicated and utilize bad technologies.
The fact that IBM and Toshiba on POS is frightening.
Doesn't appear to be merged yet, but it's coming.
I don't use RAID5, which is where I've heard btrfs is dangerous, just for the snapshots and dedupe (with https://github.com/jbruchon/jdupes/ )
(I work at fb)
That video's discussing btrfs layering, which is very neat.
I am _very_ curious what sort of experience FB has had with btrfs.
If FB has decided to go all-in with btrfs, you probably have the most accurate raw data there is to have.
Of course, any hyper-scale deployment of a technology is going to produce remarkable/exponential statistics, and these will (sadly/annoyingly) need careful parsing to normalize in a way that is generally accessible and unsurprising. (We live in a very knee-jerk world, and all. Sigh)
All this to say, I at least am looking forward to whatever sorts of numbers you end up being able to share.
--
After considering the bit about btrfs in the video, and your mention of "1 or more machines", I wonder if you aren't using btrfs on things like load balancer type systems - or, more abstractly, node configurations that can always safely be "thrown away", whether by hardware failure or explicit decommission. But then I realize the chances are you probably use such a model (individual hardware failure must be acceptable; everything important must be n-way redundant) because nothing else makes sense at scale. And then I wonder... exactly what point in that spectrum does btrfs fit in? I wonder if/how you can answer that question.