With one of the two, the filesystem was allowed to fill to 99% capacity (no automatic monitoring), so obviously that is operator error, and I'm not aware of any filesystem that can handle such situations gracefully. The system became unresponsive, with btrfs-transaction taking up an increasing percentage of CPU time. Removing files and snapshots did not see an increase in free space.
So what is super-curious is that the other server, which only ever got up to about 30% full, also started exhibiting the same symptoms: unresponsive, high load from btrfs-transaction.
I was able to mount the filesystems in read-only mode, and recover the files, and checked it against the offsite backup. So no data loss, only a service loss.
Both systems had 10 or so subvolumes, and read-only snapshots were taken three times per day for each subvolume. After 4 years, that's close to 15000 snapshots, maybe more.
I searched around, but didn't find anything particularly relevant to this issue.
Btrfs requires a regularly run administrative task via "btrfs balance". Balancing consolidates data between partially filled block groups, which will restore consumed space even if a block group is not entirely empty. Typically, one runs balance with a set of gradually increasing parameters to control the records affected (ie: bash script that starts with block groups that are more empty and proceeds towards ones with a higher percentage utilization).
In my experience with btrfs a few years ago, running a balance in some situations would result is incredibly high system wide latency.
The man page [1] does not appear to reference the free space issue directly, so I'm not sure if they've removed this need.
1: https://btrfs.wiki.kernel.org/index.php/Manpage/btrfs-balanc...
No file system should require this kind of regular hand holding. If it does, it's a bug, and needs to be reported and fixed, not papered over with user initiated maintenance tasks. It is acceptable to have utilities for optimizing the layout from time to time, but they should be considered optimizations, not requirements. If it becomes required to use them as a work around, that's both suboptimal and neat. But there's still a bug that needs a proper fix.
In the case a workaround is needed, it's recommended to do targeted filtered balances, not a full one. A full balance isn't going to solve a problem that a proper filtered balance can't.
openSuse (a distro which notably defaults to btrfs) packages these btrfsmaintenance scripts [2], and it appears they may have included it in their default install (I can't find a list of packages). Their wiki page on disabling btrfsmaintenance [5] implies that if disabled, manual maintenance (presumably via running, among other commands, some form of balance) is needed.
There's also an entry in the btrfs wiki [3] that indicates running `btrfs balance` will recover unused space in some cases. It helpfully notes that prior to "at least 3.14", balance was "sometimes" needed to recover free space in a file-system full state. (The lack of precision here doesn't inspire confidence)
Another btrfs wiki page [4] indicates that running balance may be needed to recover space "after removing lots of files or deleting snapshots".
1: https://github.com/kdave/btrfsmaintenance 2: https://software.opensuse.org/package/btrfsmaintenance 3: https://btrfs.wiki.kernel.org/index.php/FAQ#What_does_.22bal... 4: https://btrfs.wiki.kernel.org/index.php/Manpage/btrfs-balanc... 5: https://en.opensuse.org/SDB:Disable_btrfsmaintenance
Kernel 3.14 is ancient. I can't bring myself to worry about it.
I think your fourth paragraph cherry pick is disingenuous. A more complete excerpt, "There’s a special case when the block groups are completely unused, possibly left after removing lots of files or deleting snapshots. Removing empty block groups is automatic since 3.18."
Fedora 33 users will be getting kernel 5.8 from day one, and 5.9 soon after release.
I doubt my case is atypical. I haven't balanced my years old non-test real world used Btrfs file systems.
I think you could switch your concern and criticism to the fact wikis get stale.
Ah, so Fedora is limiting it's use of snapshots to avoid the need to have balances occur? Do you have some info on what level snapshot usage has to rise to before balances are needed on a regular basis? Is Fedora using snapshots at all?
There's no direct correlation between having many snapshots, and needing to balance. The once per month balance used in openSUSE is to preempt or reduce the chances of out of space error in one type of chunk (block group of extents), while there remains significant free space in another type of chunk.
Chunks are created dynamically, and are either type metadata or data (also system but it can be ignored). Different workloads have different data/metadata ratio demands, hence dynamic allocation. Snapshots are almost entirely metadata. More snapshotting means more usage of metadata.
If the pattern dramatically changes, the ratio also changes, and the dynamic allocation can alter course. Except when the disk is fully allocated. In that case, heavy metadata writes will completely fill metadata chunks, and ENOSPC even though there's still unused space in data chunks.
A filtered balance can move extents from one chunk to another, and once a chunk is empty, it can be deallocated. That unallocated space can now be allocated into a different type of chunk, thus avoiding ENOSPC. Or at least when ENOSPC happens, there's essentially no free space in either chunk type, at the same time. A "true" ENOSPC.
There are all kinds of mitigations for the problem in newer kernels. And in my opinion it's better to not paper over problems, but fixing the remaining edge cases.
This hasnt been my experience within the past few years, though I do remember it being somewhat necessary in the past. Both systems I'm running btrfs on have their free space within a few percent of the unallocated space (indicating most used blocks are fairly full and little space is wasted).
I would assume that this is not a factor with RAID-1, where all system, metadata, and data is duplicated. I could see this as being very important for the higher RAID levels.
We have since stopped deploying any sort of btrfs RAID (even RAID-1), and have gone back to using Linux MD.
https://ohthehugemanatee.org/blog/2019/02/11/btrfs-out-of-sp...
If I could mount the failed btrfs filesystem read-write for an appreciable length of time, I could try the balance operation.
I'll need to set up a cron job on the other systems to run balance on a regular basis alongside scrub (obviously not at the exact same time).
On a slightly unrelated note, I once suffered a complete failure of BTRFS where after a shutdown it just wouldn't mount anything again. Interestingly, on IRC, I was told that his is because the firmware on my Samsung NVME SSD was buggy. It might be, but ext4 has not failed me once in that regard.
An inability to mount Btrfs means metadata has been hit by some kind of corruption, and there are more structures in Btrfs that are critical. If they're hit, you see mount failure. And repair can be harder on Btrfs as well.
But it also has a lot more opportunities for recovery that aren't very well understood or discussed. In part because very serious problems like this aren't that common. And also, quite a lot of people give up. Maybe they have backups and just start over with a new file system and restore. Or through no fault of their own they don't persevere with 'btrfs restore' - which is a very capable tool but requires specialized knowledge right now to use it effectively.
One of the things that'll take a mindset shift is the idea of emphasizing recoveries over repairs. One improvement coming soonish (hopefully end of the year) is more tolerant read-only rescue mount option, making it possible for users to recover with normal tools rather than 'btrfs restore'.
Which would be Raid5/6 and 1