I'm reading through the links, and at first I thought this was a historical problem that was last reported in 2016/2017, but then gems like this from mid 2020 are popping up:
"When 'btrfs scrub' is used for a raid5
array, it still runs a thread for each disk, but each thread reads data
blocks from all disks in order to compute parity. This is a performance
disaster, as every disk is read and written competitively by each thread"
I have no words. Why would anyone ever think that this is a good idea? Who sat down at their computer and typed the code that does this!? How was this never tested?
You know what, thinking about it, I do actually have some rather choice words to describe the situation:
This boggles the mind to a level that requires further explanation, because the casual observer would likely fail to grasp the enormity of the failure that has occurred here. This isn't like, "oops, I forgot to up-shift gears in my car when going on the onramp", this is more like "the pilot forgot about the flaps after takeoff and the plane ran out of fuel.". There's a fundamental difference in the expectation of quality between, say, a random command line utility and a RAID filesystem.
To give some context: BTRFS was developed largely concurrent with, and in direct competition to Sun's ZFS. Unlike all previous SAN arrays, RAID cards, and filesystems, ZFS was explicitly designed for reliability. Sun famously had a 'test rig' where they abused each new build to death. Physically pulling disks. Randomly corrupting blocks. Running multiple operation types in parallel, while pulling disks. That kind of thing.
When I read ZFS whitepapers, I was amazed at how many fundamental flaws in RAID integrity protection they discovered, and then solved. Rigorously.
Meanwhile, BTRFS literally says, in 2020: Don't trust is, especially not for metadata, or data, or while scrubbing, which you had better baby-sit, otherwise say goodbye to your production environment!
More fun quotes:
- plan for the filesystem to be unusable during recovery.
- be prepared to reboot multiple times during disk replacement.
- btrfs raid5 does not provide as complete protection against
on-disk data corruption as btrfs raid1 does.
- scrub and dev stats report data corruption on wrong devices
in raid5.
- scrub sometimes counts a csum error as a read error instead
on raid5.
- errors during readahead operations are repaired without
incrementing dev stats, discarding critical failure information.
This is not just a raid5 bug, it affects all btrfs profiles.
You'd have to be nuts to use BTRFS for RAID 5 or 6, and I would question its use for any form of RAID.
PS: To the people downvoting this, please explain how you like people to be uninformed about catastrophic data corruption going ignored for 4 years below in the comments.