At some point after I restarted the computer and it wouldn't mount the partition saying that the path for mounting was already being used (even after restarting??), but the disk was not mounted. Searching around apparently there is some kind of bug where Linux will "cache" the BTRFS UUID [1].
At the time, the "solution" I read was to change the UUID of the BTRFS partition running btrfstune. Which I did... and it supposedly changed successfully. Except that after trying to mount the partition again, it will tell me that there was an error in the partition: the tree had mismatching UUIDs :-/. I tried updating the UUID several times, and the command ended in success, but the error remained. My BTRFS disk was officially broken...
After spending several hours trying to fix the issue following the BTRFS documentation (pretty shitty TBH), in the end, I ended up having to do `btrfs restore -iv ...` to extract the data from the disk into some other external disk (formatted NTFS this time!!). The command is still going, about 2 weeks later and I am slowly recovering my 5TB of data.
But the one thing clear to me is that I don't trust BTRFS or it's admin commands AT ALL after this experience. Once I finish recovering my data, I'll nuke the partition, format the disk in NTFS and forget about this sour experience.
[1] https://unix.stackexchange.com/questions/603528/why-am-i-get...
It's not a bug. It's just a fact that every BTRFS filesystem visible to the kernel (mounted or not) should have a unique ID. Otherwise operations on one filesystem (including mounting the partition that contains it) can end up applying to the other filesystem.
>At the time, the "solution" I read was to change the UUID of the BTRFS partition running btrfstune.
The solution would've been to find out why you have two partitions with the same filesystem ID and remove one of them. One example is that if you have the BTRFS filesystem on an LVM partition and then you clone the LVM partition, you end up with two BTRFS filesystems with the same ID. I accidentally encountered it while converting an unencrypted BTRFS parition into an encrypted one by dd'ing it into a new LUKS device - when I tried to mount the encrypted partition it ended mounting the unencrypted one.
Since you say it's an external USB drive, perhaps it disconnected and reconnected uncleanly and the kernel thought there were two of that drive connected at the same time.
Maybe it's not a code bug, but imposing global uniqueness requirements where you can't just `cp /dev/blkdev1 myfs.img` and be able to work with the new image seems rather like a design bug.
While not a multi-device filesystem, I believe XFS has the same behaviour and you have to change the UUID (or use xfs_copy in the first place) before you can mount a cloned block device.
(I wonder how ZFS handles this?)
I think the actual dynamic filesystem discovery is in user space and kernel just gets a list of devices, so it should be even easier to change that.
AFAIK the filesystem UUID, the device UUID, or both are sprinkled all around the filesystem, I suppose for reasons of identifying data belonging to the fs in case of fsck. According to my quick research device id cannot be changed. I suppose if it could be changed by some means, then btrfs rescue clear-uuid-tree could be used to fix the other uuids.
It doesn't even get much better with many drives and avoiding the docs/tooling
BTRFS arrays will consistently corrupt with my reset button -- RAID10 on gen4 NVMe drives... while LVM/dm-raid + traditional file systems are absolutely fine
YES!!! Sorry to beat a dead horse, but it has been so frustrating for me. I thought I did something wrong, like, how could the partition break for doing nothing! Where did I fucked up? I cannot imagine the experience with a RAIDed BTRFS array shudders
Sure... HDDs shouldn't do that, but it's literally the job of a BTRFS raid to protect me from such issues. The data was all still intact on the other drive. I could see it with dd, but BTRFS refused to read it.
[1]: https://unix.stackexchange.com/questions/634050/in-what-ways...
However, btrfs is a very fragile filesystem and any inconsistencies caused by the drive bugs tend to lead to it breaking and refusing to mount, whereas on other filesystems you'll just get a warning in dmesg.
It doesn't help that btrfs developers refuse to work around drive bugs.
The entire selling point of the filesystem is that it's resilient to data corruption. Drive bugs are just another class of data corruption.
In theory perhaps, but absolutely not in practice. Making things atomic is tricky.
I know you don't want to get asked this, and I wouldn't if you'd just said "I tried this once a while ago and it happened", but if it's consistent... have you filed a bug?
It's occurred to me, but I expect a fair bit of push back; 'you abuse the array and expect it to... what?'
I'm not really interested in debating that, and it's remarkably easy to get hung up on. I've floated this issue a few times unofficially and it's a constant stickler
I know expecting coherency in this situation is a little silly, but BTRFS is notably less reliable/robust than the alternatives
All in all... I'm avoiding a situation that may or may not occur, by not participating.
Not saying it's a good thing, but it's easier for me to just use What Works
I'd be kind of annoyed if I went to the trouble of reporting a bug and got that response. But I'd also not assume it would end that way...
The DRAM caches are pretty hefty compared to metadata, so I wouldn't be surprised - though that's about my extent of understanding.
Even then, though - other filesystems are certainly more durable, dealing with whatever these drives are doing.
FWIW, this was with four Sabrent Rocket 4.0 Plus drives in RAID10 with 4K sectors. I'm not sure if these are particularly well known one way or another for deception/trickery
They hold up just fine with either dm-raid or native LVM raid10 under more traditional filesystems (I tested EXT4/XFS, using the latter consistently)
I also used to run btrfs in btrfs-RAID10 configuration until apparently a flapping SATA link and fsck attempts were able to break the fs completely. Full system backups were great that day. I run https://kopia.io/ nowadays every three hours during day time and I've been quite happy with it.
Nowadays I run bcachefs.. Backups are still handy :).
I suppose the reason why you chose NTFS was to be able to access the data from Windows, at least in case of emergency? Because there are a lot of filesystems that are presumably more mature than NTFS is for Linux.
I thought of going EXT4 , but as you said, NTFS is more widely supported.
> Since you say it's an external USB drive, perhaps it disconnected and reconnected uncleanly and the kernel thought there were two of that drive connected at the same time.
That will read all of the data from the storage device(s) checking for coherency. In the case of redundancy (ie: RAID1) and corruption/rot, would use the other copy to make things whole
Beyond that and trim, most of what's in here is fairly specific. ie: avoiding ENOSPC or dealing with array changes
edit: I think running this more than ~monthly is overzealous... and only really meaningful if you have redundancy
If memory serves, the checksums are validated when you read things anyway - so I question doing passes too aggressively. I'll accept a bit for bitrot
Both SSD's and spinning metal drives have ECC bits stored with the data, and the drive can detect when the data was read easily or required lots of error correction applied. And then based on that they can make the optimal call of how often to read data to see it it is 'nearly rotted' and needs a rewrite.
The filesystem has no knowledge of any of that, so has to do dumb periodic scans.
I thought trim was enabled by default (https://wiki.archlinux.org/title/Btrfs#SSD_TRIM). Does fstrim do something more than that or is it the same thing?
Unbalanced drives can affect performance as one drive will have to do more of the work often bottlenecking. But except for extreme cases it is probably a rounding error.
One of the effects of balancing the tree is to release unused space from the allocation pools to general availability, solving that problem.
It would, if you were running a well-designed/maintained operating system. Unfortunately many Linux distributions make poor decisions.
It very well may be! Depending on your distribution... they may already have similar services/timers for periodic scrubbing
On Fedora they aren't provided; but are with Arch
Unless you run an array it's fairly meaningless, beyond being an administrative-reporting tool. It can't fix failed checksums without mirrors or parity
I think the checksums are validated on read anyway, so it's mostly to mitigate bitrot - very specific to 'glacial' storage
If it would actually affect performance, turn on autodefrag. Be aware that running a manual defrag, instead of using the mount option, will break reflinks.
[1]:https://www.phoronix.com/news/Btrfs-Discard-Tuning-Linux-6.3