I then had to manually delete the file with the I/O error (for which I had to resovle the inode number it barfed into dmesg) and try again - until the next I/O error.
(I'm still not sure if the disk was really failing. I did a full wipe afterwards and a full read to /dev/null and experienced no errors - might have just been the meta-data that was messed up)
Over the last one or two years I've experienced twice a checksum mismatch on the file storing the memory of a VMWare Workstation virtual machine.
Both are very likely bugs in Btrfs, and it's very unlikely that have been caused by the user (me).
In the relatively far past (around 5 years ago), I've had the system (root being on Btrfs) turning unbootable for no obvious reason, a couple of times.
Despite not being directly used, these blocks are kept (and cannot be reused) because another part of the extent they belong to is actually used by files.
This can happen if a large file is written in one go, and then later one block is overwritten - btrfs may keep the old extent which still contains the old copy of the overwritten block.The question was around usage, because without knowing people's usecases and configurations it'll never be usable for you while working fine for others.
As usual with all these Linux debates, there's a loud group grinding their old hatreds that can be decade old.
If 0.1% of users say it corrupted for them, and then don't provide any further details and no one can replicate their scenario then it does make it hard to resolve it
there's a feedback effect: if users know that a filesystem takes these kinds of issues seriously and will drop what they're doing and jump on them, a lot of users will very happily spend the time reporting bugs and working with devs to get it resolved.
people don't like wasting their time on bug reports that go into the void. they do like contributing their time when they know it's going to get their issue fixed and make things better for everyone.
this is why I regularly tell people "FEED ME YOUR BUG REPORTS! I WANT THEM ALL!"
it's just what you have to do if you want your code to be truly bulletproof.
No, Kent, they are not. Posting attacks like this without evidence is cowardly and dishonest. I'm not going to tolerate these screeds from you about people I've worked with and respect without calling you out.
Every time you spew this toxicity, a chunk of bcachefs users reformat and walk away. Very soon, you'll have none left.
It doesn't matter if you want to "tolerate" it if every time there's a filesystem thread there's stories about lost filesystem and unfixed bugs, and the rare times people try to report bugs or issues they go nowhere.
You're ranting about reality, and somehow I doubt you were ever a user of mine.
I've worked hard to establish an engineering community where no one has to be afraid to point out broken shit, including and especially in my code, and I wouldn't want your attitude anywhere near it.
Absolutely. Demonstrably much moreso than you over the past decade.
> I've worked hard to establish an engineering community where no one has to be afraid to point out broken shit,
You have failed. Look at this conversation: https://lore.kernel.org/lkml/fe51512baa18e1480ce997fc535813c...
Do you think Arnd, David, and Russell feel like they were rewarded for pointing out your broken shit? I doubt it. I certainly don't. You were so completely and obviously in the wrong there I have difficulty believing it was real.
I'm not replying to you again. Good luck :)
Not comparable to losing filesystems, sorry...
Also, filesystems just work. Nobody is gonna say "oh I'm using fileystem X and it works!" because that's the default. So, naturally, 99% of the stuff you'll hear about filesystems is when they don't work.
Don't believe me? Look up NTFS and read through reddit or stackexchange or whatever. Not a lot of happy campers.
I don't.
It's also the reason I am completely against the Debian "backporting" methodology to pretend they aren't using the new version of something.
Debian doesn't "backport" anything, they ship upstream LTS kernels straight, so please stop spreading this misinformation.
For some bizarre reason I can't conceive people not only believe but also keep complaining that Debian ships franken-kernels like RHEL does. See [0] for another instance.
Is it wrong to ask how to reproduce an issue?