For the record I'm using btrfs on Arch (so recent kernel) for years with no issues (including LUKS encrypted root filesystem and RAID1 arrays for backups).
Or have you been using ZFS on the same hardware?
Btw raid5/6 is still broken on btrfs which makes it a hard sell for any system with more than 2 disks. cf. raidz on ZFS
I posted this link here already:
https://www.usenix.org/conference/atc19/presentation/jaffer
f2fs (at least in its state a couple of years ago) is/was a prime example of how a filesystem can get into a barely working state with massive amounts of data and metadata corruption, and not even notice it.
God I love this site. In case of a minor disagreement with someone don't even bother to think, just press "downvote".
Btrfs RAID1 works perfectly, and RAID1c3/RAID1c4 provides additional redundancy. In place of RAID5, use RAID10 instead.
If you want more IOPS, add more raidz2 (raid6) stripes to the pool. In practice, spinning rust is the new tape. Trying to do random access under 1MB is just silly
I don't stress over rebuilds. 2 more disks failing during a rebuild is incredibly unlikely compared to everything else that might force me to restore a backup (software bugs, data center flooding, etc).
Why can't you? Granted, you need enough disks to actually have all the data - so ex. if you did RAID0 then yes you need all disks, but say if you did a mirror you can totally just yank a disk out, attach it to another machine, and `zpool import` it.
EDIT: It'd look like this: https://serverfault.com/questions/964075/how-can-i-recover-m...
* https://utcc.utoronto.ca/~cks/space/blog/linux/ZFSSplitPoolE...
Only with mirrored drives.
Any RAID-Z level would need a full export/import as data is striped, but hot-swap drives can be pulled once things are unmount.
I agree that it can be a bit daunting to operate, there are a few footguns around that, while it might not lead to data loss, but can lead to unfortunate situations.
Just the other day someone on the mailing list had managed to add a single drive as a new top-level vdev to a petabyte pool, rather than adding it as a new spare drive, simply by omitting the word "spare" from the "zpool add" command...
That said, I've been using ZFS at home here with 6+ disks for almost a decade now, and I've never lost data despite lots of various incidents, including lots of power losses and various hardware failures (like disks, mobo and PSU). So overall I'm very happy with it.
ZFS has... a lot more. It’s just very different and the way these states fit together, and worrying about how to operate on them safely, makes me more nervous, in many ways, than less-safe file systems do. I’m sure that will pass, but it’s still not fun.
For me I found it beneficial to watch the videos on how ZFS is built up, like this one[1]. Helped putting the pieces together.
When I say "corrupted beyond repair", I mean "the btrfs tools were not actually helpful".
I use zfs on everything now. I am sure at some point it will die horribly, but for now I haven't had a single problem in ~60 managed drives across 3 machines.
Btrfs has not done a good job of inspiring any confidence, many years into development. Thankfully, Ceph has moved on from FS-backed storage to its own implementation on top of raw block devices, and I no longer have Btrfs anywhere in production.
Mind you, I don't trust ZFS either; it does seem to be stabler from Btrfs, but it still suffers from the fundamental issue that all of these "fancy" filesystems do: the fsck/repair tools are never up to par, and there is next to no chance of disaster recovery (with the added drawback that ZFS is not in-tree).
My first experience with one of these "if anything fails, all your data is gone" filesystems was ReiserFS many years ago - 8 bad sectors on a disk killed my home directory and all my data was gone. Since then, I've had rather complex accidents with ext4 and XFS* where I could do manual and automated surgery and recover ~100% of my data. Btrfs and ZFS are in the same class as ReiserFS here. The repair tools just aren't there. Sure, they handle redundancy at the device level like a fancy RAID for "well-behaved" failures like devices just disappearing, but anything outside or their model, or that tickes a bug, and you can well kiss your data goodbye.
Just to give an example: I once recovered an XFS filesystem that was built on top of a RAID6 array which, due to an unfortunate sequence of events, had one drive too many drop out during a replacement, which resulted in me manually stitching together an array where one drive had out-of-date data (i.e. every block out of N was from an earlier point-in-time from the others). Fsck fixed everything, high-level checksums took care of the few files that were being written to and had become corrupted, and I lost nothing of value. On a good filesystem, fsck does its best to recover all existing data and guarantee the result is consistent.
Yes, I know, backups. I have backups. That's not a reason to neglect repair tools. Backups are one layer of defense that can also fail; they are no excuse to neglect FS-level robustness. For example, my off-site backups are bottlenecked on my 1G internet connection, which means that if I have a weird but largely recoverable soft failure, it is much more efficient to rsync data back from the backup, using checksums to avoid data transfer, rather than copy everything again.
And this is why I use CephFS as my "smart" single-host storage solution these days. It has overhead, but it works well, is much more introspectable than ZFS/Btrfs (you can dig through the stack layers if you understand how it works very easily), and I trust its ability to recover from weird failures and device states much more than any RAID solution or fancy multi-device filesystem. It is extremely well engineered.
* I don't recommend XFS either due to kernel implementation performance issues around allocations and such; it was the cause of massive latency issues on my home server for years until I discovered its antics. But at least I've never lost data to XFS. So yeah, just use ext4 if you need a normal filesystem.
Seems like a pretty easy test to run and if it found problems, they’d be well worth fixing. (And you could do the test itself pretty efficiently on a ramdisk).
https://www.unixsheikh.com/articles/battle-testing-data-inte...
[1]: https://github.com/openzfs/zfs/tree/master/tests/zfs-tests
[2]: https://github.com/openzfs/zfs/tree/master/cmd/raidz_test (run_rec_check_impl etc)
I'm using btrfs on several systems, laptop, desktop and server, on various configurations of disks.
It has served me well for years, on the server it helped me detect a bad SATA controller. It would work perfectly in light usage, but start introducing errors in heavy usage, which made one disk inconsistent with the others in the storage pool.
Btrfs alerted me to this and after moving the disk to a good controller, I ran btrfs-check --repair on the unmounted disk (after reading the warnings), which got the FS back to a consistent state, remounted the whole pool and ran a btrfs scrub to get everything back in line with itself. The whole process did take a while, but I had backups and wanted to try out the tools. In the end there was no data loss, and the pool is still running perfectly today.
And they specifically tell you to use XFS for any production deployments.
Furthermore - btrfs feels excessively complicated for simple workflows - if I want to snapshot a btrfs volume without exposing the snapshot to the machine’s view of the file system, I have to do a bunch of volume layout setup first. With ZFS I can just snapshot.
You mean like the current advice not to use anything except mirroring and striping (RAID-0/1/10)?
> Parity may be inconsistent after a crash (the "write hole"). The problem born when after "an unclean shutdown" a disk failure happens. But these are two distinct failures. These together break the BTRFS raid5 redundancy. If you run a scrub process after "an unclean shutdown" (with no disk failure in between) those data which match their checksum can still be read out while the mismatched data are lost forever.
* https://btrfs.wiki.kernel.org/index.php/RAID56
I've been using ZFS since it came out on Solaris 10 over a decade ago and it was specifically designed not to have a write hole due to its COW/ACID nature.
See this 2008 SNIA presentation from Bonwick and Moore, the creators of ZFS talking about not having a write hole:
* ACID/COW: https://www.youtube.com/watch?v=NRoUC9P1PmA&t=24m
* Integrity: https://www.youtube.com/watch?v=NRoUC9P1PmA&t=55m20s
The typical btrfs fix is (slightly exaggerated) "When doing an A while a B is pending, re-enumerate the Cs for the purpose of D unless the E is locked, in which case, reschedule the F". Where A...F are all not trivial like snapshot, out of disk space, device rebalancing, and so on.
If that level of thinking of all the things at once is required to write correct code, well, it's evidently not happening.
either way, my story is the same as yours. I have no troubles, I have at least ~10 daily snapshots of many subvolumes I can pull from if I accidentally: `rm -rf /usr` or something silly like that. It's a great FS and unlike ZFS, is actually upstream in the Linux kernel.
If you need this functionality and want to mitigate the write hole, gen an UPS.
> Seems like a feature companies would be very interested in so I wonder why nobody's ever fixed it.
Because those driving the development have no interest in RAID5/6, they do not need it. RAID5/6 is enthusiast/SOHO feature, and these groups do not take part in development, so no wonder it is neglected.
In fact, I will combine both points: it is more effective/economical for everyone, who needs RAID5/6 with btrfs to get an UPS than to pay the manhours for the development. Additionally, it would be an expense incurred by those who need it, not by those who do the development currently and do not need the feature.