Btrfs on Zoned Block Devices
lwn.net
lwn.net
I've played with device-managed SMR drives and decided I hate them. Too much I/O happening behind my back, so a low workload they can sustain before hitting a horrible performance cliff. Also, in my very limited sample, poor reliability even if you stay within the rated workload. (2/2 drives that were in continuous use started accumulating bad sectors within a couple years.)
In theory, host-managed SMR drives with the right software might have much less write multiplication. The NVR software I'm working on more or less uses them as a big ring buffer (or several big ring buffers, one per stream) which seems ideal. But I suspect getting host-managed SMR to work optimally isn't worth the effort unless you have datacenters full of drives.
So no, a single bad bit would only affect a single sector. And in fact, a single bad bit wouldn't do any harm, as all drives use error correcting codes that can handle a lot more than a single bit going bad. You'd actually lose a whole sector at time (512 bytes or 4096 bytes, depending on the drive).
This means that when writing data it overwrites neighboring tracks, and hence writing is partially overlapped like shingles.
Apart from that they're more or less the same as normal drives as far as I know, firmware aside. So a single bad bit shouldn't affect more than in a non-SMR drive.
I based my post primarily on what I recalled from this talk[1] by a HGST engineer.
They exists, and they are called host-aware SMR drives. Sadly, device-managed drives exists because drive manufacturers want to both achieve cost-cutting and increased profits (and was relatively unknown until the whole WD Red (and companions) DMSMR brouhaha last year forced every hard disk manufacturer to state if the drive is DMSMR or not).
For the record I'm using btrfs on Arch (so recent kernel) for years with no issues (including LUKS encrypted root filesystem and RAID1 arrays for backups).
Or have you been using ZFS on the same hardware?
Btw raid5/6 is still broken on btrfs which makes it a hard sell for any system with more than 2 disks. cf. raidz on ZFS
I posted this link here already:
https://www.usenix.org/conference/atc19/presentation/jaffer
f2fs (at least in its state a couple of years ago) is/was a prime example of how a filesystem can get into a barely working state with massive amounts of data and metadata corruption, and not even notice it.
God I love this site. In case of a minor disagreement with someone don't even bother to think, just press "downvote".
Btrfs RAID1 works perfectly, and RAID1c3/RAID1c4 provides additional redundancy. In place of RAID5, use RAID10 instead.
If you want more IOPS, add more raidz2 (raid6) stripes to the pool. In practice, spinning rust is the new tape. Trying to do random access under 1MB is just silly
I don't stress over rebuilds. 2 more disks failing during a rebuild is incredibly unlikely compared to everything else that might force me to restore a backup (software bugs, data center flooding, etc).
Why can't you? Granted, you need enough disks to actually have all the data - so ex. if you did RAID0 then yes you need all disks, but say if you did a mirror you can totally just yank a disk out, attach it to another machine, and `zpool import` it.
EDIT: It'd look like this: https://serverfault.com/questions/964075/how-can-i-recover-m...
* https://utcc.utoronto.ca/~cks/space/blog/linux/ZFSSplitPoolE...
Only with mirrored drives.
Any RAID-Z level would need a full export/import as data is striped, but hot-swap drives can be pulled once things are unmount.
I agree that it can be a bit daunting to operate, there are a few footguns around that, while it might not lead to data loss, but can lead to unfortunate situations.
Just the other day someone on the mailing list had managed to add a single drive as a new top-level vdev to a petabyte pool, rather than adding it as a new spare drive, simply by omitting the word "spare" from the "zpool add" command...
That said, I've been using ZFS at home here with 6+ disks for almost a decade now, and I've never lost data despite lots of various incidents, including lots of power losses and various hardware failures (like disks, mobo and PSU). So overall I'm very happy with it.
ZFS has... a lot more. It’s just very different and the way these states fit together, and worrying about how to operate on them safely, makes me more nervous, in many ways, than less-safe file systems do. I’m sure that will pass, but it’s still not fun.
For me I found it beneficial to watch the videos on how ZFS is built up, like this one[1]. Helped putting the pieces together.
When I say "corrupted beyond repair", I mean "the btrfs tools were not actually helpful".
I use zfs on everything now. I am sure at some point it will die horribly, but for now I haven't had a single problem in ~60 managed drives across 3 machines.
Btrfs has not done a good job of inspiring any confidence, many years into development. Thankfully, Ceph has moved on from FS-backed storage to its own implementation on top of raw block devices, and I no longer have Btrfs anywhere in production.
Mind you, I don't trust ZFS either; it does seem to be stabler from Btrfs, but it still suffers from the fundamental issue that all of these "fancy" filesystems do: the fsck/repair tools are never up to par, and there is next to no chance of disaster recovery (with the added drawback that ZFS is not in-tree).
My first experience with one of these "if anything fails, all your data is gone" filesystems was ReiserFS many years ago - 8 bad sectors on a disk killed my home directory and all my data was gone. Since then, I've had rather complex accidents with ext4 and XFS* where I could do manual and automated surgery and recover ~100% of my data. Btrfs and ZFS are in the same class as ReiserFS here. The repair tools just aren't there. Sure, they handle redundancy at the device level like a fancy RAID for "well-behaved" failures like devices just disappearing, but anything outside or their model, or that tickes a bug, and you can well kiss your data goodbye.
Just to give an example: I once recovered an XFS filesystem that was built on top of a RAID6 array which, due to an unfortunate sequence of events, had one drive too many drop out during a replacement, which resulted in me manually stitching together an array where one drive had out-of-date data (i.e. every block out of N was from an earlier point-in-time from the others). Fsck fixed everything, high-level checksums took care of the few files that were being written to and had become corrupted, and I lost nothing of value. On a good filesystem, fsck does its best to recover all existing data and guarantee the result is consistent.
Yes, I know, backups. I have backups. That's not a reason to neglect repair tools. Backups are one layer of defense that can also fail; they are no excuse to neglect FS-level robustness. For example, my off-site backups are bottlenecked on my 1G internet connection, which means that if I have a weird but largely recoverable soft failure, it is much more efficient to rsync data back from the backup, using checksums to avoid data transfer, rather than copy everything again.
And this is why I use CephFS as my "smart" single-host storage solution these days. It has overhead, but it works well, is much more introspectable than ZFS/Btrfs (you can dig through the stack layers if you understand how it works very easily), and I trust its ability to recover from weird failures and device states much more than any RAID solution or fancy multi-device filesystem. It is extremely well engineered.
* I don't recommend XFS either due to kernel implementation performance issues around allocations and such; it was the cause of massive latency issues on my home server for years until I discovered its antics. But at least I've never lost data to XFS. So yeah, just use ext4 if you need a normal filesystem.
Seems like a pretty easy test to run and if it found problems, they’d be well worth fixing. (And you could do the test itself pretty efficiently on a ramdisk).
https://www.unixsheikh.com/articles/battle-testing-data-inte...
[1]: https://github.com/openzfs/zfs/tree/master/tests/zfs-tests
[2]: https://github.com/openzfs/zfs/tree/master/cmd/raidz_test (run_rec_check_impl etc)
I'm using btrfs on several systems, laptop, desktop and server, on various configurations of disks.
It has served me well for years, on the server it helped me detect a bad SATA controller. It would work perfectly in light usage, but start introducing errors in heavy usage, which made one disk inconsistent with the others in the storage pool.
Btrfs alerted me to this and after moving the disk to a good controller, I ran btrfs-check --repair on the unmounted disk (after reading the warnings), which got the FS back to a consistent state, remounted the whole pool and ran a btrfs scrub to get everything back in line with itself. The whole process did take a while, but I had backups and wanted to try out the tools. In the end there was no data loss, and the pool is still running perfectly today.
And they specifically tell you to use XFS for any production deployments.
Furthermore - btrfs feels excessively complicated for simple workflows - if I want to snapshot a btrfs volume without exposing the snapshot to the machine’s view of the file system, I have to do a bunch of volume layout setup first. With ZFS I can just snapshot.
You mean like the current advice not to use anything except mirroring and striping (RAID-0/1/10)?
> Parity may be inconsistent after a crash (the "write hole"). The problem born when after "an unclean shutdown" a disk failure happens. But these are two distinct failures. These together break the BTRFS raid5 redundancy. If you run a scrub process after "an unclean shutdown" (with no disk failure in between) those data which match their checksum can still be read out while the mismatched data are lost forever.
* https://btrfs.wiki.kernel.org/index.php/RAID56
I've been using ZFS since it came out on Solaris 10 over a decade ago and it was specifically designed not to have a write hole due to its COW/ACID nature.
See this 2008 SNIA presentation from Bonwick and Moore, the creators of ZFS talking about not having a write hole:
* ACID/COW: https://www.youtube.com/watch?v=NRoUC9P1PmA&t=24m
* Integrity: https://www.youtube.com/watch?v=NRoUC9P1PmA&t=55m20s
either way, my story is the same as yours. I have no troubles, I have at least ~10 daily snapshots of many subvolumes I can pull from if I accidentally: `rm -rf /usr` or something silly like that. It's a great FS and unlike ZFS, is actually upstream in the Linux kernel.
If you need this functionality and want to mitigate the write hole, gen an UPS.
> Seems like a feature companies would be very interested in so I wonder why nobody's ever fixed it.
Because those driving the development have no interest in RAID5/6, they do not need it. RAID5/6 is enthusiast/SOHO feature, and these groups do not take part in development, so no wonder it is neglected.
In fact, I will combine both points: it is more effective/economical for everyone, who needs RAID5/6 with btrfs to get an UPS than to pay the manhours for the development. Additionally, it would be an expense incurred by those who need it, not by those who do the development currently and do not need the feature.
The typical btrfs fix is (slightly exaggerated) "When doing an A while a B is pending, re-enumerate the Cs for the purpose of D unless the E is locked, in which case, reschedule the F". Where A...F are all not trivial like snapshot, out of disk space, device rebalancing, and so on.
If that level of thinking of all the things at once is required to write correct code, well, it's evidently not happening.
https://phoronix.com/scan.php?page=news_item&px=Btrfs-Warnin...
It's kind of insane that two companies now sell a fixed btrfs-RAID5 as a proprietary commercial product (yes, in spite of the GPL). They do it by using the mdadm RAID5 code, which works, and add proprietary hooks to allow btrfs to use its checksums to figure out which spindle is corrupt in a parity mismatch situation. This (a) closes the write hole without a journal doubling the I/O load and (b) detects silent corruption, neither of which mdadm can do.
Or just upstream this patch set:
https://www.mail-archive.com/linux-btrfs@vger.kernel.org/msg...
...so we can do this ourselves with LVM (split each spindle into a small metadata device and a large device, stitch the large data devices together with mdadm-raid5/6, then add that and all the small devices to a btrfs filesystem with the small devices marked "metadata only" and -dsingle -mraid1c3).
I really feel like the btrfs guys are stuck in "shiny new thing" mode here. Last time I checked host-managed SMR devices were only available in engineering sample quantities. Even if that's changed they surely are still very rare, a tiny minority of worldwide storage device sales. The fact that btrfs-RAID5 scrubs take something like O(num_spindles^2) seek latencies is crazy... I have an 8-spindle array that scrubs in 12 hours with "-draid0" but takes 10 days with "-draid5".
Raid5/6 seems to me more popular in home use.
The fact that Synology's makes money selling their proprietary btrfs-RAID5 is incontrovertible proof of the commercial relevance.
https://www.mail-archive.com/linux-btrfs@vger.kernel.org/msg...
https://manpages.ubuntu.com/manpages/bionic/man8/mkfs.f2fs.8...
-m -z #-of-sections-per-zone
https://zonedstorage.io/linux/fs/
The f2fs section says "Zoned block device support was added to f2fs with kernel 4.10. Since f2fs uses a metadata block on-disk format with fixed block location, only zoned block devices which include conventional zones can be supported. Zoned devices composed entirely of sequential zones cannot be used with f2fs as a standalone device and require a multi-device setup to place metadata blocks on a randomly writable storage."
https://www.usenix.org/conference/atc19/presentation/jaffer
I'd be afraid to use it for anything other than throwaway test data.
This is native support for btrfs-on-SMR, without the dm-zoned layer in between. DM-zoned is meant for filesystems who are zone-unaware, and it works by batching writes and redirecting blocks into appropriate areas. Having the filesystem allocator be aware of the underlying device's zones allows for more efficient/performant use of the zones.
I reverted from btrfs to ext4 on my main desktop a few years ago, because grub couldn't remember the last selected menu item (error: sparse file not allowed), and I didn't feel like creating a dedicated /boot partition.
The 'sparse file file not found' is caused by grub, it would try to overwrite file blocks directly, but on btrfs it would cause checksum mismatch. This has been solved by storing the env block outside of the filesystem at 256K and synced back and forth once the system is booted. 256K is ok as btrfs does not use the first 1M on any device for bootloaders.