Btrfs in Linux 6.2 brings performance improvements, better RAID 5/6 reliability
phoronix.com
phoronix.com
It has gems such as:
* It won't boot on a degraded array by default, requiring manual action to mount it
* It won't complain if one of the disks is stale
* It won't resilver automatically if a disk is re-added to the array
I think the first is the killer. RAID is a High Availability measure. Your system is not Available if it fails to boot.
Of the 17 servers, only three are now remaining where I haven't had to do this. Not all of these were actual SD card failures, some where done preventively. Still, there have been several SD card failures requiring emergency repair work at inconvenient times. There have been 0 RAID failures requiring similar emergency work on the systems where /boot has been migrated to HDD based RAID.
Of course there is no replacement SD card handy with working /boot filesystem. Actually there is no such thing as "handy" in my case. I can actually reconstruct /boot partition on the HDD faster than anybody could go to the datacenter and replace the SD card. And if there were a replacement SD card, I would need to keep it up to date manually every time the kernel or initrd are updated.
I never want to see this kind of setup again, and frankly it felt insane the first time I saw it.
Of course I should have migrated all of the machines off the SD cards by now, but my excuse is that maintaining Linux on these machines is not really my responsibility (it is nobody's responsibility apparently, although many people are interested in keeping these systems online).
When I die I want to have a SD-card shaped tombstone.
I think the parent views on storage is still valid. Do not boot from the RAID volume where you store many terabytes.
I usually have a small mirror SSD for the OS (like 128G or max 256G). Those drives are super cheap, easy to replace.
For the /data I usually use ZFS or mdadm + xfs. Depending on the use case (for example do you want to add more capacity later? ) you can decide.
Using SDcards for server filesystem is negligent at the very least.
1.HA
2.Had an OS that would get fully loaded into memory (think ESX).
The idea was that SD-cards take less power and where cheap enough you could have a whole gang of spares with images ready to just drop in place. I'm not saying it was a good idea, just that I've encountered this more than once.RAID is used for two things:
1. improving the performance of slow disks (at least the read performance) 2. having 1/2 disk fail with your system remaining completely usable, till you replace the faulty disk as soon as possible.
The second point is fundamental to me: there shouldn't be any disruption of service whatsoever, meaning that only the sysadmin should notice the fault (beside maybe a reduced performance of the system since you have 1 less drive). Database transaction that were in progress when the disk did break shouldn't fail, writes/reads on the FS shouldn't fail, the only thing that should happen is an alarm triggered in the monitoring system to inform that a disk needs to be changed as soon as possible.
Having a RAID with manual recovery... it means that you could end up a Saturday evening in front of a computer to bring a system back online, and still some data corruption may have happened.
If it's your own personal box, and you're okay with that fiddling, fine. But if people are supposed to get work done while the admin gets a new disk, it's not great.
Would it randomly pick one of the bootable disks? That way you have an n-1/n chance of avoiding the bad disk?
If a disk with a boot vol failed, it would have a 50/50 chance of still booting depending on which one failed. You'd create a backup grub entry to boot off the other one manually.
Not as transparent as hardware RAID, and maybe now grub is aware enough to auto handle mdadm? I don't know - I haven't done bare metal mdadm for a decade or so.
The same way it has worked for the past 30+ years: the firmware boots from first designated boot device; if that device isn't bootable, it moves on to the next one. The next one boots since it's part of a mirror and has all the necessary data to do so, and in the rare case where the device is "half bootable" one would simply intervene and select the next good bootable device manually.
I have 2 NixOS-based NASes that run ZFS. each one has 3 equal-sized 256gb SSDs in addition to the pile o' spinning rust.
each SSD has a small UEFI boot partition, then the rest of the space is cut in half.
the root filesystem is a 3-way ZFS mirror of the first half of the SSDs. the second half of each SSD is another 3-way mirror, this time as the "special" / metadata device for the main hard-drive-backed zpool.
any of the 3 SSDs can fail, and it will boot up and mount the storage perfectly fine. I could also easily upgrade to larger SSDs, in-place and with minimal downtime (zero downtime, if my case had hot-swappable SSD bays)
it would work just as well with only 2 SSDs, but the incremental cost of a 3rd is small enough relative to the whole that I went for it.
Machines booting from SAN had only one local disk, others would have 3 local disks (small and cheap ones because application data was on SAN anyway).
Main raid would do the job for availability in term of disk failure, 3rd bootdisk did the job in case of user erroror data corruption on the main boot env. It saved our asses a few times, bringing apps back online quickly and saving us from reinstalling or fixing stuff from a tenporary live environment.
That lets you then even hotswap them if you need to while keeping the root filesystem workable.
Servers go boom, refuse to boot, or even get trigger a few alarms. Disk is replaced, server is booted, and provisioning paves the entire thing faster than you'd be able to debug.
I can't knock it, but I imagine I'd be yelling for at least RAID1 if I had to wait for a replacement.
I have used the FUSE port of ZFS to write only one member of a mirroset, then upon mounting both members elsewhere, the stale mirror was very quickly resilvered, so ZFS was able to determine only the blocks needing to be refreshed.
In btrfs, I understand that this requires a rebalance, which will read/write every used portion of the filesystem.
ZFS is far friendlier in a crisis.
"We'll even manually trigger a scrub—a procedure that storage admins generally understand to look for and automatically repair any data issues... even though we manually initiated a scrub and let it finish, our array is still inconsistent and even outright non-mountable, because it ran for a little while without a disk and then that disk was re-added. The command that we were supposed to run was btrfs balance—with both drives connected and a btrfs balance run, it does correct the missing blocks, and we can now mount degraded on only the other disk, /dev/vdc. But this was a very tortuous path, with a lot of potential for missteps and zero discoverability."
https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
There is a potential gotcha with parity being wrong following a crash or powerfail, the so-called write hole. Scrub doesn't check parity, and parity isn't checksummed. The write hole though is arguably two parts: wrong parity and propagated wrong reconstruction.
In the usual write-hole case, both happen silently. Wrong parity could happen after a power fail, crash, misdirected or torn write - whether the array is functioning normally or degraded. However, wrong reconstruction only happens if there's a failure to read a data strip (bad sector or full device failure).
On btrfs, only the wrong parity being written is possible. Upon reconstruction from bad parity, the resulting data is compared to csum and will fail, thus propagation doesn't happen.
To cause parity to be recomputed and rewritten, yes you need to do a full balance. But among all the block group profiles, that's only raid5 and raid6. Single, DUP, raid1, raid1c3, raid1c4, raid10 aren't affected and a scrub does reallocate missing blocks on mirrors.
Every discussion about btrfs immediately escalates to how the btrfs RAID support develops. But btrfs is useful on its own without that mode.
Should you wish to put a btrfs file system on software RAID (DM), you can still do that. The btrfs RAID modes are something else, where you cut the DM driver out of the picture. That's not going to be as well tested as plain old DM, even if it works, and it doesn't offer any functionality beyond plain DM RAID.
If the use case is storing real data, you want the data storage layer to be as well tested as possible. There's also more to software RAID than the block device. You want a mature resilvering daemon on a well tested schedule, and you need working monitoring to alert you of issues. RAID without monitoring will only delay the inevitable.
wikipedia: The device mapper is a framework provided by the Linux kernel for mapping physical block devices onto higher-level virtual block devices. It forms the foundation of the logical volume manager (LVM), software RAIDs and dm-crypt disk encryption, and offers additional features such as file system snapshots.[1]
Device mapper works by passing data from a virtual block device, which is provided by the device mapper itself, to another block device. Data can be also modified in transition, which is performed, for example, in the case of device mapper providing disk encryption or simulation of unreliable hardware behavior.
To expand a typical RAID setup, you have to essentially fail every single drive in the array and replace them one by one with new identical drives. Expanding an array of N drives requires N resilvers which not only takes a long time but increases the chance of actually failing those drives due to massive I/O load. It's easier to build a second array and copy the data over. This combined with the fact you have to buy all the hardware up front makes it cost prohibitive for home users. You can't buy drives one at a time and expand capacity as needed.
If you want it to behave like that then add 'degraded' to fstab. That a device is missing can have unknown reasons, the user should know better and resolve it or allow such boot. It's not automatic as there's no way to inform the user that it's degraded state.
If a device goes missing for "unknown reasons", then the machine should still work, and I'll figure out what happened when monitoring pokes me and says RAID is degraded.
It's up to you to choose at that point - is availability more important for you (add degraded to fstab), or data consistency (deal with the array first).
That's not the only purpose for it. There's three reasons I can think of that you might set up a RAID array:
* You want better uptime. (your use case)
* You want to protect from data loss. (my assumption was that this is the most common use case, but I could be wrong. This also helps with uptime because there's nothing worse for uptime than having to restore lost data from a cold backup)
* You want better performance, data integrity be damned. (RAID 0)
Booting a RAID array with a failed disk is a bad idea if you care a lot about not losing data, because now you're one less disk failure away.Booting a RAID array with a failed disk is absolutely fine idea.
How else I get access to the tools to identify the bad drive and resilver RAID on a replacement, be it in the same bay or not?
It's reasonable to want to preserve all the data you currently have—some of which probably hasn't been backed up yet—and not accept new data to be written with the durability guarantees the array was originally configured for silently violated.
Since the kernel has no way of knowing which volumes may contain important data that didn't get the chance to be backed up, it should try its best to maintain the original durability standards the filesystem was configured until some mechanism outside the kernel authorizes the relaxation of those standards.
IE (by your logic) the system should stop the writes as soon as the array became degraded.
But this is not what happens with btrfs: it would happily continue to write the data on the array until reboot.
And then suddenly it's "oh my god array is degraded!!!111 you should not write to it1111".
To add on that: I never seen for a HW RAID card to stop booting by a mere degraded state of the array. Changes in configuration of arrays, loss of more than enough for redundancy drives - yes, that would halt the boot and require the operator intervention. Array in a degraded state? Just spit the warnings to the console and boot. Nobody has the time to walk to each server with a degraded array on every reboot.
If you remove this udev rule and then add degraded mount option to fstab, it's very risky because now even a small delay in drives appearing can result in a degraded mount. And it's even possible to get a split brain situation.
Btrfs needs automatic abbreviated scrub, akin to the mdadm write intent bitmap which significantly reduces the resync operation.
Let's treat it as archived historical content it is.
These days, with UEFI prevalent, a system won't boot without EFI system partitions replicated (and possibly sync'd) on every drive.
That by itself is a complete deal breaker.
Anyone.
> Also if you boot with a degraded disk are you not asking for massive trouble
> not being able to boot until you add another disk
Great.
I'm just a [sys]admin who were given a task to do something on $server.
I jump around the red tape, claw out a 15 minute downtime because the task requires reboot, proceed with all that corporate dance with notification emails.
I do my thing, reboot the server and it doesn't return back online.
Suddenly I broke the server, missed maintenance window, amount of mails with CC and RE: in my mailbox grows in geometric progression and the most important - now I need to find out who were responsible for the server, contact him and [kick his ass] ask him to diagnose what is going on.
Bonus points if:
server doesn't have a meaningful BMC/iLO/iDRAC with KVM console
it does have it but it's broken for some reason eg requires Java 6 on Vista
server was configured 10 years ago by a greybeard who is not only retired but already died from the old age
server is 6000km away from any place with a replacement disks and the earliest time when you can send the replacement would be the next spring when the ice would break and taw enough for the ships to move. Of course you can hire a helo to deliver it, which get your ass chewed for an unplanned $50k expenses
This is all the things what I encountered in my admin days.
Including BL670 with incorrectly connected drives, so despite everything saying (and indicating) what the failed drive was in bay 2 it was actually in bay 1.
> then ridiculing btrfs
I ridicule btrfs for it's RAID mode not being a RAID mode by default.
RAID is about availability of data.
If you so hell bent on the data safety then btfrs should kernel panic as soon as one drive degrades. THAT would make sure someone would come and investigate what happened and no data loss.
Right?
> Right?
Panicking an already-running kernel would only serve to prevent userspace from handling the failure through mechanisms that are inherently beyond the scope and capabilities of the kernel alone (ie. stuff like alerting a sysadmin, activating a hot spare, or initiating a rebalance to restore redundancy among the remaining drives).
Perhaps the kernel should default to freezing a non-root filesystem when it becomes degraded, absent an explicit configuration permitting otherwise. But for the root filesystem, that would be counterproductive and prevent the failure from being handled gracefully.
Obviously, the tradeoffs are different for a system that is still trying to boot as opposed to one that is fully up and running.
It's not a backup. It doesn't protect against mistakes, bugs, or the server failing in some other way. What it's good for is for ensuring that work keeps happening if a disk fails. And disk failures happen to be fairly frequent, since spinning rust is a rather delicate technology.
If somebody has to connect to the machine and fiddle with it by hand it means that the system has been down, possibly for hours, when that was the exact thing you were trying to prevent by setting up RAID on it.
> What it's good for is for ensuring that work keeps happening if a disk fails.
Different services have different priorities. I want my work servers to keep running, but my home server under the desk to stop until I replace the failed drive. Btrfs gives you a choice.
It's an option to enable. If you needing to configure something is a dealbreaker then having RAID is probably not for you anyway.
For enterprise, if you dont have monitoring setup it's on you
It's like people justifying MySQL saying that it's okay for it to lose or corrupt data, and that transactional integrity doesn't matter as much as people say it does.
Yeah, maybe if you're website is a blog. But anyone storing real data should run away screaming from systems like this.
Maybe MySQL is sort-of-okay now? I dunno. It's possible BTRFS RAID 5 won't eat your data or crash your server regularly now.
I'm going to stay away from both anyway.
The Linux implementation is pretty awful, with it demanding that it uses the absolute `/dev/sdx` reference even if you try to build it using serials or other unchaging reference.
Replace a disk and then after the next restart it fails to bring up the RAID because `sdf` points to a different disk. At least BTRFS uses internal UUIDs so can bring up the arrays when the disk ids change.
Fairly fundamental stuff, rather than complaining about an option not being enabled by default.
Yeah, putting / on zfs was not smart. Nowadays, when the distro doesn't support it out of the box and doesn't come with kernel+bootloader+zfs tested together (i.e. proxmox, ubuntu), I wouldn't do it.
I assembled the zfs array using by-id, but after a reboot zfs reports the drives as being on /dev/sdx.
This is one of my systems using both identifications:
root@multivac[~]# zpool status
pool: boot-pool
state: ONLINE
scan: scrub repaired 0B in 00:01:33 with 0 errors on Fri Dec 9 03:46:34 2022
config:
NAME STATE READ WRITE CKSUM
boot-pool ONLINE 0 0 0
sdg3 ONLINE 0 0 0
errors: No known data errors
pool: multivac-slow
state: ONLINE
scan: scrub repaired 0B in 07:05:51 with 0 errors on Sun Nov 20 09:05:55 2022
config:
NAME STATE READ WRITE CKSUM
multivac-slow ONLINE 0 0 0
mirror-0 ONLINE 0 0 0
ad3395b6-ac43-453c-9ba2-cf65542cb710 ONLINE 0 0 0
2c423fe3-07a4-44cb-bc39-8b267c79d8c3 ONLINE 0 0 0
mirror-1 ONLINE 0 0 0
8ca9d5a2-04cb-47a0-987a-8b3d27326975 ONLINE 0 0 0
bec4c67f-fa41-4ab2-907f-0e5b1e052300 ONLINE 0 0 0
cache
bfef4d7f-c484-4c37-a0b7-81d59fa1e189 ONLINE 0 0 0
root@multivac[~]# blkid|grep zfs_member
/dev/sda2: LABEL="multivac-slow" UUID="4053361656876756561" UUID_SUB="363024978934966364" BLOCK_SIZE="4096" TYPE="zfs_member" PARTUUID="2c423fe3-07a4-44cb-bc39-8b267c79d8c3"
/dev/sdd2: LABEL="multivac-slow" UUID="4053361656876756561" UUID_SUB="10877444988141002268" BLOCK_SIZE="4096" TYPE="zfs_member" PARTUUID="8ca9d5a2-04cb-47a0-987a-8b3d27326975"
/dev/sdb2: LABEL="multivac-slow" UUID="4053361656876756561" UUID_SUB="9244362969528724588" BLOCK_SIZE="4096" TYPE="zfs_member" PARTUUID="ad3395b6-ac43-453c-9ba2-cf65542cb710"
/dev/sdc2: LABEL="multivac-slow" UUID="4053361656876756561" UUID_SUB="11658072268678819533" BLOCK_SIZE="4096" TYPE="zfs_member" PARTUUID="bec4c67f-fa41-4ab2-907f-0e5b1e052300"
/dev/sdg3: LABEL="boot-pool" UUID="13291833732043257716" UUID_SUB="15424416082895860251" BLOCK_SIZE="4096" TYPE="zfs_member" PARTUUID="b6c0b302-7599-4862-888b-be6f7a85b970"On the flip-side, I've done some of the most horrible things possible and screwed up my 30-disk ZFS array a number of times and I've never lost data. I doubt btrfs could recover from anything I've done to break my ZFS pool.
Overall I think it's a huge shame that it was determined that the CDDL was incompatible with the GPL, because it wasn't intentional, and we would be in a completely different spot for storage otherwise.
I see XFS as the performance leader (appears on TPC.org the most often that I can see), btrfs as the fullest featured, and ZFS with the strongest reliability.
XFS for databases is what is used on tpc.org, but perhaps these new improvements may help.
I think they are safe with this version:
XFS (sda1): Mounting V5 Filesystem btrfs filesystem defragment /path
There is also the mount option autodefragI've ditched RAID 5 mostly because performance sucked so much, but I have never lost any data despite going through many changes of hard drives, changes of RAID mode, rebalancing, etc.
The closest I've come to losing data is probably rotten HDD sectors, but it's always been caught and remapped by scrubs.
Yes, there is the infamous write hole, but that's not a BTRFS only problem.
As a random anecdote I've been using BTRFS for everything for over 7 years now and had no issues. Think the only FS I've ever had problems with was ExFAT.
The other reason is that the atomic snapshot feature, and ability to easily transfer diffs of snapshots makes a wonderful foundation for incremental backups. But if I can't trust the filesystem to avoid corrupting my data, how can I trust it to avoid corrupting my snapshots? Hence I can't trust my backups. So I need to go back to a more independent backup processes like rsync.
With those two features gone, my whole motivation for using btrfs over simpler file systems like ext4 is gone, so why bother.
Still seen no real proof this is true. It's always vague feelings and anecdotes. has anyone done any real testing to measure file corruption?
If your filesystem silently corrupts some data in a way you don't notice, you might just replicate the corrupted data to your backup system over time. And in that case you will still have lost it.
Data loss - e.g. due to a disk failure - is a lot more obvious.
I had recently upgraded my pool and had a bunch of old disks lying around so i fired them up in a new pool, 6x 4TB and 6x 8TB disks. I took the 4TB disks and set them in in a raid0 to get effectively 9x 8TB disks. Apparently doing this will work but you MUST make sure you properly wipe all of the newly raided disks of their old ZFS data or you can end up with ZFS trying to check that pool and breaking the mdadm superblocks.
After I properly zeroed the 4TB disks it's been working fine for weeks. This let me set it up a pool in raidz3 with them all so i can loose any of the much older 4tb disks or up to 3 of the 8tb ones. I don't use it for anything that's absolutely critical to store, just random bulk data that i don't want to worry that much about, but it's been reliable enough since then.
I've done that a couple times to make a bunch of mismatched disk fit into zfs's requirement that all are of the same size. The ability to work with a mix of disk sizes is one advantage of btrfs - if you remember to be extra careful on a disk failure and bad shutdown. Plus the ability to dynamically grow the array, but there is work to improve that in the zfs side as well.
zpool create swimming disk1 disk2 …
That has worked for over 15 years.
Say you have `1, 1, 2, 2` sized devices. A raidz requires them all to be equally sized which can be done by first adding `1 + 1 = 2` and then creating a zpool with `[1+1], 2, 2`. Can this 1 + 1 step be done within zfs? Hypothetically this would be done by creating a raid0 pool and then a raidz on top via `zfs create tank raidz (raid0 a1 b1) c2 d2`, or first running `zfs create ab1 a1 ab` and use that in the next `zfs create ab1 raidz ab1 c2 d2`.
What I suggested is not really RAID0. raidz is not really RAID5. ZFS intentionally improves on these for performance and reliability reasons.
My personal preference in this kind of situation is to still just use lvm2 to jbod the disks together if I;m fine with the risks of losing it all anyway, and just setup a good backup strategy. More because I'm familiar with it and all the tools than other setups. I think zfs might give me more early warning signs of a disk with an issue but the performance hit isn't always worth it to me there. I've got an 8tb raid0 nvme ssd array (4x2tb) setup to give me some blazing fast local storage for a few key VMs thst I just backup daily to the zfs storage. I can get 7.5GB/s reads and 6.7GB/s writes to that array doing that and using an lvm-thin pool so I can nicely snapshot during backups.
The only thing it's missing before I consider it full-featured is stable RAID 5/6. But it looks like that hasn't been forgotten.
My strategy to pull new things is to have one big feature that has ideally been reviewed and iterated in the mailinglist or there was a lot of testing already done. In addition two smaller features can be merged, with limited scope, not affecting default setup and possibly easy to debug/fix/revert if needed. Besides that there are cleanups or core updates going on so this should not touch the same code to make testing less painful. With new features the test matrix grows, code might need wide cleanups or generalizaionts before the actual feature code is merged. So this can indeed slow down development.
The raid56 is progressing but until the 6.2 pull from today there was not much to announce regarding stability/reliability. There were proposed fixes but as incompatible features, which means some changes on the user side and with backward compatibility issues. What's pending for 6.2 should fix one of the bad problems at least for raid5.
I've been running BTRFS RAID6 since 2016 and have only had one issue (arch kernel needed to be rolled back) and never suffered anything catastrophic. It's perfectly happy humming along with a 15x8TB raid array.
(It's perfectly safe to run metadata on raid1 and data on raid5/6.)
Although these drives have surpassed 50,000 hours of running time, I am fairly certain the day of reckoning is coming, but I have backups and a plan to recover.
I may recover to something not BTRFS. Not sure just yet.
Better to direct resources towards bcachefs or ZFS IMO.
As much as I want to see a wide use of bcachefs, it's still years away. As someone who actually wants to store data - why would you direct resources to bcachefs which is known experimental, rather than btrfs which plainly documented raid5 as not ready and now may decide to change it to ready... if it is? (I'm donating to bcachefs patreon, but commenting from the PoV of someone choosing a solution today)
Could you paste a link to it, please? As far as I know, all efforts now are around ZoL which is not included for legal reasons.
Why people would contribute to this mess is beyond me but everyone is free to do what they want.
You haven't glanced at a list of FreeBSD's contributors, have you?
You're probably behind a pfsense firewall right now.
You have probably seen someone playing games on a PlayStation 3 or 4.
You've definitely benefitted from the FreeBSD project without even realising.
Have you taken a casual glance at the list of FreeBSD contributors?
You haven't even looked at the list have you?
I was surprised to see that it taints the kernel when I checked the demsg.
How's that beyond useless?
Ok.
All of the arguments for why it's better to have something like zfs handle RAID rather than layering something on top of a traditional hardware RAID controller also work for explaining why you should prefer managing storage more globally rather than layering on top of something like zfs—if you have the resources to develop and maintain a true "full stack" storage solution. Which Facebook/Meta obviously does, when they can do things like publish their own spec documents that SSD vendors design around.
That whole thread above debating not booting degraded really made me lose any faith in btrfs.
No indication of any hardware issue: No recent power loss (& it's on a UPS), no SMART issues, no memory test positives. But the block tree (at least, WIP) is f*cked across all the disks (looks like two competing writers went at it) and none of the available tools can deal with it.
It wasn't a super exotic setup either: RAID10 with 4 disks (2 stripes), fairly full and regular snapshots/cleanup, but that's it.
I already converted my root to ext4 because paranoia and I'm probably going to move bulk data (what can be recovered) to ZFS.
The worst issue I've ever had with btrfs simply required zeroing the journal (by hand -- both the kernel and *fsck tools would crash when reading it). The first time I reported the issue to the mailing list and the underlying issue was promptly fixed. However, it still happened a second time, and I didn't bother reporting it. As many people say on this thread, once is too many when it comes to filesystems.
Using xfs|ext4 + mdraid is just a lot simpler and much faster, and it addresses most of my use cases.
I hope your journey ends better than mine did.
I've been managing a home NAS for years with SnapRAID and plain old ext4. Sure, it doesn't have the bells and whistles of something like ZFS, but it's simple to use and understand, scales with anything you throw at it, and there's no way to lose the whole array.
I'm currently transitioning to a multi-node setup, and will probably move to Ceph, otherwise I would stick with SnapRAID. Every so often I look into btrfs/ZFS/mdadm, but keep reaching the same conclusion that it's just not worth the risk.
The worst thing about my experiences with mdadm and ZFS is that I've had exactly one catastrophic failure in 15 years.
The best thing about my experiences with Ceph is that it's really reinforced the importance of a good backup strategy.
The combination of the two means I now run ZFS, and I have a comprehensive backup strategy that is regularly tested, but has never needed to be invoked.
I strongly recommend developing a strong backup story before you go down your Ceph journey.
If anything the fact that snapraid by not live replicating every change to every disk in the array makes it impossible to achieve the reliability offered by ZFS. It's forever the inferior cousin and in fact more complicated for layering a dissimilar technology on top of the other.
In this case, _for my simple needs_ of running a home NAS, the fact SnapRAID doesn't run in real-time is not an issue. In fact, I prefer being in control of when it runs, and what it's doing exactly. Having a short time window where some new data is not replicated is a negligible drawback to me.
OTOH, while ZFS solves this particular issue, its drawbacks of being difficult to scale, and a huge black box that claims to "just work", when in fact my _entire_ array relies on it working perfectly, 100% of the time, are a tough pill to swallow. I'm sure that with the years of dedicated improvements and stability fixes, it's a battle-tested system where the likelihood of it failing is close to 0, even on Linux. But the fact that it's theoretically possible is a deal-breaker to me. In this sense, I much prefer SnapRAID's approach that makes this literally impossible. SnapRAID could stop working entirely or disappear tomorrow, and all my data is perfectly safe.
So, yes, SnapRAID + ext4 is radically simpler than ZFS, IMO.
Would I recommend this setup in a corporate environment, where company resources are on the line? Probably not. But I would still advocate against ZFS and mdadm, and probably suggest something like Ceph instead.
We published our internal doc for how we install our Arch setup with fully encrypted BTRFS if anybody is curious. Happy to answer any questions too!
https://www.lunasec.io/docs/blog/arch-linux-installation-gui...
(We run Thinkpad X1 Extremes (+ the business P1 equivalents) as our dev boxes.)
You have a typo in your article BTW:
> aes-exts-plain64
It was a perfect storm of me being an idiot and heavy use. But I dunno if they have remotely resolved the issues
Every month or so I spin up external drives to write out monthly backups of my NAS and personal computers. Forgot I'd taken on the home video collection and overfilled the array. Deleting some old subvolumes and having the space reclaimed took hours (on a 99.9% full 5 TB RAID-1 volume). Part of that is that I use compress=zstd:10 on the external drives, but I wasn't expecting it. Writes speeds tanked to tens of KB/s. Since these are those Seagate drives from Costco, probably also SMR drives...
However I have run this same drive out of space using ext4, zfs, and xfs playing with it (it’s a drive I use for torrent etc) and none of them have crapped out like btrfs did on it.
Do/can they downstream some of the changes without changing the Kernel version?
But I do hope there will be model, especially entry model that comes with 5.10.
Synology's Linux timeline is simply too slow.
I have one of the infamous Atom C2538-based Synology there, and it is running kernel 3.10.108, built in october 2022. The device is 2019 model (the SoC was introduced in 2013).
There came a point where we were going to have to move to a larger server as there were no more slots for additional disks. New hardware procured, provisioned, prepped, ready to roll with the exception of an actual cut over. However, c-suite got involved as they tend to do from time to time, and kept kicking the can down the road.
Eventually disk space was a critical issue and a junior sysadmin decided it was a good idea to tell the COO that btrfs can be converted to raid 5 on the fly. I over the phone told them that this would guarantee data loss and refused. They literally came to my office to force my hand. Some people just have to learn the hard way.
That story aside, I'm a big fan of BTRFS and have been using in various configurations on my desktops for years. I'm glad to see this progressing.
Filesystem is light on requirements and works really well. Compression works great too.
The only thing I haven't played around with are reflinks.
For critical datasets I apply 3-2-1 backup strategy and use ZFS.
Why bare-metal admins have to spend countless man-hours on troubleshooting these?
Also note, comments on this thread are heavily biased on DYI things and not covering enterprise storage areas, which for instance can come as Blackbox with exported LUNs
Some of the problems:
* RAID cards usually have under powered CPUs, and can easily be a bottleneck. Often the best performance with a RAID card is with the RAID disabled.
* Even battery backed RAM is often pretty slow and on the wrong end of a high latency connection between CPU/RAM and disks. Generally investments in ram or intentlog/writelog/slog is a better investment.
* Metadata is often undocumented, often needs backed up via obscure device dependent methods, and is required for a RAID adapter failure
* The firmware is often buggy, especially in the handling of the numerous failure modes.
* Recovery often requires the same card with the same firmware and a copy of the backed up metadata. Not all cards can recover all metadata from just the drives.
* Often are "too" smart, and won't export raw drives for use with more advanced filesystems like btrfs or ZFS. Some require ugly work arounds like exporting each disks as a single disk RAID0.
* Often RAID cards hide the SMART info from the OS, often crippling the ability to predict drive failures. Or hides the functionality behind a weird software stack that assumes a SMTP server and integrated poorly into whatever monitoring/alerting system you use for the operating system.
* Some RAID cards require a network connection and run a buggy and insecure out of data web stack on some undocumented and rarely (if ever) patched CPU that was obsolete the day it shipped.
I have been running production servers with LSI hardware RAID cards for the past 15 years and have not experienced the items you note. The only thing I did experience was a bad RAID card (solved by replacing it with a similar LSI model).
To contrast your "problem" list:
* Compared with MDADM (mirrors, RAID-5, RAID-6), the LSI RAID cards tend to be on par as far as performance (esp with BBU)
* A failed drive won't prevent the array from coming on line (in my experience)
* All my servers have battery backed RAM; never had an issue with slowness or high-latency. Can't say the same for ZFS.
* Never, ever had a problem with metadata on the disk. Not sure why this seems to be an issue
* LSI cards have pass-thru mode and can easily be used with BTRFS/ZFS. I have done lots of tests; never an issue
* LSI cards have their own "patrol read" mechansim to scan drives for bad sectors and send alerts when something seems wrong
On a positive note: * Hard drive failures are truly plug-n-play. Send a tech out to the cabinet, spot the RED light, replace drive. Done. No OS work, no console, no crash cart needed.
* You can logically divide the drives using the RAID manager tools just like Linux MDADM, ZFS, BTRFS, etc. Easy.
* You don't need the specific card to replace a failed card. You just need a card that can read the metadata on the array to bring it back to life.
* Lots of monitoring scripts available in the wild to get RAID stats, rebuild times, etc.
I am not saying hardware cards are indestructible. But, all in all, hardware RAID cards are solid devices that have been in many, many production servers all over the world. I certainly would not dismiss them in favor of ZFS RAID (which tends to be slow for many operations). Pick the best tool for the job.Generally I (and colleagues) consider a hardware RAID failure a nightmare and a SAS HBA failure an annoyance. Doubly so if your storage design included cross connected servers and you just mount the storage from the other server. I've never tried similar with a hardware RAID. Can you cross mount and easily import/export RAID sets between controllers?
The main hardware RAID performance issues I see is with higher drive counts or NVME drives when combined with RAID6. Seen any hardware RAIDs and can manage a few GB/sec with 2 disks of redundancy? 3 disks? Even spending a few $k on hardware RAID seems to lose to a random 5 year old server using ZFS or software RAID by a large factor. Seems like even pretty old x86-64 servers manage a few GB/sec per core.
Careful on the passthru, I had many generations of LSI cards that worked, and 3ware and Areca before that. However I've been hearing that the new LSI RAID cards lack passthru/JBOD mode. I've seen reviews and complaints about the lack of it. One case had involved a LSI hardware RAID connected to a Dell JBOD array, but maybe it was specifically crippled by Dell.
I avoided ZFS for years, generally slower at the relevant workloads (used IO logs collected with systemtap and representative loads created with FIO). However once I added cache and the benchmark included multiple streams of writes, ZFS was a huge win. In particular 64 or more sequential write streams to 120 disks. In fact ZFS did better with 3 disks of redundancy that other filesystems did with 2.
In that past I had more luck with ZFS or MDADM for doing things like /dev/sd[ab]1 for a 32gb RAID1 for boot, then /dev/sd[abcd]2 for RAID5 or RAIDz2. Sounds like the hardware RAID tools are get better.
I've heard of some hardware RAID setups managing to recover hardware RAID metadata with a new card, and some not. Even cases of successes and failures from the same company. Even cases where support claimed it would work, but then decided the card and/or firmware wasn't close enough. Sure decent support can overnight a card, but nowhere near as nice as just being able to mount the drives on any Linux box.
Thus far, we don't have any high-density NVMe servers so I can't really test high-end servers. My experience is limited to 8x SSDs or 16x spinning drives in our arrays.
Finally, I avoided ZFS for a long time as well. Early on (OpenZFS 0.6, 0.7, etc), I spent an enormous amount of time trying to tweak/tune the system just to get on par with our HW RAID devices. No matter what I did, I simply could not get the server (typical Linux NFS NAS) to get any decent performance. In fact, I clearly remember HW RAID hitting 1.8GB/sec on our SSDs while OpenZFS could only get around 400MB/sec. Only until OpenZFS 2.1 mark did I see any real performance gains.
...Speaking of ZFS... I recently worked on a project to replace XFS with ZFS for Postgresql servers. I learned a lot about OpenZFS - specifically the memory latency when using compression. I had a lengthy discussion with the OpenZFS "gang" (https://zfsonlinux.topicbox.com/groups/zfs-discuss/T5122ffd3...) to identify some large latency numbers. Turns out, you need to disable ARC compression and "ADB Scatter Enabled" otherwise you will definitely hit some performance issues. Take a look at that thread - especially the very end where I publish some tuning suggestions. It was a fun learning exercise :-)
I find it useful to implement some sort of tag or link to wiki be added to alerts information [like on failed drive in your case] - makes much easier to every team member to fix the issue without guessing/invent own way during the outage.
Do you have any experience with hardware RAID provided by motherboard manufacturers like Asus? I have a spare dual xeon system I may utilize for my next NAS build and I'll pair it with two more RAID cards that can be put in HBA.
hugs zfs
... just use FreeBSD with OpenZFS for RAIDZ/RAIDZ2/RAIDZ3 pools instead.
... or even OpenZFS on Linux.
User error...both with btrfs and zfs ;)
Though I've used XFS a lot over the years. Mostly because the Debian installer gave 12yo me the choice between ext2, ext3 and xfs, so XFS it was because it sounded cooler.
Reminded me of XFS back when.
That's what redhat though ;)