Why Not ZFS (2021)
storytime.ivysaur.me
storytime.ivysaur.me
If I want no-worries and I don't want to be surprised by data loss, I use ZFS. If I want speed with data I can afford to lose, I use something else (ufs, ext4, xfs, the right FS for the job).
ZFS's integrity checking won't do a dang thing for you if you're not paying attention, don't have monitoring, or don't even run its checks. Yes, I over-provision. Yes, I'll make the reliability vs performance tradeoffs (when it makes sense, usually reliability over performance by default though).
The great thing about having options is exactly that. For me, ZFS is the right choice in most cases. For other people it's not. Being able to make an informed choice and not being forced either way is a good thing.
Zealotry/religion has no place in matters like this when there are clear tradeoffs in all the alternatives.
The hard case is databases. ZFS has a lot to offer (convenient support at FS level for replication stands out), but it is doing a lot of things that are solved problems in the design of competently designed databases and I never trust this kind of needless complexity. Of course if you really care about performance here, you should be willing to roll up your sleeves and tune the settings of the FS and ZFS is really nice in how it allows its complex features to be switched off.
I can kind of see your point, but I trust ZFS to never lose data, and I trust (in my case) postgres to never lose data, so the only issue is performance, and while that varies immensely, I mostly work on data that compresses well, so I can barely afford not to use ZFS with compression, because it saves a ton of space and actually improves I/O performance (if you're I/O bound, compressing your data lets you read and write faster than the physical disks can handle, which is still wild to me). Of course, that all depends on trusting all parts of the system; if I thought that ZFS+postgres could ever lose data, or possibly that there was a real risk of it causing an outage (say, memory exhaustion), it'd be a harder trade to make.
Very happy Ubuntu/ZFS user here, but we use only the LTS versions of Ubuntu.
Indeed. I was disappointed about the low quality of the article. A good article on why not ZFS would have been an interesting addition, to help users decide.
I've been using ZFS on my home NAS for over a decade and overall it's been a great experience, but as you say ZFS does have some limitations which makes it a poor fit for certain use-cases.
The memory cost behind Dedup explains FUD about memory cost of ZFS, arc will consume free memory, sure. But it also ejects itself properly like any memory hog and you can tune it.
The 'never in tree' thing is hardly ZFS's fault. ZFS is fully in-tree in FreeBSD.
"there are other choices" is a fine message. I chose to use ZFS for convenience of snapshots as a mechanism to drive backup to a cloned zfs disk I hold offline, as part of my 3-2-1. I also deliberately bought a larger memory device to scale to the burden. At work we use SSD to front for the cost of write, and we get good scale speed backing a DB and large filestores (large for us is still only terabytes, but I know of petabyte instances multi-zvol elsewhere in the world)
iX systems offered us support and we grabbed it with both hands. I have no complaints about maintenance and SLA on this product.
I have migrated zpools Linux-BSD routinely. It doesn't depend on RAID card specific semantic marks, CARD BIOS level config, Drive order in the frame, or "quirks" in the OS beyond conformace to a flagset. If you upgrade flags you can be stuck but we checked before upgrade.
I have lost data in JBOD, in UFS, in EXT, in ZFS. Nothing is perfect. I have lost data in soft RAID and in hard RAID. Nothing is perfect.
I do something similar with Sanoid/Syncoid and Sanoid snapshots are super easy to hook into with Borgmatic so you can have an alternate backup set that has nothing to do with ZFS.
The very first system I ever installed ZFS on had a consumer grade motherboard and very soon after installation I realized one of the SATA ports would get flaky under load because ZFS kept spitting up errors. The controller didn't report errors and happily wrote garbage to the disk from what I could tell. So, IMO, checksumming isn't worthless and I keep all my important data on ZFS these days.
I also disagree with the layering thing. I know it's the "unix way", but chaining together a half dozen independent system isn't something I find appealing. At that point I'm the biggest risk to my own data because the odds of me making a mistake are higher than the odds of hitting a bug that affects data integrity. I'd much rather have a single coherent interface to deal with.
It was the same thing with SystemD. Everyone complained about it "taking over everything" instead of chaining a bunch of existing uncoordinated systems together, but I can't imagine going back to the old way now that I'm used to SystemD. It would be nice to see a SystemD style initiative for desktop Linux.
Citation needed. I've found ZFS recovers faster and is more usable in degraded mode than an equivalent mdadm raid.
> Buggy
Compared to what? ZFS has a better record of not losing data than anything else, mdadm and ext4 included. Having a large number of bugs in your bug tracker is a poor measure of how buggy your system is.
> Scrubbing simply needs to read every file from the disk so the RAID layer notices and repairs a URE. You can simply put cat /dev/array > /dev/null on cron once a month which is enough for mdadm to notice and repair UREs.
If you do that it will swamp your disks and make that filesystem system unusable once a month, and it will take longer than it should to notice UREs. And if you reinstall your OS you will probably forget to set it up again and not notice until you have a disk failure and lose all your data.
> Checksumming is usually not worthwhile - the physical disk already has CRC checksums at the SATA level, and if you are paranoid you should also have ECC ram to prevent integrity issues in-memory (applies to ZFS too), and this should be enough. But you can easily get this if you want, either at the block layer with dm-integrity (integritysetup) below your disk or btrfs does it automatically.
Checksumming is essential to having a rebuild process that actually works. With mdadm when you rebuild you will probably get silent corrupt data in a couple of files because anything that went bad since your last scrub (if your scrub setup is even working) will not be detectable.
> btrfs does it automatically.
btrfs doesn't have stable support for raid-like modes, and has exactly the same kind of vertical integration that this author doesn't like about ZFS.
ZFS was the thing that convinced me that layering isn't actually always great and sometimes vertical integration makes sense. The file-level RAID is a lot nicer to work with in practice. The integrated tooling is a lot nicer to work with in practice than having to manage the md, lvm, and filesystem parts separately. I believe the out-of-tree thing might be an issue for Linux, which is part of why I'm much happier running FreeBSD.
I can confirm, it happened to me. A disk was corrupting my files and mdadm had no problem propagating the errors. Later switched to ZFS, and checksum error happen (defective SATA cable IIRC). By the way, the author suggests ECC RAM "if you are paranoid", good suggestion, it goes particularly well with ZFS checksumming, and I experienced faulty RAM too.
I switched to ZFS because I actually lost data with a mdadm/ext4 system, I didn't lose data with ZFS even though I went through broken hardware.
At least that is my understanding.
Anyways, see these talks from the ZFS Developer conference about how scrubbing has been improved:
* 2016 - adding batching to resilvering to improve performance: https://openzfs.org/wiki/Scrub/Resilver_Performance
* 2020 - for mirrors (not RAIDz), Sequential Reconstruction can further improve resilver speed: https://openzfs.org/wiki/OpenZFS_Developer_Summit_2020
The funniest thing is that ZFS is layered. It's just that 99% of people only see the final result, and don't notice the one place the abstraction leaks a little (zpool and zfs commands operate on separate layers with a bit of overlap).
ZFS is composed of separate layer for actual block devices, a layer of object-storage system on top of it called DMU (those two layers overlap a bit in handling data safety), and finally on top of the object storage there is implementation of posix-compatible filesystem called ZPL, an emulation of plain block storage called ZVOL, and there is optional LustreZFS which also operates on that layer.
There's a bit more segmentation when dig deep into code (whole special I/O scheduling and processing system called ZIO, for example, which provides things like encryption and compression) but yes, ZFS as a whole is quite layered - just with different APIs in between than offered let's say by Linux (you could technically build ZPL on top of now removed SCSI object storage driver, for example, but it would lack things)
Btrfs raid 1/10 is rock solid and offers similar robustness to ZFS, ie >= hardware raid.
It is a game changer being able to know not only which side of a mirror is correct (checksum), but also which side is the same age as everything else (generation).
I was mind blown that btrfs raid could handle a disk going completely offline for hours (ssd firmware bug) and then reappearing, without any kind of re-sync needed. It didn’t even ‘cheat’ like other raids by marking the disk as offline - it kept using it (automatically fixing any inconsistency) even before I ran a scrub. Thanks to the generation values it could tell that some reads coming from the desynced disk were wrong without needing to read both disks on every read
My understanding is ZFS (or any other raid) would offline the drive in this scenario, if this is true then btrfs raid is arguably better. There is also no threshold where btrfs gives up, like what LinusTechTips experienced with ZFS
There seem to be some foot guns present. See Jim Salter's article, specifically the section "Btrfs RAID array management is a mess":
* https://arstechnica.com/gadgets/2021/09/examining-btrfs-linu...
Btrfs was lunched in 2009, and there is still so much meh about it. I've been using ZFS since Solaris 10 (2006) and all the features have done exactly what they say on the tin.
The other concern seems to be the auto mounting of stale disks, which is by design and works a treat
In my experience a much bigger footgun is that btrfs raid makes files with CoW disabled completely unsafe (no sync between disks even when scrubbed), and some distros and programs (systemd-journald) selectively disable CoW by default
Maybe (though even then I've heard talk of bugs), but the article mainly talks about raid5/6-like modes and those are still marked as unsafe AIUI.
> My understanding is ZFS (or any other raid) would offline the drive in this scenario, if this is true then btrfs raid is arguably better.
ZFS can certainly handle a large number of read errors (recording them but remaining running), but if you reach the point where a device node completely disappears then it won't automatically re-add it when it comes back (you have to explicitly "zpool online" or reboot, then the pool will be imported as dirty and recover). I don't know the full details of exactly what ZFS does in every scenario but to my mind having a level of error at which you offline a drive seems pretty reasonable - once a drive is completely broken it's a waste of everyone's effort to keep retrying indefinitely.
I check the zpool status and saw that the system had been running on just 2 disk for probably a month due to write errors on one disk. Of the two remaining, another was having occasional errors that ZFS was correcting.
I pulled the bad drive, zeroed it out three times, and reinserted it. ZFS performed a resilver. After it was done, I pulled the other drive, zeroed it, and added it back.
The drives (SMART and ZFS) are no longer reporting any errors on any of the three drives and I only lost partial data on 4 files (which were replaceable).
Overall, I was surprised how resilient ZFS was and how easy the process was to replace 2 drives in a 3 drive array with minimal data loss.
Seagate drives by chance?
Edit: wait, I think I get it... "Scrubbing simply needs to read every file from the disk so the RAID layer notices and repairs a URE". So it's to avoid bit rot? I have to say, as someone who has a few TB of personal data on a NAS, bit rot is a bit scary. Data backup is mainly why I bought the lifetime pCloud+encryption package during the last Black Friday. I wonder how (if?) they avoid bit rot?
Yes,
it's also trivially to automatize in more or less all situations.
The gotcha is that you have to do it/know about it...
LinusTechTips recently did a video about how they installed ZFS on Linux and didn't have a scrub cron. They started to lose data before noticing.
Quite easy for me to setup, even though its my first NAS that I built myself and first time using ZFS. Very surprised that LTT effed that up to be honest.
The problem with by default installing a cron job when installing ZFS is that for a general purpose OS there a good default for when and how often to run it. And running it on the wrong time might even be a major problem.
Through then tbh. having a bad default is probably still better then no default in this case.
> lose data before noticing
Is a bit of an overstatement as they didn't look for quite a while, they also did not only fail to do scrubbing, they also failed to setup automated health checks and reporting.
Turning a non NAS focused Linux distribution into a well working and tuned NAS isn't easy (compared to using a good NAS OS/distribution), but making it somewhat work is easy. Which makes this a pretty common mistake for non-specialized people (i.e. like in their case).
If you don't know about ZFS scrubbing, but are using ZFS, you may wish to spend some time researching ZFS some more.
> I find it a bit scary that your have to do monthly manual work on your NAS or you'll lose.
Most distros / ZFS packages set up a cron job and e-mail out any errors. You do receive error reports from your NAS, right?
And yes, Seagate SMR drives (obviously not ideal but I'm cheap)
Anybody running 100 TB+ ZFS arrays should really have been aware of this class of problem before deploying (and doubly so if using consumer drives). ZFS protects from bitrot and hardware failures... but it's not magic. If data is written once and not read again for a long time, you won't know it's been sitting on bad sectors until it's too late, and you may end up with multiple failing drives at once if you lose a disk and try to resilver. Far better to check periodically so you can throw bad disks early.
Hopefully they had a good backup strategy!
I'm yet to see anyone make a really good set of resources to either watch or read that are not much too technically deep.
A lot of us have "learned the hard way," as it sounds like Linus himself eventually did. I think this highlights an issue with the "learn through youtube video" approach. An internet celebrity may acquire enough knowledge to do accessible demonstrations, presenting totally valid, useful, and correct information in them, and still miss crucial "unknown unknowns" that they simply hadn't encountered in their own research.
It's hard to know what to recommend for a class of issue you aren't even aware of!
[0] https://blogs.oracle.com/oracle-systems/post/disk-scrub-why-...
All that to say yes they lost all that data, no it wasn't backed up. Not critical to the business.
Let's start out with what the article spends the most time on: licensing.
This is a giant self-own and a very long winded way of saying "Linux will ignore the best implementation of a thing because we/they don't like the license."
That's a choice, not a law of the universe. It's perhaps also a canonical example of spite-induced nose cutting.
Let's see... other things... oh! Suggesting mdadm is better than well... anything... is a giant red flag to me. No thanks.
Dismissing checksums as unnecessary? Okay, now I think they're just trolling. Years of running ZFS teach you that bits flip, a lot... and you have no realistic recourse without block level checksums.
> That's a choice, not a law of the universe. It's perhaps also a canonical example of spite-induced nose cutting
It isn't because they don't like the license. The licenses of linux and zfs are incompatible, so zfs is out of tree and that leads to technical problems. No, it isn't a law of the universe, but it is the law of the United States at least, and probably many other countries as well.
> probably many other countries as well
Maybe. Enforceability of the GPL outside of the US is flaky, so whether or not this specific case would matter is always up to the courts to decide.
Well at this point, even if they wanted to change the license of Linux, they probably couldn't. Every contributor would have to agree to it, which is unlikely to happen, even if you ignore the logistics of asking everyone.
Not meaningfully; Linux can't change away from GPLv2 (too many copyright owners), and ZFS can never change from CDDL (again, too many owners, including Oracle), and those licenses are incompatible. The only way Linux could make it work is by freezing all the kernel APIs that ZFS uses, and while that is a choice, it would be an extremely expensive choice to give up being able to refactor anything that it touched.
It has nothing to do with "dislike." It's illegal to put ZFS into the Linux kernel. Just as illegal as putting a pirate copy of Microsoft's NTFS in there.
This is because of the choice of license by the owners of ZFS code!
You're blaming Linux but Linux license was chosen long before ZFS even existed. ZFS license was specifically chosen for the purpose of keeping code out of Linux.
No it's not, and you don't have to put it into the kernel.
>This is because of the choice of license by the owners of ZFS code!
Everyone else with a real free license can do it (BSD), so i would suggest that GPL is the problem and not the other way around.
>ZFS license was specifically chosen for the purpose of keeping code out of Linux
Good, it should no go into that rotting mammoth-pile of code.
And no, it's not illegal to put ZFS into linux kernel, because GPLv2 applies only on distribution.
Oh, then I guess it _IS_ in the Linux kernel already.
You're not being serious.
It's hard to determine whether GPLv2 would apply on distribution and on what kind of distribution, even.
But your comparison with "pirated NTFS driver" was completely out of whack, suggesting that it's illegal to even do it personally (as would be the case with "pirated NTFS driver"). GPLv2 applies only on distribution. Always did, Always will.
No it absolutely wasn't. It was using a symbol that saves the FPU state, and then that symbol was deleted in favor of a GPL symbol that does the same thing. There is only one abuse in the situation, and that is the flat-out lie that triggering an FPU state save intertwines you so deeply with the kernel that it makes your code derivative of it.
deliberately
As someone who _has_ operated many systems with both filesystems - for a long time - it’s malfeasance to shepherd someone in that direction. btrfs has its uses, but there’s no comparison in terms of project maturity. Sharp edges abound (behavior at high usage, RAID immaturity, the still-extant 5.16 kernel single-core max CPU use bug, etc.)
The kernel hotfix from a few days ago to break an infinite loop regression:
https://www.phoronix.com/scan.php?page=news_item&px=Linux-5....
And the announcement last year of officially giving up on RAID5/6 support, after not discouraging the totally dangerous feature for years:
https://www.phoronix.com/scan.php?page=news_item&px=Btrfs-Wa...
I would love to have options besides ZFS for check-summing and snapshots, but BTRFS isn't for people who care about their data.
Still use it on single drive filesystems for inline compression sometimes though.
I use ext4 mainly because it's the "blessed" filesystem of linux. It's simple, does what I want, and I have yet to be bitten by it. That's good enough for general use in situations where I don't care about file integrity.
ZFS marks the first time I've been able to breathe easy, knowing that my backups aren't storing corrupt files because of a raid write hole or hardware silent corruption that I won't discover for months or even years (likely long past any restore point on my backups).
Don't trust the hardware. At all. Hardware fails all the time, often silently, and can't be audited because the code and silicon are closed source so you have absolutely NO indication as to its quality (other than Backblaze reports [1]). You'll never know what demons lurk in those depths, but with paranoid software like ZFS, you don't have to care.
The only time I use non-ZFS filesystems is when I don't care about data integrity on that drive, such as pushbutton rebuildable server boot disks with backed up configurations (NixOS is great here), or cache/scratch drives. If it ever becomes easier to make ZFS boot disks, I'll probably start using them for server boot disks as well, just so that I can be alerted whenever a drive starts to fail.
You can expand ZPOOL with RAIDZ Expansion:
- https://arstechnica.com/gadgets/2021/06/raidz-expansion-code...
- https://github.com/openzfs/zfs/pull/12225
By the way the author sounds like he is in love with LVM and Linux ecosystem and does not understand ZFS at all ... while he also recommends BTRFS on LVM for the same features.
If you do now know ZFS do not read it - it will only bring false information into your mind. If you know ZFS then you can go read an laugh to make your mood better :)
It looks like you didn't read the article, either, because author already mentioned the very same links in the article
> Growing a RAIDZ vdev by adding disks is at least coming soon. It is still a WIP as of August 2021 despite a breathless Ars Technica article about it in June.
However, ~2 years ago I installed my laptop with encrypted ZFS, and frankly it kind of sucks. I don't know if this is the SIMD thing, or something else, but I can basically count on 1 full CPU running some zfs process at all times. Any apt install takes a stupid long time because it's doing something with snapshots. And I think I've only used the snapshots a couple of times. When I did though, it was handy.
But, I can't complain about my track record with ZFS: Over ~15 years I've never had data lost by it, despite having some horrible things happen to the systems that have run it.
I will fully agree that dedup is effectively useless. I've never had a system that had enough memory to reliably run dedup on any real workload.
I'll disagree that there are other tools better at doing what ZFS does though. In particular, I've had 5 events over the last year where our (normally reliable) PERC storage systems experienced corruption on LVM and RAID-6 arrays. I believe it was related to some issues with the Dell drives and Kafka broker activity on the arrays, but it sure would have been nice to have had ZFS on them instead of hardware RAID+LVM.
Based on my (limited) experience it was really bad before the whole SIMD thing got resolved. There was a huge improvement once they worked around that.
It'd explain why you get different results on NixOS.
EDIT: Seems like it was, but has been removed as unnecessary on newer ZFS versions.
2710046 root 20 0 1121M 85744 6848 S 102. 0.3 0:05.76 /sbin/zsysd
So if that is the only process spinning, that is probably specific to them.
(Not that OZFS native encryption doesn't have flaws; I am probably the last person on the planet to pick to argue that point. But I don't know that this is among them.)
(The command accepts a `-p <pid>` parameter; without that it'll profile the whole system (I make this sound much heavier than it is).)
- My primary need for it is on my backup box, which does "rsync --inplace" to keep small changes to large files from completely creating a new copy of that file every backup run. Hardlinks would cause all of those to compete if some previously similar file started getting updated (OS updates or similar starting files that then diverge.
- Each system gets it's own ZFS, so I can snapshot them at the end of the backup, keep different retention times, etc. But I can't hardlink across filesystems, I don't think.
The linked article did mention increasing dedup block size, which might help. If I could even cut down the DDT by half, that would make it more doable.
Not inline dedup, but a number of the nicer benefits, and usually cheaper.
[1] - https://drive.google.com/file/d/1csE8OuPotfhaFi9KvTGKMGy86Kx...
Encryption is unrelated to any snapshotting activity. Sounds like you may have additional hooks in place. ZFS by itself does not integrate in any way with package management or anything else, really.
Sure, as with other FS it has trade-offs and may not match all workloads or use cases, but it also has a maturity and no native Linux file systems has such a thought out and complete feature set ZFS has, albeit btrfs comes closer every release which is nice to see. The one thing ZFS is missing is rebalancing, but there are ideas out to solve that, but those are not reasons for just not singling out ZFS and never considering usage in general. Alternatively fighting with proprietary HW raid controller is risking to get one's data eaten on a simple firmware upgrade, and having no resource to introspect or salvage, because they're just a proprietary mess.
So, while installing ZFS via DKMS or the like works out, in the end they'll always be a second class citizen in the kernel and spent a lot of time on fighting changes upstream, working in a parallel universe in the kernel and thus spending much more resources and effort.
See for example the ARC, that is allowed to use up to half of memory by default, but it just nowhere shows up in the native Linux memory accounting, yes, some specialized tools like arcstat exist, but if ZFS was actually mainlined it could benefit from actually working with the kernel, not against it half of the time.
I think that if ZFS would be mainlined, and sadly that seems very unlikely to happen anytime soon, a big part of the things causing users currently trouble when using ZFS on Linux would go away pretty soon just due to being natively integrated.
There are reasons why both kernel and ZoL teams say they are happy with not mainlining.
That's only parts of the reason, and in the simplest form you can just disable the VFS page cache and render fitting arcstat properties through it.
And besides that, I see no inherent issue in solving that, at least in a much more integrated way as the status quo is now..
> There are reasons why both kernel and ZoL teams say they are happy with not mainlining.
Yes, but that's the license and not the caching/VFS differences.
As Linus said:
> And honestly, there is no way I can merge any of the ZFS efforts until I get
> an official letter from Oracle that is signed by their main legal counsel or
> preferably by Larry Ellison himself that says that yes, it's ok to do so and
> treat the end result as GPL'd.
>
> Other people think it can be ok to merge ZFS code into the kernel and that
> the module interface makes it ok, and that's their decision. But considering
> Oracle's litigious nature, and the questions over licensing, there's no way I
> can feel safe in ever doing so.
>
> And I'm not at all interested in some "ZFS shim layer" thing either that some
> people seem to think would isolate the two projects. That adds no value to
> our side, and given Oracle's interface copyright suits (see Java), I don't
> think it's any real licensing win either.
>
> [...] the licensing issues just make it a non-starter for me."
https://www.realworldtech.com/forum/?threadid=189711&curpost...And OpenZFS also wants to do so, but legal blockers and Linux maintainers not wanting to do a dual-licensed CDDL/GPL either, is their blocker:
But pagecache is still something that isn't fixable (you can't disable VFS pagecache, it's too intertwined with the whole I/O system) - I don't see a way to implement ZFS "linux way" considering things like block sizes bigger than one page being antithetical in linux.
Cf. https://www.reddit.com/r/zfs/comments/ryaaqo/zfs_on_debian_v...
> Ubuntu ships ZFS as part of the kernel, not even as a separate loadable module. This redistribution of a combined CDDL/GPLv2 work is probably illegal.
That criticism retracted, although if the author suggests Ubuntu users are under threat here, that seems kind of fanciful to me. And while Oracle are hyperaggressive, I'm struggling to see that even they would think it in their interests to go after Canonical itself.
I have compression and encryption on and the fio benchmarks were all good, similar to what I saw with the Hardware Raid 1.
Regarding monitoring: It is not that difficult. (1) Have ZFS Mail status updates to you (e.g. this is the default in Proxmox), (2) Monthly Scrubs (cron) (3) Monitor TBW with (e.g.) InfluxDB. Wrote a blog post about the last part [1] "Disk Wear (SSD) - extracted from extended Smart Attributes (Single Stat)".
All of the issues with degrading disks apply equally to other disk setups (e.g. Raid1). However, ZFS gives you better tools to predict and prevent, before disaster happens.
[1]: https://du.nkel.dev/blog/2021-05-05_proxmox_influxdb/#config...
TL;DW: Data in their massive 1PB server was suffering from bitrot because there was no scheduled scrub to repair the bad data. And they couldn't tell how bad the situation was because without any scrubs happening, the stats on data integrity were inaccurate.
about the data loss in the video, mistakes easy to spot: 1st: using seagate. 2nd: installed by us and never updated 3rd: insufficient reading of docs before going all in on zfs. 4th: buying more seagate drives ;))
I think they had a way higher chance of losing their data going the usual stack mdadm/lvm/ext4/luks/btrfs, I think mastering those is harder than mastering zfs.
Although everyone else can learn the important lesson that RAID / ZFS isn't magic and you need to have stuff setup correctly and monitored. The fact that RAID isn't a backup as well[1].
[1] Although if the LMG servers affected are just for data hoarding raw footage that is unlikely to be needed again, it's possible the risk / cost balance pushes away from backups and just relying on RAID, but that's a niche case (and they lost the gamble...).
I'm certain at least that Linus knows that RAID isn't backup. And I'm Linus is going to try had to get the data back, but it seems to me that this isn't some devastating failure for him.
If you are building a storage array, do not do this. Ensure that you are using a variety of drive types (obviously same size and interface technology). Doing so guards against the danger of too many drives going wrong at the same time (within the same time window) causing a failure from which it is impossible to recover.
btrfs filesystem defrag -r /
You can't do that with ZFS.People can really hurt you with that.
https://www.usenix.org/system/files/login/articles/login_sum...
Well, this article is fascinating, but maybe the reason why btrfs has built-in defrag functionality is because without it, performance degrades incredibly steeply?
https://www.usenix.org/system/files/conference/fast17/fast17...
Glancing through, it seems ZFS is about middle of the pack of the analyzed filesystems (BetrFS Btrfs ext4 F2FS XFS ZFS) in terms of performance aging caused by change history. And only BetrFS did not show any aging from these tests and was surprisingly stable throughout. I'm curious if that's changed in recent versions of any the analyzed file systems, I know ZFS has had a major release recently.
Thanks for the USENIX link!
[1] https://www.percona.com/blog/mysql-zfs-performance-update/
Saying to be cautious of using zfs because some day there will be an in-tree file system to replace zfs is absolute nonsense.
What happened with btrfs? People waited for a decade and still no wide deployment.
And no one wants to use zfs-fuse. Everyone says the performance is a joke but yet, it's not mentioned in there.
Word of advice, try zfs than read this article.
Isn't Facebook all in on btrfs? https://www.networkworld.com/article/2367229/smooth-like-btr...
I did have some problems in early versions that seemed plausibly related to its contiguous memory needs, and later I saw a release quite drastically changed this (0.7), and indeed, it was stable from then on.
I found it more cohesive to work with than other options available at the time, e.g. l2arc for caching. I didn't have need for pool dynamism, but I found it rather easy to use, though I also have experience with lvm and mdadm.
That said, some of the source code I found more incomprehensible than average, and I'm fairly used to reading linux to answer certain questions. Maybe I just needed to study it longer.
E.g. Borg was worthless because the retrival of data was so slow from a dedicated hetzner SX6_ Host that getting data back with less than 10MB/s was disastrous for a 10TB+ repo. (Not a bandwidth problem.)
Been there, done most of the suggestions. Most of them are unpractical -- regardless what benchmarks at phoronix say and how often one jumps from solution to solution for certain aspects.
Still use ZFS, still IMO overall the best package for a lot of data management.
"Unless you have ECC", of course I have ECC, it's for my file storage.
That said, there are valid points and something to be considered, especially the licensing stuff is a bit scary.
Funding bcachefs development would help that as well: https://www.patreon.com/bcachefs
2. relicensing would kill the project
To expand on (2) - CDDL was specifically designed to make it easier to include in other projects while also keeping certain protections for all involved. The issue is that GPLv2 has somewhat complex case of whether it applies ("derivative work") and also tries to push you to distribute code under GPLv2 terms and doesn't play ball if some part of it has requirements that exceed GPLv2 ones.
The specific incompatibility is, iirc, due to patent litigation protections in CDDL, which do not exist in GPLv2, meaning you can't just redistribute CDDL code under GPLv2 umbrella
Unfortunately, it's IIRC the part that, according to some, makes it incompatible with GPL
I was able to get it working (with a few setup issues) with a RAID 1 mirror: https://github.com/microsoft/WSL/issues/6711#issuecomment-10...
(Forgive linking in the middle of a Github issue. Everything ended up working with the latest updates and being an administrator.)
Have had zero hosed drives with ZFS. Although with ZFS not being supported in the kernel, it's fairly hard to do a ZFS-on-root setup unless you're using Ubuntu.
YMMV.
The reason I don't use btrfs is I don't want to deal with problems I currently don't have. ENOSPC with space available, performance issues when space available, quotas.
I want evolution from where I am, not jump ship.
I want to depend on others on what to spend my time on and using in kernel on whether to "jump ship" is a filter I've decided. YMMV. Like I said, I don't use btrfs either.
I'm not aware of any distros except NixOS which will then proceed to roll back the entire transaction.
To the end user the functionality is identical whether or not it's in the kernel as opposed to just being a kernel module
https://openzfs.github.io/openzfs-docs/Getting%20Started/ind...
It is as easy as "Copy and Paste", but not "point and click" easy
My impression is that zfs is generally considered to be less buggy than btrfs, which is the filesystem that is most often used when directly comparing to zfs. Though I admit I mostly use ZFS on freebsd, not linux (and I think the article is mostly focused on ZoL)
ZFS compressratio: source: 1.40x (Separate pool) output: 1.68x-2.49x (Multiple datasets to make management easier)
The output compression allowed me to use a 1TB drive as a 2TB drive effectively, allowing me to store a lot more output and not have to wipe away build output from x to build y.
That alone makes zfs worth it for me (and I know many other Android devs who use it in a similar fashion)
Disclaimer: I work at Red Hat, though nowhere near filesystems.
I also think the article underestimates the value of compression, its insane for the file formats to include it at all, its absolutely a filesystem property. How much and how hard a file should be compressed isnt known on creation, its known by the server that is currently providing it to users, meaning precompressed files are always either to strongly compressed, or not strongly enough, and often using custom shitty compression. Its also great for many kinds of research and engineering workloads where data does not have built in compression, or generates intermediate files.
> No disk checking tool
I don’t understand this complaint. Is this not what you get by running scrub with checksumming enabled?
> High memory requirements for ARC
This section is simply wrong or at least very misleading: the main motivation for ARC today is not that ZFS can’t use the page cache but that its own strategy can perform better and more optimally (if tuned properly). Things generally do not get double-cached in both page cache and ARC at all. The memory usage can be tuned with various parameters including min and max sizes (though I do wish these parameters could be set per pool as opposed to on a system level). This section too seems to be in bad faith as a reader who is not deeply familiar with ZFS already will have an incorrect understanding.
> causing long periods of unavailability until a workaround can be found.
Has it really? The one situation I’m aware of is the one they’re mentioning. Which was not experienced any differently than the usual expected delay for things to bubble down to distro repos.
> btrfs
Is simply not ready for casual users yet. It will hopefully get there before too long but at this point I find it irresponsible to recommend it for non-expert users and production workloads.
OpenZFS is not perfect and there are valid reasons not to use it. You don’t have to misrepresent it to find them.
However linux pagecache isn't compatible with ZFS for many reasons (XFS was bit by it as well in the past)
For most of my usecases, I typically default to xfs over mdadm raid10 in f2 mode.
I love the ideas behind bcachefs but suggesting it's anywhere close to being ready for primetime is just silly.
https://www.phoronix.com/scan.php?page=news_item&px=Bcachefs...
> Why not bcachefs?
To which, "it isn't stable or anything like prod-ready, and you still have to build it as an out-of-tree module" is a perfectly good answer. Hopefully, the answer will become "yes, bcachefs is the solution" soon, but we aren't there today.
> Bcachefs can currently be considered beta quality. It has a small pool of outside users
>However, given that it's still under active development backups are a good idea.
With regular backups, of course.
RAIDZ expansion and disintegration of ARC and kernel page cache is valid problem what I wish solved.
The biggest downsides I can think of are:
1. You have to do an initial initial initialization of checksums on the disk, which takes forever.
2. Automatically bringing up the dm-integrity layer of the plain disks at boot time is a bit of a dark art involving udev rules.
If you skip the checksum initialization step and add the device to the RAID immediately, it _mostly_ works, because disk blocks are rewritten by the resync/reshape, but then on reboot it fails because some part of the system tries to read an otherwise unused sector (probably to find metadata or a partition signature of some kind?), and then determines the device is unusable because of read errors (dm-integrity errors show up as read failures to the next layer in the stack).
I've set up my RAID to not haeve a "bad block list" so I _think_ it will try fixing it a few times, and drop the drive if there are too many errors in a row.
Should be easy enough to try with a "test" RAID made out of a bunch of files (or even USB sticks).
Last time I checked btrf _seemed_ like a mess with many hard to reproduce subtil bugs and for was basically a no-go for any usage requiring any reliability. Like it looked bad to a point that I wondered if it only still exists due to the sunken cost fallacy.
But neither then nor now I was familiar with either btrf nor zfs, so I might have been wrong.
That's coming from someone who has been using Ceph for a couple of years now. I know it's apples to oranges, but still. (For one, zfs is non-distributed, whereas Ceph is).
Not minor quibbles or terminology or preferences - but flat out incorrect.
I post this for posterity / search engines / future readers: Ignore this post.
[0] https://www.youtube.com/watch?v=-zRN7XLCRhc&t=33m (the meat is closer to 35 minutes in, right around "What you think of Oracle, is even truer than you think it is.")
I've always figured it's at least partially because someone was too stubborn to give up their work creating btrfs. Though relicensing it to proprietary is Oracle.
See section 4: https://opensource.org/licenses/CDDL-1.0
Another issue is that what is specifically a problem with GPLv2 is that CDDL-1.0 provides patent litigation indemnity, losing which would be a big problem.
I mean, with an open source license, who could get sued, anyway? The current steward of the main repo?
CDDL has provisions that mean that Oracle can't sue you for (hypothetical) patent infringement in ZFS code because Sun gave automatic license to those patents when it released code under CDDL
> Out-of-tree and will never be mainlined
So the article is heavily Linux-focused. Fine. To an end user, having a driver that's not part of the mainline kernel sources is hardly a deal breaker. It still ends up running in kernel space with all the advantages that comes with.
>> Ubuntu ships ZFS as part of the kernel, not even as a separate loadable module. This redistribution of a combined CDDL/GPLv2 work is probably illegal.
Canonical lawyers have obviously disagreed with this assessment. So far so good.
>> Red Hat will not touch this with a bargepole.
Red Hat also doesn't touch btrfs. Or basically anything that's not ext4 and XFS.
>> You could consider trying the fuse ZFS instead of the in-kernel one at least, as a userspace program it is definitely not a combined work.
No, you really should not. zfs-fuse has not been maintained in over a decade, doesn't even remotely come close to supporting the features of modern ZFS, and frankly... it's FUSE. It's slow as molasses.
> Slow performance of encryption
>> ZoL did workaround the Linux symbol issue above by disabling all use of SIMD for encryption, reducing the performance versus an in-tree filesystem.
Only partially true, but the damage is limited to some metadata structures and the bulk of encryption code does use SIMD instructions (eg, the parts that encrypt your file data).
> Rigid
>> This RAID-X0 (stripe of mirrors) structure is rigid, you can’t do 0X (mirror of stripes) instead at all. You can’t stack vdevs in any other configuration.
Hard to say why it's useful. Hardly anyone ever chooses that configuration even in the non-ZFS world. It's the same capacity tradeoff for a different approach.
>> For argument’s sake, let’s assume most small installations would have a pool with only a single RAID-Z2 vdev.
Okay, not an item that can actually be refuted, but the idea that "most small installations" have only a single raidz2 vdev is a stretch. I'd wager there's a heck of a lot more single-disk and two-disk mirror configurations than all the other types combined.
> Can’t add/remove disks to a RAID
Everything here is accurate. Part of it is because of ZFS's original target audience, part of it is really hard math problems that hadn't been solved in a manner where you can actually pull it off before the death of the universe. Mind that mdadm isn't exactly magic either -- it'll frequently refuse some operations to reshape a RAID array, and it's not quite as upfront in the documentation about what those scenarios are.
> RAIDZ is slow
raidz has to make sure all the disks have the same set of data synchronized. This is a major reason people are generally recommended to use mirrors instead of raidz.
> File-based RAID is slow
ZFS does not use file-based RAID, it uses block-based RAID. Yes, it knows what blocks are used and will only need to scrub/resilver those blocks.
>> Sequential read/write is a far more performant workload for both HDDs and SSDs.
Yes it is. Which is why ZFS 2.0 introduced sequential scrubs.
>> It’s especially bad if you have a lot of small files.
ZFS isn't file-based, it doesn't matter if you have a single 2TB file or a million files. It's the same work either way.
> Real-world performance is slow
Comparing to ext4 isn't the most fair thing to do. ext4 is a dumb file system that will give you raw disk performance, every time. If this is the utmost priority, use ext4. ZFS adds compression, checksums, and redundancy to the mix. Extra protection means it's a bit slower.
> Performance degrades faster with low free space
>> It’s recommended to keep a ZFS volume below 80 - 85% usage and even on SSDs. This means you have to buy bigger drives to get the same usable size compared to other filesystems.
Basically every file system gets bad when in high-utilization, and probably should be taken as a sign to either upgrade the storage or start deleting.
The threshold for where ZFS starts getting painful depends on which anecdote you listen to. I've heard from people running up to 95% utilization on an SSD without feeling the burn.
>> ZFS’s problem is on an entirely different level because it does not have a free-blocks bitmap at all.
ZFS has had a spacemap feature since 0.6.4 and an improved version since 0.8.0. If the output of the "zpool list" command has a FRAG value other than a hyphen, your pool is using this feature already.
> Layering violation of volume management
Basically "but muh Unix philosophy" argument. It's tiresome :)
In order for ZFS to do what it does so well, it has to incorporate all these features that used to be in different layers. The reason that resilvering a ZFS pool is so much faster than the whole-disk RAID solutions of yore? Precisely because it knows exactly what blocks are in-use and what are not.
>> If you use ZFS’s volume management, you can’t have it manage your other drives using ext4, xfs, UFS, ntfs filesystems.
You can create volumes and store ext4, xfs, ufs, ntfs filesystems on top of ZFS just fine.
>> And likewise you can’t use ZFS’s filesystem with any other volume manager.
You can, but you probably shouldn't.
> Doesn’t support reflink
Some work in progress is around to do it, but not in a release version yet. Doesn't mean it never will have it.
> High memory requirements for dedupe
Indeed, and there are research projects at a new dedup algorithm to drastically reduce the memory requirement. Don't use dedup unless you really really need it.
> Dedupe is synchronous instead of asynchronous
Dedup is meant to work on live data so it does what you want immediately.
>> (By comparison, btrfs’s deduplication and Windows Server deduplication run as a background process, to reclaim space at off-peak times.)
Sometimes "off-peak times" don't exist, and this just highlights the limitations of the two mentioned technologies: they don't have an online dedup mode. I know at least for btrfs, you have to completely take the file system offline to do a dedup pass after-the-fact (and the end result is the same as the aforementioned reflink).
> High memory requirements for ARC
Basically a long rundown of Linux's own page cache fighting with the ARC. It's actually a fair point, but it's probably overblown. It's nowhere as dire as the section makes it out to be.
>> Even on FreeBSD where ZFS is supposedly better integrated, ZoL still pretends that every OS is Solaris via the Solaris Porting Layer (SPL) and doesn’t use their page cache neither. This design decision makes it a bad citizen on every OS.
FreeBSD's technical implementation of how they ported ZFS doesn't really matter, and this is the second time the article has said "supposedly better integrated" -- it's not supposed, it's literally as well-integrated as ZFS on Solaris is. (Guess what? Solaris still has UFS too, it's pretty much on-par with FreeBSD)
> Buggy
The flimsiest argument of them all :)
Bug trackers track bugs. Some of them aren't even bugs (as the nature of user-submitted bug reports are). It goes more to show the popularity and widespread use of ZFS than anything else.
I'd be far more concerned about a software project that has no bug reports on display at all.
> No disk checking tool (fsck)
>> Yikes.
There is, it's called scrubbing. See "man zpool-scrub" for details.
>> In ZFS you can use zpool clear to roll back to the last good snapshot, that’s better than nothing.
"zpool clear" is an administrative command to wipe away error reports from storage devices. It should only be used when an administrator determines that the problem is not a bad disk.
ZFS pools can be made of many file systems with any snapshots you desire. There is no "the last good snapshot". Maybe zpool checkpoints are what the author is thinking about, but I doubt it.
>> merely rolling back to the last good snapshot as above does not verify the deduplication table (DDT) and this will cause all snapshots to be unmountable
That really should be impossible. I've never even heard of such a thing happening.
>> coupled with the above point (“Buggy”) if ZFS writes bad data to the disk or writes bad metaslabs, this is a showstopper
This is an error that is detected and provided as part of the "zpool status" command (and as mentioned, "zpool clear" can even clear the errors).
>> and so it should have an fsck.zfs tool that does more repair steps than just exit 0.
You could replace it with one that does a scrub, but that can take weeks on some pools :)
> Things to use instead
>> The baseline comparison should just be ext4. Maybe on mdadm.
If you think mdadm+ext4 is comparable to ZFS, you are waaaaay off. Even btrfs can't hold a candle to ZFS and it comes closer than mdadm+ext4.
Here in this section, the author comes across the term "scrubbing" but doesn't really apply it in the way ZFS uses it.
>> Compression is usually not worthwhile
Hard disagree :)
>> Checksumming is usually not worthwhile
Disks lie. All the time. The claims of "physical disk already has CRC checksums" was around over 20 years ago when the ZFS project started, and the fact that disks lie or do not do strong protection is a huge reason that ZFS was created in the first place. The problem in 2022 remains the same as it was in 2000.
> Summary
>> you can achieve all the same nice advanced features
You really can't. It's not even close. ZFS is so far ahead of the game, that even if some alternatives (eg, btrfs) offer a few similar features, they don't even approach it.
>> ZFS also has a lot of tuning parameters to set.
Having tuning parameters and requiring them are two different things.
>> In the future we’re waiting to see what stratis
stratis is a dead-on-arrival joke.
>> bcachefs
Probably the only thing that has a shot at competing with ZFS.
My approach is a login script that tells me the health of my zpool. My crontab has a "0 2 * * 0 /sbin/zpool scrub tank" and 6 variants of "0 2 2 * * /usr/sbin/smartctl --test=long /dev/sda &> /dev/null". I've learned to resilver a dead drive recently, it's on a UPS, and I have automated iterative backups elsewhere for critical data that I've practiced restoring from to verify my solution works. Never had much luck with email alerts unless I want to get all of crontab's emails sent to me.
Could also look into using zed for getting email notifications about pool problems as soon as they happen.
I'm using https://habilis.net/cronic/ to make sure I don't mess up the email notification part of the cronjob. It's a simple wrapper script that sends an email in a readable format if a cronjob fails.
I use Sanoid to create snapshots on my home server, and use Syncoid to push those to a cloud VPS with a beefy network drive as an off-site backup. Both tools are available here: https://github.com/jimsalterjrs/sanoid
The free tier of https://cronitor.io/ makes sure I'm alerted if a cronjob fails, or fails to run on time. Especially that last bit is interesting: that way I'm sure cronjobs aren't silently failing for days/weeks on end.
I have 4 monitors set up in Cronitor: snapshot creation, zpool status on the local and backup machine, and send/receive with Syncoid. This is how that looks on the Cronitor dashboard: https://img.marceldegraaf.net/v6IpNAyxrZ54vgLqIJpY
Let me know if you want more info or examples, happy to share whatever I can to help :-)
EDIT: feel free to reach out via email as well, my address is in my profile.
Could you elaborate a little bit on this point. I've used both btrfs and zfs and they have both been fine, but I'm not using them on a large enough scale to see problems.
It can go on and on. btrfs makes an attempt to compete with some of it, but after 14 years of development, it's still not even close to ZFS's first public release (which was 4~5 years after internal Sun development)