Linux 5.10 BTRFS performance regression
reddit.com
reddit.com
[0] https://bbs.archlinux.org/viewtopic.php?pid=1943906#p1943906
Which triggered an immediate 5.10.1 release: http://lkml.iu.edu/hypermail/linux/kernel/2012.1/07005.html
Thanks for your hard work, Greg!
This isn't "negativity", this is people not understanding how the process works :)
And you're welcome!
Merry Christmas/Isaac Newton's Birthday!
Bugs can always be fixed, mental health can't.
Thanks for all the awesome work Greg!
Thanks for all the hard work!
Anyways, Linux needs some more CI so that such bugs can be found during the RC phase.
I'm a software engineer who's not involved in Linux Kernel Dev... but I've got a stack of old laptops that I'd be happy to set up to run automated CI if that'd be helpful.
Is there a webpage or doc somewhere I can look at?
(I'm not trying to snark - the fact that you're you and you're here asking for help is making me want to dip my toe in).
Second-simplest thing to do is to run the linux-next branch/tree on your machines and report any build warnings and runtime issues you find. That's what will be the "next" kernel releases and is where all of the developer/maintainer trees are merged together before they are sent to Linus.
Both of those should be very easy to do, and any problems found there should be easy to fix and resolve before they get to a "real" release.
Back when I was building kernels for embedded hardware (Sheevaplug) in the 2.6.33 timeframe, I found a USB audio regression between 2.6.33.7 and later versions. If there were a semi-turnkey way to set up a testbench that could automatically reboot hardware in every new kernel, run through some basic tests, and report any deviation, I probably would have been more likely to do so. At the time I was working solo trying to release a polished consumer product (sadly though the product was released the business didn't work out) and didn't have time to dig into and report bugs.
We have the 0-day bot from Intel that runs so many things on all developer trees. We have kernelci running on many many different hardware platforms, and we have Linaro test systems also running on many different branches and hardware platforms.
If you want to tie your own hardware into the system, kernelci is the best place to start, I recommend looking into that.
thanks!
Despite appearances, "the kernel" is not a single monolithic thing. There is a about a 100 kloc core (but I haven't looked up that number in years). The rest, hardware drivers, network protocols, file systems, crypto, raid ... bolt on as modules.
Those modules are maintained separate teams. They are as related to the kernel as the phone dialler app is related to Android. The quality of each module is the responsibility of that team, not "the kernel" team. And that applies to testing the module as well.
In a sense, "the kernel" team is more like debian or redhat than developers. What they have done is develop a framework that lets them take bits created and maintained by a cast of thousands, and bolt it together into what appears to be a single coherent thing from the outside. So the answer to "how is the kernel tested" is "it's complex, and not centrally planned".
The other answer is what you are seeing is in fact part of the testing process. Most people use kernels packaged by their distribution. kernel.org releases are more like Microsoft's pre-releases of Windows. Most Debian users for example won't see it until it gets to Debian testing. To get there it must pass through Debian experimental (which is where 5.10 sits now) then sit in Debian unstable without bug reports for a while. Those release names should give you a hint about the anticipated stability of the kernel version. I personally won't use it until it takes another step, which is from Debian testing to Debian backports (which is when it because available to Debian stable users who are willing to risk compatibility issues).
This means that for for most users, 5.10 it's done yet as it has barely begun it's testing regime.
Both do checksumming and regular scrubbing to help detect bitrot, but both need sufficient redundancy configured to actually be able to repair the corrupted data when it is detected.
apt install dkms spl-dkms
apt install zfs-dkms zfsutils-linuxIt’s like security - it’s done in layers.
Using a checksumming filesystem without ECC RAM is an improvement on not using a checksumming filesystem at all. Do not consider ECC RAM to be a requirement for ZFS / BTRFS.
It does add another layer of improvement though, so if you can, you should!
And even if you can’t, and even if you can’t provide the redundancy to enable bitflips to be resolved, still use a checksumming filesystem if you can!
I have problems with USB packets getting corrupted on their way to my audio interface. Signal integrity is really underappreciated in the PC peripheral world.
The SMART thresholds on the drive were strange too. A small but non-zero amount of read errors was not considered a health problem yet. I only realized that after the scrub reported problems.
How did you avoid the crippling performance penalty with storing VM disks on btrfs? The usual workaround suggested to btrfs users is to disable copy-on-write for VM images, but doing so also turns off checksumming thus disabling the very features that make btrfs, well, btrfs.
My next storage setup will definitely use a checksumming filesystem.
ZFS has similar features and has eaten 0% of the data I have stored on it. would suggest.
I think ECC RAM and a good FS are complementary tools to achieve storage high reliability. One without another may not be that useful.
The CIGAR is not a checksum, even though you'll notice if it's corrupted, because it won't match the read length anymore. However, BAM is zlib-compressed (even if you asked samtools not to compress it). Each block therefore contains an actual checksum, and corruption will show up during decompression. Therefore, if you can decompress BAM, but the CIGAR doesn't make sense, you've got a software problem, not a hardware problem.
From my experience in beerinformatics, you got some software that... uhhh... interprets the BAM specification differently than you do. It's depressingly common.
Those bits could still flip in memory, which isn't BTRFS's fault. That's still unlikely, because you'd typically get segfaults, not silent corruption. So, it's probably bad software.
> those algorithms should be deterministic, right?
Combine bad code (samtools) with multithreading, and you will quickly become convinced that your computer is possessed by an evil spirit.
Currently not using btrfs by the way. But you are positive that BAM files should be self consistent? And when would you catch a bitflip then? What would it look like in a compressed file lake BAM? A much larger effect than just a cigar length mismatch supposedly?
BAM is a sequence of zlib-compressed blocks. Each of them has a checksum, and any corruption gets past it with probability 2^-32. Any sane software will give you an error. When in doubt, just decompress the stupid file using gunzip or zcat, those will verify the checksum. Don't guess, check.
But if you insist on using software that just ignores the checksum (I think, not even samtools, which segfaults(!) on misspelled file names(!) is that awful), all bets are off. Decompressing a corrupted file yields nonsense that doesn't even have the same length as the original data. BAM would also quickly get out of sync, and subsequent records would contain complete nonsense. You'd know if you saw that.
However...
> you are positive that BAM files should be self consistent?
That is a loaded question. In addition to being syntactically correct, BAM files should maintain a few variants that are hard to check. They rearely do. BWA(!) produces inconsistent BAM files, and I've seen many, many scripts in the wild that produce complete bullshit that only happens to work in a very specific setting. Like I said, you're probably using some awful software.
I've used both ZFS and Btrfs extensively. Only Btrfs has lost data and behaved badly. I would highly recommend evaluating ZFS.
It will be investigated: https://marc.info/?l=linux-btrfs&m=160869337604422&w=2
Would Rust solve the non-speed issues? Rust-in-kernel discussion from August: https://lwn.net/Articles/829858/
In kernel, you can run a privileged cache or mmu instruction or a write to some magical memory position and all the sudden the "normal" rules don't apply anymore.
(But I think there are other parts of rust that are nice to have in kernel or any complex software).
I didn't know this requires certain features which are not available inside the kernel. I only knew all existing interfaces may be unsafe because they are in C though. Rust does not seem as useful then.
Thank you for your input.
I didn't think about the hardware issues, hmm. I can't see how to do that, when the compiler guarantees get invalidated by hardware. Checks are also needed like in C? (assuming there are checks which do not get compiled out..)
EDIT: also see panpanna's comment
This are tools. A screwdriver is not the right tool, when you need a power drill. And vice versa. Yes, sometimes you can use both and stick with your accustomed tool.
Anyway. C and C++ and the tool chain are constantly improving like others.The moern memory sanitizers in GCC and LLVM are awesome.
Are they not solvable because the kernel does not give enough guarantees as it gives to userspace, because the c-interfaces of the kernel have to be wrapped in unsafe or because of other reasons (architecture, data model, kernel constraints, ...)?
The issue will probably be some sort of pathological case in an algorithm being used, or perhaps from a poorly chosen algorithm. The point being, it’s not clear yet, and to solve that requires understanding the problem, not effecting a needless rewrite in a new language.
Agreed.
> The point being, it’s not clear yet, and to solve that requires understanding the problem, not effecting a needless rewrite in a new language.
Which is why the first question is why btrfs has issues others do not have. Some mention its CoW architecture, system design, too many features and not limiting storage to 90% and so increasing complexity. Others mention usual kernel issues.
Others mention that they had no issues to begin with and that its mixed reputation is unwarranted. I'm not clear who is right, but they are data points.
Thank you for your input in cautioning of rewrites to avoid needless work, I appreciate it.
As someone that has built infrastructure on BtrFS for years, the scary stories are mostly just hot air and the stability of other filesystems is really not significantly better.
Bugs like this happen, this is why Linus releases many release-candiates every kernel, this one got through as 10 was a rather massive kernel and there were several regressions. Including one that caused a new release just hours after the supposedly final one. Distros wait a bit longer before shipping an updated kernel and none of these hit actual users.
As far as I know there is no reason to abstain from using Btrfs. When Fedora talked about not using it, they had as reason that they had no in-house expertise.
I've lost 2 root filesystems to btrfs, on a laptop with only 1 drive (read: not even using RAID). Have you considered that you're just lucky?
I can't speak to if there are other foot-guns waiting around or how common problems like this are because we migrated back to FreeBSD and ZFS shortly after that experience. I do know they have since updated BTRFS to make that scenario less likely (but still not impossible).
Btrfs is the C++ of file systems: it’s powerful and works for a great many people. But the tooling is intimidating to new comers and unless you know exactly what you’re doing, it’s a ticking time bomb due to the plethora of foot guns and hidden traps.
This is why some people claim to have success with it while a great many other people, rightfully, claim it’s not yet ready for prime time.
ZFS, on the other hand, has not only protected me against failing hardware but it also has sane defaults and easy to use tooling thus protecting me against my own stupidity.
We expect filesystems to work robustly. We do not expect them to fail after an arbitrary time interval merely by being used. Even terrible filesystems like FAT don't do that. They might get fragmented and slow, but they don't just stop. I find it incredible that this is often minimised by people; it's a complete show-stopper irrespective of the other problems Btrfs has.
I made exactly the same migration you did. ZFS has been solid, and it does exactly what it says on the tin.
Mostly because it has lots of features and as a consequence, is pretty large and complex. Closer to ZFS than ext2.
Btrfs suffers from a initial bad rep, which is difficult to overcome.
Granted, there are some limitations to the design, but it doesn't affect my use cases, so whatever...
The ZFS vs. BTRFS choice, I think, depends more on whether you need specific features like offline deduplication or L2ARC / SLOG cache devices. And which one you're more familiar with (can troubleshoot better).
Btrfs tries to fully use the space and gets all the associated complexities. Additionally, because data/metadata ratio is not fixed one can get into situations where the file system is full and there is no more metadata space. For every action it needs to carefully check if there is enough space to actually perform the action even if the file system is nearly full. Improvements in this area caused this regression.
And no Rust wouldn't help. How often do you get a kernel Oops, dead lock or memory leak? Rust would help with those.
So Rust does not decrease the complexity, but only removes certain kinds of errors which the compiler can detect. Neither logic errors nor speed regressions.
Thank you for your input.
Decide what? Cow is a fundamental part of how btrfs works, and Rust didn't exist for the majority of Linux's life. (Although if you're into that, look at Redox)
Thanks fr mentioning Redox!
Why? There are several reasons, but if you go right back to the beginning, there's a single reason which caused all the other problems: they started coding before they had finished the design.
All of the other problems are fallout from that. Changing the design and the implementation to fix bugs after the initial implementation was done. Introducing more bugs in the process. And leaving unresolved design flaws after freezing the on-disc format.
When you look at ZFS as a comparison, the design was done and validated before they started implementing it. Not unsurprisingly, it worked as designed once the implementation was done. Up-front design work is necessary for engineering complex systems, it really goes without saying.
This isn't even unique to Btrfs, but filesystems are one thing you can't hack around with without coming to grief; you have to get it right first time when their sole purpose is to store and retrieve data reliably. Many open source projects are ridden with problems because their developers were more interested in bashing out code than stopping and thinking beforehand. Same with a lot of closed source projects as well for that matter.
In the case of Btrfs, which was aiming from the start to be a "better ZFS", they didn't even take the time to fully understand some of the design choices and compromises made in ZFS, because they ended up making choices which had terrible implications. Examples: using B-trees rather than Merkle hashes; this is at the root of many of its performance problems. Not having immutable snapshots; again has performance implications as well as safety implications, and is rooted in not having pool transaction numbers and deadlists. Not separating datasets/subvols from the directory hierarchy; presents logistical and administration challenges, while ZFS datasets can freely inherit metadata from parents and the mount locations are a separate property. ZFS isn't perfect of course, there are improvements and new features that could be made, but what is there is well designed, well thought out, and is a joy to work with.
Thank you for your input!
For work I'm involved in relating to safety-critical systems, we use the V-model for concepts, requirements, design and implementation, with extensive validation and verification activities at each level. Tools are used to manage all of the requirements, design details and implementation details and link them all together in a manner which aims to require self-consistency at all levels. When done correctly, this means that the person writing the code does not need to be particularly creative at this stage: the structure is completely detailed by the formal design. But it does require significant up-front effort to carefully consider and nail down the design to this level of detail. But it does avoid the need to continually revise and adapt an incomplete or bad design in a never-ending implementation phase.
This approach is definitely not for everyone, and there are many things one can criticise about it. But if you are willing to bear the financial cost and time costs of doing that detailed design work up front, the cost of implementation will be much lower and the product quality will be much greater. There is a lot to be said for not madly mashing keys and churning out code without thinking about the big picture, and Btrfs is a case study in what not to do.
How to decide whether such meticulous design is necessary or not? In hindsight Btrfs may have benefited, but how to decide when to and when not to in the future?
I would also be interested to know what tools are used for this. The ones I looked at seemed quite dated.. :-)
Thank you for answering! This is very interesting to learn about
In terms of deciding if meticulous up-front design is necessary (again my own take), it depends upon the consequences of failure in the requirements, specifications, design and/or implementation. A random webapp doesn't really have much in the way of consequences other than a bit of annoyance and inconvenience. A safety-critical system can physically harm one or multiple people. Examples: car braking systems, insulin pumps, medical diagnostics, medical instruments, elevator safety controls, avionics etc. It also depends upon how feasible it is to upgrade in the field. A webapp can be updated and reloaded trivially. An embedded application in a hardware device is not trivial to upgrade, especially when it's safety-critical and has to be revalidated for the specific hardware revision.
For filesystems the safety aspect will relate to maintaining the integrity of the data you have entrusted to its care. Computer software and operating systems can have all sorts of silly bugs, but filesystem data integrity is one place where safety is sacrosanct. We set a high bar in our expectation for filesystems, not unreasonably, and after suffering from multiple dataloss incidents with Btrfs, it's clear their work did not meet our expectations. We're not even going into the performance problems here, just the data integrity aspects.
I can't say anything about the tools I use in my company. There are specialist proprietary tools available to help with some of the requirements and specifications management. I will say this: the tools themselves aren't really that important, they are just aids for convenience. The regulatory bodies don't care what tools you use. The important part is the process, of having detailed review at every level before proceeding to the next, and the same again when it comes to validation and verification activities.
Often open source projects limit themselves to some level of unit testing and integration testing, which is fine. But the coverage and quality of that testing may leave some room for improvement. It's clear that Btrfs didn't really test the failure and recovery codepaths properly during its development. Where was the individual unit testing and integration test case coverage for each failure scenario? Where the V-model goes above and beyond this is in the testing of the basic requirements and high-level concepts themselves. You've got to check that the fundamental premises the software implementation is based upon are sound and consistent.
Then a minor issue did happen, some files got missing after a system crash. That's not a problem in itself, that happens.
But then I discovered that the tooling of btrfs is awful to use. Tools to repair, restore, check, scrub, ... I think I got it to finally print some filenames of missing files after a huge amount of hoops. Just see how complex btrfs restore is: https://btrfs.wiki.kernel.org/index.php/Restore. Ideally those tools would know how btrfs works internally including all indirections and be able to translate that into recovering files to the user, not have, for example, a user manually puzzle together objectid numbers from a huge list.
But anyway, that made me simply format everything back to ext4 and go with that again for now.
So, I was a happy btrfs user until I discovered how bad its tooling is. No issues with bugs or anything though.
I really want btrfs to become better given that it's present in the kernel and has checksums, cow, .... I hope it'll keep improving!
The most annoying issue I've run into (and I really hate it) was being low on space and it being really hard to free some since it's very intransparent what's actually using it - but with btdu I mostly got that under control as well.
That said when repairing the boot drive I did experience some confusion due to there being multiple different ways - the documentation could be better on that front. For simple things at least (subvolume management, snapshot creation, checking free space) I've been quite happy with the btrfs command line tool.
Write holes during resilvers are healthy and natural, but the same in a software RAID 5/6 configuration on BTRFS is unacceptably risky.
Brought to you by the same people who don't think SMR disks are used with ZFS in enterprise.
The linked post explains that there is a performance regression in a very rare and specific usecase: creating 100k files in a short span of time. And I'm not talking about unpacking boost, that's not enough files.
Additionally, it will be fixed before any distro packages it. (even Arch is still on 5.9.14.)
On the other hand btrfs got rather impressive speedups for fsync https://www.phoronix.com/scan.php?page=news_item&px=Linux-5....
Why are people making so much noise instead of just reporting the bug upstream when stuff goes wrong with btrfs, and why do we hear so little about the positive impact of new improvements. Almost like the people behave like customers of commercial operating systems. :-)
If that is the case, then your machine maker obviously does not support linux. So, it's no surprise that your machine had problems with linux.
If you buy a machine that supports linux, you would not see those problems. Nothing to do with linux, mostly to do with interaction of linux and your machine manufacturer.
On desktop, Fedora has been by far the most reliable. Ubuntu LTS, as far as I can tell from using it for work, just means that annoying bugs that have been fixed in the upstream never get backported.
Case in point: Bluetooth on Ubuntu LTS has always been much less reliable for me than on Fedora.
Disclaimer: This is only my anecdotal experience, use whatever works best for you.
> you can't get only the bug fixes without the new bugs
> if you know how experiment with getting a new kernel or video driver
I would agree with you if I was on my old machine with an nVidia graphics card which for what it is worth my new machine is a much better experience with AMD processor and integrated graphics but still how do you explain audio in failing to work as soon as I upgraded to Fedora 33 / kernel 5.8 and it (mostly) resolving itself when upgrading to 5.9 kernel? Is this a kernel bug or not? What caused it? What can we do to prevent it from happening in the future?
Then when I want to try the next LTS years later I install it on a different partition and check if it works or not for my use case.
For my work I use Intellij , when there is a big update I get the .tag.gz and try it. If something goes wrong I still have the working version and I can go back.
Updating to latest and greatest was fun when I had the time and when I knew how to format the disk using "fdisk" and it was pleasurable to read and tweak stuff, this days I don't care about shiny stuff and my work does not require latest libraries and I am not using any first gen of hardware.
So IMO if you want stability and no surprises and the ability to maybe upgrade the kernel and a driver Ubuntu is a good solution(not sure about Debian or SUSE) and also I don't have experience with Wayland so maybe that invalidates things and is impossible to get a stable working Wayland setup.
In JavaScript world that's really not that much. A project I've been working on has 120k+ files in node_modules, and `yarn install` creates the whole directory tree in a few seconds. It's not some obscure issue nobody ever runs into.
The main competitor to btrfs is zfs, which is harder to set up, arguably more complex and has licensing issues so you can't really ship things with it built in.
I like zfs on my server, but am also happy with btrfs on my laptop.
I've been bitten by btrfs back when it was the new hot FS, so I'll just stick to ZFS until a distributed (multiple machines) equivalent of ZFS comes along.
Ceph?
However, I am really not sure if the facts they provide are still relevant today. Like always with commercial documents, that must be taken with a grain of salt.
This means BTRFS isn't able to heal itself because there is no copy of the data that BTRFS could use.
I built my own NAS so I can just use ZFS with raidz, and when I ever have a silent data corruption the repair is done in a few seconds, not hours.
At the time, Matt told them off because it's wasn't "ready for porting yet" and wanted to save them useless work.
On the server side, the tendency is now towards vertical striping (directly addressing disks on multiple servers through the network with software).
Recently with super powerful multi-core CPUs being the norm especially in places where RAID is a consideration, hardware RAID makes less and less sense. You just create additional complications like when moving your RAID array between 2 otherwise identical cards but with different firmwares fails. So you end up always buying spares and testing the migration before putting any data, then you freeze the config in place. On the other side you may be stuck forever with a buggy firmware.
The reason I was given is twofold. 1) If your RAID card dies it's nice to be able to plug drives into any controller and be able to access the data. 2) "look at how often MDADM fix bugs. Now look at how often your RAID card firmware gets fixed... use MDADM".
Modern filesystems like btrfs & zfs both prefer to deal with the RAID aspects themselves as it gives them far more control. Even if you're a hardware raid afficianado, I wouldn't recommend formatting (eg) ZFS on top of a hardware raid pool.
It was updated plenty, from 1.58(B) (earliest I can find) to 6.64(B)
https://support.hpe.com/hpsc/swd/public/detail?swItemId=MTX-...
I'm not recommending some abstract hardware RAID, I'm recommending this particular card for personal use, with backup obviously. Although if you just use RAID 0 or 1, then data is perfectly readable outside of RAID with a normal SATA controller.
+ Raid 5 requires BBWC + Raid 6 requires licensing + HPE requires a support contract to download most firmware + When used in servers, they report non-hp drives as being in a constant fault state.
why wouldn't you use it if you have it?
> Raid 6 requires licensing
Keys are out there and can be googled, in any case you probably don't need RAID 6 or any extra features for personal use
> HPE requires a support contract to download most firmware
Um, firmware for this card is on the page I linked, free as in beer
> When used in servers, they report non-hp drives as being in a constant fault state
I wouldn't know, haven't seen anything of the sort personally
Nice to see firmware available - in 2014 HP started requiring support contracts or warranties to download firmware updates for the servers which really turned me off them. I didn't realise RAID firmware wasn't included in that (for some reason iLO firmware isn't either, but I think that's probably because it includes OSS). For this reason alone I'd avoid HP on principal.
I'm hoping one day I can find a reasonably cheap RAID card that still supports using hte BBWC in JBOD mode...
hardware raid has the hw complexity / pickyness you mention, but also is safer in crashes for checksumming raid levels (e.g. 4,5,6) since the stripe checksum happens from controller->disk in hardware rather than os->controller->disk in software, so you're less likely to lose a stripe due to OS issues (all else - like bugs - being equal)
also, IO wise, any mirrored raid is going to require N*mirrors of IO on the system bus, so you're more likely to saturate it
that said, with faster tech (NVMe), individual drives can easily saturate a single card, so its 'worth' paying the multiple IO pentalty multiplexed onto multiple individual lanes since the card is more likely to be the bottleneck than the bus
also, hw raid controllers even in jbod mode might be needed to get enough device fanout, though you're not using the controller for raid in that case
Is that new?
I have always been told to avoid hardware raid when possible.
It's great that you've had good experiences with this particular hardware RAID controller and are recommending it to others, but it's also the case that there are legitimate reasons to avoid all hardware RAID cards in some setups.
Too bad it's not possible to boot from it in HBA mode too.
- don't use it because you think the name is "cool". Every time I go to a linux meetup, some idiot starts talking about the name.
- it takes 5-10 years for a filesystem or database to mature after it's released. Otherwise you will lose data and cry.
- the recent BTRFS performance improvements are against itself, not other filesystems. EXT4 is an excellent fs, as good as XFS performance-wise overall.
Source: DBA and storage engineer.
Great news! BTRFS was introduced in the mainline kernel in March 2009! It's been ten years.
I never do a separate home partition though, so I ended up with the entire disk in Btrfs. I wasn't paying as much attention as I should have when setting up partitions during installation, and didn't totally understand the implications of using Btrfs back then.
Eventually though, it filled up my entire drive with snapshots, causing me much confusion and many disk full errors until I could figure out what had happened.
That experience turned me off from openSUSE for several years, until I finally started using it (with a more conservative choice of Ext4) again for a few machines a couple years ago.
If you're on lvm, you're still getting snapshots now, even with ext4.
That said, I'm going to continue to opt for Ext4 or XFS for the time being, even on distros where Btrfs is the default. Probably in a few years I'll get around to upgrading to Btrfs.
(Nb, avoid SMR hard drives - the rebuild took more than a week!)
It does however require a bit of care to maintain performance, and you need to know your expected workload going in. Otherwise you can find yourself in a situation with a very poorly performing pool where the only realistic route to recovery is a send/receive to a fresh pool and back.
If you do not have a lot of sync workload (VMs, DBs) and don't have super-high performance needs, say a home NAS, then mainly you just need to think about not filling up the pool too much.
A nice way to do this is to create a root dataset where you set a quota to say 75% of capacity, and then create all other datasets below this one. You should not go above 85% space usage, as ZFS switches allocation strategy then to one which can significantly increase fragmentation.
If you do have a lot of sync writes, a SLOG device is basically mandatory. The SLOG device does not have to be large, it only stores about 5-10 seconds worth of writes, so 10-20GB can be plenty. I've partitioned up my SSDs and created a mirror out of two small partitions, using the remaining SSD space for other things.
One thing to keep in mind is that while an L2ARC device sounds like a great thing, depending on your configuration you can actually slow things down with one. A bunch of disks has a lot more bandwidth than a single SATA SSD. An L2ARC device also requires some memory overhead, so reduces your primary ARC. Again depending on load this can be detrimental.
And finally, don't ever think about using deduplication, unless you've read about the consequences, measured the performance benefits and ensured the memory overhead is acceptable. It sounds great on paper but has a lot of associated downsides that can ruin pool performance, and disabling it does not make it go away.
At least that's what I've picked up so far.
I think recreating the entire thing is about my only option at the moment.
My pool, 2 vdevs each a 4-way RAID-Z1, has been used and abused for almost 7 years now. I've gone over the 85% mark but it still works fine, performance too.
Some VM stuff but mostly media and similar.
Do you mean slow IOPS or also sequential performance?
But yeah sadly that's the one area where ZFS is less stellar. Once it's fragmented it's hard to fix. Easiest is to send/receive to another pool, but as you note that is not always feasible.
Assuming your free space fragmentation is not too bad (check output of zpool list, "frag" column is free space fragmentation level), you could just move files back and forth between the two. Assuming you don't have snapshots holding on to the files, this can help reduce the fragmentation.
Otherwise you're stuck with the send/receive.
Though I'd try the mailing list[1] to see if any of the gurus can help identify what's going wrong before attempting random ailments.
ZFS on FUSE should be deprecated by now, and ZoL I believe merged with OpenZFS for the 2.0 release.
FreeBSD ZFS is fine, and just check your versions if using it on Linux. Ie- don't use anything pre-2.0.
I'm currently on Bcachefs, which has erasure coding with similar promises and works better for me.
ZFS is more rigid, I have to have all matched disks for best performance, plus RAM unless I like bad performance (I don't). This could be solved by striping disks into 1TB partitions, but ZFS doesn't like living on a partitioned disk nearly as much.
Mixed disk has a bunch of cool properties for homelab users. Bcachefs also solves another problem; caching and speed. I could designate my 1TB NVMe as fast and my SATA 2TB SSDs as slow. I could also define that /tmp requiers no Erasure Coding and everything else requires 2 replicas. Then I would be able to combine the performance of my NVMe with the capacity of my SATA SSDs while gaining a simple redundance solution. My NAS has a dozen of mixed disks that I could combine more easily than ever.
The issue with ZFS is not that it's not ready to use. The issue is that it's born from enterprise and requires a costly enterprise setup to run, rather than a cheap homelab setup.
Don't know how true is this, since it's not even possible to create a zpool on the whole unpartitioned device on linux. It automatically creates GPT label with zfs and a small efi partition.
Every ZFS pool I have is composed of GPT partitions, and this is also recommended in ZFS books as good practice to make it easier to identify a failed drive. Since you see the GPT partition name in the "zpool status" and other tools' output, it's handy.
You can't unfortunately tell it that different disks have different speeds.
It also works just fine w/ a small (1GB or so) ARC if you want, just at an obvious performance penalty compared to having a larger cache.
For instance, a common threshold for "ready to use" is "included in a released version of the upstream Linux kernel, and not under CONFIG_BROKEN or CONFIG_STAGING". Under that definition, neither ZFS nor bcachefs are "ready to use".
ZFS has been quite feature-full and production ready for more than a decade... why do you think it is not finished?
Yes. When it does what it advertises, properly.
In what way has your experience differed?
ext4 had at least two critical data corruption bugs in the past 5 years in stable kernels (personal anecdote: one ruined the root filesystem on a server that I used, after which I stopped using ext4).
No it hadn't. Those bugs were lower in the stack in block layer.
The one you're referring to is https://lwn.net/Articles/774440/ which is a different bug.
I should add that as a user, what matters to me is the reliability of my data. If a bug exists outside of fs/ext4 in the kernel but affects only ext4 and not other filesystems (such as https://www.phoronix.com/scan.php?page=news_item&px=MTIxNDQ which was caused by an ext4-related commit, and similar others in prior years), this makes ext4 unreliable for me.
> which was caused by an ext4-related commit
Now that's blatantly wrong. It wasn't caused by an ext4-related commit. The commit to blame was scheduler code in the block layer. Nothing filesystem specific.
The point is, ext4 isn't a monument of reliability like you delude yourself with, and it had, and will continue having data corruption bugs.
> Now that's blatantly wrong. It wasn't caused by an ext4-related commit.
Yes it was. This is very clear from the context: https://patchwork.ozlabs.org/project/linux-ext4/patch/500F1C...
You can try and weasel your way out with the semantics, say that it didn't change any files under ext4/ (which does happen with xfs/btrfs/... related patches too BTW, simply because fs/ contains common fs code), but the reality is, it was an ext4 related fix, it appeared in "linux-ext4" mailing list, and Ted Tso, the ext4 maintainer, signed off the patch.
Lastly, idolizing a piece of code is nonsensical. Yes, btrfs had its data corruption bugs, but you can't pretend that ext4 and other filesystems didn't.