XFS Metadata Corruption on Linux 6.3 Tracked Down to One Missing One-Line Patch
phoronix.com
phoronix.com
I have enough on my plate just dealing with the issues arising from using stable code, I think it’s admirable that people find the time raising their glance to future releases and helping us all enjoying a less panic-inducing experience.
And even if you perfer stable, the latest will become stable eventually. Not trying your workload out on the next releases has pretty much the same risk profile of just running latest.
Many problems can only be found by running your particular workload.
Running on "latest commit from master" from many projects (not Linux) will just get you code nobody even tested and so a lot of bugs fixed quickly.
Running on "latest stable" (whatever that means for project) means fixes from time to time when it updates, but in vast majority of cases not that much work.
Anything behind that like LTS releases ? Extra work.
Now any doc you find might be about never release or feature that changed. "Bugs" might not get fixed if they are not big enough to backport.
Upgrade to new LTS version will also get you years of changes in app that you then have to apply to the system, vs having to do it "change by change" when keeping up to date.
If you use configuration management that also often means multiple different configs to manage at the very least till previous LTS version gets finally upgraded
In fact, many of these bugs were on stable releases too.
We used to run -stable, and update every few years, like from FreeBSD 9.x to FreeBSD 10.x. We found that when we did that, we would often encounter some small subtle bug that was tickled in our environment, and which was incredibly hard to track down. That sort of bug was hard to track down because the diff between branches was enormous, and because there were thousands of commits to sift through, and because the person responsible for the bug may have committed it months or years ago, and has forgotten about it.
We eventually decided to track the main branch, updating frequently. This means that while we find more bugs, but they are far easier to fix because they were introduced more recently, and there are a lot fewer commits to look through to find where they came from.
For a lot of companies, inspecting source code and filing bugs directly is just not a capacity that exists, which is where LTS of a Linux distribution makes a bit more sense — and without throwing any shade on FreeBSD (I love it), maybe the smaller amount of users globally compared to Linux means that “stable” isn’t quite as stable, especially if you’re doing bleeding edge stuff anyway.
I guess you could say the same thing is true for a company like Cloudflare considering their network related patches to the Linux kernel.
Thanks for the perspective!
If you're just a consumer, then it makes a lot of sense to consume the LTS branch. Whenever I've run ubuntu (which I do not contribute to), I run LTS for that reason.
In our case, we have a team of kernel engineers who make frequent contributions and are very familiar with the FreeBSD source code. So we're in a good position to inspect the frequent merges from upstream.
The other benefit to tracking the main branch closely is that it makes it far easier to contribute changes. When tracking the main branch, its easy to test a change in our tree, and then pick it up almost unmodified as a patch to the FreeBSD main branch. That makes it much easier to get the code into FreeBSD. In fact, for most smaller changes, we try to push them upstream first and bring them back with our frequent upstream merges This is much harder when running a several years old branch, as then the patch needs to be forward ported to the FreeBSD main branch. As such, its very hard to integrate and test changes suggested by reviewers, as the patch needs to be ported back and forth. And things get worse for large changes (like new TCP stacks, kTLS, etc), which are harder to port back and forth.
Also, do you folks tend to review the freebsd code before upgrades or only after the fact (like, if there's a show-stopping bug or two)? Thanks.
But in reality what's happening here is folks are getting access to bleeding-edge kernel development snapshots who choose to run these kernel versions, and are lucky to get such quick access to patches even before the scope of new bugs are entirely understood by the developers. Note there's nothing preventing these affected users from simply running a prior known-stable kernel version until the bug is better understood, they're opting in on the chaos.
It's unfair to assume Dave Chinner et al won't be running the issue seemingly fixed by this one-line change fully to ground.
If you're not interested in playing the role of kernel QA and interacting with the upstream devs when things break in not yet understood ways, don't run bleeding edge kernel versions. LTS and -stable releases are offered for a reason.
Relative to a kernel version you'd encounter in something like rhel or debian stable however, tracking mainline's "stable" branch is still pretty damn aggressive.
If you want to run hardware released in the last year or so, LTS kernels have a funny tendency not to ever work. I've yet to buy a laptop that didn't at least need latest mainline.
It shouldn't have to count as 'aggressive' to want a kernel that will work with my hardware, but I guess that's what we get when there's no HAL.
If there’s a metadata corruption bug in a file system, I think I would prefer a rapid kernel crash over a magic change that may or may not fix the issue.
You claim it can be done. Have you ever actually done it? I bet not.
ref: various 'this rocket exploded because of one line of code' headlines
This doesn't even consider the fact that issues related to things like concurrency are usually difficult to properly unit test at all unless you already know what the bug is in advance. If you have a highly concurrent system and it requires a bunch of different things are in some specific state in order to trigger a specific bug, of course you CAN write a test for this in principle, but it's a huge amount of work and requires that you've already done all the debugging already. Which is why developers in C/C++ rely on a bunch of other techniques like sanitizer builds to test issues like this.
The fact that it would be hard to test certain edge cases does not in any way excuse the fact that the overwhelming bulk of functions in Linux are pure functions that are thread-hostile anyway, and these all need tests. The hard cases can be left for last.
But note that not only are there no silver bullets, as sibling comments note kernels (or anything that touches hardware, and "the real world" (ex. getting back packets with random reorders and duplicates and drops) really) have trouble using unit testing. And even in those cases where it might work it's not universally applied, I think.
Edit: it looks like 6.3.3 is the first broken version. I would _guess_ 6.3.4 is also affected. Presumably 6.3.5 will not be.
said you know who
First, they do have unit tests (KUnit). However, I suspect the "real" tests that result in a mostly-working kernel are massive integration tests run independently by companies contributing to Linux. And, of course, actual users running rc and release kernels who report problems (which I suppose is not unlike a stochastic distributed integration testing system).