Sometimes it actually is a kernel bug: bind() in Linux 6.0.16
utcc.utoronto.ca
utcc.utoronto.ca
It seems to me that for a long time now stable releases haven't been trying very hard to follow the stated policy [1], in particular the parts that say
> It must fix a real bug that bothers people (not a, "This could be a problem..." type thing).
and
> It must fix a problem that causes a build error ([...]), an oops, a hang, data corruption, a real security issue, or some "oh, that’s not good" issue. In short, something critical.
[1]: https://docs.kernel.org/process/stable-kernel-rules.html
The last three 'stable' releases have contained an annoying number of refactors and fixes for amdgpu in particular
Play with PowerPlay tables in staging, I haven't been able to upgrade
Stable seems an odd place /shrug
Those 'fixes' may have fixed what they looked at, but I'm still broken. Feels more like rearranging chairs than actually fixing
It should fix an actual problem, even if not found until now.
Test the damn thing!
Well, that's all in my interpretation :) for more broad background see https://lwn.net/Articles/863505/
That is, I think the stable-release maintainers should update Documentation/process/stable-kernel-rules.rst so that it tells the truth.
(I think this is about normal stable kernels, not 'longterm' ones. I don't think 6.0 was expected to become the next LTS release.)
I wish this was true. It would have made my job of writing compilers so much easier.
I learned that Gcc has a policy not to fix performance regressions on any but the development branch.
Back in the late '80s we spent fully half of each working day tracking down the trigger and workaround for cfront crashes. Kids today have it easy.
I say, it’s “rarely a compiler bug.” When teaching new developers I remind them that you should assume the bug is in the code you just wrote before jumping to the bug being a compiler, framework or OS issue. It’s just a matter of probabilities.
At the end of the day, you want to optimise for debug time. There's the probability that it's a compiler/developer bug (respectively very low/very high) and the time it takes to rule it out. It's of course best NOT to start by investigating the compiler bug.
1) With with an esoteric environment that has relatively few users.
2) Have specific objectives that require a lot of edge case testing or ridiculously thorough fuzzing of the binary.
3) Go out looking for them by crafting nifty things that aren't much used.
Otherwise it would imply large companies (Google-scale) experience thousands of compiler bugs per year...
This seems likely to be true, from my experience.
Fortunately, most compiler bugs I run into (mostly with gcc and llvm) are not code generation bugs (which can eat months of debugging), but just segfaults / rejecting correct code / other broken stuff.
These features are much more likely to have buggy edge cases in the compiler. And only a small select group of programmers will run into those bugs over and over again.
Eg complex template metaprogramming - which had known broken corner cases in all C++ compilers up until a decade or so ago.
Large companies with compiler teams want to use them, so they do update their compiler, so they will find bugs in it.
Btw, what do you think this prints with clang? (Whether the answer is a compiler bug is debatable…)
printf("%#x\n", 1 << 32);(Spoiler: it prints uninitialized stack memory. A bit closer to demons flying out your nose than you'd expect!)
On Darwin/x86_64, this actually means that it prints out bits from an "uninitialized" register, specifically %rsi in this case, since that's where the second argument should be (even for variadic functions).
Certainly a jarring failure mode! However, if updating your compiler causes you to run into this, you already have a bug in your code (one that clang -- arguably not forcefully enough -- already warns you of).
I find them because of a combination of : - a very large codebase - compiling with optimizations (O3) and targeting recent archs - yearly compiler upgrades - an extremely extensive testsuite
I have found all kinds of bugs (frontend, middle, backend) - the codegen ones tend to be very nasty to diagnose.
Networks can and do go down of course. But in my experience, the vast majority of these issues are actually the result of something extremely mundane like a typo in a hostname. I've resolved an incredible number of issues over the years by simply reading the error message someone sent to me and asking them to check the exact thing that the message says is wrong.
Not that most problems still aren't in my code, but I've increasingly run into framework/library level issues and very rarely compiler issues. Can't say I've triggered any kernel issues yet.
Compiler bug or CPU issue, take your pick :)
Basically there is a system called KUnit for writing white box unit tests and there is code coverage to determine coverage as you would expect, then there are systems for static analysis and verifying assertions.
https://kunit.dev/ - unit tests for kernel
https://docs.kernel.org/dev-tools/testing-overview.html - entire testing overview.
You can usually trust that select() is not broken, but select.js is deprecated, select-kitchensink.js is not compatible with your toolchain and unicyclect.ts leaks memory like a sieve
I don't think it is a bug in the kernel. More likely the userspace or the host system needs to be reconfigured to work well with the bleeding edge kernel. Still I am curious, what exactly may cause this dramatic drop?
Back in the days of ISO-8859-1, you could make gcc crash by having the string literal "ä" in the source code...
Which commit caused it?
https://lore.kernel.org/stable/CAFsF8vL4CGFzWMb38_XviiEgxoKX...
A patch was backported to the 6.0 branch from the main branch, but they forgot a line of code, leading to a buggy behavior.
If you don't have that constraint yeah, not much reason.
There's a tradeoff between the risk of running an older kernel that has known bugs, upgrading to the latest new kernel which has bug fixes but may introduce new bugs, and getting backports for known bugs to your known working kernel. Most of the time the last option is reasonable but it definitely depends on your use case and what you're optimizing for.
Why aren't the processes either manual or automated in place to check for things like this?
Aren't there some tests in place to check for such basic functionality errors?
Doesn't kernel development process mandate much facilities? I'm sure the NSA, Unit 8200 and GCHQ have tests like this in place but don't share their findings.
Is it a matter of funding or leadership philosophy and priorities?
For Linus's releases this is easily solved by slowing down progressively the pace of development towards a release, so that cross-subsystem issues where maintainer A breaks maintainer B's subsystem become progressively less likely over the two months of the release cycle.
For stable releases this is much harder to do because of the short cycle. The stable branches in the end are a mostly automated collection of patches based on both maintainer input and the output of a machine learning model. The quality of stable branches is generally pretty good, or screwups such as this one would not make a headline; but that's more a result of discipline of mainline kernel development, rather than a virtue of the stable kernel release process.
The issue is this bug is not hardware related. Its a pure software issue.
Hardware bugs are an entirely different kettle of fish.
BTW is that bonzini of GNU Smalltalk fame?
The stable kernels pre-release queue is posted periodically to the mailing list and subsystem maintainers _could_ run it through their tests, but honestly I don't believe that many do. Personally I prefer to err on the other side; unless something was explicitly chosen for stable kernel inclusion and applies perfectly, I ask the stable kernel maintainers to not bother include the commit. This approach also has disadvantages of course, so they still run their machine learning thingy and I approve/reject each commit that the bot flags for inclusion.
> BTW is that bonzini of GNU Smalltalk fame?
Yes it's me. :) Did we meet?
I'm a fan of Smalltalk and I used to follow your development of GNU Smalltalk.
What's happened to it? It seems to have fallen by the wayside.
You can get a kernel from Red Hat that has been through Red Hat's release process. Red Hat has their own test suite/labs and will also pay attention to test results from elsewhere - including Fedora, their evergreen distro for putting new software into the wild ahead of its incorporation into Red Hat Enterprise Linux.
Substitute the distro of your choice.
> As 6.0.y is now end-of-life, is there anything keeping you on that kernel tree?
Uh, several distributions. It wasn't EOL enough to prevent breaking it, so fix it.
Don't even technically need their input, Git and all.
I'll buy this EOL thing if they revert the change that caused this and stop releasing under 6.0. There were at least two more after this
That is the kernel bug tracker, not the distributions bug tracker.
>It wasn't EOL enough to prevent breaking it, so fix it.
It wasn't EOL at the time the patch was backported. It's EOL now.
>I'll buy this EOL thing if they revert the change that caused this and stop releasing under 6.0. There were at least two more after this
Not sure what "this" in "two more after this" is, but there have been no 6.0 releases since it was EOLed.
The point is that the contributors reasoning for being on the tree is irrelevant. Like you said, they just made it EOL
Distributions are the continuous/constant answer as to why countless people will be. This isn't an ancient release, something from the grave.
Is the expectation, then, that distributions would have to patch out the regression - or take on a more major upgrade (6.1 / 6.2), likely breaking something else?
Neither of these are particularly tenable. I'm glad GKH was willing to accept further changes to make it correct, but reverting is also applicable.
Breaking something, calling it EOL, and not fixing it is closer to dead than end of life.
Correct.
>Neither of these are particularly tenable.
Yes they are.
>Breaking something, calling it EOL, and not fixing it is closer to dead than end of life.
You're awfully confident about how things should work, even though you don't understand how they already work.
This is really just a pedantic criticism on the handling of 6.0, and the 'ignorance' (I hate the connotation of the term) of why people don't run latest.
I'm not asking them to bend over backwards, here.
Things are going more or less the way I want, 6.0 will get fixed [edit: upstream]. Please don't take this the wrong way.