Follow up to “I booted Linux 293k times”
rwmj.wordpress.com
rwmj.wordpress.com
That's funny because when I was reading through this I kept thinking, "Wow, this is a heroic effort."
It was a heroic individual effort that led to a heroic team effort. The team wouldn't have been assembled had the individual not raised the alarm and focused the attention. Then the team was able to handle it through their collective efforts. Doubly inspiring.
It's the Bystander Effect in code.
> So the commit I had bisected to was the wrong one again – it exposed a latent bug simply because it ran the same code as a sleep.
The fix was by someone else, and is here: https://lore.kernel.org/stable/20230615091830.RxMV2xf_@linut...
But the root cause of the bug was in another earlier commit. In a complex system such as the linux kernel, there could be multiple contributing factors. Both the "introducing a sleep" and the "root bug" are contributing factors, but we as engineers tend to think of the "root bug" as the actual problem, because we like more elegant explanations. Also, the "root bug" is likely something that goes against the ethos of the software project, and "common sense best practices", to a much larger degree than "introducing a sleep" is.
808 points, 3 days ago, 266 comments.
There are devs and then there are kernel devs.
The relentless pursuit of bugs is truly impressive.
> Paolo Bonzini suggested to me that I boot the kernel in parallel, in a loop, for 24 hours at the point immediately before the commit
Otherwise I don't disagree that perseverance is important to get to the bottom of this stuff.
https://scholar.harvard.edu/files/mickens/files/thenightwatc...
Awesome read! Thanks for sharing.
The amount of attention this received probably led to a way faster fix.
Basically, you know commit 0 is fine, and change 52,363 is bad, so you run 1000 tests on change 26,180 to cut in half the number of commits that could have introduced the problem. If the bug is still there, then you run 1000 tests on change 13,090; if the bug is not in change 26,180, then you run 1000 tests on change 39,270. Eventually (after 15 or 16 iterations) you'll narrow down the problem to a single change which introduced the problem. Unfortunately code is complicated, so the first time they ran that procedure, their tests didn't find the right bug, so they did it again but slightly differently.
(many details omitted for simplicity)
Unfortunately, when the problem is in kernel boot software, that is not a practical solution for obvious reasons, and you're left using more basic techniques like running a binary search on the commit history.
The long bisection process identified the second commit but it took further work to identify the true cause of the bug.
Would it have been possible to generate a list of commits on timeline and perform bisects in a method similar to binary search? Or that was already done and still 293k boots were needed? I am genuinely curious.
That's what bisect means. You do a binary search for the breaking commit
The human brain struggles when it has to execute almost, but not quite, the same exercises over and over again. It becomes a blur and after a while you can’t recall if you ran all of the steps this cycle (if you can’t automate the test well) and all it takes is typing “bad” when you meant “good” to end in failure. Which doesn’t sound like a big likelihood but when you’re testing a negative it’s much easier to get your wires crossed.
https://news.ycombinator.com/item?id=36325986