Keeping a project bisectable
andrealmeid.com
andrealmeid.com
Mozilla has a tool (command line or user-friendly-ish GUI front end) called “mozregression” that Firefox users can use to bisect regressions without compiling Firefox. It downloads and runs Firefox Nightly builds from an archive of years of builds, bisecting good and bad builds until the regression commit is found. An indispensable debugging tool for a very complex application!
I will reject any such contributions and ban anyone who persists with them. I’m not sure what lies below the fold, but if this is a baseline assumption then you’re already fucked.
But... this approach clearly exists. A while ago I was talking with a coworker about some open source upstreaming they wanted to do. They were struggling to get their code out there because they were fighting against all sorts of broken test suites. They were told that this was simply the norm in the community they were contributing to and the test suites are always broken so don't worry about it. Baffled, they made sure that their new tests passed and sent the code off for review and integration.
So there is perhaps a need for software engineering principles that apply to this situation. Maybe that's wrong. Maybe the solution is like adopting source control - it is a fool's errand to try to handle the case where it isn't present so just tell everybody to adopt it or they aren't serious. But maybe there is a population that needs tools that help them even in a world of fucked integration tests.
A feature I’m still missing is to bisect using a new test. That is, I have discovered a bug in the program that there is no test for, so I author a failing test. Now I want to find out when this failure occurred but I can’t bisect because even the previous commit is green as the test didn’t exist. It’s possible to work around by placing the test outside the repo root and bisecting with a script that “authors” the test (e.g copies the file) at each step, but it’s cumbersome and feels like something that could be possible to add to the bisect command, e.g an option to merge a commit (the new test) before testing each step of the bisect.
You could just copy that entrypoint into another untracked file, have its list include (or be only - quicker) the new untracked test file, and execute that instead of the normal one while bisecting.
Before moving on, a quick `git reset —-hard` will get back to a clean flip
1. Create new test (e.g. "spec/reproduce-bug_spec.rb")
2. Ignore it locally: `echo "spec/reproduce_bug_spec.rb" >> .git/info/exclude`
3. Run bisect (something like: `git bisect run rspec spec/reproduce-bug_spec.rb`)
If running the test gets more complex (e.g. installing dependencies as they might change travelling through history), I usually create a wrapper script (and ignore it) to bisect-run-it.
$ git bisect [good|bad]
$ git cherry-pick new-test
$ ./run-the-test
$ git reset --hard HEAD^
$ git bisect [good|bad]
I sometimes use the stash for that: cherry pick the test as a working tree uncommitted change. "git stash" it before every bisect step, then "git pop".git bisect run --cherry-pick abc123 [....]
to do a regular bisect but at each step cherry picking the commit abc123
- find bug
- write new test to expose it
- bisect commits, running new test each time
- find culprit commit
- stop bisecting, fix bug
- when the test passes, you know bug is fixed.
If a test started failing, I could annotate it as an xfail with an open-ended version pattern. If I noticed that an xfail started passing, making it a upass (I think these were blue in our tests results), I put an upper bound on the xfail version pattern.
The nice thing about this was that I could always run the latest and greatest suite against any version, although it made it a little more complicated to run them as they existed when they were first run.
In general, I want to store one or more services and a first/last valid commit for each one as metadata for each test so that my framework knows which ones to run based on the environment under test. Haven't added this feature yet though.
- write new tests to expose it, but not commit the tests
- bisect commits, running new tests each time
- find culprit commit
- stop bisecting, fix bug
- Look for similar bugs nearby, and ways to improve the new tests, now I know the root problem.
- Figure out the best series of commits for the tests and fixes together, that will be easiest to review by my colleagues/future me/alien archaeologists.
- Send PRs/commit to main.
Disclaimer: Google used to pay me to write the coolest features in Bazel like its system for downloading files. https://github.com/bazelbuild/bazel/commit/ed7ced0018dc5c5eb... But I stopped supporting Bazel because they kept breaking git bisect history permantently with backwards incompatible changes.
Well I hope people do this
Not necessarily have a downloaded carbon copy of every dependency, but at least enough information to recover it. package-lock.json file for example.
NB. I added a script to backup node_modules of every CI build. Just for paranoia.
[1] https://www.joelonsoftware.com/2000/08/09/the-joel-test-12-s...
Greatly cuts down on the "tests pass on my system, not sure while it failed in CI"
The result is a linear history, that is much easier to bisect.
It also encourages developers to make small commits, since they are easier to rebase.
The main page of a repository configuration in GitHub let's you enable/disable:
- Allow merge commits
- Allow squash merging
- Allow rebase merging
for PRs.
And under branch protection there is a "Require linear history" option.
https://github.com/isaacs/github/issues/1017
I really, really like semi-linear branching/merging. I.e. always rebase-merging, but with a merge commit.
Reasons, in comparison to Github's "rebase merge" which doesn't produce a merge commit:
1. It makes it clear which commits were part of one PR
2. It makes it clear who did the merge
3. It's okay to not have every commit build. but the one being merged will.
4. Still pretty bisectable. You'll narrow things down at least to the PR that caused an issue, and from there it's usually quite simple.
5. Looks very tidy in gitk & Co
I find commit history is important if you have a long running client project with evolving context, so you can figure out why something was changed, not just what.
As well as linear log, cherry-pick and scores of other tools which do not have intelligent merges handling.
I find that quite surprising, given how I see `git bisect` referenced by the Linux community as useful for tracking down the commit that introduced a problem, and Linus uses merge commits (including `octopus merges`) very heavily.
That’s what you need to avoid in your mainline. Mandating what developers do in development branches is like having a dress code at work and then complaining about your colleague’s choice of underwear.
Git log also has the first parent option and that is really helpful for helping you cut through the noise and see the high level overview of the history without seeing every commit.
But IMHO if swash merged lead to a situation that relevant history is lost the PR is too big.
However, squashing everything down isn't the only or best answer. One can also create a series of commits, each of which is individually passing all the tests, but isn't one big bang. Of course sometimes you have no choice, some refactorings just fundamentally involve huge atomic changes to the code base, but it's still better to try to break these things down into multiple steps if at all possible so you can come back later with a new test and bisect.
I don't use bisect often, but when I do, it can save me days of effort trying to find problems, so it's worth it to maintain the discipline of keeping the tests passing with all commits. Plus it's already a net benefit; maybe you are a way better programmer than me but I find that my changes that couldn't possibly break anything else in the code base break other things in the code base in modules I wasn't expecting with some frequency. Keeping all the tests passing as much as possible is a critical component in making sure that I make what I think of as "monotonic forward progress". There are things more demoralizing that breaking half the code base every time I make the slightest change and it not being caught until QA or deployment... but making a habit of this is up there for sure. Definitely a great way to add some burnout to your life. I've seen multi-decade projects basically strangle themselves to death this way; nobody was willing to touch any of the core code because it had no testing coverage and every time they tried it broke everything, so they just lived with some real garbage code at the core of it.
If you squash it into one commit on the other side, git will have a very hard time figuring out that the file got renamed and is not a brand new file, as both the rename and the fixup happened at the same time.
The main impediments are external dependencies: the things contributing to the behavior change you're looking for which are not in the git repo you are bisecting:
- components from other git repos it depends on which were changing in parallel at around the time the problem showed up
- platform/toolchain and other dependencies; maybe the problem showed up because of a compiler upgrade.
Plus other things:
- difficult reproducibility of the issue being tested.
- multiple unrelated problems masquerading under the manifestation that you're actually testing for.
Regarding buildable commits; if you want a buildable commit to stay buildable, it has to reference the exact toolchain that was used, which has to be available. A commit that built with GCC 4 may not build with your GCC 12 today.
I've done bisects going back 10, 12 years.
Bisect is a useful tool. Making commits that don’t compile, don’t pass tests, or don’t run is useful. Squashing commits is useful. Not squashing history is also useful.
There’s no technical reason we can’t have all these things.
I'm not sure I agree with you there on that one.
To prevent this problem, I wrote this push script: https://github.com/nblockchain/fsx/blob/master/Tools/gitPush...
For example, you may want to get your tree in one of the intermittent states and play around with it. Run some tests, try how some object behaves, to better understand the different commits comprising the PR. I do this quite a lot.
I aspire for this. I don't always succeed; in C++ code, often older commits will rely on transitive stdlib header #includes without including the proper header for what is used. This works, up until a compiler stdlib update will remove these implicit includes, breaking all past versions of your code (and requiring you to manually patch in the include when bisecting in these versions), and requiring you to add the necessary includes moving forward.
If you are going to bisect you either
1. Write unit tests
2. Run the app.
For 1, the problem is you are bisecting, so your test for HEAD may not compile with earlier versions
For 2. unless it is a nice command line app you can script, you have to wait for it to start up and manually do the steps to reproduce the problem.
Ideally of course if you are always aiming to be bisectable you can mitigate these with fast as possible compile and run times. Probably you would try not to build a monolith in that case.
it's a fantastic tool, but if you are chasing a kernel bug in an OS, its intensely time consuming to do, without the local clone of the repo to walk back the atomic(ish) changes and my lived reality is the kernel re-make is a lot slower than you want because imputed dependency chains typically re-compile a LOT more than you think.
I would not be surprised if duck debugging and printf() yields the outcome faster than what is (forgive loaded language) blame-finger pointing, sometimes.
It unquestionably works. It is not low-cost. The goal of preserving it is worthy. This is sensible.
The solution that I found was fixing the bug at the commit that created it, squashing, and then rebasing on top.
But it's far from ideal, and completely non-viable if there's heavy use of the repo from other developers.
commit 011bf12e41fea384c06873505b25cd44f995926d (HEAD -> main)
Author: Tom Jakubowski <tom@crystae.net>
Date: Sat Aug 6 09:27:05 2022 -0700
initial commit
diff --git a/README b/README
new file mode 100644
index 0000000..3b18e51
--- /dev/null
+++ b/README
@@ -0,0 +1 @@
+hello world† My company's not on Github Enterprise and I don't have a repo and a second user handy to double check on public Github.
You need better tools. Even in GitHub you can click on "force pushed" and see the diff.
> there’s no guarantee they would all pass CI
There wasn't any guarantee before, either. If you want that guarantee, make the CI build every commit.
Guarantee is the wrong word. The point is that if you always squash merge after passing the entire test suite you don’t have a bunch of potential garbage commits in history that you have to wade through when bisecting.
As for building every commit that’s probably a tough sell and a poor use of money for what benefit?
Second, you might have more than one commit you want to merge to master, not just a single commit.
Also rebasing is trivial and very quick once you have done some basic reading about it.