DiffDebugging
martinfowler.com
martinfowler.com
I can't speak to 2004, but these days "bisect" is the standard term in my circles (even when we're not actually using git and its bisect command).
It would be interesting to know whether git was the first to associate that word with this activity... especially given it's atrocious naming of everything else :P
It's also possible to do this is a Bayesian fashion.
e. g. recently I was debugging a problem, after two halvings, I saw there's a big refactoring in the remaining interval, so I picked the revision before/after.
https://github.com/ealdwulf/bbchop
But it seems likely that heuristically picking a multiple of times to run the test produces similar results
Let's say you've been working on an "algorithm", which is a combination of cleverly-crafted SQL to fetch the right data, some ad-hoc data processing code in Python, getting predictions from a model, and some additional mathy stuff to translate the model prediction into something useful. You probably built it interactively in a REPL or notebook environment. You might have some ad-hoc assertions scattered across your scripts, but nothing resembling a test suite.
That's just how things go, and if you did it differently you'd probably never get your work done. But now you have a problem: you have what you believe is a mostly-working/mostly-correct implementation, which you built incrementally and interactively, but now needs to be put into production. Without any real tests, your script has immediately become "legacy code": https://understandlegacycode.com/blog/what-is-legacy-code-is.... If things went well, you had time to extract some common utility functions and write unit tests for them, but you probably don't have more than that.
So now what do you do? You want to refactor it into testable units, so you can build out a proper test suite. But you can't refactor code without tests, lest you risk breaking something. You did extensive manual checking and validation of your outputs, but you can't keep doing that over and over.
In this case, the best option that I've found is to do pretty much what Fowler is advocating for here. Extract a set of known inputs that don't take too long to run through the algorithm, and save their known-valid (or known-close-enough-to-valid) outputs. Then, piece by piece, start refactoring, re-running at each stage. If the outputs differ by more than floating-point roundoff error, you've either found a bug in your original implementation, or you introduced a bug in your refactor. Once you're sure that your refactor is OK, you can "lock in" the changes by adding new tests for the refactored sections. Repeat until satisfied, or out of time.
Bisecting is useful when you find a bug which your test suite did not, and which may have been introduced a long while ago. So you search for the first offender between the last known-good commit and the most recent, usually with an automated test that reproduces the bug, and which should be introduced in the suite to prevent it from happening again in the future.
A snapshot test is basically a “let me know if this unit stops giving this output for this specific input”. And sometimes the fix is to acknowledge that this new output is now the correct version by overwriting it to the snapshot.
I added some tests like you described. Some known price inputs and configuration options and the observed return value from a presumed good state. It didn't verify anything was correct, but it did catch “drift.” The expected object would be updated when changing the calculations, if the difference conformed to expectations based off the changes.
I got some pushback because the tests werent testing correctness but it was trivial to implement and it did catch cases where we changed something and something else that should have been unrelated changed as well. A shit test for even shittier code I guess.
If you're unfortunate enough to have an impoverished "linear history", you're stuck figuring out whether it's "bad" or "skip", sorry.
With linear history you squash everything into a single, working commit which is applied to the tip of the shared branch. The rationale is that each change is like a published piece of writing: it will have been through multiple drafts and final versions as well as being edited and peer reviewed, but ultimately the only thing any of your peers care about is the finished work. None of the incomplete, broken, or unreviewed intermediate versions belong in the shared history and they can be thrown away.
Two counter points to this. First, the code review discussion is often recorded forever but in a social tool that’s kept separate from the code itself.
Second, if you end up being as famous for your code as Austen, Thackeray, Shakespeare or Da Vinci were for their literature then your “work in progress” commits and v0 drafts are very valuable and worth keeping. The amount of people this could apply to is not a large number.
An opposite which I could accept is asking people to carefully rebase their drafts, but I see it as a higher effort with a bigger risk of screwing up, with a dose of bikeshedding an additional topic, the commit history: "why do you split in 5 commits when 2 would be enough?"
I guess you can keep the drafts if you don't delete feature branches. But I don't see it as extremely useful, unless it's a very talented individual AND very methodical with his commit history.
- reading the affected code - reproducing the problem step by step with a debugger
Which I guess works for all kinds of problems and not just for regressions, therefore people are more used to it?
- data dependencies cause regressions, so you. need some way to factor this in - production changes cause regressions - debugging at the source code level can be laborious past a certain point
The way I approached doing this at another (maybe same) T$ compnay was tooling that could take anything observable -- e.g. logs, program traces, signals from services, service definitions, data, stack traces from debuggers for side-by-side running -- convert it to a normal form (e.g. protobuf or JSON) and then diff that to look for regressions.
Could be useful but parent mentions a megacorp, so maybe it's just some bureaucracy or ball-breaking.
I find the exact opposite to be true.
When you merge 20 commits from a branch, you add to the history a bunch of unfinished-state checkpoints, which may not even compile individually, and that may make the search longer than it should. Also, those intermediate commits probably never were deployed as individual units.
In a merge-squashed main branch all commits:
* were reviewed as a single unit, but may have originated from multi-commit branches
* are guaranteed to have passed the test suite
* by definition, are at a ready-to-deploy state, even if they contain feature-flagged components
These are much easier to bisect through. And their changeset shouldn’t be any larger than any regularly-merged branches. It’s the same unit of work, merged differently.
If the PR is too large, then you’re probably lumping too much stuff together, which also leads to slower and worse code review cycles.
In my team, we merge squash our PRs, having a guideline that they should be a small increment. Here's what happrns:
- when you checkout a commit, you are guaranteed to have a version that was reviewed by a person and validated by ci/cd - most commits have chsnegs with substance, instead of "log", "debug", etc
Without squash, a person can push 3 commits to a feature branch, ci/cd says the last commit is OK, it gets merged, and you may get 2 commits that don't compile or cause a very basic runtime error due to some typo or whatever.
(not that it could not be solved with due diligence, but that comes with higher effort too)
How do you assure working commits and avoid noise without either squash or rebase?
[1] Yesterday, my program worked. Today, it does not. Why? PDF link: https://www.cs.purdue.edu/homes/xyzhang/spring07/Papers/p253...
It’s magical to experience a short but sweet tip and you just _know_ that (programming) life got a bit easier. Good bang for buck blog post! A few minutes well spent.
(To be fair I think the first time I heard it explicitly from a teacher was Jr high shop class)
Edit: at what age do people usually learn to play 20 questions?
Even if people know / use it intuitively, there's a value to describe it more formally and give it a name.
(BTW sccs was 1972; no doubt it had predecessors. Before disks were large enough to hold multiple versions of a source tree we kept them on tape)
In my circles, it's not ubiquitous even today.
> BTW sccs was 1972
Many projects didn't use a reasonable versioning system until 10 or 20 years ago.
Before decentralized versioning system, this technique was highly impractical - in one SVN project I worked on, checkout of a branch / revision took ~30 minutes.
CVS tracks commit on individual file basis, checking a global state of the project at a particular timestamp wasn't that trivial / necessarily correct.
Stop doing it.
When you're looking for a smoking gun commit, you want it as small as possible, otherwise it'll take longer to fix.
sorta. this really varies strongly from person to person and i wouldn't generalize. but you can construct small commits after you're done with your work.
> contain intermediary trains of thought that aren't relevant
no. delete these, they should not be present in the final history.
> decompose the [...]
i don't know what this is supposed to mean.