Bifurcate the Problem Space
potetm.com
potetm.com
—
[…] Imagine the frustration of the Stanford University instructor with an especially thick repairman who kept showing up to fix the same computer. He’d stay for a few hours, poke around, announce the machine fine, and leave; but he never knew what he was doing, so the instructor kept calling him back because the machine was never fixed.
Finally the instructor got so mad that he said “Let’s do a binary switch. The problem has got to be one of the circuit boards.” The repairman didn’t know what a binary switch was, but he agreed because he had an angry customer.
A binary switch involves taking half the circuit boards from the machine having trouble and switching them with identical boards in another machine. If the problem’s been transferred to the new machine, you know that something’s wrong with the switched boards. Then you reswitch smaller and smaller number of boards until you pinpoint which board is defective.
They switched half the boards and, sure enough, the problem moved to the other machine. The repairman said, “So what do we do now?” So they swapped half of the changed boards. The trouble was that now both machines worked. (Perhaps shaking the boards around had put a loose contact back in place.)
“That’s great!” said the repairman. “I’m going to use this procedure from now on. I never knew you could fix a machine just by swapping boards!”
—
The best way to troubleshoot these machines was to split the problem in half again and again until you found the part that needed to be cleaned, repaired, or replaced. The way to do that was to run different known chemicals with their own known signatures through the machine again and again. For example, you could run a single chemical (maybe ethanol - I honestly don't remember which ones at this point) and see if it goes straight through the GC and gives you the peak you expect. If it doesn't you could look at the result and see if the timing was off (indicates something with the GC or gas flow) or if the peak was wonky (indicates something with the MS). And then you just keep going. (Sounds a lot like unit and integration tests right?)
Applying that same strategy works wonders for data pipelines as well. Is it something with the extractor or loader (or god forbid Airflow)? Break it down and go from there.
My only nit on this is that while bifurcate is technically accurate (or is it actually?) it feels unnecessarily complex sounding for folks learning this skill.
Bifurcate is commonly used by Indian colleagues, and perhaps programmers are familiar with the term, but it's not something used by US business people in my experience.
If one has pretty good automated tests, it’s possible to automate and pinpoint which commit has the last working version and which commit had the test failure by using git bisect and using the test results as input.
He told them that even if the CCTV log went back to the earth's formation, they would only need to take about 57 stills to pinpoint the second the bike was stolen (apparently they were not persuaded).
(In fact, you don't even know if a bike was ever there without already having identified a frame with the bike in it. After all, the physicist could be lying.)
Given that they knew when the bike was there (presumably the physicist knew when they locked it up or near enough) they could have found that point and looked forward, and it would have taken far fewer than 57 frames to identify where it disappeared if you're only interested in getting down to the second.
So if they knew the bike was present at 1pm, and it was gone by 3pm (hypothetical since not enough information is given) then they can do a binary search on that 2 hour window, that's only 7200 seconds worth of frames. Start at 2pm, is it present? Flip to 2:30, else 1:30. Repeat. Even a week is only 604k seconds, which would require no more than 20 frames to be checked.
(Btw, I realize it's just an anecdote, I'm just being extremely nitpicky.)
There's a summary set of slides here that gives a great overview: https://courses.cs.washington.edu/courses/cse474/18wi/pdfs/l...
My original training was as an electronic technician; specifically, an RF tech.
We would debug through a signal path.
The very first thing they taught us to do, was stick a probe in the halfway point, and see if the signal was what it was supposed to be.
A brute force (non-targetted) binary search may always be an option if you have a problem where you don't even know where to look.
Effectively, we define some function to map N -> B, where N is the natural integers and B is in [0, 1]. (We can then apply some arbitrary function on B for our actual parameter value). We find the initial values as 0, 1, 0.5, 0.25, 0.75; that is: the extrema, all halves (1), all missed quarters (2), missed eighths (4), missed sixteenths (8), and so on. I used some binary representation of N to get B; there are probably other ways to do it.
I found this pretty useful at bifurcating a highly dimensional parameter space – it meant that made sure at all times I could get a decently equal view of all points (with bias towards B=0) in expanding detail. It also meant I didn’t have to think about allocation time, which was useful on a somewhat packed (but use-as-needed) compute cluster.
A greyscale camera can tell you white or black and you could differentiate day from night. But a color camera can tell you red white blue black and so now you can derive sunrise sunny cloudy night. The set of observables is what you partition into signals.
Enhancing your observable space lets you derive stronger signals to act on.
I'll note that, when pair programming, the people who sit and reason about the code often outperform me, so the way I look at it I'm trading time against complexity.
Bifurcation means divide, split, or fork something into two.
'bifurcate the problem space' sounds so Silicon Valley.
* Add lots of logs (e.g. "entered function X with parameters a, b, c", "exited function X, return value W"). * "Make things explode in a useful way". This is, cause an error on purpose, which halts the program. This is useful for debugging server-like environments where many different things might happen at the same time, so the logs become insufficient. It is way slower than having a debugger doing step-by-step scrutiny, but it can be done.
Your debugger may not work with a multithreaded program. Or it might cause the bug to disappear because it changes scheduling/execution order. Printf is more reliable in this context.
Or they may be debugging hardware, which the debugger can't step into. Or debugging object code without symbols.
Or they may be debugging a distributed system with remote code execution and asynchrony.
Or they may be debugging a language that doesn't have an execution stack.
Or they may be debugging an issue that only happens in production or customer machines that they can't access directly.
Or they may be debugging an issue that only happens very rarely, longer than the patience of a programmer to wait for a breakpoint.
Debuggers are one tool we have available, but hardly the only tool, and definitely not a complete tool.