Some ways to get better at debugging
jvns.ca
jvns.ca
I only know like 10% of the tool and it’s just indispensable. Use strace!!!!!!!
Except not on MacOS. Lots of tracing is denied by System Integrity Protection; you can use Instruments.app [0] but it doesn't lend itself as well to quick exploration in my experience.
[0] - something like: `xctrace record --template 'System Trace' --launch -- /bin/ls`, then open the trace with Instruments.app.
Mind=blown. Of course! I've been neglecting this thing for so many years!
Question: any recommended reading for using strace effectively?
`run_my_build.whatever` -> hangs, damn.
`strace run_my_build.whatever` -> awesome i can see all the system calls so i know what `.so`s it's pulling in, and what config files it's trying to read, and if it's hanging on a network socket or whatever. This is usually where `strace` just immediately solves my problem.
If you've got like a background daemon that's misbehaving you can do `strace -p <PID>` and it will attach and start printing out the syscalls, this can also be really useful.
`strace` (on all my systems at least) logs to `STDERR` by default, so sometimes you want some combination of like `2>/tmp/log.strace.blah` or to interleave the `STDOOUT` of the process so it's just the usual shell stuff `strace whatever 2>&1 | rg -C ...`.
My use of the tool is very basic, but that's part of what makes it such a great tool, a few simple invocations will just save your bacon. This is especially true on a new team/company or whatever where you run the thing from the Wiki and it doesn't build or start or...
I wonder if there was formal education for this, a lot of people would sleep better.
I studied chemistry, and I'm not sure if it's a causal relationship, but I find I am much better at debugging (in development) and "firefighting" (in production) than my peers.
I guess my conclusion is CS students should be required to take more science classes, or, preferably, software engineer hopefuls should be pushed to study other disciplines.
My longest debugging session were too long because I spent so much time trying to validate a hunch when I should have been trying to invalidate it. This might seem like common sense but it's taken me years to figure this out and it's really made me much more efficient at finding root causes.
Maybe during debugging I consider the strategy of writing a minimal reproduction, but decide it's too time consuming. Then after reviewing the video, I can see that it would've been much faster if I had just written the reproduction to begin with.
Taking breaks does interfere with this, but I still think it's often a good idea.
Another strategy is to add assertions for your hypotheses, which can sometimes be preferable to logging.
I really learned to appreciate an abundance of asserts; the sooner you catch the issue, the less time you have to spend chase it around. Of course, asserts are in essence filtering the allowable state at a point and it's always more powerful when you can express that in the type system, which is why I love languages with rich static type systems, like Rust and Haskell.
This is some combination of: content cached by Cloudflare / Vercel / node server (e.g. using "node cache") / the browser (figuring out SWR, why it's working, not working, working too well, etc. etc.) / cookies/local storage / PWA settings...
Aside from using browser tools like Chrome's Network and Application tabs combined with switching to different browsers and outputting a lot of console logs... what tools should I be using to get to the root of these issues?
A lot of the time the answer is there if you read it :)
I've noticed this while coaching / working with junior developers, is that this is a still that some haven't quite developed yet, and others clearly have.
Unless you're using golang. There's no stack trace by default. It's left as an exercise to to programmer to wrap and add context to errors.
"In order to know how to solve problems get more experience solving problems so you know how."
or maybe
"To know the answers to your questions learn what you need to know to answer them"
One of my fave quotes (Attributed to "Nasrudin," but usually referenced from Will Rogers or Rita Mae Brown).
1. Write things down. Have an assumption? Write it down. Then write down how you will test your assumption and the results.
2. Being able to systematically enumerate possibilities and rule them out one by one, often times indirectly. Yes, one needs to understand the system but one also needs to have a good strategy to use that knowledge. By "indirectly," I mean being able to reason like so: If P, then Q. Not Q. Therefore not P.
- Set a goal
- Experiment
- Visualize
- Undo
[1] Ellnestam, O. Brolund, D. (2014). The Mikado Method. Manning Publications Co. 1. See
2. Plan
3. Do
4. Check- Observe
- Orient
- Decide
- Act
- Orient: check code that maybe involved, check in environment related variable values.
- Decide: decide what you will change in your code to try to fix bug.
- Act : implement fix.
Repeat all steps until you found fix.
Of course you can use different approach to debugging but this one works quite well too.
1) how to reset frame 2) how to set conditional breakpoints 3) do binary search for interesting breakpoint candidates. Going down from a single point isn't very efficient, is better to put a break before and after, then restrict the instruction space by moving them both. That is for non obvious bugs of course
Bonus point: use an ide or something to display objects in a custom way (if your language is OO) like IntelliJ allows.
And another thing to keep in mind is Sherlock's rule: once everything possible is ruled out, it remains the impossible. Your bug is there
Were that could be done without altering production code or restarting production services that don’t have any facility to change the logging levels on the fly.
Was investigating an issue today in some code that only logged that an HTTP request was made to a specific endpoint: no details on why it decided to make that request where it was making the request from, what the request is expected to do or if it succeeded. The unit test approach is the only viable option at that level; isolate as little data as possible to replay until it succeeds. But that costs a lot of time and may not be safe depending on the side-effects. Maybe knowing how to snapshot and restore system state should be on the list too?
Looking at the error and trying to deduce where the problem can work.
But normally, you need to run the code and observe what happens. When what does happen differs from what should happen, you can zero in on the bug empirically.
This process is typical a binary search, which finishes in a few steps, as opposed to the intellectual riddle solving.
WHERE IS THE BIT ABOUT THE DUCK????
Just assume the friend is a rubber duck.