How so? Can't you just attach to each process? Do you mean at production scale?
How so? Can't you just attach to each process? Do you mean at production scale?
Debugging a show-stopper synchronization flaw in the bootstrapping of a parallel job, I had to tell the job system to launch each of 64 nodes wrapped in gdb wrapped in xterm with remote X display back to a laptop. It was something that had "worked in test" reliably, but that was always on a smaller number of nodes or simulating a larger number with time-sharing. It seemed to need real parallel hardware allocations to show up.
So, I was able to launch the job, allow all nodes to run until distributed deadlock, then interrupt and show all thread stacks on one debugger after another until I found the odd process which was out of phase with the rest. As soon as I saw the stacks, I had no further need for the debuggers. Just knowing the "impossible" state configuration happened was the necessary clue.
I have seen lots of satisfactory use of logs for diagnostics, but this is a case where I think the debuggers were almost essential. The debuggers gave a distributed state snapshot due to the deadlock-induced quiescence of the whole system. To have logs show the same snapshot would require a perfect arrangement of unbuffered logging so that the final state would be visible and not caught up in RAM of some or all of the stalled processes.
I've used that a lot black-box debugging java processes in production.
I also wrote such a signal handler for a ruby app.
There are some of cute hacks to get a remote debugger which I haven't used in years.
https://github.com/sassoftware/epdb can start and connect to a remote debug instance
But worse, we were debugging our library linked into someone else's application. So, we would have needed their cooperation to add signal handlers. And, we naively thought our library had already been sufficiently tested and did not anticipate the need for last-minute debugging when our user moved their application to a different computing resource that was not available during prior months of preparation and prototypes.
So, the ugly app-in-gdb-in-xterm with X over Internet addressed all that with adhoc instrumentation and communications that could be folded into the existing parallel, distributed job on short notice.
This was also in the days of pthreads in C. The other valuable use of gdb I recall was for memory watchpoints to help track down unexpected changes to certain data. The other tool we got a lot of use out of in those days was Purify to help audit for use of uninitialized memory and for memory leaks.
I was working on a cluster and often needed to do the same thing on all the nodes, run some database query or os command, and gather and integrate the results. I wrote yet another Python stream-objects-instead-of-strings shell (https://geophile.com/osh), which included features for distribution: Run the same command on all nodes, bring back the results, merge streams, gather files, distribute files, etc.
That was several years ago, and it wasn't a full shell. Since then, I've done an improved system that actually is a shell (https://marceltheshell.org, https://geophile.com/marcel).
You could, I suppose, but I'm not sure why you would want to. Going back and forth to two debuggers (at least) seems like torture. And with multiple processes, timing and synchronization issues could make interactive debugging a real nightmare. Why? Why would you do this?
Interactive debugging doesn't scale in any dimension.
All the usual problems, but doubled. Oops, I stepped too far, start over. Oops, the other one stepped too far, start over. Oops, start over, oops, start over.
Never again.
I start with basic lifecycle logging, at an INFO level: The FooBar is created, the FooBar is destroyed, and major states in between. If a FooBar manages a set of things, then I might also add DEBUG level logging for operations on those things. But mostly, detailed logging comes later, as I investigate problems. You really can't anticipate too much logging, because you don't know what's going to need it.