Monte Carlo profiling
stackoverflow.com
stackoverflow.com
Back at Electric Cloud I was working on debugging an apparently live locked build process using our tool (this was back in 2004) where make would run forever with the CPU pegged.
It turned out that the particular build tree structure (at the file level) was somewhat pathological for our tool and we were traversing the tree a ridiculous (exponential) number of times relative to its size. We were doing this to constantly calculate a value that depended on sub-tree size.
After ages trying to narrow down the bug I just went into the debugger and broke in to look at a particular data structure. After doing this a few times I noticed that I was always in the same function.
Memoizing the function fixed the problem.
"Profiling also involves watching your program as it runs, and keeping a histogram of where the program counter happens to be every now and then."
"The run-time figures that gprof gives you are based on a sampling process, so they are subject to statistical inaccuracy."
-- http://www.cs.utah.edu/dept/old/texinfo/as/gprof.html#SEC11
It's a pretty good technique, actually.
Combine with watch
$ watch pstack PID
On RHEL6 pstack is in gdb RPM package.
(bonus: shows all threads)
As a variant, if you want a language-level backtrace in an interpreted language, you can:
- set up a signal handler which dumps a (perl, python, etc) level backtrace
- hit the program with the signal a few times
My Python server processes during testing handle 2 signals SIGUSR1 & SIGUSR2.
SIGUSR1 : runs guppy memory usage & leak profiling (http://guppy-pe.sourceforge.net/)
import guppy
_hpy = guppy.hpy()
log(...,_hpy.heapu(),...)
log(...,_hpy.heap(),...)
log(..., <other _hpy info>...)
SIGUSR2 : open an interactive console using rconsole (http://code.google.com/p/rfoo/) and let me log and muck around with live object while the server process keeps running.[NOTE: don't use on production systems, it opens an RPC port that lets anyone control your process remotely!] from rfoo.utils import rconsole
rconsole.spawn_server()
Then on the console: $ rconsoleCompared to that, firing up the debugger and hitting break is pretty easy. If the hotspot is like core meltdown hot, it's probably faster.
Now, I'm sure there are special circumstances where this method is the right thing. But that's totally different from it apparently being suggested as best practice to newbies.
So I wrote an Alt-Sysrq handler that dumped the stack of the running process. The failed one was usually the one running so it was pretty easy to find, and then the stack trace told us enough about why it was deadlocked to make a fix without needing more investigation.
It was essentially just a one-sample sampling (monte-carlo) profiler.