Strace – My Favourite Secret Weapon (2011)
zwischenzugs.com
zwischenzugs.com
mkdir /tmp/empty
LD_LIBRARY_PATH+=:/tmp/empty # or PATH, or PERL5LIB, etc.
strace -f program |& grep /tmp/empty
I had a need to track what files are used by what processes spawned by a program (to infer dependencies between the processes), and it seemed (and probably was) simpler to use ptrace directly rather than to parse strace output, so I wrote https://github.com/orivej/fptrace/ . Soon it turned out useful to dump its data into shell scripts that reflect the tree of spawned processes and let you rerun an arbitrary subtree. I mostly use it to debug build systems, for example to trace configure and examine what env var affected a certain check in a strange way, or to trace make and rerun the C compiler with -dD -E to inspect an unexpected interference between includes.LD_DEBUG=libs command
You can check out all the valid options with:
LD_DEBUG=help cat
or check out `man ld.so`.
On mac, there is something similar. Lookup all the DYLD_PRINT_* variables in `man dyld`.
strace -e trace=file -f program
is much better in this case."perf trace" has improved since I wrote that, and may well be mostly strace-equivalent (but without crippling overhead) in the latest Linux.
I often run ftrace, perf, and eBPF on our production instances for syscall tracing. If I ran strace, the instance would suddenly be very slow, and it would trigger Hysterix (and other) timeouts and be removed from the ASG and auto terminated. Our environment is fault-tolerant, so yes, we can run strace -- you just don't get much output, and the load vanishes from the instance you are looking at.
Do newer versions of `perf trace` expose these?
Are you referring to that it makes syscalls slow, or is it about something else?
Yours is a very superior treatment of the strace topic.
This sounds like a really interesting job - consultancy? How did you get started in it
Strace was always my go-to on Linux for solving error messages that fail to be a complete sentence: "connection refused" (to what?) "file not found" (where did you look?) and so on. Since I've moved to Windows the best alternative seems to be Process Explorer.
I'm not OP but I do the same thing, solve the problems and bugs that the dev team can't. In my world, this is simply "Operations" or, in later years, "Systems Engineer." I'm really good at doing that but I am not at all skilled at large-scale software development even though I can read and understand already-written code as part of my troubleshooting.
Descriptions like the author's are why I'm dismayed that Ops has such a bad reputation in the modern computing industry and why I'm not so keen on the combination of the two roles into "DevOps." Troubleshooting a working, in-motion system requires a different set of skills, in my experience, from writing and debugging code. It's also a set of skills that doesn't seem to overlap very often. The two roles go hand-in-hand, of course, but asking someone to do both in a large environment hasn't ended will (again, in my experience).
In most large non-tech companies operations is a cost center budgeted alongside facilities and basic costs of business and also usually run by managers less frequently with a software background as much as an IT or traditional business background focused around cost efficiencies. I’m not sure whether this designation or Taylorism is fundamentally the root causes for why “devops” from the cultural sense is not really happening.
I'm not convinced of this - or rather, if someone can't troubleshoot a system in motion, they're going to struggle to actually develop anything substantial. Debugging is a very underestimated skill of programming.
I think the conflation of "devops" is nothing to do with skills and everything to do with culture. When you have them as separate business units or teams it becomes adversarial: every deployment is a potential headache for Ops, so Ops try to prevent deployments or set up staging barriers to minimise this. Whereas developers may not appreciate the difficulties involved in deploying a large system vs the toy one they're developing on.
Devops can also mean that you're developing automation to reduce manual maintenance or even remove some ops responsibilities entirely.
Think of self healing clusters, auto provisioning and so on.
I've also come across the the definition of devops as a traditional ops dude that just worked for the dev team, keeping their development environment maintained. I'm still confused why they called that a devops position...
Admittedly ProcExp will make you seem amazing as it is but harnessing ETW power will make you unstoppable
You can read more about my background here IYI:
https://zwischenzugs.com/2017/10/15/my-20-year-experience-of...
> /usr/bin/perl -mblahlib
Can't locate blahlib.pm in @INC (you may need to install
the blahlib module) (@INC contains:
/usr/lib/perl5/site_perl/5.26.1/x86_64-linux-thread-multi
/usr/lib/perl5/site_perl/5.26.1
/usr/lib/perl5/vendor_perl/5.26.1/x86_64-linux-thread-multi
/usr/lib/perl5/vendor_perl/5.26.1
/usr/lib/perl5/5.26.1/x86_64-linux-thread-multi
/usr/lib/perl5/5.26.1 /usr/lib/perl5/site_perl).
BEGIN failed--compilation aborted."Zug" is also a form of pull(ziehen). So it could also be interpreted as 'while pulling'. But that's pretty far fetched tbh
I'm not the parent, however, and can't answer your question.
The only thing I'd still like to have is something like strace but for general function calls. Often, the issue is within the application and does not show in syscalls. I guess this should be possible with gdb, but I haven't looked into it yet, also because any meaningful names are often stripped from the binaries.
$ whatis ltrace
ltrace (1) - A library call tracer
That seems nice indeed. It seems to make things a lot slower, but I guess that's to be expected from tracing at this level! strace -o 1.log strace -o 2.log ls
and observe in 1.log how the second strace uses ptrace call.The limitation is that a process can not be ptraced by multiple processes at the same time. If you add -f to the first strace, it will start tracing the fork meant for ls before the second strace has a chance to setup its tracing. That setup will fail, and the second strace will kill the fork instead of running ls. You can read this from 1.log!
If I'm not mistaken, a ptrace tool may untrace its grandchild right when it intercepts a child attempt to trace that grandchild to make the attempt succeed, but I don't know if strace can.
ltrace is a program that simply runs the specified command until it exits.
It intercepts and records the dynamic library calls which are called by the executed process and the signals which are received by that process.
It can also intercept and print the system calls executed by the program.
Also noteworthy is that you can simply press s key in htop to immediately attach a process and inspect it, which was handy many times for me.http://www.brendangregg.com/blog/2016-03-05/linux-bpf-superp...
csrutil enable --without dtrace
Otherwise you'll get a permissions error due to recent MacOS code signing protection... unless you copy the program you want to analyze to /tmp as a work around.Also, on OSX you have access to (mostly) the full power of dtrace, so you can do diagnostics on closed-source, debug-stripped running programs that make strace look small by comparison.
I already allowed myself to debug my own system by disabling "System Integrity" in recovery mode.
https://www.w3.org/Daemon/User/Installation/PrivilegedPorts....
fuser is useful too, and available on many Unixen. I've used "fuser -k" many times to find and kill rogue processes.
https://en.wikipedia.org/wiki/Lsof
https://en.wikipedia.org/wiki/Fuser_(Unix)
Edited for grammar.
https://bitbucket.org/twic/devtools/src/1b7a8f9ab849b36de70c...
Which, given a specification of a filehandle (eg an IP address or path) uses lsof to determine the PID of the owning process and the numeric value of the filehandle, and then uses gdb to attach to that process and close it. It's a crude but simple way of exercising error-handling code.
Are you running Elixir in the last line? the iex?
sudo gdb -batch -n -iex "set auto-load off" -p $TARGET_PID -ex "call close(${TARGET_FD})"
The -iex is a flag which executes a command on startup: -init-eval-command command
-iex command
Execute a single GDB command before loading the inferior (but after loading gdbinit files). See Startup.
[1] https://sourceware.org/gdb/current/onlinedocs/gdb/File-Optio...