The Magic of strace
chadfowler.com
chadfowler.com
man 2 read
This should probably be mentioned somewhere.Otherwise, great writeup. Thanks for sharing!
(edited)
This might be a distro-specific thing.
What is the number after the command called? For example, when I look up 'man sed', I find a manpage for 'sed(1)'[0]. When I look up 'man kill', I find manpages for 'kill(1)' and 'kill(2)'[2].
Can somebody tell me what that number is called so I can look it up? Thanks..
0: https://developer.apple.com/library/mac/documentation/Darwin... 1: https://developer.apple.com/library/mac/documentation/Darwin... 2: http://man7.org/linux/man-pages/man2/kill.2.html
Under windows, strace is an SSL/TLS monitoring tool (also hella useful). It shows payloads passed to CryptoAPI/CNG libs so you can easily troubleshoot explicitly encrypted protocols like ldaps. Especially useful if you use client authenticated TLS where is is not possible to use a TLS mitm proxy to snoop the layer 7 data.
VMware is using SpyStudio for creating and troubleshooting application virtualization packages, this is, for example, a twitter post from a VMware escalation engineer: https://twitter.com/DooDleWilk/status/428562701313662977
[1] http://www.nektra.com/products/spystudio-api-monitor/
[2] http://www.nektra.com/products/deviare-api-hook-windows/devi...
Running version 7 on my work laptop. :(
I wrote a short article about stap using Rubys probes as an example: http://www.asquera.de/blog/2014-01-26/stap-and-ruby-2
Thats why I decided to focus on Linux > 3.5, which will be available with Ubuntu 14.04 LTS, where installing gets much easier if you know the right packages. I definitely wanted to make sure that people can start playing around with it in a few minutes.
Also, UPROBES are in my opinion one of the most interesting features to grab with strace. It allows you to easily combine detailed kernel-level tracing with tracing of your application.
(Not that I want to suggest that compiling your own kernel is hard to do, it just takes the fun out of "let's trace!")
strace / dtruss just spit out kernel syscalls of any user land program wo modifications.
Also dtrace, as seen on smartos
That's useful and all, but what if you want to instrument arbitrary parts of a program, not just the syscall interface? By function or instruction? Either in userspace or in kernel? With statistical functions? And speculative tracing? And extensive control flow (except loops, which prevent certain safely guarantees DTrace makes). And a lot more.
Don't be fooled by the single-letter change: strace is to DTrace what edlin is to emacs. Or something else ridiculously extreme. They're barely comparable.
>What is this "DTrace" thing? It stands for "Dynamic Tracing",
>a way you can attach "probes" to a running system
>and peek inside as to what it is doing.It can be used to write tools like strace (see "dtruss" on OSX), iotop, topsyscall, etc.
Disclaimer: LTTng developer :)
So I think the comparison, despite the superficial similarity and similar mechanism of action, isn't really fair.
We've got a fairly complicated bioinformatics pipeline that calls about 100 other programs, and creates or reads about 100 different files. I'd love a way to create a picture of what's going on. Which files each program uses, etc.
If such a program doesn't exist, would that be worth building? Could it be something I could potentially sell?
#!/usr/bin/perl -w
$|=1;
use strict;
my (%pidmap, @order);
while ( <> ) {
chomp;
if ( /^(\d+)\s+(\w+)(.*)$/ ) {
my ($pid, $syscall, $args) = ($1, $2, $3);
if ( $syscall =~ /(^clone$|fork$)/ and $args =~ / = (\d+)$/ and $1 > 0 ) {
my $clonepid = $1;
$pidmap{$clonepid} = { -parent => $pid };
push(@order, $clonepid);
}
elsif ( $syscall =~ /^exec/ and $args =~ / = (\d+)$/ and $1 == 0 ) {
my $exec = $args;
@order = ($pid) if !@order;
$exec =~ s/^\("([^"]+?)",.*$/$1/g;
push( @{ $pidmap{$pid}->{-exec} } , $exec );
}
}
}
foreach my $pid ( @order ) {
my $spaces = walkpid($pid);
print " " x $spaces . join("\n" . (" " x $spaces), map { $_ . " ($pid)" } @{ $pidmap{$pid}->{-exec} } ) . "\n";
}
sub walkpid {
my $pid = shift;
my $c = shift || 0;
if ( exists $pidmap{$pid}->{-parent} ) {
return walkpid($pidmap{$pid}->{-parent}, $c+1);
}
return($pid, $c);
}particularly when you don't know which process is calling all the syscalls.
Mix "perf record" and "perf trace" & you have the next generation of strace tools.
perf --help
trace strace inspired tool- allowed exploring forking behavior of daemons, in particular the nitty-gritty of gunicorn's prefork behavior, and understanding the rationale behind single- and double-fork daemons generally (very important to understand for job control e.g. writing upstart/init.d jobs)
- isolated hot reads to memcache in situ, by identifying the socket associated with the memcache connection, and finding which key was read the most by a process (we built better logging after the fact, but sometimes there's no substitute for instrumenting prod during tough perf/stress problems)
- let me explore the behavior of node.js's several threads, and find one of them sending "X" over a socket to the other (still not quite sure what this is, some kind of heartbeat/clock tick?)
- helped understanding "primordial processes" and the exact details of how forking/reparenting work on linux
It's a great tool and one that every ops/infrastructure engineer should be familiar with.
I don't know about node.js specifically, but this is a common pattern to wake another thread that uses a select()-style event loop.
strace -f <command> 2>&1 | grep ^open
Really useful to see what config files something is reading (and the order) or to see what PHP (or similar) files are being included.There's normally other ways to do this (eg using a debugger) but sending strace's stderr to stdout and piping through grep is useful in so many cases it's become a command I use every day or 2.
More modern tools such as dtrace for the solaris and systemtap for linux addresses similar problems but with a broader coverage.
I'd also like to point out that a key to using strace successfully is the result column... Programs that fail often make system calls that fail right before they exit... You can often tell what the program is trying and failing to accomplish...
I just thought I’d let people know that it can be a lot easier to read strace’s output if you read the output log file using Vim as it contains a syntax file which can highlight PIDs, function names, constants, strings, etc. Alternatively, if you don’t want to create an strace log file, you could pipe the output to Vim and it will automatically detect it as being strace output, e.g.
strace program_name 2>&1 | vim -Tried my luck with gdb. Sure enough...there was libQt5DBus pointing to the old libs leading to the crash. If you are feeling particularly adventurous, you can step one instruction at a time after starting. Even without debug symbols, there is quite a lot of info that be used while troubleshooting.
also remember also useful 'ltrace' - libraries tracing
ftp://86.0.252.89/pub/release/website/tools/trace-20140126-x86_64-b95.tar.gz
This is a tool called ptrace - which does everything that strace does and a lot more. You have working binaries in there, and most of the source - I havent extricated the full build dependencies so it all builds, but this includes extra facilities like reporting summaries of process trees, showing only connections or files, and shlib injection into a target process.
If people are interested more on this, contact me at CrispEditor-a.t-gmail.c-o-m
So now we know his ISP too. Useful, eh?
Elsewhere in this discussion: There's a difference between man page section 1 and 2 - and read was quoted as an example for a potential ambiguous result if you invoke "man read" (opens man 1 read here, when man 2 read was the syscall I might want to look at after running strace).