Show HN: Sysdig, a tool for Linux system exploration
github.com
github.com
I like the filter syntax - would be nice for perf_events to pick this up. Although, if it did, I hope that the stable filter fields API can be extended with unstable arbitrary expressions as needed, for when dynamic probes are used.
What perf_events realy lacks is a way for custom processing of data in kernel context, to reduce the overheads of enablings. Eg, lets say I want a histogram of disk I/O latency. sysdig has chisels, which look like they do what I want, but from the Chisels User Guide: "Usually, with dtrace-like tools you write your scripts using a domain-specific language that gets compiled into bytecode and injected in the kernel. Draios uses a different approach: events are efficiently brought to user-level, enriched with context, and then scripts can be applied to them." Oh no, not user-level!
I tested this quickly, expecting DTrace's approach (which is the same as SystemTap and ktap) to blow sysdig out of the water. But the results were surprising (take these quick tests with a grain of salt). Here's my target command, along with sysdig and DTrace enablings, and strace for comparison:
Target: dd if=/dev/zero of=/dev/null bs=1k count=1000k
sysdig: sysdig -c topfiles_bytes
DTrace: dtrace -n 'syscall:::entry /execname == "dd"/ { @[probefunc] = count(); }'
strace: strace -c dd ...
sysdig slowed the target by about 4x. DTrace, between 2.5 and 2.7x. strace (for comparison), over 200x. This is a worst-case test, and if I'm willing to slow a target by 2x then taking that to 4x doesn't make much difference. With what I normally trace, the overheads are 1/100th of that, so DTrace is negligible. The take-away here is that the overheads are closer to the "negligible" end of the spectrum than strace's "violent" end. Which I found surprising for user-level aggregation.The Sysdig Examples could do with some sanity checking. Eg:
"See the top processes in terms of disk bandwidth usage sysdig -c topprocs_file"
I saw:
Bytes Process
------------------------------
134.65M dd
4.82KB snmp-pass
603B snmpd
332B sshd
220B bash
107B sysdig
That's while my dd between /dev/zero and /dev/null was running. No "disk bandwidth"! :)edit: formatting
Good catch on topprocs_file, we'll have to find a better name for it.
In terms of overhead, we put a lot of effort in it and, as you pointed out, we're already extremely optimized. But we think we can do even better. For example, we don't have any kind of kernel-level filtering yet. Coming soon! :)
On my workstation the plain dd runs in 40 ms, with systemtap one-liner instrumenting/aggregating those 1 million write(2) syscalls in the kernel extended that tiny runtime to about 50 ms for a 1.25x slowdown. But such small numbers are hardly meaningful.
I'm curious to what extent userspace perf script postprocessing is deemed a technological equivalent to this; or why a new kernel module was deemed necessary versus the perf_event_open(2) ring-buffer abi.
Plus the packaging is top-notch; its kernel modules are rebuilt automatically on kernel upgrade via DKMS (which I wish other vendors like FusionIO would do).
edit: Nevermind, I found it. It's a kernel module and user app that uses Lua scripts for interpreting data. Sorry about my harsh tone before, but jesus I hate it when there's more gloss than content.
To answer the question "what it does and how", sysdig captures system calls and other system level events using a linux kernel facility called tracepoints, which means much less overhead than strace.
It then "packetizes" this information, so that you can save it into trace files and filter it, a bit like you would do with tcpdump. This makes it very flexible to explore what processes are doing.
We also pack it with a set of scripts that make it easier to extract useful information and do troubleshooting.
Dump system activity to file, so that sysdig can be used to process it later.
* sysdig -w trace.scap
Print process name and connection details for each incoming connection not served by apache.
* sysdig -p "%proc.name %fd.name" "evt.type=accept and proc.name!=httpd"
See the files where apache spends the most time doing I/O.
* sysdig -c topfiles_time proc.name=httpd
Show the network data that apache exchanged with 192.168.0.1.
* sysdig -A -c echo_fds fd.sip=192.168.0.1 and proc.name=httpd
Show every time a file is opened under /etc.
* sysdig evt.type=open and fd.name contains /etc
http://www.ktap.org/doc/tutorial.html#faq
Is Sysdig design similar?
From the architectural point of view, sysdig is closer to tcpdump/wireshark than to systemtap/ktap.
systemtap/ktap work similarly to dtrace: - a script is loaded into a user level process - the process compiles the script and dispatches it to a kernel module - the kernel module hooks the script into specific places in the kernel - the kernel module sends the results back to userspace where the user can see them
sysdig works this way: - the kernel module hooks into specific places in the kernel (using tracepoints), captures everything, and puts it into a shared memory buffer - the buffer is accessed from the user-level sysdig process that reconstructs state (so it knows that fd 23 means /etc/passwd) - filtering is applied - scripting in Lua is applied - the whole thing is optionally saved to disk so you can analyze later
Both approaches have pros and cons. We think that the sysdig approach creates a more natural workflow, ideal for troubleshooting and system administration tasks. Plus, writing scripts in Lua, with access to its rich libraries, is quite fun. :)
I want to give more details in a future blog post, so stay tuned.
Interesting enough, ktap vm is also based on luajit, if you want to add kernel filter scripting functionality into sysdig in future(without GCC needed), then just integrate ktap into your solution. :)
Looks nice otherwise. Too bad it needs a kernel module.
I would also like to point out that the sysdig workflow is quite different from the drace one. In addition to supporting real-time investigation, sysdig lets you take a rich "snapshot" of the machine activity that you can analyzer later. From this point of view, I don't think sysdig is less powerful than dtrace. Quite the opposite. But we're eager to know what you think.
And you have the option to install it manually if you want https://github.com/draios/sysdig/wiki/How%20to%20Install%20S....
Does appear to build fine with cmake/make, though.
An https link should be the default IMHO.
edit: the tool looks really useful though.
https://s3.amazonaws.com/download.draios.com/stable/install-...
We'll update the page in a second
http://drj11.wordpress.com/2014/03/19/piping-into-shell-may-...
Gerald Combs is the creator of Ethereal/Wireshark: https://www.wireshark.org/about.html
Less Dwight please.
https://github.com/draios/sysdig/issues/39
Has anyone encounter with this error before? Any help would be appreciated.
# sysdig fd.type=ipv4
error creating the process list
Has anyone seen this one before? Any help would be appreciated.
sudo sysdig -w file1.log
file1.log contains lots of junk characters (fix this) ^@^@^@^@^@^@^@^@^@^@^@^@^
Better alternative
sudo sysdig > file2.log
file has proper logs
"sysdig -w" switch will generate a binary dump (in a pcap format) containing the "raw events" coming from the kernel (plus a snapshot of information gathered from /proc), so it's not supposed to be human-readable, you have to use "sysdig -r" on the dump file to get the output.
If you're used to tcpdump, it's the same thing.
They're located in my town O.o