Pin - A Dynamic Binary Instrumentation Tool
pintool.org
pintool.org
Second, PIN gives you very low-level access, sure, but it also requires quite a bit of work to get that data. You can get a lot of the same data (as seen in the examples) with libpfm4 (http://perfmon2.sourceforge.net/). CPU Performance counters can give you a lot of this data with much lower run-time overhead, and a lot less work.
Also, PIN's a little hairy. If you just want to generate code for run-time execution, LLVM's your best bet. If you want to diddle with a running executable, you can always use libelf(3) and ptrace(2) to read and diddle with the running process. It may be useful for specific sorts of analyses you want to run on an executable, but it's messy. If you're doing performance instrumentation, dynamically modifying the code is going to alter your results that can be hard to compensate for.
Performance counters can tell you quite a bit, and they cost very little to set up. Snapshots and a little differential analysis can get you more comprehend-able data without transfer/storage problems.
Want to profile dynamically generated code? libelf won't be much help there. Or do you want to run memory accesses through a cache simulator? ptrace won't be much help. There are plenty of uses when DBM is the right (or only) tool.
What's the preferred way of using libpfm4? I've ended up using it's sample program as a way to convert from readable counter names to hex to put into a perf command line. I've found several defunct patches to give perf this functionality directly, and am confused why this seemingly essential functionality is left out.
What I'd like is the ability to measure just sections of code, and access to all available counters without needing to copy-and-paste hex. I'm getting the sense that perf is not the tool for this:
http://lwn.net/Articles/441209/
http://www.mail-archive.com/linux-perf-users@vger.kernel.org...
Is Likwid a viable option? https://code.google.com/p/likwid/
A couple other uses that I thought it would be good for were marking instructions with unaligned loads and stores and flagging switches to and from 256-bit VEX code. Although probably you can do these with other tools as well.
I posted this here because it seemed potentially useful, and I was surprised I'd never heard of this project. I'm mostly familiar with the ones Lally mentioned, but Thomas brought up DynamoRIO which is also new to me. Are there other niche optimization tools I should know about?
A specific question would whether there is anything more useful than IACA and gut instinct for determining optimal instruction order. Even just something that would generate an easy to parse data dependency graph for a short section of code?
[1] http://www.offensivecomputing.net/?q=node/1687 [2] http://www.offensivecomputing.net/?q=node/492
The main focus of the DBT tool is Linux kernel modules, but let me know the kinds of stuff you need it for and I can a) figure out if my tool is applicable, and b) perhaps share the code.
I need to know everything about that binary. How it works, what ports(unix domain & network sockets), files it opens on the harddrive, libraries it's linked to, how it decides what to do. Anything & Everything there is to know about it. ^_^
1) What ports it opens:
- netstat shows you what ports a program has running
- DTrace shows you all syscalls a process makes (among other things). dtruss is a convenient wrapper script included in OS X which shows you all the syscalls a process makes (including opening sockets.
2) What files it opens
- Again, DTrace's syscall provider lets you introspect all syscalls, including open(). There's even a handy wrapper script included with OS X called opensnoop.
- Alternately, you can use the fs_usage command line tool to tap into the xnu kernel's trace mechanism. This shows all sorts of filesystem events, including what files are opened.
3) What libraries a binary is linked to
OS X binaries use the Mach-O format, not ELF like most other Unixes. So you have to use OS X's binary introspection tools to understand that format rather than the standard GNU binutils. What you're looking for here is otool, which lets you introspect Mach-O binaries. Specifically, "otool -L /Applications/Mail.app/Mail" for instance shows you which libraries Mail links to. Run this recursively to get the transitive closure of all dependencies a binary links against. Another way to do this is to run "vmmap -v <pid>" to show you the vm layout of a process, which includes the __TEXT/__DATA segments of all libraries the process links against.
And of course, gdb/lldb is included with the developer tools, you can just attach to whatever process you care about and set breakpoints, type "info sharedlib" to see what libraries are in the address space, etc. Also, for better or worse, Objective-C is an extremely dynamic language, so you can even do things like write a shared library with code you want to inject into a process (potentially monkey-patching existing methods using ObjC categories) and dlopen it from gdb to insert it into the target process's address space.
This blog talks about the tool for automated testing. It's kinda "complicated" to setup. In the blog, he eventually gets to a commandline:
instruments -t /Applications/Xcode.app/Contents/Developer/Platforms/iPhoneOS.platform/Developer/Library/Instruments/PlugIns/AutomationInstrument.bundle/Contents/Resources/Automation.tracetemplate "/Users/jc/Library/Application Support/iPhone Simulator/5.1/Applications/C28DDC1B-810E-43BD-A0E7-C16A680D8E15/TestAutomation.app" -e UIASCRIPT /Users/jc/Documents/Dev/TestAutomation/TestAutomation/TestUI/Test-2.js
But there's a lot of setup, so =/
Huh. What are the licensing conditions and price for commercial use then?
* Valgrind was designed to support rich analysis plugins (like Memcheck, which keeps a shadow copy of every bit of data) and performance was a secondary concern (on Valgrind, applications run on average about 4x slower, threads are serialized, etc).
* DynamoRIO and PIN are designed not to make much of an impact on performance (usually a few percent) and are more suitable for running in production, but it's somewhat more complicated to write plugins for them.
Both DynamoRIO[0] and Valgrind[1] maintain lists of publications which go into much more detail.
[0] http://www.dynamorio.org/pubs.html [1] http://valgrind.org/docs/pubs.html
Pin, on the other hand, dynamically rewrites the binary to inject instrumentation. This allows it to inject instrumentation code at a higher granularity (individual instructions). It's useful where the application might not define dtrace probes, or, for example on OSX, where the application has "opted out" of it to protect it from being screenshotted (a la iTunes).
Pin is more like a scripted debugger than an instrumentation tool.