Incurring a 1 us overhead on each function call is very steep if you are doing a function entry/exit trace and nearly totally smears the profiling information you could get. In contrast, efficient recompilation-based instrumentation should only incur maybe 100 ns down to maybe around 10 ns depending on how aggressively you instrument and how much overhead you are willing to incur in the logging disabled case. In aggregate, a efficient recompilation-based approach should only incur a whole program overhead in the low double digit percent range when enabled and at most a low single-digit percent, if even that, when disabled. As a corollary, if 1/10th the per-invocation overhead results in say a aggregate 30% overhead, then we can reasonably assume the full overhead case is around 10x as much overhead resulting in 300% aggregate overhead, or a program taking 4x as long to run. That is a qualitatively different amount of overhead.
[1] http://dtrace.org/blogs/brendan/2011/02/18/dtrace-pid-provid...
Most of the overhead comes from the fact that it's using kprintf() to print the tracing info, since I'm happy to spend a few extra nanoseconds having more elegant code. So it could totally be improved further. Another thing is that right now it's only line buffered. So if it buffered between lines, it'd go faster.