Linux Performance Analysis: New Tools and Old Secrets
brendangregg.com
brendangregg.com
It only got better once they added trace points to perf events. Now I can simply trace all futex syscall enters to look for lock/synchronization contention.
One extra tool I recommend is dynamic logging which you can enable from kernel build. Create any log mask (file/like me/module) set a level and you're off to the races. Invaluable in certain scenarios (like debugging bugs in fail path of fscache module).
Does anyone else spend hours tweaking builds/configs only to then just assume you have made an improvement without confirming it?
I wish I had learned of Knuth's "premature optimization is the root of all evil" quote sooner.
Edit: I should add -- there are three examples in the original post that use ftrace. You can't currently do any of these with sysdig.
My company has a performance engineering team, but I imagine smaller companies don't have any performance staff, or time to deal with understanding this. Hence the value in creating canned tools for others to use, provided their warnings are made clear.
Is there an authoritative beginners guide to linux performance analysis? A cursory glance at trace-cmd tutorials shows I have a lot to learn. Or maybe it would be more beneficial to start playing with a tool like sysdig?
Not a ton on ftrace in it, though. Maybe we'll see a new book more focused on Linux specifically with some ftrace stuff?
I've been thinking of how best to cover my new ftrace and perf_events content in book form...
As for learning how to do all the stuff... If you're a programmer, a good approach is to write your own simple poor performance programs to anaylze. Eg:
- burn CPU in a function
- do too much disk I/O in a function
- do lots of large size network I/O from a function
- do heavy memory allocation from a function
Now analyze these using the Linux toolset. Quantify the behavior (how much per second), show its effect on system resources (%utilized, or IOPS, or whatever), then drill down and identify the code path responsible: the function.This sounds easy, since you wrote the program to start with, and already know the answer. But there's often a lot of quirky behavior with perf tools, and assumed knowledge to learn, that you want the target to be as easy as possible.
A mistake beginners make is to aim performance tools at complex production workloads, and then get lost and confused. It's like trying to run before knowing how to walk.
I think Julia Evans is going to cover this approach in an upcoming talk, so that should become a good reference.
You'll also want to develop a good understanding of software, systems, and kernels, so that you can reason about performance. That may take years to pick up, and the study of various text books. I tried to summarize all the content in my last book, systems performance.
Edit: formatting.