KUtrace: Low-overhead Linux kernel tracing facility
github.com
github.com
It gave me appreciation for the amount of knowledge one accumulates over a career, and what a loss it is to an organization when one so knowledgeable retires.
KUtrace is absolutely one of the most powerful tools I've used for deeply understanding performance bottlenecks (after isolating issues) such as poor scheduling behavior. I would highly recommend reading his book "Understanding Software Dynamics" [1] if you are interested in learning more about KUtrace or performance bottlenecks/optimizations in general. The book is quite dense and dives deep into the performance characteristics of many examples of the five fundamental resources (according to Dick): CPU, Memory, Disk/SSD, Network, and Software critical sections.
[1]: https://www.oreilly.com/library/view/understanding-software-...
[1]: https://www.intel.com/content/www/us/en/docs/vtune-profiler/...
How applicable to the general cases is it? I’m deeply interested in the topic, but unlikely to actually be running KUTrace, fwiw.
"Measurement" delves into understanding and measuring four fundamental resources: CPU, Memory, Disk/SSD, and Network. This section is quite dense and explores both the depth and breadth of understanding performance of programs. For example, there is a chapter on optimizing code to use caches more efficiently. Though I will say this section is obviously not a complete exploration of all aspects of performance as there are many many many more things which can affect the performance in such complex systems like modern computers. But what Dick does is in this section is to give you more tools in your toolbelt to understand performance better.
"Observation" looks at existing tooling (so profilers, tracing tools, etc.) and discusses where they are useful or where they fall short.
"KUtrace" introduces KUtrace, its kernel module, and its timeline visualization tool. It discusses its design and implementation and why is it so fast and low-overhead.
"Reasoning" has case studies that looks at particular kinds of performance pathologies such as "waiting for CPUs" etc. Dick uses KUtrace here to tease out the underlying inefficiencies in the analyzed programs.
So the first two sections are essentially orthogonal to if you want to use KUtrace or not, but the last two sections are about KUtrace and how to use it to understand performance bottlenecks. Even if you don't use KUtrace, the "Reasoning" section can still be insightful imo as KUtrace is just a tool at the end of the day, and the real insight is why or what is causing the performance issue.
Nice handle btw. Grinding for that was unforgettable...
The upside to perfetto of course is the much much richer tooling, infrastructure, and ease of use since it comes pre-installed on your phone.
> Nice handle btw. Grinding for that was unforgettable...
Haha thanks. Symphony of the Night is easily one of my favorite games -- I can pick it up any time and play it until 200.6% completion ;)
- Hacker's Delight
But it works by patching the kernel, not just using eBPF like many performance tools recently. So it needs active maintenance all the time considering the current velocity of internal kernel changes. And I would not be surprised if it didn't build or work correctly if you have a heavily patched and customized kernel.
On the positive side at a first glimpse the maintenance to adapt to new kernels looks very active.
Also, if the overhead is negligible, maybe the author can try to merge this into mainline with the use of static key to make the incurred overhead switchable. In spite of the static key, the degree of the accompanied inteferences on cache and branch predictor might be an intriguing topic though.
$ sudo wc -l /sys/kernel/debug/kprobes/blacklist
783 /sys/kernel/debug/kprobes/blacklist
Edit: Perhaps an alternative approach would be to attach probes to relevant (precise) PMU events. There's also this prototype of adding breakpoint/watchpoint support to eBPF [1]. But actually doing stuff within this context may get complicated very fast, so would need to be severely limited, if feasible at all.[1] https://ebpf.io/summit-2020-slides/eBPF_Summit_2020-Lightnin...
Note also that this work emerged within Google a decade before eBPF was really useful.