Llvm-mca – LLVM Machine Code Analyzer
llvm.org
llvm.org
(The code is from the stream vbyte repo; see https://lemire.me/blog/2017/09/27/stream-vbyte-breaking-new-... ).
Edit: Example interesting fact: the first iteration takes 35 cycles before it finishes; but the reciprocal throughput of the loop (assuming it executes a few times) is 5.3 cycles.
April 2019: Intel® Architecture Code Analyzer has reached its End Of Life. Users may want to try LLVM-MCA. This is NOT a recommendation to use LLVM-MCA nor a comment on its accuracy or usefulness. Thanks for being faithful users of Intel Architecture Code Analyzer throughout the years. We hope it was useful for you.
From: https://software.intel.com/en-us/articles/intel-architecture...
Article: https://pdziepak.github.io/2019/05/02/on-lists-cache-algorit...
Understanding where the limiting factor is can help optimize and tune small chunks of code.
The gist of it is that your CPU loves to multitask (because that's so much faster), and if you like performance you want to maximize how many things it's doing at once at any time. You want every part of the CPU to have work on its schedule all the time so it doesn't sit idle. This shows you what the schedule for your code looks like so you can optimize it.
--
In more details, this tool is going to read your code instruction by instruction, compute what kind of schedule the hardware will be able to make for executing each instruction, and tell you which circuits might be overworked or going unused. (Caveat: the CPU makes its own schedule on the fly as best as it can — this tool is just an approximation in software — but your compiler's best guess is still a pretty good guess in general!)
For example (actual numbers! [0]) your CPU could have 4 "ports" (circuits) that can start doing math on integers (with their ALUs) at any instant in time, or 2 of those "ports" could also multiply floating-points numbers, but each circuit can only be given one kind of task at a time to keep things manageable. Well it turns out making a good schedule is a surprisingly hard problem, since if you send int work to the first two ports and you didn't foresee there'd be float work coming after, the first two ports will have a full schedule while the other two will be doing nothing. A better schedule could have had all four of them busy in parallel!
You don't directly have a say in the schedule — the hardware does its best — but with LLVM-Mca (or Intel's IACA, which inspired it) you can write code that you know will be easy to schedule in parallel, and that's already some pretty awesome tools to have!
[0]: https://en.wikichip.org/wiki/intel/microarchitectures/skylak...
That's not quite true. The order you present the instructions influences the actual execution order greatly, and it's why instruction scheduling remains an important part of compilers.
You're right that there are many many things you can do to the code to influence the scheduling — like the reordering the compiler is doing — and at the end of the day that has a predictable impact on the scheduling. Don't get me wrong I love my compiler, and the fact that we can impact scheduling is why LLVM-Mca is useful in the first place.
What I meant to write is that your x86 isn't some kind of mostly statically scheduled VLIW. the behavior of the hardware is only partly predictable, and even IACA has to make some tragic simplifications. Tweaking the alignment to play with fetch boundaries has an effect, vectorizing obviously does, picking a different mix of instructions can help, artificially loading a port to prevent a bad scheduling decision down the line is not always entirely stupid, etc...
I feel it's important to keep in mind that the hardware scheduler keeps dynamic statistics on port usage, so in a sense it's more like a JIT than a compiler, static analysis is only an approximation. What little experience I have told me it's always a good idea to compare IACA and LLVM-Mca's predicted schedule with a real profiler's output :)
(Thanks for giving some nuance, it's appreciated. I'm not actually a compiler engineer or doing low-level magic for a living, so if you see anything wrong I would love to be corrected!)
Which targets does that include?
On x86, they have a scheduling model for most recent Intel Tock cycles going back to Sandy Lake, and AMD Piledriver, Jaguar, and Zen microarchitectures.