Precise timing of machine code with Linux perf
easyperf.net
easyperf.net
Wish I could get more time to work on micro-optimisations, in this age you often get funny looks proposing to dig down and optimise sluggish code. Spin up more compute and call it a day they say.
While that's completely understandable from a business perspective, it's somewhat unsatisfying as a developer, the expertise that comes from it surely has lasting benefits too.
The sample program is doing 2^7 = 128 NOPs, 4 at a time for 32 cycles, and then it is doing a memory access.
The address of memory that is going to be accessed is known right before doing the 32 cycles of "work", so a prefetch can be issued at that time.
The meaning of the 'prefetch windown' term is the number of cycles that you have between when you issue the prefetch to when you issue the instruction that accesses the address that was prefetched. So it is based on the structure of the program being analyzed.
I would better go for analyzing not the whole function (all basic blocks of the function) but only the Hyper Blocks (typical hot path through the function). Here is the example of how to do it: https://lwn.net/Articles/680985/ chapter "Hot-path analysis".