Actually, I thought of another thing: when optimizing some module, I often find myself making a test harness that puts a lot of load on the module in question. Then I tweak it until it is fast. The problem is that I've optimized it in a state where the data is likely to be cached and the branch predictor has learned how the branches go. But then when you put that module back into a full program, the rest of the program might have overwritten our module's state in the cache and branch predictor, so the performance we get is much lower. In which case, Getting Accurate Results requires flushing the cache and branch prediction state in our test harness. I'd be interested to see some ideas on how to do that.
Flushing the cache is easy: just read 4MB (or whatever your L3 cache size is) of some dummy data and pretty much everything you had cached before would be kicked out.
At the same time, resetting the branch predictor AFAIK can only be done by powering the CPU off and on again.
About two decades ago, I was fighting to optimize a thing where the right answer depended on the icache pressure that the rest of the system was creating. More icache pressure -> inlining bad.