Actually, I thought of another thing: when optimizing some module, I often find myself making a test harness that puts a lot of load on the module in question. Then I tweak it until it is fast. The problem is that I've optimized it in a state where the data is likely to be cached and the branch predictor has learned how the branches go. But then when you put that module back into a full program, the rest of the program might have overwritten our module's state in the cache and branch predictor, so the performance we get is much lower. In which case, Getting Accurate Results requires flushing the cache and branch prediction state in our test harness. I'd be interested to see some ideas on how to do that.