[1] http://www.cs.umass.edu/~emery/pubs/stabilizer-asplos13.pdf
[1] http://www.cs.umass.edu/~emery/pubs/stabilizer-asplos13.pdf
Care to explain why? I find the argument in http://blog.kevmod.com/2016/06/benchmarking-minimum-vs-avera... for using the minimum to be pretty compelling.
Also, FWIW, Julia's automated performance tracking determines likely regressions based on the minimum. Their reasoning is explained in http://math.mit.edu/~edelman/publications/robust_benchmarkin....
Imagine, for example, your program has 100 possible ASLR states: in 99 of those states it takes 1s to run and in the 1 remaining state it takes 0.8s to run. Using the minimum, you will think your program takes 0.8s to run even though there's only a 1% of chance of observing that performance in practise. That's bad, IMHO, but the problem compounds when you compare different versions of the same program. Imagine that I optimise the program such that in those 99 slow states the program now takes 0.9s to run, but in the remaining state it still takes 0.8s. The minimum will tell me that my optimisation made no difference (and thus should probably be removed) even though in 99% of cases I've improved performance by 10%.
The problem with the existence of 100 independent states is that you need a very large number of trials to get a performance measurement of each—and you're still left with the problem of statistic to reduce each to (min? mean? and how??). Attempting to take the mean with the CST should eventually work, but it might need a very long time. Instead, the minimum lets to pick whichever mode was fastest and compare the time of that mode. Sure, that might not be great for all cases, but if that's the worst problem with the benchmark suite, I'd say you're in pretty good shape. I've run into issues where the CPU appears to have memorized the random number sequence in the test data—how are you supposed to pick the better algorithm when the CPU won't let you run the 99% case under a benchmarking harness...
Yeah, minimum isn't perfect, but it's at least pretty clear what it says, as long as there isn't incentive to abuse it (I might say the same about p-value vs bayesian).
disclosure: worked with the author of the Julia paper cited above.
You’ve lost me. What is the single, obvious meaning of the minimum measurement, and what are you benchmarking against, and what do you intend to get out of the benchmark? Typically the minimum just means “program ran x input as fast as y”, which doesn’t actually help with most questions about performance.
If X is constant, and time is the only thing you care about, then maybe that works.
well, from a marketing standpoint it means that you are able to write
OBSERVED MAXIMUM SPEED : 789634 gigabrouzoufs per second on a core i5-xxxx
on your brochure, and you wouldn't be lying, and it would be better that a majority of products which don't even give an observed speed but just a theoretical one.
[1] https://soft-dev.org/pubs/html/barrett_bolz-tereick_killick_...