Statistically rigorous Java performance evaluation
blog.acolyer.org
blog.acolyer.org
For example, see Table 2 in a recent paper I am a co-author on (page 8 of the pdf, page 73 using the proceedings numbering): http://www.scott-a-s.com/files/debs2017_daba.pdf In this paper, we care about latency, and we report the average latency along with the standard deviation. Here, a tighter standard deviation is more important than confidence that the mean falls within a particular range. And the variation in latencies is caused by both software and hardware realities of the memory hierarchy.
I sometimes have used standard deviations for errors myself, but I'm not really sure I can defend it.
This is, I think, a philosophical difference from, say, measuring the charge of an electron. The electron has a charge. The mean of independent measurements is, we hope, very close to that true value. Deviations from that mean are indeed error. But when measuring performance in a computer system, all observable values are valid. We're trying to figure out not what the "true" value is (there isn't one), but what the range of values your performance can take on, and where in that range you're likely to fall. (I almost wrote something about excusing "real" errors such as memory corruption, disk failures and segfaults, but sometimes you want to include that! If you're doing performance analysis of a distributed system that scales to hundreds of thousands of compute nodes, real runs of an application are likely to encounter such things, so your system better be resilient to them, and they will have an impact on real performance.)
Things like cache locality, branch prediction, prefetch etc. should be flushed out with a good microbenchmark prologue.
A value of microbenchmarks is to establish a stable environment when a baseline can be set, where predictions can be tested. It's a development tool, not a qualification tool. You will have anyway to measure the real performance impact in the complete environment.
Deficiencies of mean and std are easy to demonstrate. Just draw multiple samples of different sizes from a Pareto distribution (some other long tailed distribution would work too) and observe how drastically the mean and std varies among those samples.
Gaussian distribution based methods, for example ANOVA, had their moment of glory in centuries past (good stats has since moved on). They are still fantastic tools when the assumptions they make are true (this is rather rare) but they break badly when those assumptions are violated a tiny wee bit. One can easily end up making a wrong decision.
At the very least one should verify that the Gaussian assumption holds before believing those statistics. There are many post-Gaussian statistics that one can use, if it turns out that the data is not Gaussian.
I agree that trimmed-means are L-moments are better at characterizing a distribution. But! I had to look them up just now, and I have never seen a computer systems paper which used them. I don't know if most readers would be able to intuit meaning from them. My claim is not that mean and standard deviation is the best way to characterize computer systems performance data, but that for my work, it's better than mean and a confidence interval.
This is great to hear, hope everyone follows this.
> I have never seen a computer systems paper which used them. I don't know if most readers would be able to intuit meaning from them.
That's rather unfortunate, is it not ? Current practice is broken. So should one not strive to fix it and introduce the community to better tools, champion better tools. I am aware I am being harsh here but, correctness be damned, lets follow the crowd does not show the community in stellar light. If trimmed mean is too exotic, even median and inter-quantile distance would be fine.
> My claim is not that mean and standard deviation is the best way to characterize computer systems performance data, but that for my work, it's better than mean and a confidence interval.
I have two comments to make (i) Its not the case that mean and std are mostly OK but rather suboptimal. I would be fine with that. The thing is these can be horribly wrong and unrepresentative. Seemingly innocuous looking distributions have infinite variance. Since you use a lot of Gaussian based tools, you would be familiar with a t-test and the t-distribution. The variance of a t-distribution with 2 degrees of freedom is infinite, any finite number one reports for the std would be wrong by an infinite order of magnitude.
> it's better than mean and a confidence interval.
(ii) I am having trouble with this one. Assuming that confidence intervals have been computed using the Gaussian assumption, width of the confidence interval would be a scaled up version of the standard deviation, nothing fundamentally different. Instead of std, you get, say, 3x of std. When they are bad, both are equally bad.
I am not against the use of Gaussian, sometimes it is very reasonable, but if correctness is a goal, then it needs to be demonstrated that it is not a wrong thing to use. There are standard techniques to verify that using Gaussian based stats would be fine.
I would be sympathetic if it was the case that we don't know how to handle departures from Gaussian, or it is very expensive to handle that. Neither of these are true, and haven't been true since the 80s. in fact earlier.
The point of confidence intervals in TFA is to determine whether there are statistically significant differences - but I think they must be making their own assumptions about the underlying distribution to come up with the CIs. So not sure how much I'd trust it.
"Quantifying performance changes with effect size confidence intervals" Tomas Kalibera and Richard Jones Technical Report 4-12, University of Kent, June 2012.
https://www.cs.kent.ac.uk/pubs/2012/3233/
Kalibera, Tomas and Jones, Richard E. (2013) "Rigorous Benchmarking in Reasonable Time"