Because warmup is probabilistic, there must be cases where it goes off the rails.
So yeah, I’m dismissive.
Because warmup is probabilistic, there must be cases where it goes off the rails.
So yeah, I’m dismissive.
First, VM authors I've discussed this with over the years seem roughly split down the middle on microbenchmarks. Some very much agree with your perspective that small benchmarks are misleading. Some, though, were very surprised at the quantity and nature of what we found. Indeed, I discovered a small number had not only noticed similar problems in the past but spent huge amounts of time trying to fix them. There are many people who I, and I suspect you, admire, in both camps: this seems like something upon which reasonable people can differ. Perhaps future research will provide more clarity in this regard.
Second, for BBKMT we used the first benchmarks we tried, so there was absolutely no cherry picking going on. Indeed, we arguably biased the whole experiment in favour of VMs (our paper details why and how we did so). Since TCPT uses 600 (well, 586...) benchmarks it seems unlikely to me that they cherry picked either. "Cherry picking" is, to my mind, a serious accusation, since it would suggest we did not do our research in good faith. I hope I can put your mind at rest on that matter.
- Academics don’t publish results that aren’t sexy. How many people like you ran the same experiment with a different set of benchmarks but didn’t publish the results because they confirmed the obvious and so were too boring? How many times did you or your coauthors have false starts in your research that weren’t published? You’re cherry picking just by participating in the perverse reward system.
- The complexity of the data analysis sure makes it look like you’re doing something smart, but in reality, it’s just an opportunity to cherry pick.
- These results are not consistent with what I’ve seen, and I’ve spent countless hours benchmarking VMs I wrote and VMs I compete with. I’ll believe my own eyes before I believe published research. This leads me to believe there is something fishy going on.
Anyway, my serious accusation stands and it’s a fact that for large real-ish workloads, VMs do “warm up” - they start slow and then run faster, as designed.
To that end we work in the open, so all the evidence you need to back up your assertions, or assuage your doubts, has been available since the first day we started:
* Here's the experiment, with its 1025 commits going back to 2015 https://github.com/softdevteam/warmup_experiment/ -- note that the benchmarks are slurped in before we'd even got many of the VMs compiling. * You can also see from the first commit that we simply slurped in the CLBG benchmarks wholesale from a previous paper that was done some time before I had any inkling that there might be warmup problems https://github.com/ltratt/vms_experiment/ * Here's the repo for the paper itself, where you can see us getting to grips with what we were seeing over several years https://github.com/softdevteam/warmup_paper * The snapshots of the paper we released are at https://arxiv.org/abs/1602.00602v1 -- the first version ("V1") clearly shows problems but we had no statistical analysis (note that the first version has a different author list than the final version, and the author added later was a stats expert). * The raw data for the releases of the experiment are at https://archive.org/download/softdev_warmup_experiment_artef... so you can run your own statistical analysis on them.
To be clear, our paper is (or, at least, I hope is) clear to scope its assertions. It doesn't say "VMs never warmup" or even "VMs only warmup X% of the time". It says "in this cross-language, cross-VM, benchmark suite of small benchmarks we observed warmup X% of the time, and that might suggest there are broader problems, but we can't say for sure". There are various possible hypotheses which could explain what we saw, including "only microbenchmarks, or this set of microbenchmarks, show this problem". Personally, that doesn't feel like the most likely explanation, but I have been wrong about bigger things before!