I’m surprised that the author of the post (who I generally find writes interesting, thoughtful and seemingly correct things) feels that the research talked about in this article was pushing the state of the art in reliable benchmarking and understanding vm warmup, but that your comment seems mostly dismissive. I would especially expect a VM developer to be interested in more reliable benchmarking. Do you think that that assessment by the author was incorrect or that these things are widely known among vim developers?
I also don’t really understand your comment, eg you claim that VMs win on average but the numbers in the article show that, even for a case which ought to be good for a JIT (I might be wrong here?) with a small very hot section, VMs on average fail to end up in a steady state with better performance than the interpreted case. Is the idea that the good steady state performance you get some of the time will average out with the poor or inconsistent performance you get other times to get good performance? I can sort of see that argument being made about a big system with many instances but I don’t think it’s a correct one because in practice many systems like that will block on the slowest part and so have worse tail performance than a single component. I guess another idea could be that a real program has many parts and if 10% reach a steady fast state while the rest are inconsistent or slower, then on average the program can still be faster. But I don’t really buy this because I expect there are a few places that will matter far more and so you don’t average out very well, but my prior could be badly wrong here.