All these things make the inter-language comparison as presented nearly useless. Crucially, the site itself doesn't even state what it is that it purports to measure or compare, like a paper without any hypotheses, claims or exposition. That alone is enough to place it in the low-quality benchmark category (regardless of the merits of the problems themselves, which I haven't looked at). But that's where most non-professional benchamrks are. ¯\_(ツ)_/¯
Now, I would guess that there is some meta-hypothesis to the game pertaining to markets, but that's not what's being analyzed or even presented.
> Which programs?
I tried to look at programs that report big performace differences and I'm looking at C++'s regex-redux, and I see it uses a library that isn't in the standard library. Or look even at "Java" vs "Substrate VM" pidigits. These are two VMs for the same language, yet the programs are completely different. I'm not sure what I could learn from that about HotSpot vs Substrate VM.
Or, the example that caught my eye when I reached the game when looking for SBCL benchmarks -- SBCL's reverse-complement, that is reported to outperform C by more than 10x. Well, not only does it use a completely different low level algorithm (doesn't use parallelism), it doesn't outperform C or Java at all, but crashes, and what appears to be some junk result is reported. Had it not crashed, it would still have used a very different algorithm.
> So why haven't you told us how few 1/10ths of a second?
Because that would be misleading. Now, I admit that what threw me off at first was Java's performance compared to SBCL's reverse-complement, but that's just a mistake and has nothing to do with compilation. Real programs can take many seconds to compile, but looking at the benchmarks, they do look quite small, so you are correct that they will be compiled rather quickly (less than a second). However, I see that Java is run with tiered compilation, and obtaining good optimization, even for such small programs, can still take many seconds (say 10 or 30). As a general rule of thumb, we always warmup for at least a few seconds, even for microbenchmarks.
It is true that the difference would often not be large compared to the inter-language differences reported by the benchmark -- and as I said above, it's unclear what it is they compare -- but they can certainly be in the 5-10% range, and maybe more. This is a big difference for what I call "high quality" benchmarks, benchmarks intended to really give a precise performance measurement. For example, Twitter has a whole machine learning system that tunes and re-tunes their JVMs just so they can get a 7% performance improvement. Even the game's own examples that you linked to and seem to say, see, warmup doesn't matter, report differences of 10-200%.
> Which programs?
Now I'm looking at Haskell's fannkuch-redux
> And don't you expect code written to be fast to be different?
Yes, and if the comparison was for the same language (and compiler) then this would be interesting, but what does this tell you when comparing different languages?