However, the benchmarks game (and microbenchmark-based comparisons in general) is hard to take seriously if you are thinking about application performance. Some languages microbenchmark very well, but don't translate that to system performance (C is the poster child of this effect), and some microbenchmark poorly but work very well in practical systems (Go is the most popular language with a big gap here, but some functional language like OCaml or Haskell probably has the biggest gap).
The reasons for these gaps can include things like it being harder to use the optimal data structure for your application (eg C code using red-black trees instead of btrees in 2023) and large code size causing terrible caching behavior (heavily templated C++). I also remember seeing something here about some non-optimal calling convention in the Rust compiler, which would be another thing that shows up in a system that doesn't in a microbenchmark.
Microbenchmarks are not a good replacement for system-level comparisons.