The submitted code snippets for Java do not need OSR, though... yet they can be improved further, e.g. should drop the use of String entirely (which would not feel very Java). The other attempt to convert int -> String (byte) uses the naive way to divide by 10 on each iteration, Java.'s Integer.toString does it way better.
Edit: on a 2nd thought, having a dedicated direct buffer [same allocation, different slices] per all the 8 out of 15 ares for numbers and NOT converting int->String each operation but adding 15 would be a pretty boon as most of the time the change would be only the last 2 bytes, and there won't be any 'div' to be had. Div is generally slow (compared to L1/L2 cache misses, and L3 hit), there is not algorithm to parallelize it, and there is one (few) unit that can perform div, unlike 'add')
In theory, yes, but in practice I've never seen it happen. Best I ever saw was matching C speed at toy benchmarks. Even in this benchmark here, Java is decent, but does not beat even the naive implementations in C/Rust.
Also, AOT compilers can do PGO as well, so they can use the same techniques. But they also have way more time and resources, so they can do things like whole program optimization, which is something JITs cannot do because they have much smaller computation and memory budget.
It happened already many times to me that the first naive version of a C/C++/Rust program/function I wrote was already faster than a carefully tuned Java equivalent. AOT compilers for "fast languages" got really good these days. The design of language also influences how well it can be optimized by the compiler. E.g. it might look impressive JVM can devirtualize dynamic calls at runtime, but C++/Rust often don't have to do this at all as programs in those languages tend to have very few virtual calls if any at all.
Java doesn't beat C in this benchmark but beats Rust with ease.
Also, technically there is nothing C can do that Rust can't.