Maybe Array.sort() isn’t that frequently used, as data sorting is often done by the database?
Most Java application code doesn't, but it typically uses libraries that do. A sorted array is a priority queue, a binary search tree, etc.
If you have an application that is truly bottlenecked the performance of number-sorting in any measurable way, then you most probably didn't write it in Java. It's not really a number crunching language for a variety of reasons.
You end up using libraries like fastutil, which is "generic code" templated by C preprocessor macros,
This is planned as phase two of Project Valhalla [0]:
The second phase will focus on generics, extending the generic type system
to support instantiation with inline classes (which will include primitives),
and extending the JVM to support specialized layouts.
[0] https://cr.openjdk.org/~briangoetz/valhalla/sov/01-backgroun...(Scala has specialization and value classes and opaque types so that covers a fairly big range but they make interop with Java tricky.)
This includes JEP 401 "Flattened Heap Layouts for Value Objects" [1, JEP 402 "Enhanced Primitive Boxing" [2], and I think also "Value Objects" [3] and "Universal Generics" [4].
It is a huge task with many dependencies and requires careful design. It feels like it might finally make it in the next long term release.
[0] https://openjdk.org/jeps/12
[1] https://openjdk.org/jeps/401
[2] https://openjdk.org/jeps/402
But then when you look at the actual code of JIT compilers, such JIT-only optimizations seem extremely rare. The JVM, surely one of the platforms that have had more than enough dollars and top-level CS talent thrown at it, even after ~20 years it still apparently lacks many optimizations that you'd expect to be in there if you've only read the introductory texts about it.
The linked MR about AVX optimizations for sorting is one such example IMO. AVX512 was first announced 10 years ago, and (not being deeply into JVM development myself) I might have assumed its use would be more prevalent in the JIT output than actually seems to be the case.
Nonetheless, you might find the Graal compiler doing a better job at autovectorization than C2.
OTOH, another way in which JITs "should" generate better-performing code is by tailoring their output to the platform on which the program is currently being run. With AVX being quite prevalent on the server-grade CPUs on which many big JVM programs run on, I don't think it would be unreasonable to expect the JVM to have more support for AVX512 in its code generator than it apparently does. Is low hanging fruit like 10x speed improvements in sorting something you'd expect out of a very mature platform like the JVM?
I don't mean to harp on the JVM devs here, JIT development is a Very Hard Problem. It's just that I can understand why GGP is disappointed, JITs in general don't seem to quite deliver on the excitement they generated when they were new.
But it doesn't unless your "almost" is very generous. Java is pretty consistently 2-10x slower than the major performance-focused AOT offerings (C, C++, Rust)
Now maybe you call 2x "almost", but let's phrase it in terms of CPU performance over time. That's equivalent to 10 years of CPU hardware advancements.
To me that's a lot of overhead. Depending on who is paying for the CPU time vs. the developer time it's regularly a cost worth paying, but at the same time don't pretend it's "almost native speed", either. It is a cost and a rather significant one at that. Just, so are engineers. They also aren't cheap.
Also, which 10 years of CPU advancement do you mean? It is definitely not a linear graph, we have reached an almost plateau on single-core performance.
And even code that does have an overhead is not as simple to judge. Could you write that same code in a lower level language that it will still remain correct and safe? Is the algorithm actually expressible in Rust’s much more restrictive style (to stay safe)? If you do locks and ref counting everywhere, will that code actually be still faster? For example, a compiler might very well be faster in Java/haskell/another managed language.
Because it's not really interesting to debate. Nobody writes Java code like that, and even when they do there's still overhead to it. The exact amount of overhead is kinda irrelevant since the language very obviously doesn't want you to write code like that.
> Could you write that same code in a lower level language that it will still remain correct and safe?
Rust is safer than Java, so yes :)
But you're drifting into the productivity argument anyway, which I already pointed out is a reasonable reason to pay runtime overhead to get.
Rust has data-race freedom, while java has “safe” data races (it is tear-free, so even in case of a data race you won’t corrupt the memory, which is not true of rust with any number of unsafe parts). So I don’t really buy the argument that Rust would be safer.
Likewise if you code in C, C++ or Rust with allocations everywhere, bad algorithms or data structures, being AOT won't help.
GraalVM does it much better than OpenJDK, then there are OpenJ9, PTC, Aicas, Azul, ART (yeah not really, but close enough), microEJ, and a couple of nameless others from Ricoh, Xerox, Cisco, Gemalto,... on their devices.
Even if we stick to OpenJDK, the distributions based on it aren't all the same, for example Microsoft's fork has additional JIT improvements (-XX:+ReduceAllocationMerges).
JITs do regularly have a lot more specialization optimizations, though, but is that really because it's a JIT instead of an AOT or is it more because JIT'd languages just often tend to also be more dynamic ones as well?
Nonsense. What evidence do you have for this claim?
> JITs often have preset heuristics for figuring out where to split the code that's practical to implement rather than being the most performant possible. After all, the code needs to hot-swapped in without much disruption.
You seem to be talking about JITs without on-stack replacement. So not state of the art high performance JITs.