How?
How?
The submitted code snippets for Java do not need OSR, though... yet they can be improved further, e.g. should drop the use of String entirely (which would not feel very Java). The other attempt to convert int -> String (byte) uses the naive way to divide by 10 on each iteration, Java.'s Integer.toString does it way better.
Edit: on a 2nd thought, having a dedicated direct buffer [same allocation, different slices] per all the 8 out of 15 ares for numbers and NOT converting int->String each operation but adding 15 would be a pretty boon as most of the time the change would be only the last 2 bytes, and there won't be any 'div' to be had. Div is generally slow (compared to L1/L2 cache misses, and L3 hit), there is not algorithm to parallelize it, and there is one (few) unit that can perform div, unlike 'add')
In theory, yes, but in practice I've never seen it happen. Best I ever saw was matching C speed at toy benchmarks. Even in this benchmark here, Java is decent, but does not beat even the naive implementations in C/Rust.
Also, AOT compilers can do PGO as well, so they can use the same techniques. But they also have way more time and resources, so they can do things like whole program optimization, which is something JITs cannot do because they have much smaller computation and memory budget.
It happened already many times to me that the first naive version of a C/C++/Rust program/function I wrote was already faster than a carefully tuned Java equivalent. AOT compilers for "fast languages" got really good these days. The design of language also influences how well it can be optimized by the compiler. E.g. it might look impressive JVM can devirtualize dynamic calls at runtime, but C++/Rust often don't have to do this at all as programs in those languages tend to have very few virtual calls if any at all.
Java doesn't beat C in this benchmark but beats Rust with ease.
Also, technically there is nothing C can do that Rust can't.
"GC pauses" are greatly exaggerated in terms of impact and frankly for the vast majority of uses cases GC simply doesn't become an issue.
The JVM is really really good because at one point or another they had basically every luminary in the field working on it.
And it's not even that slow anymore compared to, say, starting the JVM in 2010.
Some numbers for those curious...
On modern hardware launching the JVM to run a program immediately exiting takes less than 100 ms (so does starting Emacs complete with its GUI and running some elisp code exiting Emacs: 80 ms on my Ryzen 7000 series including reading from the M.2 NVMe PCIe 4.0 x4 SSD).
It's once you start loading lots of classes that JVM startup time can be slow.
One example would be a Clojure program doing nothing besides exiting: thousands of Java classes being used and you get into 1.2 seconds territory to do nothing. 12x slower than a Java program doing nothing.
As a sidenote for both Java and Clojure there are now ways to reduce startup time, like using GraalVM (which, for example, Babashka, a natively compiled Clojure interpreter, making Clojure startup so fast it can be used for scripts).
> The JVM is really really good because at one point or another they had basically every luminary in the field working on it.
I agree. The JVM is an impressive piece of machinery and it'll even give you, say, an AIOOBE (ArrayIndexOutOfBoungsException) instead of an exploit if you fuck up.
I'm with you all the way and am a daily emacs user but I'm not sure I'd point to emacs as a good performance comparison, haha.
Emacs has always been a bit of a dog in my experience.
It's similar to SWAP in Linux. It was implemented somewhat meh in older kernels, and usually if swap got hit, the system died anyhow. So there was no real difference between random processes being OOM-killed or the system grinding to a halt swapping. Modern kernels in the 4+ line have received quite a bit of work on the swap handling and swap is used a lot and very cleverly to eek out just a bit more available memory more quickly.
Old habits die hard though and it takes time for old knowledge to change.
Low pause GCs (ZGC, Shenadoah) for Java are not generational yet.
Also, even if they were, a very high temporary object allocation rate increases young gen GC frequency and thus increases the number of objects pushed to old gen eventually.
And it is not like those young gen GCs are free either. They burn quite some CPU time and they cause micro-pauses - the GC has to scan parts of the heap to learn which objects are reachable and then it has to copy the survivors. In practice, high allocation rate requires a decent amount of overhead RAM to make that process efficient. It doesn't matter if pauses are only 10 ms short if you do 50 of them per second. ;)
In non-GCed languages those trivial short-term objects are typically allocated on stack and their allocation/deallocation is trivial and doesn't pause at all.
They are. GenZGC is merged into mainline and got into the latest JDK, GenShenandoah didn't manage but will be in next release.
>The problem of GC pauses is solved only if you can afford to waste 5x-20x more memory than the app is really using
I don't think that is true. ZGC uses multi-mapping in order to dereference its colored pointers, this causes some tools to report excessive memory usage but nothing actually hits RAM.
Hmm, for generation GCs, it should use Card marking, the pointers in the tenured gen should have means to be trivially determined where they belong to. With 64bit pointers, there is space for quite a lot of metadata, incl. the Class (or most commonly allocated/used classes).
I feel like I've been hearing this same line for 20 years now
It works by allocating all or most needed memory at program start, instead of asking the operating system for it every time. But, as soon as you don't use heap memory, and use the stack, C++ is again much, much faster than Java.
It all depends on memory management.
To your second point about raw speed, one of the extraordinary things we found in long-running processes performing computation using real-time marketdata at Goldman was that the HotSpot JVM was able to optimize java programs through the day due to their usage, so if you started them each day they would actually end up faster than the C++ versions at the end of the day even though they would start off slower. That's not due to memory allocation it's due to things like inlining of functions.
The implication of that is that if you very carefully inlined all the functions appropriately in the C++ version based on profiling actual usage you would be able to achieve the same result, but for the JVM it just happens automatically without you doing anything.
[1] Search https://en.cppreference.com/w/cpp/language/new for "Placement New"
Can you point us to some benchmarks?
I don't doubt it is faster, but I doubt the "ridiculously" part. Last time I measured it was only a tad faster than jemalloc (~20%) if you allowed the benchmark to run long enough for GC to start cleaning up. Unfortunately I haven't saved it (it was a very informal benchmark). I'm curios to see that esp. on modern low-pause GCs.
Anyway, don't use new in C++ is my opinion.
You could of course FFI into e.g. C for those parts, but that is usually harder to maintain than a few well optimized java classes.
It's a quite common myth developers believe about performance. Hotspots do happen sometimes, but once they have been optimized you quicky end up with a flat profile and an "everything is slow" problem. And in some types of apps, the majority of code is performance critical.
If the majority of code is performance critical, the tradeoffs are of course different.
Hopefully Foreign Function & Memory API [0] makes FFI so much easier that we get to drop down to C without much fuss.