For instance, you can’t use String on hot paths. Also, NPEs can happen almost anywhere, and the compiler doesn’t enforce thread safely (by which I mean “statically checked to be data race free”)
Also, the 10-50% slower rule of thumb only applies to throughput. Tail latency is generally much worse than that.
There is nothing comparable in C as performant as all of the streaming tooling built up around Kafka for example.
So your can benchmark percentages all you want, but there are a lot of things you just can't do in C that you can on the JVM (without 100 developers and three years)
This is an ill-chosen example.
librdkafka is as or more performant than the Java stuff.
As a server, Redpanda blows Kafka out of the water (but this is rarely any bottleneck with brokers).
I was referring more to the general ecosystem of Flink/Spark/KStreams/KTable etc.
Then it's still a bit of an untruth since all of those use RocksDB for their high-performance storage layer, which isn't in Java. You could maybe (maybe!) argue for some ergonomics (though especially hard to argue for Spark!) but performance still goes plainly to C (/ C++).
Not to mention you still pay a real performance penalty for being unable to make direct use of RocksDB features, because Kafka is shoving a million "abstraction" and cache layers in there and you just want your your damned merge operator to apply.
I don't agree Kafka off the JVM is subpar. Kafka Streams has a lot of sharp edges from "abstractions" that aren't, Spark is a godawful developer experience and tuning it for performance is black magic for 99% of data engineers, and Flink keeps dithering about how consistent/stable baseline features like queries will be. You can stand up very good apps for a huge set of common cases just as quickly in Go or C++, and not have to deal with Kafka Streams's funny opinions about threads/tasks/consumers.
Now if it means I can run fewer smaller brokers that's awesome.