So then, would Rust be better than a JVM language for a distributed compute framework like Apache Spark?
Based on what others said in this thread, these are the primary arguments for Rust:
1. JVM GC overhead
2. JVM GC pauses
3. JVM memory overhead.
4. Native code (i.e. Rust) has better raw performance than a JVM language
My take on it:
(1) I believe Spark basically wrote its own memory management layer with Unsafe that let's it bypass the GC [0], so for Dataframe/SQL we might be ok here. Hopefully value types are coming to Java/Scala soon.
(2) Majority of Apache Spark use-cases are batch right? In this case who cares about a little stop-the-world pause here and there, as long as we're optimizing the GC for throughput. I recognize that streaming is also a thing, so maybe a non-GC language like Rust is better suited for latency sensitive streaming workloads. Perhaps the Shenandoah GC would be of help here.
(3) What's the memory overhead of a JVM process, 100-200 MB? That doesn't seem too bad to me when clusters these days have terabytes of memory.
(4) I wonder how much of an impact performance improvements from Rust will have over Spark's optimized code generation [1], which basically converts your code into array loops that utilize cache locality, loop unrolling, and simd. I imagine that most of the gains to be had from a Rust rewrite would come from these "bare metal' techniques, so it might the case that Spark already has that going for it...
Having said that, I can't think of any reasons why a compute engine on Rust is a bad idea. Developer productivity and ecosystem perhaps?
[0] https://databricks.com/blog/2015/04/28/project-tungsten-brin...
[1] https://databricks.com/blog/2016/05/23/apache-spark-as-a-com...