Blaze: Fast query execution engine for Apache Spark
github.com
github.com
RDDs will be worse, so it shouldn't matter. No vectorization, no column processing, lots of serialization and de-serialization. They're basically always slower than DataFrames barring some strange use case.
[1] https://docs.databricks.com/en/clusters/photon.html [2] https://dl.acm.org/doi/10.1145/3514221.3526054
One of the issues was that we started experimenting with Delta Tables and EMR was horrible in leveraging that.