> datafusion
> blazingly fast
I’m going to need to see a citation for that. Last I checked, it was being beaten by Apache Spark in non-memory constrained scenarios [0]. This may be “blazingly fast” compared to Pandas or something, but it’s still leaving a TON of room on the table performance-wise. There’s a reason why Databricks found it necessary to redirect their Spark backend to a custom native query engine [1].
[0] https://andygrove.io/2019/04/datafusion-0.13.0-benchmarks/
[1] https://cs.stanford.edu/~matei/papers/2022/sigmod_photon.pdf