[0] -- https://www.pola.rs/
[0] -- https://www.pola.rs/
If you use an r5d.24xlarge-like[1] instance, you can skip spark/dask for most workflows as 768 GB is plenty enough. On top of that, polars will efficiently use the 96 available cores when you are computing your join, groupby, etc.
Also polars is getting more and more popular[2]
[0] -- https://h2oai.github.io/db-benchmark/ [1] -- https://aws.amazon.com/fr/ec2/instance-types/c6a/ [2] -- https://star-history.com/#pola-rs/polars&Date
Polars runs orders of magnitudes faster than pandas does. Which means EDA can be completed quicker.
A statistician/data scientist wrangling data and making plots would not have cared whether loading a CSV file takes one second or one microsecond, because they may only do it a handful of times for a project.
A data engineer has different requirements and expectations. They may need to implement an operational component that process CSV files repeatedly for billions of time a day.
If your use case is the latter, then pandas is probably not for you.
Polars excels when pandas operations take 30 seconds or a minute to complete. Bringing that time down to the second or ms mark is really amazing.
I've definitely procrastinated doing some analyses or turning prototypes into dashboards because of the potential for small slowdowns to turn into big slowdowns, so it's nice to have other options available. I'm very interested in Dask but have also been apprehensive about doing something stupid and incurring a huge bill by failing to think through my problem sufficiently.
[0] https://modin.readthedocs.io/en/stable/
I don’t have much experience with this though.
IME xarray and pandas have to be used together. Neither can do what the other does. (Well, some people try to use pandas for stuff that should be numpy or xarray tasks.)
I still use pandas for tabular data, but anytime I have to deal with ND data, xarray is a lifesaver. No more dealing with raw, unlabeled 5-D numpy arrays, trying to remember which dimension is what.