Pandas has also moved to Apache Arrow as a backend [1], so it’s likely performance will be similar when comparing recent versions. But it’s great to have some friendly competition.
[1] https://datapythonista.me/blog/pandas-20-and-the-arrow-revol...
[1] https://datapythonista.me/blog/pandas-20-and-the-arrow-revol...
Spark tended to do this and it makes complete sense for distributed setups, but apparently is still faster locally.
Only for SQL databases, so not really. Source: have been running dplyr since 2011.
https://arrow.apache.org/cookbook/r/manipulating-data---tabl...