Agreed RE GroupBy being challenging, especially compared to dplyr.
As I've worked on a port of dplyr to python over the past year, though, I've realized the dtype issue (like you said), indexes, and GroupBy being difficult are likely connected. Basically,
* dplyr can chop up a dataframe into 50,000 groups and apply arbitrary functions to it--no problem.
* custom pandas grouped applies are very slow
There are basically three reasons for slow pandas apply methods... 1. creating an index for each subgroup is slow (will not be a RangeIndex)
2. initializing a series for each group is slow (mostly due to type inference being re-run; could be avoided)
3. AFAIK more type inference is run when concatenating results
This leads to a world where grouped calculations can't be run using arbitrary expressions (e.g. lambdas), but have to go through specific SeriesGroupBy methods.I wrote a bit on how I tried to work around that, to enable fast dplyr-like syntax over grouped data in python. Would definitely be interested in your take! There are other libraries, like ibis that do a good job with it, too!
https://siuba.readthedocs.io/en/latest/developer/pandas-grou...