I know that isn't how dplyr works, but it feels more Pythonic, and this solution isn't entirely analogous to dplyr anyway.
This approach seems a little complicated, though I'm sure with some use I could learn to enjoy it.
I know that isn't how dplyr works, but it feels more Pythonic, and this solution isn't entirely analogous to dplyr anyway.
This approach seems a little complicated, though I'm sure with some use I could learn to enjoy it.
Right now, a compromise I've been exploring is just attaching siuba's DF functions to a pandas DataFrame, e.g. df.siu_mutate(...). This seems to be what pandas wants people to do [1]! One obstacle here is that the DataFrame has 300+ methods, which can be overwhelming to learners.
I've spent a lot of time wondering whether the piping syntax feels like too much vs chaining. It's still an open question in my mind, so it's really helpful to hear what feels most natural!
https://pandas.pydata.org/pandas-docs/stable/development/ext...
Edit: oh that's pretty much what the linked decorators do.
I think I agree. Pandas has a lot of overhead/baggage that I don't want 90% of the time. Being able to chain _simple_ verbs on a data frame would be great -- something like mtcars.groupby(cyl).summarize(avg_hp = hp.mean())
from siuba import _
from siuba.data import mtcars
# mtcars is a pandas DataFrame
mtcars \
.groupby("cyl") \
.siu_summarize(avg_hp=_.hp.mean())