Why Pandas feels clunky when coming from R
sumsar.net
sumsar.net
library(data.table)
dt <- fread("purchases.csv")
dt[, .(total = sum(amount - discount)), by = "country"]
and much faster to run q = (
pl.scan_csv("docs/data/iris.csv")
.filter(pl.col("sepal_length") > 5)
.group_by("species")
.agg(pl.all().sum())
)
df = q.collect()I often dont want to manually write out column names, but programmatically specify them, and similar for a lot of other of these examples. I don’t want to manually configure them.
I haven’t seen examples of that higher level programming in these various R python comparisons. It’s always tediously manual examples.
The examples usually feel like manual analyst query type tasks. Even the tone strongly reinforces that with text like “oh and Maria asked me to xyz”