Discovering copy-on-write in R
franklin.dyer.me
franklin.dyer.me
Here's some completely not-the-point of the article code review since I can't help myself. If you can set up earlier steps give you a named list for `col_grouping`, and use `lapply`, the code is a little more concise:
efficient_flow_agg <- function(dat, col_grouping, gpcol_name="GroupMembership") {
make_postproc <- function(gp, groups) {
gp$preproc(dat[gp$which_cols]) |>
lapply(collapse::BY, groups, gp$aggfun) |>
gp$postproc()
}
col_grouping |>
lapply(make_postproc, groups = dat[[gpcol_name]]) |>
as.data.frame()
}
* I had previously written here that `tapply` is probably faster, but apparently `tapply` does exactly `unlist(lapply(split(x, g), f)))` anyway? wtf R. Strange there's not something like `collapse::BY` in base R.Good advice on `col_grouping` as well, accessing those components of an aggregation rule by index rather than by name is a bad code smell and decreases readability for sure.
Pandas has a global option to turn on copy-on-write.
https://pandas.pydata.org/docs/dev/user_guide/copy_on_write....
To be default mode in Pandas 3, but seeing as how long it took them to pull the trigger on Pandas 2, that could be a while.
A lot to be said for not defaulting to data frames, in both r and python. Or, if you must, using something like r data.table or python’s polars if you don’t think in other data structures easily or just want convenience.
Thanks for mentioning polars, I hadn't heard of it before but it looks neat.
I would even add especially in Python. The main issue I have found is that pandas heavy code is just not as easy to integrate into other Python tools/features/abstractions as code using mostly numpy, dictionaries and various comprehensions to do the vast majority of your work.
As a heavy pandas user for several years, I decided about a year ago to not import pandas by default and instead treat most data problems like regular python problems. I've been genuinely surprised as how much easier it is to create useful abstractions with the code I've been writing, and also how much easier it's been to onboard non-DS devs into the code base.
There are a few obvious cases when Pandas is very helpful, and I'll pull it out in those places, but I've been able to do a tremendous amount of data work in the last year and used very little pandas. The result is that I have an actual codebase to work with now rather than a billion broken notebooks.
(Note that in general, I'm the biggest pandas hater I know)
This is the biggest part. Giving yourself permission to make real abstractions, rather than forcing yourself to go directly from data-on-disk to pandas (or whatever) makes it that much easier to test, repeat, modify, and extend whatever analysis you're working on.
https://github.com/smithzvk/modf
With modf, you use the existing place syntax to refer to part of an object. It looks like you're mutating that object, but in fact it will return a clone of the entire containing object, with the modification, while the original remains untouched.
x<- c(1,2,3)
y <- c(4,5,6)
z<- c(1+4, (1*2+4*5)/(1+4), sqrt((1*(4+9) + 4*(25+36))/(1+4) - ((1*2+4*5)/(1+4))^2))
But your z[2] is just an elaborate weighted mean, I would use the following built-in function - weighted.mean(c(2,5),c(1/(1+4),4/(1+4)))
Similarly, your z[3] is just the weighted standard deviation, available in library modi. I was wondering, isn't it better to store the data in some vector x & the weights in a different vector y, and compute the weighted mean & weighted variance in a straightforward fashion like above, or am I missing something. Thanks.