dplyr 0.7.0
blog.rstudio.org
blog.rstudio.org
"R makes it easy to create DSLs thanks to three features of the language:
* R code is first-class. That is, R code can be manipulated like any other object (see sym(), lang() and node() for creating such objects). We also call expressions these objects containing R code (see is_expr()).
* Scope is first-class. Scope is the lexical environment that associates values to symbols in expressions. Environments can be created (see env()) and manipulated as regular objects.
* Finally, functions can capture the expressions that were supplied as arguments instead of being passed the value of these expressions (see enquo() and enexpr()). "
dplyr is a testament to how these nonstandard features of R can be combined to produce a language for interactive data analysis whose convenience and clarity cannot be matched by traditional languages.
[1] http://rlang.tidyverse.org/articles/tidy-evaluation.html
The key feature in this case is tidy evaluation, which will be slightly different to use since it favors functional programming paradigms. (pull() is neat for certain use cases as well)
I'm genuinely curious, because I just started using pandas in a new job a few weeks ago, and it seems robust enough so far. I glanced through your link and didn't see any key differences.
My gripes thus far about pandas are that it seems a bit verbose sometimes, eg doing groupby's.
And I haven't quite grokked the indexing. As in - I never use indices, I always reset indices after groupby's to get the groupbys as columns. And I find multi-indices a hassle. Eg, if I group by a column and want to get a sum and count or max and min in one go, without the multiindex that requires using a tuple to access the column afterwards.
Oh, and one actually annoying one - grouping by a column that contains NaN's silently drops those rows. Not the behaviour I'd expect, and requires ensuring all groupbys are preceded by fillna's, which adds to the verbosity.
And just thought of another annoyance. Integer columns are silently turned into floats if a row has any NaN's. So your column of integer ID's turned to floats won't join with another table expecting ints (I've had to workaround and turn to strings)
Besides that, pandas seems pretty reasonable. I've found its use of masks to be pretty powerful, for instance.
df %>% mutate(x = y[2] + z[3]) %>% filter(x > 4)
I haven't found anything like this in pandas; I haven't found any of the pandas dplyr emulators able to do either of these transformations cleanly.
If anybody knows a way, I'd love to hear it. (But afaik dplython etc can't do it.)
df.assign(x = lambda x: x.ix[2, 'y']).query('x > 3')
Get the last n characters from a column string, where n is determined by another column. Then filter to rows where those characters are 'AB'.
( df .assign(last_n_chars=df.apply(lambda x: x['name'][-x['n_chars']:], axis=1)) .query("last_n_chars == 'AB'") )
df %>% mutate(last_n_chars=str_sub(name, start=-n_chars)) %>% filter(last_n_chars == 'AB')
But in pandas they are highly optimized and proof tested, and a breeze to work with once you get a hang of it. They make merging dataframes easy, pivoting easy, data tidying easy, and etc.. However, since you are still learning the api, it can be a pain to use them.
my_column %>%
gsub(" ","",.) %>% # Remove whitespace, then...
gsub("[a-zA-Z]+","",.) %>% # Remove letters, then...
strsplit("-") %>% # Split on dashes, then...
lapply(as.numeric) %>% # Make each vector in the list numeric, then...
lapply(mean) # calculate the mean for each list index
The alternatives here would be gigantic function calls that need to be read from the inside out, or multiple variable assignments.Now days I would do the same with the tidyverse equivalents:
my_column %>%
stringr::str_replace_all(" ","") %>%
stringr::str_replace_all("[a-zA-Z]+","") %>%
stringr::str_split("-") %>%
purrr::map(is.numeric) %>%
purrr::map(mean) df %>%
select(var_1, var_2, var_3) %>%
left_join(df2 %>%
select(another_var1, var2),
by = c('var_2', 'var2'))
This is really over using `%>%`. However, I sometimes take this pattern way too far... :D