The Grammar of Data Science
technology.stitchfix.com
technology.stitchfix.com
http://nbviewer.ipython.org/github/davidrpugh/cookbook-code/...
BTW: For pandas-dplyr dictionary: http://nbviewer.ipython.org/gist/TomAugspurger/6e052140eaa5f...
This is my biggest beef with R. It is constantly changing the dimensions and types of your data without telling you. Want to grab some subset of the rows of a matrix? Better add some extra post-processing in case there's only one row that satisfies your query, or else R will change its type!
The solution is not to make the programmer memorize obscure edge cases.
> x <- data.frame(foo=1:5, bar=1:5, baz=1:5)
> dim(x[,'foo'])
NULL
> dim(x[,c('foo','bar')])
[1] 5 2
> dim(x[,'foo',drop=FALSE])
[1] 5 1
compared to > x <- dplyr::data_frame(foo=1:5, bar=1:5, baz=1:5)
> dim(x[,'foo'])
[1] 5 1
Although I think these are more reasonable (I've got multiple commits at work with messages bemoaning drop=FALSE), this can ironically also mess you up if you got used to the old defaults :)I wish Hadley had developed these tools for some other language, such as Python, or in a language-agnostic way. Hopefully that is the direction things will go in the future.
Thinks like function parameters being promises make it far easier to deal with functions like optimizers where there really are 10+ tuning parameters or things you may want to tweak. Iterative languages are far easier to understand for people who don't want to be programmers.
You cannot develop plyr or ggplot in a language agnostic way, because they need the purpose built syntax R has. Contrast to eg the fight in python to get an infix matrix multiplication operator.
[1]: Here is a simple example of that for +, note the very strange overloading of quotes.
R version 2.15.2
> "+"(7,5)
[1] 12math:
S = ( H β − r )^T * ( H V H^T )^{-1} * ( H β − r )
python, ugly mess S = (H.dot(beta) - r).T.dot(inv(H.dot(V).dot(H.T))).dot(H.dot(beta) - r)
python, better: (although @ is an ugly matrix operator) S = (H @ beta - r).T @ inv(H @ V @ H.T) @ (H @ beta - r)
the latter is an order of magnitude easier to understand, and looks just like the math. Having one layer of indirection: math to code, is far better than two: math to code to obfuscated code because you won't create infix operators.Edit: examples stolen from the matrix operator pep
Moreover, if it's hidden in a library then why does the user care if it's ugly? It's the library designer's job to test it and make sure it's right.
That said, the other big area of complaint in R is the type system. We are too often having to coerce types, but I'm not exactly sure of the solution for that.
The %>% operator alone (which to be fair was originally from magrittr) is a great help. Not sure if this is my personal biases, but I always find it easier to read calls chained postfix-style.
I agree with you about dplyr + ggplot and was pretty much gobsmacked at the obviousness of "this is the way it should be taught" and am glad I'm in the position to help review such a text!
I wonder if eventually this is the future of standard R.