The data.table package (https://github.com/Rdatatable/data.table/wiki) does make progress on some of these - I'd say #1, #3, maybe #7, #8. Dplyr has a query planner too, fwiw.
Point 10: R has lazy evaluation, which means here that a function will not be evaluated when you define it, it will be evaluated when you call it (maybe not quite the same as some other language's lazy eval). I'm not aware of any built in feature for query planning, if you ask for nrow(some_func(myframe)), it will evaluate the some_func(myframe) function and then count up the rows. You could always write your own query planning function I suppose.
Point 11: R has several multicore/cluster libraries, and some are actually decent. If you are like most R users, you use StackOverflow a lot, and you'll end up with one algorithm that uses the snow package and another algorithm that uses multicore, one that uses parallel, and so on. A few very well written packages have hooks that make going to multiple cores easy, but most do not and you typically have to roll your own.
Out of SAS, R, and Python or even C++ and Java. Data frame are native.
So personally, the syntax especially for data frame is beautiful compares to python.
Python don't even have missing value built in or subsetting data frame.
It seems unlikely you meant this as stated. How is it possible to "evaluate a function" when you define it? Certainly you need to give it arguments, and that can only happen when you call it.
I can pass an argument but R won’t try to evaluate unless it needs it. This can be beneficial when only some of a function’s branches need the argument. You can pass a solver for the traveling salesman problem but R won’t waste CPU cycles until it reaches a point where it has to solve the TSP to get an answer. Maybe the first branch of the function is a feasibility check, and the TSP will be skipped for something else.
This is a little harder to explain for data frames, but you can create R functions that act a little like generators in python. This can help with memory management where instead of a gigantic matrix you have a function that generates the part of the matrix that you need.
Welcome to R!
If you look at attempts to do this stuff in python---e.g. patsy, which emulates R's formula DSL, and there's another project that emulates dplyr I don't recall the name of---you see they have to resort to parsing and eval'ing strings instead of working on expressions (language objects that represent ASTs), which is not nearly as nice or safe.
Edit: But just to emphasize your surprise -- yes, you can definitely be surprised by delayed evaluation in many contexts if you're used to more traditional languages.
> y <- 10
> wat <- function(x=10*y) { y = -y; x }
> wat()
[1] -1000
But (1) good library writers don't play these kinds of tricks, so it doesn't come up too often in practice; and (2) when writing/debugging my own code, I've not found it too hard to reason about, anticipate, and avoid these effects. > y <- 10
> less_wat <- function(x=10*y) { force(x); y = -y; x }
> less_wat()
[1] 1000Wow. Didn't knew that.. strings is the one thing I always minimize in my datasets, due to speed and memory considerations..
BTW, I wanted to thank you for your 2012 slides on how you used hashes to group and join data. It led me to learn more about categoricals and I ended up implementing a Factor() object [1] in the other tool I use (Stata) that ended up being a life saver. In fact, once you have a powerful and fast categorical type, with a set of key functions, you can do anything from group the data, to count distinct categories, to run fixed effect regressions in no time.
[1] http://fmwww.bc.edu/repec/scon2017/Baltimore17_Correia.pdf
factor(x = character(), levels, labels = levels, exclude = NA, ordered = is.ordered(x), nmax = NA)
[...]
levels
: an optional vector of the values (as character strings) that x might
have taken. The default is the unique set of values taken by
as.character(x), sorted into increasing order of x. Note that this set
can be specified as smaller than sort(unique(x)).
labels
: either an optional character vector of (unique) labels for the
levels (in the same order as levels after removing those in exclude), or
a character string of length 1.
This way, you can do something like that: > x <- 1:3
> factor(x, levels = 1:2, labels = c("foo", "bar"))
[1] foo bar <NA>
Levels: foo bar
But this actually is: > factor(as.character(x), levels = c("1", "2"), labels = c("foo", "bar"))
[1] foo bar <NA>
Levels: foo barMy impression is that factors in R are borderline-deprecated, especially in the tidyverse, in favor of just using the equivalent non-factor vector.
Could you please provide a real world example where this actually is a problem.