How much of R is written in R?
r-bloggers.com
r-bloggers.com
Well written R code tends to be incredibly compact, because the functions available in base R are plentiful and the language is both heavily functional and vector-oriented. The amount of manual memory management and explicit looping required by C easily inflates the lines of code.
A better metric, perhaps, would be to count the number of functions written in each language. Of course, there are issues of style there, but I think that would lead to a more comparable estimate. I don't know of a tool that does that for C - anyone know of one - or should I go the parser route?
But in terms of code quality within R itself, I was surprised to find that a number of its .c files are actually machine-translated Fortran, so I'm guessing the author's statistics are not far off.
I discovered this when I decided to confirm my suspicions that my (the?) most frequently used function, t(), which computes the transpose of the matrix, was implemented about as naively as possible. If R developers were really concerned with speed, this is probably the first place to start optimizing.
However R has the advantage that it will have support for every obscure statistical analysis routine you can ever think of. It also has better support for reading in data from all kinds of sources and handling things like missing and invalid data. So if your goal is to quickly read in a bunch of data sets (that are small enough that performance isn't a critical issue) from arbitrary sources, run a bunch of statistical functions on that data and turn those results into pretty graphs, then R is pretty great.
However, the vast, vast repository of every statistical analysis under the sun - not just 'core R' but every thing that any statistician has hacked up - is unparalleled.
My 'coping with R' strategy is to do all the heavy lifting data manipulation in C/C++/Python, then do one-shot things in R. I just pass csv files around but there are tighter integrations of R and python if you want to look into that.
For example, I used the free R package "earth" to confirm that something like MARS(Multivariate Adaptive Regression Splines) is a good approach to a particular analysis. For my client that initial test justified paying Salford Systems for their great, but expensive, CART/MARS software.
Must be some pretty dense code!
http://www.kdnuggets.com/2011/08/poll-languages-for-data-min...
1. Use the functional style stuff rather than loops ( i.e. especially the apply family of functions)
2. For large data sets avoid the default memory management which involves loading everything into memory all at once. The sqlite dataframe stuff is probably a good default for larger data sets :)
It's still slow though.
It'd be cool to see R evolve more quickly as a language (implementation wise). We're getting hints of it with the new byte code compiler though.
That would be ideal.
I quite like the middle ground of JMP (kind of SAS lite), and its a damned sight cheaper than SPSS too.