Renjin: a JVM-based interpreter for the R language
renjin.org
renjin.org
The thing that makes this even more awesome is inter language optimization - inline your C code into your R code at runtime, which can result in zero overhead inter language calls.
Looking forward to poking around at their implementation. If it’s a JVM friendly implementation, that could open up some nicer visualization options from other languges like clojure, scala, etc.
Nowadays, my R usage has gone full tidyverse. Anyone know if dplyr and the gang work here? I know dplyr has a lot of C++ code, and don't know how well the "Renjin-specific" Rcpp interop story is.
http://packages.renjin.org/package/org.renjin.cran/tidyverse
For clarification, the JVM JIT-compiles the most frequently used blocks of bytecode to machine code, so the bulk of the heavy lifting is not interpreted.
The benefit of doing it this way is that compilation can be optimised based on the runtime-characteristics of the code, not compiler judgement. Cases of JVM code running faster than equivalent C/C++ code are not uncommon.
Aside from that, access to Java libraries (eg for integration) may be important to some shops.
I've only used R for a short time, so I can't comment on it myself.
I see that the propaganda machine never rests, no matter which language. Either that, or they haven’t heard of Vertica’s R, which is excellent and counters these claims.
The rest with the “fragmentation of R” is a poor marketing tack: people who run R are perfectly capable of thinking with their own brains, thank you very much; Renjin has no place patronizing. We’re not sheep.
I personally think it's great that there are people who are working on making R faster/better/popular, even though I may disagree with the methods they've chosen to achieve those goals.
Oh, GNU R is also garbage collected of course and compared to the JVM's garbage collector it is pretty primitive. This is an area where TIBCO have also improved their own R interpreter called TERR [1].
The language statistics show a good deal of FORTRAN (%24.5) [0], however that is largely skewed by the included LAPACK code [1], which accounts for 221,921 / 259,773 lines of FORTRAN in R.
[0]: https://github.com/wch/r-source [1]: https://github.com/wch/r-source/tree/trunk/src/modules/lapac...
which is where most of your number crunching will be happening..
Looks like Renjin can also be used from GNU R as an R package[2]. I'd assume any installed R package should also work in conjunction with it.
Admittedly, I haven't tried Renjin, but only just found it and thought it was quite interesting.
[1] http://packages.renjin.org/package/org.renjin.cran/Rcpp
[2] http://docs.renjin.org/en/latest/package/index.html#using-re...
"As a service, BeDataDriven provides a repository with all CRAN (the Comprehensive R Archive Network) and BioConductor packages at http://packages.renjin.org. The packages in this repository are built and packaged for use with Renjin. Not all packages can be built for Renjin so please consult the repository to see if your favorite package is available for Renjin."
[1] http://docs.renjin.org/en/latest/interactive/index.html#prer...
While I don't have direct experience of Renjin, or any other alternative R interpreters, I'm inclined to believe Hadley has and if he believes they improve speed, I'll take him at his word.
Makes sense. Why, if you want numerics that were traditionally written in Fortran, of course you have to go to JVM. :)
Does anyone know of a similar tool for c# ?
1. x = rnorm(1000) 2. plot(density(x)) --> not work 3. stem(x) --> not work 4. summary(x) --> works
For R data handling, I always use data.table for its efficiency and power.
1. library(data.table) 2. x = data.table(x=rnorm(1000)) --> not work > x = data.table(x=rnorm(1000)) ERROR: Exception calling Calloccolwrapper : Unimplemented GNU R API function 'DUPLICATE_ATTRIB'
However, I would agree with you that at this point in time it's fair to assume that R is no longer a language in and of itself - it's more akin to a collection of APIs around years of optimized C++ code, somewhat separated into a few universes (bioinformatics, time series analysis with xtz, data.table aficionados and tidyverse acolytes).
And thus making a faster R alternative should focus on 100% package compatibility along with speeding up native R code. Renjin, pqR [1], et. al. all seem to be compromises that work for some workflows but not all of them.
The project that I'm excited for is fastR [2] - bringing R, C, C++ and FORTRAN code together into one Truffle/Graal environment would be absolutely huge for this ecosystem.
However, I've tested fastR on my employer's [3] workflows and it does not work yet.
[1] http://www.pqr-project.org/ [2] https://github.com/graalvm/fastr [3] http://syberia.io/