R for the JVM: 60% complete, help wanted!
code.google.com
code.google.com
There are currently > 3500 packages in CRAN, many of which can be found nowhere else. Being able to access java libraries is great, but I an access java libraries from all sorts of different languages. Until it is seamless to install and use these CRAN packages, I'm sticking with the standard interpreter.
But others -- like the survey package -- are pure R and run well on renjin. (Just do library(survey, lib.loc='/path/to/R/library') )
The central goal at this point is to support embedding packages like 'survey' in web apps or larger java apps. If you're looking for a seamless user experience for ad-hoc analysis, we're still quite a ways off!
There are a few features of the R-language that make direction translation into byte code daunting for a muggle like myself:
1. Computing-on-the-language: R code expects to be able to access and modify the AST and frame of itself, its caller, and other closures. 2. Impure call-by-need argument-passing semantics.
The compiler that's in the trunk is experimental but evolving fast, I think the next steps will probably to start compiling simple but performance-critical basic blocks to byte code at runtime, and then slowly expand the scope of language that can handled from there... (Expert advice welcome!!)
I'd hardly count myself as an expert, but I think the best win we've had is in thinking carefully about callsite caching strategies and having a eureka moment about just how insanely powerful MethodHandles.exactInvoker can be.
Does the JVM environment have any tools that would facilitate building a better R debugger? That's one area of the R ecosystem that could use a serious upgrade IMO.
One of the projects for 2012 is to integrate Renjin into StatET, including a line-by-line debugger. Any takers?
http://heuristically.wordpress.com/2010/01/04/r-memory-usage...
As for memory usage, I believe object.size() will double-count your input data when it is referenced by the resulting model objects. Better to check memory.profile()
At present, Renjin benefits from the JVM's state-of-the art garbage collection, so you may see some improvements even at present, but I expect the big difference will be once we roll out non-memory-backed stores for R Vectors. Then your input data could be stored in a database and only partially loaded into memory as needed.