FastR: An implementation of the R language in Java [pdf]
oracle.com
oracle.com
There are many other R implementation efforts going on right now -- Radford Neal lists a few (as well as his own) here: http://radfordneal.wordpress.com/2013/07/24/deferred-evaluat...
The presentation focuses on the R programming language, which they nicely show has all sorts of misfeatures that impede rapid execution. If you're going to not try to have compatibility with R and CRAN, you might as well start from scratch with design and performance in mind, as in Julia: http://julialang.org/
We've been working hard and now systematically to get Renjin (also R on the JVM) to run CRAN packages: http://packages.renjin.org. Renjin also compiles C and Fortran code to JVM bytecode, though there is still some work to do there as well.
Regarding the hand-tuned matrix math libraries, there's nothing to stop you from using them with Renjin - you can drop in MKL or Atlas as desired, or fall back to pure-Java versions in a pinch.
As for Java, Oracle together with AMD, are in the process of making the GPU trasparent to Java developers as part of the Sumatra project.
So this is one are where R could benefit of running on Oracle's JVM. It remains to be seen if other Java vendors would adopt such feature.
As of 2002 (JDK 1.4) Java has excellent integration with native memory (you can freely pass pointers from C/FORTRAN to Java and vice versa[1]).
There are numerous Java math libraries that use BLAS/LAPACK already[2]. In fact, AFAIK, most Java matrix math libraries use FORTRAN code (at least as an option).
[1]: Java side: http://docs.oracle.com/javase/7/docs/api/java/nio/ByteBuffer... C side: http://docs.oracle.com/javase/7/docs/technotes/guides/jni/sp...
[2]: For example, https://github.com/fommil/matrix-toolkits-java, http://mikiobraun.github.io/jblas/
(Personally, when programming Java I find it more convenient to use primitive arrays as opposed to matrix libraries, but that might be dependent on the operations I tend to do: lots of increment/decrements and only occasional linear algebra. I guess this isn't exactly relevant to the R replacement question.)
Direct byte buffers are very common in high-performance Java code. Reading/writing from/to those buffers can be made just as fast as plain Java arrays.
I had used memory mapped buffers (which should be equivalent in performance), and there was no way to make the JIT inline access to these arrays. It was all calls (indirect, not properly branch predicted at that). Equivalent code to C++, running 10 times slower, with no way to speed it up.
(And the reason I was using memory maps, if you insist -is a 2GB read-only dataset used by multiple processes at the same time - I went to C++ eventually, because there was no way to get reasonable performance from Java, either memory use or speed. This is circa 2010)
For example, the Java Chronicle library[1] uses these techniques, as well as memory mapped files, to implement fast persistent message queues.
The sun.misc.Unsafe class is used extensively by JDK classes, and is meant for internal use. It provides intrinsics that are translated to a single machine instruction. Normally, you don't use the class directly. For example, you use, say, AtomicInt for CAS operations (which, internally uses s.m.Unsaafe) or the ByteBuffer class (which internally uses s.m.Unsafe for direct pointer access). The JDK classes add all sorts of protection (like range checks) around s.m.Unsafe, but if you know what you're doing, using s.m.Unsafe directly and eschewing some of those protections (usually adding ones more pertinent to your domain), you get some performance gains which may be significant depending on your use case.
2. The math in FastR, if I understand the presentation correctly, is performed by FORTRAN libraries anyway. Using battle tested FORTRAN libraries for matrix computations is common practice in C, Java, Julia, Matlab and most other environments. They basically all share the same underlying matrix math code.
Although it is standard in high-level systems to call out to a BLAS library [1], for some inexplicable reason it seems that both R and NumPy use the reference BLAS by default, which is quite slow – around 4x slower than better BLASes. Matlab ships with Intel's proprietary MKL, which includes a very fast BLAS implementation, while Julia ships with OpenBLAS, which is a similarly fast open source BLAS implementation derived from the legendary GotoBLAS [2]. Since all BLAS implementations share a common Fortran ABI, it's easy to swap them out, but it's not quite true that all of these systems are using the exact same Fortran code.
[1] https://en.wikipedia.org/wiki/Basic_Linear_Algebra_Subprogra...
GNU R, for example, is implemented in C, and the implementations of R's basic arithmetic functions are actually quite complicated because they take care of so many of the edge cases cited in the cited post. For example, the round() function casts its argument first to a 64-bit before calling the C library's rint() function to preserve precision. [1]
http://download.oracle.com/technetwork/java/javase/community...
http://openjdk.java.net/projects/graal/
https://wiki.openjdk.java.net/display/Graal/Publications+and...
If you don't want to use R, Pandas in Python provides very powerful data frames (which are likely faster for many cases). However, it depends on NumPy, matplotlib, and a few other libraries, which probably total more than 60 MB.
See http://r.789695.n4.nabble.com/R-in-the-browser-td4667985.htm...
But the JS blob ends up being like ~15mb!