Faster R with FastR
medium.com
medium.com
Well, I can't really use it in my day to day work, since that almost always involves cleaning and munging via one of those two packages. And it's not like ggplot2 is where my R code is most delayed, usually I'm working on aggregate data or perhaps a very much smaller analytical dataset which requires much less speed for plotting. My hang-ups are in initial munging phases where the data is still very large, which often calls for data.table over dplyr due to the latter's much slower performance.
In the event that the data doesn't fit into memory, it's better to preprocess w/ SQL at the data-store level. There hasn't been a case where I'd need to feed massive amounts of data into a ggplot2 visualization unaggregated.
I spend hours cleaning up data and only have to run the code once (I normally save the output to a feather and then work with a separate file from there).
I still believe that the 'tidyverse' is hands down the best thing that has happened to R and is the whole reason why R has grown so fast.
Big barriers to adoption here: not a truly drop-in replacement, R people have an aversion to Java (we've all spent hours debugging rJava; luckily most of those packages have been rewritten in C++ now), and nobody likes Oracle.
I think the best-case scenario here is that progress on FastR pushes the R-Core team to improve GNU-R.
This claim is made about a lot of things, Ruby, Python etc. I think the important point is it that there is no trade going on. It just that these things are all slower / less efficient than they need to be.
If I was Lord Of Computing I wouldn't let languages out of beta until they had a high quality compiler or JIT. Turns out I am not though.
Julia gets very good performance without the massive manpower that has gone into Javascript VM's.
I'm not even sure GNU-R is the most important comparison (although it is an important comparison). How does it compare to R with Intel MKL? How does it compare to other (faster) languages?
We didn't want to include comparison to R-3.5.X, because FastR itself is based on the base library of 3.4.0, but the results for GNU-R 3.5.1 almost the same as for R-3.4.0.
AFAIK ALTREP is not used that much yet inside GNU-R itself. They can now do efficient integer sequences (i.e. 1:1000 does not allocate 1000 integers unless necessary), which would save a little bit of memory in this example, but that's about it. FastR also plans to implement the ALTREP interface for packages. Internally, we've been already using things like compact sequences.
https://github.com/QuantStack/xtensor-r https://github.com/QuantStack/xtensor
Disclaimer: I'm one of the core devs.
We've mainly gone through RCpp for the R language, and that has been working great. I don't know about changes in R 3.5 or ALTREP. Is there something we should know/change for it?
https://www.youtube.com/watch?v=HStF1RJOyxI
It's a little disappointing, because the conclusion is that R will probably never "run fast", but very interesting nonetheless.
Great work, thank you for sharing.
data.table is a different beast and we will probably provide and maintain patched version for FastR. They do things like casting data of internal R structure to byte array and then memcopy it to another R structure. This is very tricky to emulate if your data structures actually live on Java side and you're handing out only some handles to the native code.
Context ctx = Context.newBuilder("R").allowAllAccess(true).build();
Value rFunction = context.eval("R",
"function(table) { " +
" table <- as.data.frame(table);" +
" cat('The whole data frame printed in R:\n');" +
" print(table);" +
" cat('---------\n\n');" +
" cat('Filter out users with ID>2:\n');" +
" print(table[table$id > 2,]);" +
"}");
User[] data = getUsers();
rFunction.execute(new UsersTable(data));
The example above combined with "JEP 326: Raw String Literals" and an IDE that understands Java with embedded R code would be cool to play with.Hm... looks like the issue may have been fixed. I'll have to try again.
Thanks!
That, or something like Cython where, instead of writing inline C++, you translate a restricted subset of R to C, which is then compiled.
E.g. can you quickly spin up a REST-like HTTP interface for your goods?
There's also opencpu (https://github.com/opencpu/opencpu), though the pros/cons of one vs the other has never been clear to me.
Unfortunately, not even close to flask. Plumber doesn't even handle concurrency...
https://www.rplumber.io/docs/runtime.html#performance-reques...
I'm curious to see how https://github.com/thomasp85/fiery performs and if anyone has used that. May be higher performance than plumber (re: concurrency) because I get the sense from the docs that it's closer to libuv.
So I'd say the answer is yes, and you'll have a good time as long you only need the HTTP interface to do certain things (responsive dashboards; and do them well!).
Webserver implementations exist in R, but don't have near the time / attention put into them as with Python.
On the contrary, it started life as a Bell project called S, more or less a math/stats DSL. It was implemented in GNU as R, and R became one of many competing "stats packages" you may or may not be familiar with: SAS, Stata, SPSS, etc.
While it can be used for general purpose programming, its main advantage is that it is still primary a math, statistics, and data analysis DSL at heart. The concept of a "data frame" (which you are familiar with if you've used Pandas) as a data structure originated, as far as I can tell, in R. Data frames are built into the language, and the language offers custom syntax support for them.
Also, the standard library is full of high-quality statistics tools. Fitted model objects have handsome, human-readable string representations. The formula DSL is elegant and convenient. Manipulating data (replacing missing values, etc) is easy and relatively concise. Math and linear algebra is similarly and it is linked to BLAS so it's pretty fast. Plotting is built into the language and it's pretty intuitive, even if the defaults aren't that pretty. The language is also fully homoiconic and wildly dynamic, allowing you introspect and modify pretty much any chunk of code.
And all that's just in the standard library. The package ecosystem is downright enormous. You can write R packages in C/C++ just like in Python if you need something to go fast, aided by Rcpp. There's Shiny, which is a self-contained HTTP server for data-driven web applications. GGPlot2 was a minor revolution in elegant data visualization. The Tidyverse package collection was similarly mold-breaking by letting users write organic "data pipelines" instead of imperative code. Caret is at least as good as Scikit-learn for general-purpose machine learning. XTS takes the pain out of time series manipulation and modeling. Data.table can efficiently join and subset billion-row datasets in memory using indexes. The list goes on.
Long story short:
- domain-specific niceties
- batteries-included standard library that mimics features found in big monolithic stats packages
- has general-purpose programming capability
- extensible in C for speed
- built-in plotting that's not perfect but it's pretty good
- huge package ecosystem.Oh how I wish this was true! Luckily RStudio hired the author of Caret to develop a family of smaller tidy modeling packages (https://github.com/tidymodels), and with recipes we're finally close to having something like sklearn's Pipelines, which IMO is one of the best parts of sklearn.
With R? Why would you want to do that with R? R is not suitable as a web server. May be you can write a package for that using C. There are 13170 packages for R. ın fact 99% of R consists of packages. You don't sit and write web server with R.
R is used for statistical data analyses. I was using R to get the most occurring error in Apache/PHP error logs, only with 2 lines of R code. https://cran.r-project.org/web/packages/ApacheLogProcessor/i...
Even then, if the stats being done in the background were hard to reimplement, I suppose plummer & R could still work with the right cloud / load balancing infra. Might end up being more expensive than it needs to be in the final iteration, but in the meantime money could be flowing in and customers gettin’ happy.
You get to use the best language for the task at hand and don't have to worry about performance penalties for doing so.
Doing something like that is definitely possible, all the parts are there and work well. Shiny gives you a lot out of the box, is great for prototyping and can be customized. I’ve been working on a less opionated package that isn’t ready for anything but gives an idea of what would be possible:
http://adv-r.had.co.nz/Computing-on-the-language.html
Functions in R are not referentially transparent, so replacing an argument with its value is not necessarily the same. That is a clear restriction on optimizations. If you would want to choose a restricted subset of R to speedup, then this would be a good candidate to cut out since the standard place to compile is at the function level (Numba, Cython, and Julia all do it at functions).
I suspect pass-by-value is a much bigger barrier to speed in R than non-standard evaluation.
However, it is true that the warm-up and memory usage are something we need to improve. We're working on providing native image [1] of FastR. With that, both the warm-up and memory usage shold get close to GNU-R.
[1] https://www.graalvm.org/docs/reference-manual/aot-compilatio...