Boost Your Data Munging with R
jangorecki.github.io
jangorecki.github.io
The vignette linked above describes some compatability between dplyr syntax and data.table, but I admit I have not tested that.
Nowadays, if you are working with multigigabyte datasets, it may be worth looking into Spark/SparkR (https://spark.apache.org/docs/latest/sparkr.html) for manipulating big data in a scalable manner instead, as there is feature parity with data.table + analytical tools [GLM] as a bonus. (The syntax is still messy, to my annoyance)
I was a data.table acolyte for years, but after being forced to learn dplyr I can't imagine going back. I'm hoping if I need update-by-reference in the future I can use tbl_dt but I haven't played with it yet.
I find that some tasks are frustrating with one package but easy with the other, so more frequently I just have some "dplyr lines" and some "data.table lines" in my scripts.
And the compatibility between the two has improved a lot, so you don't get so many "invalid selref" warnings and unexpected behavior when you combine them in pipelines
If you prefer Azure, a G5 instance with 448GB RAM is available for $9.65/hr.
Not a lot of need for Spark/SparkR for the vast majority of data sets given the cheapness of compute these days.
Such an approach worked to great effect for me recently, when I needed to perform a `zoo::rollapply`[3] across a time series with tens of millions of rows. The speedup when throwing more cores (7, on my laptop) at it is roughly linear. If I ever need to scale up the analysis to hundreds of millions of rows, the 128 vCPUs of an x1 EC2 instance would be well worth the $$/hour.
[1] https://cran.r-project.org/web/packages/foreach/index.html
[2] https://cran.r-project.org/web/packages/doFuture/index.html
https://gist.github.com/stillmatic/fadfd3269b900e1fd7ee
if your function has a well-defined rolling form, e.g. rolling mean or standard deviation, you should use a pre-optimized function for that. package `catools` has a bunch of useful ones.
data.table transformed the way I do analysis. Honestly probably 75% of my lines have are dt[] calls. It allows me to analyze on 500M row data.sets routinely on a powerful workstation interactively.
my beef is with the new tibble-ish format of dplyr. the format doesn't support lists as cells, unlike data frames. right now, `tbl_df` is a good in-between, but since dplyr is converging to the tibble/feather format, it's gonna be rough
The article mentions Reshape2. I also love TidyR. It's much easier to use than Reshape2 when dealing with a large number of variables (i.e., columns).
ETL is just one step of the pipeline.
I have tried yhat's rodeo and also yhat's port of ggplot for python but it just isn't there yet in my opinion.
I use both heavily in my daily work.
On the flip side, I've done a fair swath of work in the machine learning arena, and in particular the deep learning topics before it was called 'deep learning', and it was nearly hopeless to use anything except Python or MATLAB (both strongly tied in with C++/CUDA libaries). I think this is still largely the case.
As others have mentioned here, for me the biggest pull away from R was that it's not general purpose. Hard to ship someone R code. Hard to throw a web framework in front of your code. Hard to build a rich desktop GUI on top of it. I know you can do most of those things in R, but last time I dealt with it there was a massive ravine in usability/maturity. I'd also be insincere if I didn't admit that the pervasive R coding style just drove.me.fucking.crazy. That and I found that while there was a mindbogglingly large pile of libraries, documentations was usually very lacking.
/rant