> Native data science structures.
DataFrames are often easier to use than Pandas. However, in production workflows we're often using more datatypes than DataFrames, and R is weaker there. For example:
* Lists must be accessed with `[[ ]]`, instead of `[ ]`. I've seen many silent bugs slip through due to this. * There are 3 competing implementation of classes. This results in classes being mystical and rarely understood. * R is a Lisp 2. Variables and functions may share the same name. This leads to confusing errors. * Catching specific types of errors can be awkward. * Adding elements to a list iteratively is slow [1]
> Non-Standard Evaluation
This can be handy while quickly working in RStudio, but it's not easy to maintain. I've seen code that failed because it specified `f(!!variable)` instead of `f(!!!variable)`. I like R's formula notation, but I'm happy enough with sklearn's API that I don't miss it.
> The glory of CRAN
CRAN is not set up for production. It makes pinning versions very difficult [2]. Many people resort to using MRAN, which is a Microsoft supported snapshot of CRAN at a specific time, so a dev can just pretend they are installing software as if it were 6 months ago. I have seen MRAN go down multiple times [3]. Not to mention, the owner of CRAN is notoriously prickly [4] and packages will not be accepted to CRAN unless the maintainer ensures their software runs on Solaris [5]. Hadley Wickham has done so much for the community with `devtools` and his books. He gets a lot of praise, but it's not misplaced.
> Functional programming
Okay, this is actually pretty great. Hooray functional programming! Not totally related, but R has a great polymorphic dispatch of functions, which really can't be undersold (the way automatic documentation generates for this is kisses fingers)
Ultimately, R is a cool language. In interactive settings, I would rather work in RStudio than Jupyter any day. I like RMarkdown better than Notebooks for sharing analysis, too. If there is a specific Bayesian model necessary only available in R, that's fine, wrap it in a container. But the rest of the ETL and pipeline code feels easier to write and maintain in Python.
[1] https://stackoverflow.com/questions/17046336/here-we-go-agai... [2] https://stackoverflow.com/questions/17082341/installing-olde... [3] https://github.com/Microsoft/microsoft-r-open/issues/51 [4] https://www.reddit.com/r/rstats/comments/2t5oqp/dont_use_the... [5] http://www.deanbodenham.com/learn/r-package-submission-to-cr...