R vs. Python for Data Science
github.com
github.com
I disagree with the "learning curve"; if you've learned other programming languages Python has a pretty simple and familiar core, and Pandas (while the API is an inconsistent mess) is well documented. Base R is quirky compared with modern programming languages, and the API is pretty inconsistent.
I also strongly disagree with the Tidyverse bashing. I'd say it has the shortest learning curve (especially for someone familiar with SQL), and is one of the main reasons I still use R today outside of deep learning - I find it much more friendly to work with than any alternative.
Your joking right? Eggs, wheels, virtualenv, venv, pyenv, poetry, conda, python2.7, 3.4, 3.5, 3.6 (yes, we have all 4 versions installed in my current company's "production environment", not to mention 3 different versions of python 2.7 but no 3.7)...
Fortunately conda came along and pretty much solved all my problems.
I don't have much experience with packrat - but as opposed to pip it's another thing you need to discover and install. And so people don't do it by default when releasing code, and I've had to bisect versions of dependencies to get a working version of code. This can happen in Python too, but is rarer.
I've deployed both R and Python for completely junior datascientists team, on top of a poorly managed infrastructure. I'd say they both have pros and cons and are actually both pretty bad. But R's packrat makes it slightly better than python. Python is a mess when you want to reproduce a working environment. Conda and pip both have huge issues. R's package management is pretty poor too with completely misleading errors, but at least it's unique and once you know your way around the most common errors you can build and run different projects quite consistently.
I've managed both RStudio+Shiny for R and Jupyter for python and overall my experience is better with the R stuff too. Things look a bit standardized while Jupyter needs tons of dependancies and (I felt) lacks a clear opinionated way of doing things.
I have 0 opinion on the actual languages though, as I'm not a developer.
I had a colleague try to set me up with their R-studio project recently and we gave up because getting the packages installed was such a mess. So I'm not currently a huge fan of either. I don't do much data science or machine learning, but I work with and support people who do.
There are many aspects here:
* You simply use it (in interactive mode)
* You integrate it into your application as static lib
* You integrate it into your application as a dynamically linked lib
* You use it via some kind of (remote) API
* Do you integrate via source code, static lib, dynamic lib, API
R is a rubbish general purpose language and should not be used as a general purpose language. But it’s damn awesome for anything that considers matrices and vectors as the primary units of manipulation. Further, CRAN has everything ...
I'm being facetious here, PyPi is a few times bigger but less curated.
You don't miss anything by skipping this.
>By contrast, just now I tried to find nearest-neighbor code for Python and at least with my cursory search, came up empty-handed; there was just one implementation that described itself as simple and straightforward, nothing fast.
> The following searches in PyPI turned up nothing: log-linear model; Poisson regression; instrumental variables; spatial data; familywise error rate; etc.
This is not how you search for things. Usually I search on google for "poisson regression scipy" if I'm looking for poisson regression.
I'm not sure I follow this, you can just set the interpreter.
I've started using rMarkdown more heavily, with reticulate & python for most data munging and r for plotting. Partly because I already know how to solve the problems I have in python more quickly than in R. The only thing I have against it at the moment is the debugging story isn't very nice by default, though I've not looked into how to improve this.
edit - if you've not looked into rmarkdown, I heartily recommend it. It is to me what the final output of notebooks should be. I can easily interleave code and descriptions, hide what I want, run it from scratch entirely as a default, and produce a range of outputs including interactive static webpages. Once web packaging is finally sorted, it'll be near perfect.
* R has Better statistical correctness based on "some dude"?
* R has better OO programming because you can print functions to the command line?
In my workplace R and Python are both well represented and I always hear from the R users that there is no "real support" for classes in R as there is in Python and that they miss it. I can't judge for myself though.
I'd be interested in seeing it included in the comparison, though (I'm afraid it would lose on multiple points if only due to lack of funding, but it is very usable).
I currently have a fairly big project written in Octave, but will likely rewrite it in python for maintainability (and would rewrite it again in something else if it grew too much for Python).
there's also Sage, which is an interesting contender, but I do not know enough about it to know how it compares. Arguably, though, Sage and Octave are more geared towards numerical computing than data science, and I think that's where they shine. So, depending on your data and the processing you need, those could be more adequate.
Sage is not really aimed at numerical computing, even though it can be used for that. It's primary use case is more towards computational algebra and number theory (and related areas) and is generally more focused on features needed by researchers and academics.
I think ggplot2 is a decent library, but in my opinion it gets more praise than it deserves.
I mostly use python and prefer it to R, but putting things into production is not a strength of python and R wins that comparison a thousand times over.
- Learning curve
- Machine Learning
- Parallel computation
- C/C++ interface
- Object orientation
All of these are wins for Python, some of them like the learning curve are wins by a huge margin. I should probably do a point by point rebuttal later but so many of his points are incorrect and/or poorly justified.
> For instance, though functions are objects in both languages, R takes that more seriously than does Python. Whenever I work in Python, I'm annoyed by the fact that I cannot print a function to the terminal, which I do a lot in R.
I assume "print a function to the terminal" means print its source? If that is the main complaint, it is available out of the box in IPython (%psource), which if you are doing data science you are probably using already.
https://github.com/search?q=poisson+regression
I've not used Rcpp, but Pybind11 is pretty mature, and works well, so I'm not sure why he's saying it's under development; by the same measure there was an update to Rcpp last week, so that is too. He mentions Cython which allows you to compile Python code, but my main use case for that is exactly what he says - wrapping up C/C++ libraries, which is very easy in it.
"Python is currently undergoing a transition from version 2.7 to 3.x. This will cause some disruption, but nothing too elaborate."
This is pretty out of date; in Jetbrain's survey, 84% of devs had transitioned.
R
- The tidyverse ecosystem had given a huge boost to R. It had brought intuitiveness and consistency to R which was much required especially if you are a programmer coming from other languages. Also there are other ecosystems like Bioconductor which are also very mature.
- Rstudio and especially Rmarkdown notebooks are much better for reproducible analysis than Jupyter.
- It is very difficult though to develop standalone tools with R. For example it doesn't have a good argument parser.
Python
- The language is much more intuitive and more ideal for developing standalone tools.
- The ecosystem is many cases very fragmented though with a lot of libraries doing similar things.
- It lacks a good plotting system. Matplotlib is very powerful but has a very steep learning curve. In comparison ggplot2 in R is very intuitive.
If you're writing code that's deployed in any way, it's best to avoid the tidyverse as much as possible. This is also acknowledged to some extent by the main developer - https://www.tidyverse.org/articles/2018/06/tidyverse-not-for...
Someone noted below that python is better for deploys: true. R is better for interactive use, and has a better package universe, though quality control for packages is vastly lower than something like scikit learn. R also completely dominates in classical stats, which is generally bread and butter compared to having the latest goofy neural thing. Assuming you actually do data science.
There are lots of regression options. Scipy, Scikit learn, Pymc3, PyStan.
Metaprogramming in Python is easier than in R, and arguably more predictable and consistent.
In any case, I've worked extensively in both environments and I don't think the author has considered every aspect. Below or my two cents.
- Elegance: slightly disagree. R might look more concise, but the language comes with many strange aspects (quoting, non standard evaluation) that can put a wig between novice and experienced team members. Python is more verbose, perhaps, but cleaner overall
- Learning curve: disagree. Even when working in R, modern practice would ask you to learn the tidyverse or data.table first instead of sticking with base R. Good tutorials are available for both
- Libraries: depends, the notion of "libraries" is too broad anyway, better to split it up according to the subcategories below. Both come with lots of packages, so I'd agree with it being a tie
- Statistics: agree with R. R is still the statisticians language, and many implementations of some more obscure techniques are only available in R. This being said, most ML shops today would be more interested in e.g. a good GBM implementation or deep learning rather than some robust statistics package. In R: think regression, ANOVA, significance tests, time series and niche subfields like bioengineering. In Python: think RF, GBM, t-SNE, deep learning
- Parallel computation: I'd say both are lacking, and you'd need to look more towards tooling such as Spark anyway. I'd also say out of memory computing becomes your first concern more often. Dask and Pandas on Ray are very nice on Python
- Foreign interface: kind of disagree. I think Python has matured better here
- Object oriented programming: disagree. The problem with R is in fact that is has about 4 (or more) OOP ways
- Interop: agree that you should avoid it, at the moment, it will only make deployment more cumbersome
Some other concerns I'd consider.
- Pipeline approach to ML ("model dev / model run"): better in Python. E.g. the clear approach of scikit learn to consider both preprocessing as the model itself as part of the fit-transform-predict pipeline with clear methods is way better than R. I've seen many novice R users fall into the trap of preprocessing a data set before splitting in train/test, for example. This has been one of the biggest drivers to push me towards Python coming from R. Most established libraries in Python commit to a shared, best-practice way of thinking whereas every package in R seems to come with its own ideas in terms of pipeline and usage
- Deployment: also a win for Python. Better package management / reproducibility, though it is possible in R as well
- Data exploration: I find this easier in R. Packages like dplyr help a lot here. Pandas' API is somewhat cumbersome
- Charts / visualizations: ggplot2 in R is still a champion, though good dashboarding tools exist for Python as well. Still, I find this easier to use in R
- Spatial analysis: both come with very solid libraries, though I find whipping up a quick visualization easier in R
- Deep learning: clear win for Python. Tensorflow, PyTorch and even Keras are not fun to use in R
- Reports authoring: possible in both, though R's markdown functionality combined with RStudio is fantastic. Nevertheless, Jupyter notebooks can be made to act as a reporting tool for both languages
I can’t claim to have tried everything out there, but so far it’s been my experience that matlab beats the socks off everything else when it comes to data exploration. I’m mostly looking at time series like data (but mostly not statistics). The ability to do things like click on a datapoint and export it and it’s index back to your workspace sound trivial. But in practice it’s a huge convenience and I’ve been unable to find a plotting package for python or Julia etc that can do things like this.
> Given R's magic metaprogramming features (code that produces code), computer scientists ought to be drooling over R.
The benefit of this... is very debatable.
And it's a pity. Like a lot of "old" technology and "old" languages, also SPSS is far easier to use for normal people. Asl usual, IBM is just wasting potential.