Why Clinical Laboratorians Should Embrace the R Programming Language
aacc.org
aacc.org
R is a data analysis DSL that also happens to be a full programming language.
- Rmarkdown. I prefer a text document over a web notebook for exploratory research
- The standard library is for statistics and data: dataframe, lm, anova, etc. are builtin
- A huge range of probability distributions are built-in. I don't need to import extra libs to do simulations.
- Between functional programming techniques and vectorization, I can write very clean and concise code
- Tidyverse and data.table are lovely and coherent approaches to data management. Data.table is fast and memory efficient.
- Advanced models are trustworthy: For example, glmnet, mgcv, nlme, rms are authored by statistical heavyweights, and have accompanying books that are excellent. I don't have the same confidence in python's statsmodels.
- CRAN is easy to use, I can access it from my R session, and there are rarely problems (big thanks to Uwe Ligges)
- Libraries for design of experiments and surveys are available. R supports the entire design -> data management -> model cycle.
- base graphics/lattice/ggplot2 are excellent for plotting. If I need something advanced, I can use grid. If I need vector graphics, I can use tikzDevice for latex.
- Rstudio is a an excellent IDE, and Emacs Speaks Statistics is an excellent Emacs plugin
- It is very easy to get help without going to google. (?foo, ??bar, etc) Documentation is well organized, and the documents often contain citations and relevant links.
- Lots of advanced models can't be found outside of R. Today I fit a splines-on-a-sphere model using mgcv (https://stat.ethz.ch/R-manual/R-patched/library/mgcv/html/sm...)
- Rapid iteration in modeling using Wilkinson notation formulas. Built-in formulas are the actual killer app of R, IMO.
- Things are generally fast, but if you need extra horsepower, plugging into c++ is easy using Rcpp.
- R feels like lisp. Experimentation is easy, and I don't feel forced into any particular paradigm while using R. I have a lot of ways to evaluate code (https://ess.r-project.org/Manual/ess.html#Evaluating-code)
Compared with Python the language ergonomics of R are confusing and inconsistent.
I guess momentum and establishment is also a feature in itself though I’ve personally never felt that, one of the selling points, the esoteric statistics packages at the edge would be of any use to me.
The use I’ve seen is reminiscent of Spyder and Notebooks: tangled, unreadable mess of line-by-line execution where people are prone to re-running stuff out of order.
They are different tools for different purposes.
That being said, R is amazing for exploratory work, while Python is better for integrating with the rest of the world.
R is far more of a Lisp than Python, and in a field that heavily relies on DSL abstractions (which definitely includes clinical laboratories), R is going to fight you a lot less than most choices.
In regards to packaging, you have MRAN snapshots, you have conda (which will give you binaries on Linux), you have renv, roughly in order of preference. The situation is not ideal, but it's certainly not worse than Python, this is not a hill I'd die on!
Julia might be an exciting and welcome alternative to both; from where I'm sitting, it hardly even enters the conversation currently (in data science where I'm at, it's all Python and R, with Python unfortunately taking by far the larger slice of the pie), but I wish it a bright future, it's a great language.
That said, R had come a long way in recent years and is enjoyable to use. It is very complete as far as statistics go.
If you like to write lots of code, Python might be better. But if you use it in clinical research, R has probably better packages for whatever you need.
Python gets scrutiny but most packages are on github and feedback can be received. Although, you should read the source code regardless.
EDIT: And I re-emphasize -- never trust the source code, even if the company you work with indemnifies. One commercial product we used in particular had an insidious bug in one of the new time series packages that got corrected with in later versions, but we never would have found it if model testing requirements didn't also require implementing in R or Python. Since the package didn't exist for Python, and we wrote it ourselves, we found the performance issues.
In R you more packages that do what you expect (mathematically) but the implementation is inelegant and slow. Written by someone who knows exactly what they need and what the package should do, but has difficulty of writing it down.
In python you many well implemented neat packages where the code is well implemented and performs well, but is not exactly doing what user need or skips important features because they are conceptually difficult.
I disagree with this point explicitly. Many packages are not only poorly written programmatically and systemically, they also produce bad results in common cases and fail silently. This has been discussed well before on our very own YCombinator.
https://news.ycombinator.com/item?id=17308554
> is not exactly doing what user need or skips important features because they are conceptually difficult.
This is why classes are cool. You can extend them, modify them, and so on. Perhaps we should teach this more to fellow data scientists.
You may be unaware, but there are vast similarities between the object system of R and Python, mostly due to their common inheritance from the Art of the Metaobject protocol. They look very different (generic functions vs classes), but they are equally extensible.
The trouble with R's systems is that there's three of them, and people use whichever works without really understanding any of them, but all the tools are there.
You said, in the context of ensuring conceptually difficult parts of a model/method were implemented:
> This is why classes are cool. You can extend them, modify them, and so on. Perhaps we should teach this more to fellow data scientists.
I pointed out that both R and Python have similar object systems.
> ou can implement what you want in Python relatively easy.
I don't get why this is necessarily true, but I might be missing something. Can you clarify?
IME, for production Python rules and for data exploration, graphics, stats and adhoc analysis R rules.
It's problematic because people abuse it for everything (300 line pipes are common, sadly) but it's a really useful tool in moderation.
* I can't quite get htmlwidgets and docx formats to work together in bookdown without using separate commands for the interactive tables (DT and flextable).
I'd say the R dev community was then actively hostile to a culture change to support any model other than individual contributors working at their desktop.
It's also important to note that much of the original core of R is based on S, which was developed around the same time as C, so some baggage would be expected.
> Unlike Excel and many other graphical user interface (GUI)-based programs, R’s reliance on text-based structure makes it straightforward to review at any time the commands used in a data processing pipeline to ensure that the correct steps were taken.
> Furthermore, the ability to view the underlying commands facilitates transparency and reproducibility of analyses.
The article seems to be targeted at people with zero programming knowledge. The arguments here are valid, but don't rely on R.
The title could just have been ".... Should Embrace a Programming Language".But it's very powerful, it's exactly right for these use cases and its ecosystem is mindbogglingly huge. Also, it tends to be easier to grasp for folks who don't have prior programming knowledge (anecdotal, but I've seen people pick it up very quickly who struggled a lot with, say, Python. And Python is the only language/ecosystem that comes close to R.)
So, yeah, a lot of languages could be used for the use cases in TFA, but R is uniquely suited, weaknesses notwithstanding.
I'd recommend python based on the slightly-saner tooling. I've found that python with conda/pipenv/poetry results in mostly reproducible installs of the tools needed to run a computation.
This is all about tradeoffs. Fundamentally, if your package doesn't compile on the latest version of R, it gets removed from CRAN. This means that each version of R has a consistent set of packages that (mostly) work together.
Contrast with Python which does facilitate reproducible builds, because you can hack together ancient versions of Python and make them continue working. I could go into a massive rant here about pip, but it's trending in the right direction now and I don't want to discourage any of the people working on it.
R is better in terms of being able to ensure that for a given R version, any package you install will be compatible, Python is better for making sure that that one application built three years ago keeps working in the same fashion.
Also, it sounds like you're running a nix based system, have you considered (I'm sure you have) using the system packages. For example, the Debian/Ubuntu ones are pretty comprehensive, at the cost of using older versions. I believe that R-studio also have pre-built packages for Linux (but have not tested this) so that could also work.
To be fair, conda is pretty good as a package manager, because it handles the C++/C dependencies. But to your Docker point, that's how I handle the insanity that is python packaging, especially in the data science space, so it may just be an issue with the field itself.
[1] https://www.amazon.com/Book-First-Course-Programming-Statist...
I’ve taken courses on statistical computing in R and statistical computing in SAS in my statistics degree. We were always told that SAS is the standard for anything health care, pharmaceutical, or where regulation and publication comes into play.
Anecdotally, my friends who did PhDs in biochem and immunology all used SAS for their data analysis.
Have I been misled or is this up to individual preference?
The first guide is for clinical settings, doing data analysis for inflammation in patients who have been given a new treatment for arthritis [2]. Another is a general introduction to R for non-programmers using gapminder data for reproducible scientific analysis [3][4].
[1]https://software-carpentry.org/lessons/index.html
[2]http://swcarpentry.github.io/r-novice-inflammation/
Conventional programmers seem to be somewhat reluctant to learn R's syntax and adjust their programming model.
Non-programmer types think in maths even less so they like the python "straightforwardness".
R is hard to reproduce library setup. Unable to compile and static validation.