Enhancing R: The Vision and Impact of Jan Vitek's MaintainR Initiative
r-consortium.org
r-consortium.org
It's the language I'm most proficient in. Tidyverse is the most human-friendly way for explorational data analysis. Data.Table is blazing fast. RStudio is the best IDE I've ever seen and so tightly coupled with the mostly amazing documentation that it's a pure delight. CRAN's quality control next to none.
That being said, I prefer Python for my production use cases.
Why? What's missing in R, imho, is:
a) a decent interface to the web.
b) a decent way to use async processing / utilise the CPU fully.
I've found ways around both, but compare `Shiny` to `FastApi` + `Jinja2` and `future` to `asyncio`.
Shiny just feels clunky, no matter what you do. And R `future` (which saved my butt in 2019 when I was processing millions of geospatial data points for a project, so not throwing any shade) can be a total mindfuck and bug out (at least back then).
(Man, I miss working more in R, though.)
Of course this is partly attributed to R's great DSL capabilities and making documentation first class. But I've definitely seen terrible APIs in R too.
Wonder if anyone else has had a similar experience with another ecosystem? (Regarding API design)
R is also thriving in bioinformatics with Python trailing behind as an afterthought. I think the main reason why R is fading is simply why some other languages fade: fashion, and prestige of the new masters (read: professor X in field Y likes esoteric language Z, so his grad students write out libraries A/B/C to corner the Z library space, and professor X gets more clout than he ever would in a default language).
Maybe this is sub-field dependent, I'm a bioinformatician who hasn't touched R in about 5 years, and everything is now in python.
It took a long time for Perl to go away from daily use in bioinformatics, solely dependent upon how common it was twenty years ago.
In our lab, there is me, a (mid career) polyglot who mixes Python, Go, Java, and R daily. I use tab delimited text files to transfer data. I also grew up coding in C++ and like learning new languages.
We also have a (mid career) staff scientist who grew up in Perl, but switched 100% to R and an (early career) postdoc who has always used 100% R. For both of these people, if work can be done in R, it is. If it can’t be done in R, they figure out how to do it in R anyway (even if that is shelling out to another program).
We also have a (young) grad student that is 95% Python. They try to keep to the Python tooling, even though they are quite aware of the R ecosystem.
There is a generational shift in the field, and it is more apparent each year. I find it interesting that Python took over from Perl first and now it’s trying to take over from R.
Both are under active development and are used in several transcriptomics atlas projects, as far as I can tell.
Look at the packages now for integration, pseudo time, pseudobulk - R (and therefore Seurat) dominates heavily
And apparently that's the problem, FTA:
> Fortran isn’t as popular as it used to be, and we encounter issues when compiling it with modern compilers like LLVM. Ensuring Fortran compiles across all desired architectures and operating systems has been a persistent challenge.
I don't compile any old C or fortran code, so I can't say anything about if and what kinds of problems arise from using modern compilers against modern targets.
Even if most people use R interactively, having contributers working on compiler has many positive spillovers for the language.
Also note that the R code running behind the scenes of your scripts (powering the functions of your favourite packages) is quite a different language, using less dynamic features. This is where a better compiler would always be appreciated.
I think you're right about fads and the appeal of a new sexy language. No dispute on that point.
However, some of what constitutes fads are really more like a coincidental convergence of advantages. So for example, something gets picked up in field X because it's more convenient and all the people learned that in college, and then the same thing happens in another field, and then when libraries in the two fields interact, it's like multiplying the reasons. These network effects happen everywhere in tech. It's not necessarily good, but it happens.
With R in particular there's a long arc to reasons why it might fade. I've heard about R fading before and then it picked up again, so who knows, but it will probably fade and there are reasons why.
If what you're doing involves mathematical or computational fundamentals, all that wrapping around C and Fortran gets annoying really fast. Not everything involving heavy lifting has been coded in fortran or C already, and sometimes passing back and forth between those heavy lifting routines becomes a huge bottleneck.
R is slow as hell, and yes you can write things in C or Fortran (or Rust etc?), but it turns something that should be fairly straightforward into a library project on its own almost. It's just easier to be able to write all the underlying stuff and IO/API stuff all in the same language and have it perform optimally. In fact, I'd probably rather just write it all in C or Fortran than write parts in one language and then wrap it in R — the R wrapping would mostly be to make it accessible to others (which is important, but there's the library bit).
R too has become horribly fragmented in my opinion. A lot of things like ggplot and tidyverse are great, but it's led to this kind of fracturing of syntax in what was already a kind of fuzzily defined syntax in some ways.
For what it's worth, I'm not sure I greatly prefer python. It's more general-purpose than R so has that advantage, but also has the same performance issues as R, and doesn't seem quite as well suited to statistical and numerical computing to me. Maybe it's just the object-heavy structure of python or something — maybe I prefer something closer to either lisp or C in the end — but python has never felt quite right to me. I'm looking forward to seeing what happens with Mojo, because that could be a real game changer, but am not holding my breath.
Julia is appealing to me and seems to check all the right boxes, has a lot of the fun I had with R early on, but I agree with some of the criticism it's gotten for library interdependence problems due to type flexibility and conversion-type issues. I also think error handling needs a lot of work. It's great when it works, but sometimes just becomes intractable to debug.
For me, statistical/mathematical computing is exciting in the last several years in a way it hasn't been in a long time, but also feels not quite "there" yet. I could still see a lot of currently used languages be superceded pretty dramatically by something new.
I think similar issues with syntax is what drove many students away from perl, preferring python with its one-way-to-do-it philosophy.
I think you’ve proposed a real reason a language’s popularity changes. But let’s add a few more: First, as languages grow in use that in itself leads to more use (and vice versa). Second, users of languages use said languages as protection of their jobs, livelihoods, culture, and so on by simply not learning less popular languages or allowing the use of less commonly used languages. There is power in numbers.
IMO R’s problem has nothing to do with R’s limitations (which when limited are easily expanded). To the contrary, if anything, R’s ability to enable users to do so much is more of a problem than its limitations. The norms (culture) of a company say using many different languages for different organizational roles can clash with an R user who can succinctly work cross-functionally. Employees could be interested in preserving siloed roles and employers could prefer limited scope for employees. Instead of one person building something and presenting to users, you could have many employees serving in many roles using standard languages for each role. R is basically a wrapper language, and the ease at which third-party libraries can be built and installed for R, and R’s flexibility is inherently “useful” but less dogmatic. Less dogmatic but also less standard and less common.
… So I just don’t think the problem for R is that “it’s not deemed useful” and I actually think such an argument is disingenuous on the part of people who want to limit the power of R users. Granted, I think, particularly in large organizations, using the most popular languages in itself has valid justification. But the reason is to attract employees and to build siloed expertise; not to enhance “usefulness”.
Especially anything using multi-threaded algorithms or from the recent algorithms research world, I really get a kick out of hacking into a usable R module. It is almost like doing computer science again (as opposed to the real day job, which is data munging for reporting tools).
This description seems to fit Julia really well. What other languages are like this?
The all-in-one ecosystem of R is nice, but text encoding is still a major pain point (e.g. people try to put emoji into RMD to translate into pdf via tinytex, and fail miserably, of course).
R had a good run, and IMO they still have some packages that are just too good to see it languishing into the future, notably the package data.table. I have not come across a better library for data manipulation, in R or Python. The syntax is excellent and it is faster than most alternatives, especially the popular ones.
I think an important factor that contributed to R's downfall is the decreasing hype around data science, as well as the fact that the core base of R users do not have a background in the STEM fields, but rather the humanities and social sciences. I do believe that the R community of developers are dedicated and perhaps, in relative terms, are more involved with their projects than a typical Python developer. But that's not enough, the sheer numbers of Python developers eclipses R's and that is too hard to overcome.
But in the academic world (I'm in bioinformatics)... There seems to be nothing but R, in my experience. I don't really like that, because we have Snakemake and a lot of ML stuff, all Python, and the R people have a barrier to get started. I myself associate R with "just scripts and notebooks". Never do the R people seem to make anything into a well maintainable module. They make the notebook, use the build-in R functions and then their work is done. It seems to be different in the Python world where I see people writing modules that are "re-useable assets", and then those are used in notebooks for data science. This is probably my industry bias and perhaps Python-using academics also never make packages.
I guess also that there is no such a thing as poetry in R? I'm not entirely sure...
Exactly, it is an industry-driven change IMO. In fact, R has gained a lot of popularity in academia, especially in the Social Sciences --though perhaps Python has gained even more.
But in industry, R's falling hard. I also think that the growing popularity of cloud analytics platform solutions such as Azure Synapse (now Fabric) is a significant factor. Though SparkR is a decent R-native API to Spark, Python has so much support in those cloud analytics ecosystems, it's hard to keep doing things in R.
Your guess is (largely) correct. For most users, the workflow is exactly as you described:
> They make the notebook, use the build-in R functions and then their work is done.
R is a highly functional language modeled after Scheme. In many ways it's more powerful than python, especially due to its metaprogramming capabilities. It's possible to write maintainable, readable, high-quality code in R (just look at the Tidyverse[0] libraries). The issue is that the user base is mostly scientists and statisticians, and they just don't.
The bioinformatics space seems especially dire with dumpster fires like Bioconductor.
“R is for statistics, but if you want to do machine learning you have to use Python”
Which is ironic given that pythons success as a data science tool is built on the back of a matlab clones.
I think I like that actually. The core is quite conservative, meanwhile the packages that the average R community member uses break their own backwards compatibility every few months for silly reasons like renaming an argument `linewidth` instead of `size` in a function that was otherwise backwards compatible for a decade.
R is still unbeatable for wrangling tabular data, visualising, and fitting complex models on tabular data (complex up to GAMs let’s say).
It's not just a matter of manpower and funding, which Python has more of. Python is actually a less capable language in some important ways. It's not that expressive, and has basically no affordances for domain-specific languages.
I'll be happy when we have a successor language for R, but it won't be 2024 Python.
> I'll be happy when we have a successor language for R, but it won't be 2024 Python.
When it comes to statistics, and especially having to program stats analysis instead of just calling libraries, R is a much better language than Python or Julia. When it comes to other things, they are better than R. I use all 3 daily.
And as someone who uses both languages daily, Python is much better at some things, R in others. R is much better in coding stats analysis.
Competition is great, may both thrive.