R: Introduction to Data Science (2019)
rafalab.dfci.harvard.edu
rafalab.dfci.harvard.edu
For those of you who us R still what is your use case?
We found R has a really hard time integrating into data pipelines and was best used as a standalone tool by individuals, which doesn't really work in our particular professional setup where everyone works collaboratively together.
What we found was that R had alot of packages but most haven't been touched in years and when you contact the owner you find they've often moved onto the python/pandas/scikit eco system
The unique qualities of R that allow this are that it's so easy to use, extremely reliable for package installation (problems occur approximately never), and the tidyverse makes it incredible easy to translate ideas into code, not only in its broad, easy to understand and powerful vocabulary, but in there being little 'nesting' required; instead working left to right and top to bottom (via the magrittr pipe) - i.e. your code, for the most part, is like reading a page in a book.
So if you have tabular data, it’s a no brainer to use R.
Getting the data into table form is often better suited for python. Fitting models that leverage autograd are also better with python.
Totally agree. If it's in a dataframe, I prefer R. If I get involved pre-dataframe, I prefer Python.
I was an R user from about 2003-2010.
We didn't have DPlyr at the moment though ggplot2 was coming around about that time I think. That helped alot for easy to develop visualizations.
But in our specific cases, the distributed libraries we used were written in python and integrated well with native python code. Pandas was just coming out around 2010, I think, and I think multi threading was also an issue then, but I can't really remember.
So our issues was partially our infrastructure tooling was going to python, but also we had a far easier time hiring people who were proficient in python and harder to find the same for R.
And once you start writing more code in python it starts to become harder to justify two separate code bases that can do the same thing so the R code got phased out and rewritten in python so we could have a single code base and not have to duplicate functionality in two languages.
Also a slight push for python came from the programmers who thought python represented a better language to know for their careers. Which looking back it does seem like python is used more often these days in general.
So I guess there isn't much you could have done in this case.
And as a side note, thanks for all the work you've done with R!!
- when we wanted to build a web app that processes data, it was a lot more straightforward to build both in python, so we can process data within the web servers instead of having to manage multiple stages of infrastructure and different languages. There's no Django for R.
- R will often do something instead of explicitly failing. This is the wrong tradeoff when running a production system, as if you're returning the wrong results to users you may not realize it unless there's an error
- R reproducible builds are worse than python. That's saying something because python is a pretty low bar. But running production systems you can't have builds suddenly fail week over week because one of a hundred packages was updated
There's renv that addresses that point already: https://rstudio.github.io/renv/articles/renv.html
> There's no Django for R.
Nowadays you can integrate R with WebR (WASM) in a web app: https://docs.r-wasm.org/webr/latest/
And of course on the web side you have shiny (https://shiny.posit.co), which now also comes in a python flavour.
Aside -- we tried using dash in our production app and then had to remove it after a month, because these types of frameworks that spit out front-end code are almost never flexible enough to do what you actually need to do in a full app context, and you end up doing more work to fight the framework versus the time-savings from the initial prototype.
But the fact that there's no Django for R means shiny's a dead-end for a production web app.
ambiorix might be what you're looking for.
check it out: https://ambiorix.dev/
it provides: - routing - api generation - templating - web sockets
I mentioned exception handling above, but this is more specifically the problem.
I think it's a hard problem to solve, because the behaviour of older libraries is so varied.
I have sometimes thought that something like a try catch wrapper which pattern matched or tested the value returned would be useful.
(this all prefaced with a massive thank you for tidyverse, without which R is very crusty).
I love R for interactive work and quick analyses, but I'm currently trying to integrate various bits of R code into a large document-building pipeline and wishing I could use Python for it:
- Exception handling and error processing seem a pain in R. Maybe I'm doing it wrong, but if feels like a mess and not nearly as ergonomic as python. Trycatch seems to have gotchas related to scope because the error handling is in a function. The distinction between warning, stop etc seems odd. The option to stop on warnings isn't useful because older packages seem to abuse warnings as messages. I have just discovered `safely` which is helpful, but then you have to unwrap lists in pipelines which feels clunky.
- Related, I _really_ wish we could just drop model objects or other tibbles as single objects directly into a tibble cell rather than as list(df). Unpacking lists and checking objects inside them exist is much more of a pain (e.g. can't just do `filter(!is.na(df_col))`)
- I really miss defaultdict from python, and dictionaries generally.
- Passing variable names as strings to dynamically generate things seems clunky compared with python. Again, it may be because I'm doing to wrong but I end up having to wrap things in !!sym the whole time and the nse semantics seem hard to remember (I only use R about 20% of the time). I liked cur_data() for passing a df row to a function but this now seems deprecated.
- String formatting -- fstrings are just great. Glue is OK, but escaping special characters seems more tricksy. Jinjar is OK, not quite jinja.
- purrr is nice, but furrr just isn't a drop-in replacement. Making http requests in parallel seems non-trivial compared to doing it with python. Is there an easy way to do it without creating multiple processes? Why can't I just do something like `. %>% mutate_parallel(response=GET(url), workers=10) %>% ...`?
- 5 different ways to do wide to long and long to wide over the years even in the tidyverse. - A lot of dependencies to connect to DBs and difficult programs. Rstudio/Posit does have some premium libraries but they should be made free and bundled with the tidyverse to really promote the ecosystem. - Shiny support to save interactive charts and tables. This is a massive problem for me. If I have a heavily stylized HTML table with a bunch of css, I need to rely on webshot, webshot2 which are both alpha or beta versions and they are poorly documented. How can I evangelize R if my deployments cannot be used properly by my community?
I'd love to hear more why you're using webshot etc to talk screenshots of your shiny app. A more typical workflow would be to generate a separate HTML/PDF with quarto/RMarkdown.
The packages I think are the dependencies of some DB connectivity libraries. https://www.rstudio.com/tags/databases/ - these are the ones I was referring to.
Re webshot my use case is: I have a heavily modified DT table in a shiny app. Users log in, play around with the DT table, update ggplots etc and then download the snapshot and send it to a WORD file. I can't move away from word and use html or pdf because we need the word file formatted by editors for publication and they need to follow the corpo guidelines. So, I am having to use webshot to grab a screenshot of the tagged html instead of natively handling it. I tried using officedown and a few other methods and it just didn't work.
ps: I hope the rebrand goes great and I am rooting for you.
Hmmm, I'd still try generating the table with quarto (since you can output word documents), or try gt (https://gt.rstudio.com), which I know has much greater control over output, and supports RTF output (https://gt.rstudio.com/reference/as_rtf.html) which should import cleanly into word.
Use suppressWarnings() to silence misbehaving functions or withCallingHandlers() to stop or handle specific conditions.
> Passing variable names as strings to dynamically generate things seems clunky compared with python.
Can you give me an elegant example in Python? Because I don't understand what you want to generate dynamically.
That said, I dislike the tidyverse solution as well. Too much abstraction for not enough benefit over a base solution with substitute()
my_plot_script.R --plot_col=g_max --output_type=pub_quality data_file1 data_file2 data_file3
It's possible to use optparse/OptionParser() to get that information (but you have an option for every argument, no --param1 X --param2 Y file1 file2 file3) but it is much more difficult to fit those arguments into the RStudio environment. I want an RStudio to be able emulate reading command line arguments (since they do not exist in RStudio). Right now, I have to check to see if there are commandArgs(), and, if not, do something else to get the information to the RStudio script.
(2) There needs to be an option that says STOP if something doesn't make sense. I have dozens of beautiful data plots that look great, but in fact do not in fact plot what I think they do, because factors have not been properly assigned to colors, shapes, or linetypes. (And it can be really hard to recognize that the data has not been plotted properly.) Give me an option that says, if I did not explicitly declare a column a factor, and I did not specifically associate colors/shapes/lines with factors, then the data will not be plotted.
(2) I think part of that is in scope for strict (https://github.com/hadley/strict). You might also be well served by adopting some more data validation tooling, e.g. pointblank (https://rstudio.github.io/pointblank/).
Maybe that’s the problem?
In my org we have several 100% R teams (including mine) that have been developing and maintaining business-critical, data-intensive applications for a decade now. We don't find R difficult to integrate into data pipelines. We write our data pipelines in R, and we find it very efficient to do so. They talk to databases, APIs, command line tools, etc without issue.
Doing what we do in Python is unimaginable, especially if pandas is the tabular lingua franca in the team. I vehemently agree with this article on the clunkiness of pandas from a sister comment: https://www.sumsar.net/blog/pandas-feels-clunky-when-coming-.... Compared to dplyr and the tidyverse, pandas very noticeably gets in your way rather than being a tool of thought. (For what it's worth, there are other teams in my org that use Python for entirely justified reasons, and they use polars these days, not pandas.)
If I had to complain about anything in R these days, it would be the increasing complexity and illegibility of error messages. Tidyverse tracebacks are often dozens or hundreds of lines. This is made much worse if you have a web app in the Shiny framework, as Shiny seems to mangle and garble what little useful information you can get (my kingdom for an error with a file name and line number). Even outside of advanced packages like Shiny, the reporting of error messages suffers from some clunkiness and irregularity.
As an expert user, I can usually squint at the error barrage and infer what is really going on, but it's probably quite confusing and off-putting to newer users.
Overall though, I'm not seeing any competition for R in our space. My fondest hope is that in the coming decades there arises a new, thoughtfully designed language with the Lispy flexibility of R, but also optional type safety and static analysis affordances. I'm not sure if that's even possible, but I hope the computer science geniuses figure out a way.
(The intersection of tidyverse and shiny tracbacks are a known pain point that's hard to resolve. Unfortunately shiny and tidyverse did a bunch of parallel work that took us in slightly different directions and now it's hard to re-align.)
One thing we are missing is a guide to reading traceback for newer users. Often experts can get a good sense of where the problem is, but we've failed to teach newer users how to get the most value from a traceback.
The biggest one I have right now is a little niche, but probably useful to address. Moderately complex dbplyr pipelines on wide tables have a tendency to generate very long queries, and if there's an error, the generated SQL returned tends to overflow some text or line limit allotted to show the error at the command prompt. My workaround is to use sink() to dump the error to a file, which is a little painful as the sink() API and documentation are not the most straightforward or intuitive. (Hmm, I wonder if a withr wrapper would help me make something simpler to use...)
I filed an issue so I don't forget about this: https://github.com/tidyverse/dbplyr/issues/1471
I think many of us saw Julia as the successor to R. Unfortunately, the package ecosystem---one of R's strongest points---still has a long way to go.
My sniff test for a successor language to R is whether it can replicate the tidyverse API with 100% fidelity. The API is already optimal for tabular data analysis, especially the dplyr core. It can be thought of as a specification for other languages to implement.
There is a great deal about how R works that is negotiable. But if the language can't implement dplyr to spec, or somehow doesn't "want to", it's not the language for the audience served by the tidyverse.
Here's how dplyr-style chains look in their system:
using TidierData
using RDatasets
movies = dataset("ggplot2", "movies");
@chain movies begin
@mutate(Budget = Budget / 1_000_000)
@filter(Budget >= mean(skipmissing(Budget)))
@select(Title, Budget)
@slice(1:5)
end
Not a character-for-character match to dplyr, but gets much closer than most other attempts!Can't speak to abandonment, but it seems a lot of recent devel is occurring inside the the tidyverse, which is deprecating a whole bunch of other stuff.
What Hadley Wickham has done is very impressive.
The drum about not fitting into data pipelines... if you're literally using a bash pipe its true most R programmers have no idea how to do that. Otherwise, that is where Docker and k8s shine.
On packaging. R's package authority runs tests and ensures that all packages work with the latest version of their peers. The dependency heck is much less deep as a result.
We use R at my employer still because we put statistical data science into production. Our experts come to us comfortable with R. Reimplementation would be absurd.
Typically those applications are not the sort of line-of-business enhancements ML in Python is more tuned to. I.e. recommender systems, NN models, and so on.
- R Markdown is just great for static reports. We use PowerBI or ArcGIS for interactive stuff.
- GIS is a breeze. My work provides licenses for ArcGIS, which has a Python library for scripting. Despite that, it is so much easier to do stuff in R, which can read and create ArcGIS shapefiles.
- Exploratory data analysis is easy. Often, before meetings, I'll connect to the database in R and make a few basic tables. Then I can query, aggregate, or plot data sitting the meeting. I have custom ggplot themes in a package, so even my happy hastily created plots look nice.
- RStudio is amazing. What it lacks in editing tricks, it more than makes up for in simplifying R-specific tasks. Showing plots is automatic, rendering and viewing markdown reports (of any type) is two buttons, testing and building a package are each two buttons.
- I spent a lot of time evangelizing R (team-wide presentations, being the "R guy" for troubleshooting, organizing an R User Group with members from different teams, creating an internal package repository). Some became happy converts, the rest begrudgingly accepted it as a tool we would use. I don't know if I could do it again with another language.
I'll admit my work doesn't get incorporated into pipelines. We get the data, analyze it, create reports, and share the reports by email or on our public website. The statisticians are segregated from the developers here. State government resists change, especially role changes that don't match grants' or laws' wording.
Quarto (also supporting R) is a good replacement for rmarkdown (with a saner syntax) and I say this as someone who has extensively used rmarkdown over the years.
As a "bilingual" R & Python user, I've found this to be true for the latter language as well :)
I don't have much to add on top of what other useRs have mentioned, except another testimonial that our company has successfully used R in production for 6+ years, from data "pipeline" stuff you mentioned to dozens upon dozens of predictive models of varying complexities.
When faced with a new data analysis ask, 99%+ of the time I reach for R (although without the tidyverse, that number would be much lower). Like another commenter said, the ease by which you can plot in R blows Python away. Seaborn seems like a decent compromise in my limited experience, but plotting in "base" matplotlib makes me want to die.
The ecosystem is simply better. The folks who maintain CRAN do a fantastic job. I can’t remember the last time a library incompatibility led to a show stopper. This is a weekly occurrence in Python.
Oh, it’s very common unless you basically only use < 5 packages that are completely stable and no longer actively developed: packages break backwards compatibility all the time, in small and in big ways, and version pinning in R categorically does not work as well as in Python, despite all the issues with the latter. People joke about the complex packaging ecosystem in Python but at least there is such a thing. R has no equivalent. In Python, if you have a versioned lockfile, anybody can redeploy your code unless a system dependency broke. In R, even with an ‘renv’ lockfile, installing the correct packages version is a crapshoot, and will frequently fail. Don’t get me wrong, ‘renv’ has made things much better (and ‘rig’ and PPM also help in small but important ways). But it’s still dire. At work we are facing these issues every other week on some code base.
I think that most problems are ultimately caused by the fact that R packages cannot really declare versioned dependencies (most packages only declare `>=` dependency, even though they could also give upper bounds [1]; and that is woefully insufficient), and installing a package’s dependencies will (almost?) always install the latest versions, which may be incompatible with other packages. But at any rate ‘renv’ currently seems to ignore upper bounds: e.g. if I specify `Imports: dplyr (>= 0.8), dplyr (< 1.0)` it will blithely install v1.1.3.
The single one thing that causes most issues for us at work is a binary package compilation issue: the `configure` file for ‘httpuv’ clashes with our environment configuration, which is based on Gentoo Prefix and environment modules. Even though the `configure` file doesn’t hard-code any paths, it consistently finds the wrong paths for some system dependencies (including autotools). According to the system administrators of our compute cluster this is a bug in ‘httpuv’ (I don’t understand the details, and the configuration files look superficially correct to me, but I haven’t tried debugging them in detail, due to their complexity). But even if it were fixed, the issue would obviously persist for ‘renv’ projects requiring old versions.
(We are in the process of introducing a shared ‘renv’ package cache; once that’s done, the particular issue with ‘httpuv’ will be alleviated, since we can manually add precompiled versions of ‘httpuv’, built using our workaround, to that cache.)
Another issue is that ‘renv’ attempts to infer dependencies rather than having the user declare them explicitly (a la pyproject.toml dependencies), and this is inherently error-prone. I know this behaviour can be changed via `settings$snapshot.type("explicit")` but I think some of the issues we’re having are exacerbated by this default, since `renv::status()` doesn’t show which ones are direct and which are transitive dependencies.
Lastly, we’ve had to deactivate ‘renv’ sandboxing since our default library is rather beefy and resides on NFS, and initialising the sandbox makes loading ‘renv’ projects prohibitively slow — every R start takes well over a minute. Of course this is really a configuration issue: as far as I am concerned, the default R library should only include base and recommended packages. But it in my experience it is incredibly common for shared compute environments to push lots of packages into the default library. :-(
---
[1] R-exts: “A package or ‘R’ can appear more than once in the ‘Depends’ field, for example to give upper and lower bounds on acceptable versions.”
For those whom want to use both R/python, I have notes on using conda for R environments, https://andrewpwheeler.com/2022/04/08/managing-r-environment....
It's a bit of faff but that seems like it should work (but maybe I'm missing something).
As a SWE I much rather inherit and maintain R services than Python services.
Take something like adegenet, where the manual itself is approaching 200 pages:
https://cran.r-project.org/web/packages/adegenet/adegenet.pd...
It’s a language which feels like it has a lot of magical incantations you need to remember - the default namespace is much more crowded. Functions like sapply vs mapply are tricky to reason about from the documentation alone. The values NA vs Null vs integer(0) are all used as standins for real thrown errors and knowing which one to check for after calling a function can be tough.
But after using it for a few hundred hours to do data processing and statistical regression it’s hard to imagine python or Julia being faster to use. But in all honesty for the pharmaceutical industry it’s mostly momentum that keeps R on top same reason they use a lot of FORTRAN90.
I can’t agree with this: especially in PK/PD, R is only just now taking over from the previous (closed-source) systems. Momentum would keep R out, not in.
Could you please expand on that? It's unclear what you're referring to.
> The values NA vs Null vs integer(0) are all used as standins for real thrown errors and knowing which one to check for after calling a function can be tough.
`checkmate::assert_numeric()` (or similar)
with base R you want isTRUE():
`stopifnot(isTRUE(is.finite(x)))` (or is.na or anything else) will error on empty values.
Traditional stats
Very fast iteration for data exploration in REPL (vs code or R studio).
Prefer pipeline workflows (Tidyverse/maggrittr).
Prefer functional
Prefer array based.
Prefer 1-indexed arrays (yes there are some of us).
I always understood that it was its open source nature being a successor to the S language that gave it traction
Python could be so much better with some minor syntax extensions.
What about R's language makes it better for Repl driven development?
For anything else, we use Python.
1) tidyverse makes prodding and plotting my data faster and more enjoyable. when I am prototyping a model I'll sometimes do the groundwork in R and then migrate the production version to python
2) I can't seem to write data wrangling code in py that is as aesthetically pleasing and easy to reinterpret later. could just be that I started in R, but while the methods in pandas "work" I don't always totally understand why they work the way they do. with tidy it works the way I expect and feels easier to read back and iterate on
R is so much better than python in many areas concerning data pipelines: connecting with external database systems through an unified API, superior data munging utilities, as well as plotting, a more comprehensive (obviously) statistical analysis toolset.
I even find rmarkdown vastly superior to jupyter.
But IMO the best reason to use R rather tha python is that its tools will make you approach the problem as a statistician rather than a programmer.
if you are doing Bayesian stats, fitting hierarchical models, or using Stan in any serious capacity, R/Stan is so much more ergonomic than Pystan. Here’s a long list of pros-cons:
https://discourse.mc-stan.org/t/various-observations-on-rsta...
At my FAANG company, there are teams that use it for econometrics. I think that’s Rs sweet spot, still in 2024.
However - I'd love to learn your ways; specifically - what are your best recommendations for python over R?
Specifically, even though my R skills are weak - I think that RStudio is pretty darn amazing - what do you recommend over Rstudio?
I'd truly like to hear what a good toolbox looks like from your perspective these days (especially now this little GPT toddler is bonking into everything in my domain)
Still the best replacement for EDA and reproducible analysis that used to be done in Excel.
and package management is much, much more reliable in R than in python.
I consider it a bad relic of the 70's. It doesn't have a "learning curve" -- it has a "learning straight line." Even when you're experienced and semi-competent at it, it's still difficult and surprising.
If you have to use R, use the tidyverse.
I like R and use it often as it find it more concise to work with than Python for simple statistical purposes. I forced myself to use R instead of spreadsheets and don't regret it.
This is one the reasons why (thanks, Zed Shaw) https://web.archive.org/web/20110702162929/https://zedshaw.c...
For example SAS makes R look beautiful and consistent. And that’s more a comment on SAS than R. And this isn’t to say python is perfect either, but I prefer it.
I've got a decade in with Python numeric computing, and I'm interested in Julia and all of the cutting-edge stuff.
I've only dabbled with R until now, and I haven't researched it enough to know if rumors of it's inevitable demise have any substance.
There are a lot of interesting math problems other than training gigantic neural networks on NVIDIA gear, and I've got some Computer Algebra System / ergonomic linear modeling needs on a current project:
I need the best tool for someone who is messing with Black-Scholes type stuff, who is still building the fidelity with tricky antiderivatives by hand, but I have enough fundamentals to check the computer's work.
What role should R play here?
So, if I wanted to dabble I'd easily use R and if I was in the quant developer world I'd be doing C/C++
Something like intraday momentum/sector rotations can easily be done entirely in Python/R, from what I've seen.
In your other comment, you said you are looking to price "weird derivatives". How weird are we talking? If its OTC I won't be able to help anyway, if its standard then I can at least try to point you in the right direction. The fact you mention Black Scholes makes me think it might be something closer to "vanilla" than the other way around.
I have a vague intuition that transaction costs will be sort of cumulatively symmetric: participants who get in quickly will pay a lot per unit time, but conversely, people who VWAP in will get zero-rated on the way out.
There’s a legitimate underlying switching cost, there’s a stability premium thereby, making that equitable for all participants is an interesting problem.
The reality is closer to an option on an FX forward, with a very nasty empirical MC as Q* for the payoff equivalence.
I’m not fancy enough, I know when to sub-contract!
If what you have is FX-like I wouldn't be able to help beyond that anyway, FX modelling is its own thing and I haven't done anything there since the obligatory uni courses(in equity space myself). AFAIK the general way to do things in rates/FX is SABR for vanilla and then PDE/MonteCarlo for exotics, but I was never on an FX desk so don't want to point you in the wrong direction.
But your reminder to think of SABR/implied-vol is useful: I think there’s a convexity argument that can be made around how fat the tails would need to be.
I’m not sure anyone is going to be thrilled at “anywhere between one hundred dollars and one hundred million dollars”, but my job is to figure out the bounds.
You do any consulting on non-adjacent areas.
+ extra points for using quarto
[0]: https://gist.github.com/mine-cetinkaya-rundel/03d7516dea1e5f...
I don't know what R does now, but this was a deal breaker for me at the time because I was dealing with really large integers that regularly broke this limit.
A lot of the stuff above was complaining about issues where Python is a lot worse than R, about non-issues or with a fundamental misunderstanding of the language. I'd given up hope of seeing a real weakness named as such :)
There is bit64 and doubles being used as 53bit pseudo-integers - but if I needed 64bit integers, R wouldn't be my first choice, definitely.
Of course that doesn't make it a DSL. It simply means that R was designed with a particular application in mind. So was Perl (regular expressions as first-class citizens) or Javascript (DOM manipulation). Not to mention PHP.
"R Programming for Data Science" https://leanpub.com/rprogramming