RStudio: Integrated development environment (IDE) for R
github.com
github.com
VS Code/Python has made some major improvements in the past couple years but it’s still very clunky compared to the ease of running R code line by line without having to start up a debug instance. And now with copilot the most frustrating parts of R (such as remembering all the Tidyverse syntax) have been abstracted away.
There is something to be said about running and processing large CSVs and keeping that in memory while running other parts of the program as well as having clickable access to all the dataframes loaded into memory.
The problem with a lot of end-user R code is that it is written by statisticians, not programmers. They'd write the same garbage and huge scripts in Python (trust me, I know).
[0] http://adv-r.had.co.nz/Functional-programming.html
I have to deal with getting code from data scientists into production, and simply getting it to run outside of their mutant local environment can take days. Things are starting to get a bit better with packrat initially and now renv/pak/rig and the like, but most DS haven't heard of them, and major breakages between minor library versions are still commonplace, as are undocumeted system library dependencies. Then there is the whole stringsAsFactors nightmare, thankfully slowly on its way out but still around causing occasional catastrophic breakage.
There are lots of nice things about R, but it makes it very easy to shoot yourself in the foot.
All that said, I still greatly prefer it over Python for DS work.
I've had exactly the opposite experience. For R, I download R and install it, and download Rstudio and install it. Then when I need a new package I just install.packages("coolnewpackage") and it just works (TM). Occasionally I get info messages about packages being built in newer versions of R, and once a year or so I eventually get around to looking up how to use the updateR() function, but in five years of doing biostats in R I can't remember a single time I had a dependency issue.
Python, on the other hand, is a nightmare. Conda makes life a lot easier, but it is not easy to learn if you are not a software engineer (remember, R was made not just by statisticians, but for them as well). For many projects, my Python flow was something like...
Try creating a new conda env with the packages I think I need. Try starting the project, oops I don't have spyder-kernels installed. Oh, and my environment isn't compatible with it. How about just running it in VScode? Well now I don't have my variable explorer. How about Jupyter? How do I get Jupyter to find my conda env again? Oh wait I need this other library it's only on conda-forge, and then the conda environment solver fails. I guess I'll start from scratch with a new conda env, and maybe after several trial-and-error sessions of carefully composing the correct "conda create -n ..." incantation in a text editor before copy-pasting them to the command line, I might get the environment I need up and running, after conda finishes its 10-minute compatibility search and downloads 80 GB of python libraries.
And using conda is the easy way of doing it! Don't even get me started on pip and venv...
At work, I've been giving out about pip to one of our DEs for a while, and when he needed to upgrade a bunch of DS packages he finally started coming around to my opinion.
It doesn't have modules or namespaces, and the current fashion is for packages to use non-standard evaluation which adds friction to user's writing their own functions.
Note for many R packages, the NAMESPACE file is autogenerated from roxygen docs: https://cran.r-project.org/web/packages/roxygen2/vignettes/n...
Which are all dumped into the one single global namespace regardless if you want everything or not.
I can't remember the exact number, but tidyverse package imports literally thousands of things into your global namespace on package load, coupled with any other dependencies and you have a hell of a time figuring out where any function or constant came from.
I guess I should take offense as a statistician. But its a fairly common complaint. The reality is, most of us statisticians are trying to compute a result. Like once. Or sometimes twice. For a paper. Or a task. If someone comes to me with a time series and asks me to test it for stationarity, or find the p lags to make it MA(p) stationary, they aren't asking me to write a program. The goal is not reproducibility. The goal is a fast answer. I've used R at trading desks & financial institutions - the goal has seldom been "run the same program again, but with this new input". If that was the case, I would write a function & stick it in a nice library with documentation. But these aren't tech firms. We aren't shipping software. The goal is to compute something fast so you can get on with life & make the trade, or draft the next paragraph in your paper, or... Like if they give me a set of bespoke mortgages with some hairy constraints & ask me to compute the value at risk, there is not much point in building some VaR function. Because its a once in a while thing. Next time it will involve a different set of args & they'd be different constraints & so forth. So just write some 10 line script & get the number & move on. Yeah, sometimes I would stash the script in some repo & write a 1-line comment on how it works - but its kinda pointless, it doesn't get much play/reuse. We aren't programmers in that sense, we are just trying to solve problems.
My kid knocked on my office door yesterday. He's in some AoPs course where they use generating functions to count stuff. So he had a problem about the number of ways to add three odd numbers to make 1001. He had worked out the algebra & gotten some number, but before he hits Submit, he wants to doublecheck with me because wrong answers have a penalty. Now, I don't have the time to go back to school and learn what is a generating function. And I don't want to write lots of for loops & if statements & fight with syntax errors & so forth. So my 1-liner in R
dim(subset(expand.grid(a=seq(1,1001,2), b=seq(1,1001,2), c=seq(1,1001,2)), a+b+c==1001))
tells me there are 125250 ways. He says he got the same number with generating functions. Boom done! So that's what R is for. Quick & easy.
I recently had to inherit someone's R stuff and I had to learn R and fix it all. It now runs from a makefile repeatably.
Anyway it could be worse. It could be Minitab.
That's not really RStudio's fault. It is just how many people use R and were taught.
> code is run out-of-order which makes the code organization and flow of a program a complete disaster.
In my experience, with R Markdown, this is untrue. I see Jupyter Notebooks with cells run out of order much more often.
That's more a REPL issue than specific to a particular language. It's the tradeoff you make. I write my R programs in Geany and then run the whole thing using Rscript. That gives me a clean environment on every run.
Is there a good demo or video you can point to that shows this? I have no experience with R, RStudio, or data science, but you've piqued my interest.
Just open a .py file, then select the snippet of code you want to run and cmd+enter
It will open a new REPL for you (using your selected interpreter) the first time, and after that all commands are run in that same one.
The python interactive window has pretty much fully replaced my use of jupyter, since it gives you notebook-style output without the annoyance of the notebook format. My usual workflow is highlighting lines of code and shift-enter to execute (there's also a cells syntax).
I'm surprised by this because it _is_ possible to use R in Jupyter (although I never really liked the experience, R Studio was far superior).
Yes it does.
The support for R looks a bit different (to me at least?): https://code.visualstudio.com/docs/languages/r
In the screenshot the window on the right does not look comparable to the output in a jupyter notebook. It looks more like a standard terminal. e.g. does it support interactive charts, html tables etc?
The Python interactive window uses the ipykernel package to allow rich outputs like that.
I still might be wrong and would like to be corrected on this, since it would mean R support in VS Code is now better than I thought (I haven't tried it fora. while)
See my other comment in the main thread with more info.
I composed almost all my homeworks in grad school using RMarkdown in RStudio. You get LaTeX whenever you need it, code (I usually use it for R or Julia), and markdown for ordinary text. The kable function renders tables nicely from data frames and ggplot2 creates beautiful plots.
Mathematica and Jupyter have a few advantages, but overall I'm very happy with RStudio.
For my uses, it replaced RStudio 100% of the time.
For knitting, you can use Markdown image links.
There isn't much advantage to using it over RMarkdown for R, IMO.
https://posit.co/pricing/individual-products/
If you want a Rstudio server to host for a research group containing more than 5 people, talk to their sales Rep.
Otherwise each person will need to host their own Rstudio server side-by-side on the same machine.
Jupyter and JupyterHub is the way forward.
Especially if they get multi-kernel notebooks mainlined (read: what Org-Mode has been doing for decades)
Between the standalone desktop app, and the convenience of running JypyterLab in the cloud thanks to https://mybinder.org/ links, there is now a smooth path for beginners getting into stats/ML/data science: (1) read notebook on github or nbviewer, (2) run notebooks in the cloud via mybinder links, (3) install JupyterLab Desktop app, (4) learn to install Python+env-manager via command line. Previously, new learners were forced to jump straight to (4), but now there are logical steps along the way!
[1] https://github.com/jupyterlab/jupyterlab-desktop?tab=readme-...
[2] https://blog.jupyter.org/jupyterlab-desktop-app-now-availabl...
Yeah for sure when I use RStudio it seems much more polished, but I guess my attachment to (and comfort with) Python still makes it worthwhile to use JupterLab rather than switch to RStudio.
There are a couple of things I would love for the R ecosystem: project scaffolding to do bulk data generation (e.g., from continuously generated data sets). What's the best way to do this: makefiles, or what? I have a relatively short entrypoint R file that sources other leaf files to run specific analyses, but it makes the software engineer inside of me want to curl up and die.
Rstudio is nice but lacks a lot of nice things from something bigger.