Announcing RStudio v1.0
blog.rstudio.org
blog.rstudio.org
Most of our students use RStudio for their work. It's convenient and easy. For developing standalone scripts or functions, rather than notebooks or R Markdown files, the typical workflow is to write code in a file, then run it in the current R session by selecting pieces and hitting "Run".
But this encourages terrible practices. After a few iterations, the current environment does not reflect what's written in the file. If students write tests in a separate file, they neglect to source in their functions, because they're already in the workspace and so the tests run fine. Datasets that were loaded in the R console are used in code without being explicitly loaded there. Code gets changed without being run, or variable definitions are changed but old copies in the workspace accidentally used instead.
We end up getting many homework submissions that simply don't run if you start them in a new R session. Pieces are missing, code is out of order, tests only run if you manually select the test code and run it. It only worked in the R session of the original author.
R Markdown is a decent step, since rendering the HTML should start a session from scratch, but when we're asking them to write well-tested and modular algorithmic code, using R Markdown doesn't really fit.
I've begun to appreciate DrRacket's approach, where there is a "Definitions" pane and an "Interactions" window. When you Run the definitions, the current workspace is blown away and everything is defined from scratch from the Definitions, so there's no lingering state from REPL interactions. (Unfortunately your REPL history is lost, which can be annoying.) You can't run into the same inconsistent state as RStudio actively encourages.
As someone who had been a student in statistics for more than 10 years, I confess I had never written a single test for my homework. Frankly I just didn't have the time or interest (too much homework, and becoming a professional software engineer was not the goal of the homework assignments). That said, when I put on my software engineer hat now at work, I'd definitely do what you advertise here and write tests carefully. If you want your students to enjoy the benefits of both R packages and R Markdown, I wrote some thoughts here a couple of years ago: http://yihui.name/rlp/
Don't get me wrong. I'll all for teaching students good practice of software engineering. I just want to speak from my own memories and experience as a student. Sometimes I feel teachers are like parents: they want kids to learn all possible right things, no matter if they are practically able to swallow all the good stuff (sometimes this has bad psychological consequences, like rebellious children). If I were an instructor in statistics, I'd only require students to submit an R Markdown document. Other things like tests can earn extra credits but not required.
It's the nature of learners to make mistakes. The more things they have to remember to do, the less cognitive power they'll have to focus on what they're trying to learn.
Perhaps one way to make a teachable moment of it is to help them set up a baby CI environment. Then every time it catches something the value of good practices is driven home.
This would be a problem in a straight-line script where you're doing a bunch of data munging and analysis, since reloading from scratch might mean redoing expensive computations. But if you're building clean, reusable functions to implement interesting algorithms -- building a package and not a script -- then it's exactly the behavior you want.
Our course very much focuses on software engineering. Major topics include writing modular code, object-oriented design, thorough testing, and version control. We don't cover statistics concepts in the class -- it's computing for statisticians, not computational methods in statistics. We believe that teaching statisticians to compute like software engineers will, in the long term, dramatically improve their work, since they'll have a stable base of robust, modular, well-tested, reusable code.
One recent project, for example, required students to write a pipeline of scripts: one script takes the name of a CSV file as a command-line argument, processes and filters the data, and dumps it on STDOUT so the next script can read from STDIN and load the data into PostgreSQL, so another script (an R Markdown document) can do some queries and generate an automated report on the new batch of data. The processing and analysis stages have to be written as functions, not just top-level scripts, so they can be thoroughly tested.
A future project will involve using dual k-d trees for fast approximate kernel density estimation, or building R trees to efficiently query spatial data. These are definitely more like packages than scripts.
We take reproducibility very seriously. The fact that RStudio's Knit button uses a new R session, instead of the current R session, to compile R Markdown documents was a deliberate choice to make sure your output is produced from a clean R session. But if you are doing EDA, it may not be very pleasant to click this button over and over again every time you update your code (you can if you want).
If your course is focused on software engineering, everything you said makes perfect sense. Statisticians can learn the good principles in CS, but they are statisticians after all. There must be tradeoffs.
My difficulty is more that for most of these students, this is their first time using R or a real programming language. I had hoped that R Markdown would make the process simpler and more sensible. But the tendency of a student is to try different things, and not have a strong understanding of the logic behind those processes. So they create R Markdown files, and then gripe about how the code works in their IDE, but doesn't knit, and so knitr must be broken!
I mentor learning data scientists and my advice is always to start using RMarkdown as soon they're remotely comfortable with RStudio. Not only does it avoid issues with an easily polluted global namespace, but more importantly encourages literate programming from the early stages. In stats/data science literate programming is vital to having any idea what you were working on a few months ago. It also makes writing reports much, much easier.
RStudio makes it pretty easy to put together R packages, and the package structure for R does a great job of enforcing proper documentation and testing. Sourcing R files should primarily be used to quickly play around with ideas, or for exploratory data analysis that doesn't fit well inside an RMarkdown document. Any code you intended on reusing between projects should end up in a local package.
I do think it's a problem that R has no intermediate method of organizing code like simple modules in Python. But this means if you're serious about writing clean R, you just have to bite the bullet and teach students to write packages.
But in general you really should run it from the top from time to time and learn best practices for error handling and such.
http://stackoverflow.com/questions/480083/what-is-a-lisp-ima...
One of the reasons I switched to using Jupyter over R/RStudio directly was the native rendering of notebooks on GitHub, which made it pretty (example of mine: https://github.com/minimaxir/stack-overflow-survey/blob/mast...)
The addition of native R notebooks may make me switch back, although I'll have to experiment on the differences between Jupyter Notebook rendering and .Rmd rendering on GitHub. (and since the notebooks are theoretically language agnostic, it might be fun to experiment with Python code too!)
Native sparklyr is something I'll also have to research/experiment with, since according to the official Spark documentation, although R has first class support with Spark, there is not API parity with Python + Spark, for example. (although, sparklyr has most of the important transformers/models so it is definitely worth a look: http://spark.rstudio.com/mllib.html)
I'd been waiting for the official release, instead of preview because of some issues, the official release seems to have ironed out the issues.
Excited to code directly into notebooks as reproducible code for fellow workers.
Sparklyr is really also changing the game for me in how I can integrate new users into R. It used to take a little work to get people off SAS EG or whatever statistics package they looked for. Not so much anymore.
Not a bad thing, it's just R Notebooks kind of seem like old news to me.
Addendum: I looked it up and it appears that Jupyter Notebooks actually have an R kernel (https://irkernel.github.io/).
There wasn't migration to python because there's a total lack of anything to migrate to in python. It has only textbook examples as far as statistical models go. That's around 10% coverage and pretty much 0 coverage for all the abstractions and utility packages in R. Pretty much the same with machine learning (minus neural networks where it is strong). Python took some share from matlab and data processing tools but I don't see it retaining that. IMO it will lose it to Julia for the engineering & science people and to Scala for the data processing people.
The "second best language for everything" doesn't really work too well in the "data science" world.
Most of the nice stuff from RStudio is making its way into https://github.com/jupyterlab/jupyterlab which is starting to look nice.
I am really excited about using R Notebooks.
R Notebooks allow you to prepare something for presentation, not just for running/testing. I.e.: You can run code and then choose to present the code and the output, just the code, just the output or none of them (for instance just putting a sentence in the presentation where you say you applied some mathematical function to the data, but show the mathematica function (latex) instead of the code.
This is really important to publish and to present data and it's not possible to do in Jupyter/iPhython.
Also, it allows you to export directly to HTML/PDF. Something Jupyter doesn't.
R and RStudio just installed and worked for 3 of the 4 companies. The 4th required a tweak to one environmental variable and everything installed/worked after that.
Corporate IT restrictions can make or break software.
[1] https://en.wikipedia.org/wiki/Affero_General_Public_License
It's certainly not our intention to deceive and then submarine our customers, and I sincerely hope that's not what you were implying. IANAL but RStudio users have nothing to fear from the AGPL, as the copyleft provisions are for derivative works of RStudio itself.
If OTOH someone is trying to build an R editor interface for their commercial SaaS data science startup, and want to leverage our code to do it, then yeah--the AGPL is going to apply and if that's a problem then we try to work something out.
(BTW I'm a fan of the work you're doing with Julia!)
Congrats on the release! RStudio pushes the envelope on very many data analytics UI features. Excellent work.
While I am a big fan of Julia, in my opinion it has a long way to go to be ready for our version of prime time.
Python on the other hand just doesn't have the breath of statistical models that R has. Generalized Additive Models is just one example.
Did you mean to use "may" instead of "will"? The latter would seem to indicate that the R-Studio makers actually plan to do this.
Agreed. However, the R programming language and many of its libraries are GPL licensed.
Can someone elaborate how R can be legally used in to develop proprietary or commercial software? RStudio's website [1] lists all of this large companies that apparently use R.
What you cannot do, however, is modify GPL software, as you're then creating a derivative work which must legally also bear the GPL license.
If you create a program that uses any of the GPL licensed R libraries, that program must also be GPL licensed [1].
Additionally, the GPL license prohibits using GPL licensed software in proprietary systems [2].
[1]: https://www.gnu.org/licenses/gpl-faq.html#IfLibraryIsGPL
[2]: https://www.gnu.org/licenses/gpl-faq.html#GPLInProprietarySy...
AFAIK RStudio is now being lead by JJ Allaire, same person who did Coldfusion and stuff. Also in there is Hadley Wickham, of ggplot2 / dplyr fame.
I dont see us moving back anytime soon, because production code in python is orders of magnitude better than R.
I just wish there was a decent dplyr for python though :(
It explains it better than i can
Oh and it's fast as shit (the tight loops are in C++).
"DSL for figuring out things aa quick as possible " is a religious statement. I suggest you spend some time on the Python side to understand how good it is.
However I can quantify some of what you said - that tge ecosystem of analytics libraries is bigger in CRAN. I agree with that... however at this point, the only library I really miss is dplyr.
Markdown is extremely limited (what we need is something more like knitr), there are quite a few bugs and it's difficult to access documentation of the packages.
I appreciate the work of the author, but all in all, it's nowhere near RStudio and - except for the fancy interface - actually behind Spyder.
It would be great if Spyder would add a plugin to edit these formats or something similar.
A modern interface would also be nice.
And native support for Vim (I know there is a plugin but it's quite limited).
But mostly, the mardown/knitr part is the big thing missing for data exploration/presentation.
How Spyder looks in each platform depends on Qt, but we have some additional customizations to make it look better at least on macOS.
We'll try to add something like knitr as a third-party plugin in the future. We didn't know it's so important to have it.
Also, since you mention you didn't had idea about knitr importance: It's normally one of the strong points presented when discussing Python vs R for data science.
Knitr makes it much easier to present (and now the RStudio notebooks make it also much easier to explore data) your work and to adopt a literate programming approach since you don't have to keep a Jupyter Notebook and then hack a Markdown/Latex document on the side where you are inserting just some plots and just some code and just some output from your Jypyter notebook that's actually relevant for the presentation or publication you are doing (and then having to remember to change parts of it whenever you change something in the code and the output changes).
But you're probably aware that you can also generate PDFs from Jupyter notebooks. So I guess the advantage of knitr over notebooks is the ability to easily version control markdown docs.
Are there other advantages you'd like to share?
But the big selling point is that you can choose which parts of your notebook to include in the HTML/PDF output. Input, output, plots, you can write a document for a scientific paper, presentation or even a book and not have to show all the code, or all the output or all the plots like you must do with iPython. That's the really strong point with knitr.
There have been long discussions about hiding output or code from Jupyter notebooks throughout the years but (unfortunately) they haven't derived on a final decision from the Jupyter team.
In any case, your feedback is a very strong motivation to start working on this in Spyder. Thanks a lot for it!
We have several things planned already for the next six months, but we'll try to have an initial implementation of a knitr equivalent before the summer of next year :)
I would still suggest that you would look into knitr before, since it already supports other languages besides R (although I can't understand if so extensively as R) so perhaps it would be easier to just build from that.