One Year with R
github.com
github.com
To which I mean R is a highly optimized, well-oiled machine if you're using it for its highly-optimized, well-oiled purposes. I tend to have notebooks full of tiny fragments like this
dat_min %>%
group_by(ymd = make_date(year(date), month(date), day(date))) %>%
summarize(vol_btc=sum(vol_btc), vol_usdt=sum(vol_usdt), tradecount=sum(tradecount)) %>%
ungroup() %>%
pivot_longer(cols=c(-ymd)) %>%
ggplot(aes(ymd, value)) +
geom_line() +
facet_grid(name ~ ., scales="free_y")
It's madness if you're not familiar with the tidyverse, but 3 dozen fragments like this is enough to eviscerate a fresh data set. Almost any question you can dream of is a 3-20 line set of transforms away from a beautiful plot or analysis answering your question. Very notably, this includes some of the finest modeling tools available today.Terseness here is a huge advantage as well because in many data analysis workflows you are rerunning that same 10 line snippet over and over, making small changes, adjusting to eventually visualize the thing you're looking for perfectly. Having all of that in the same small block is ideal.
Finally, for the non-trivial number of folks in this specific scenario, the integration between Stan and R/RStudio is top-notch and makes using both tools very pleasant.
You can replicate all of this in Python, but optimal Python/Jupyter is still a far cry away from R/RStudio for these specific sorts of tasks.
My skin writhes every time I need to type:
table.loc[(table.column > 2) | (table.column2 < 3)].reset_index(drop=True)
when I want to subset a table.
0: https://marketplace.visualstudio.com/items?itemName=ms-tools...
If you know the class of an object, let's say Class, but the object has yet to be "constructed" so that IPython can correctly infer its type, you can type `Class.[TAB]` in IPython and look at its methods.
For example, in Sympy, you have a matrix type called Matrix. You can do `(A * B).diagonalize()`, or alternatively you can do `Matrix.diagonalize(A * B)`, which has some advantages because doing `(A * B).[TAB]` does nothing useful because Python can't infer types.
You can also do the same for modules. `ModuleName.[TAB]`
To be honest though, I found the experience smoother in R for some reason.
`table.query("column > 2 and column2 < 3")`
table.loc[lambda df: df["column"].between(2, 3, inclusive="neither")]
this is useful when your dataframe has a long name, or when you have some long method chain and you need to subset at the end: table.foo().bar().baz().loc[lambda df: ...]it is still more verbose, but I actually prefer always providing column names as strings. it's more explicit. I don't like R's environment-manipulation metaprogramming magic where you can give column names as symbols.
as for resetting the index all the time, this is something of an antipattern. if you set up your index right beforehand it isn't necessary so often.
I would say the most frustrating part about RStudio is that it is a workbook where you can execute code based on your cursor. For my wife, these workbooks become a total mess because things aren't necessarily run sequentially.
1. Always run from the top, using the "run previous chunks" button. When this gets too slow, you know that it's time to think harder about your workflow. For a more extreme version of the same idea, regularly restart R using Ctrl-Shift-0, and run from the top. It'll ensure your code is working right.
2. Have a setup chunk that always gets you to the same state. Make sure every other chunk works directly after calling the setup chunk. Then just alternate between "run setup chunk" and "run current chunk".
> the question is how much code you should write in the notebook, versus having it in a more organized set of functions and libraries. It's very easy to end up with a huge bloated document which contains thousands of lines of spaghetti.
But this isn't correct:
> your options are more or less use a notebook, or copy and paste your results.
What's wrong with writing scripts that write images to disk? That's how millions of academic papers were written before the advent of notebooks. You could use Makefiles if you like, or you could even use a technology such as Sweave to automatically mix images with LaTeX output.
I mean this as politely as possible but the fact that you think that the options are "use a notebook or copy and paste" I think shows that you've caught a notebook mentality disease! The fundamental point I'm trying to make is that you we don't need to do everything interactively from REPLs. REPLs are great for trying things out, but when it comes to producing the images for your paper, those should be produced by scripts, not by commands entered into a REPL, or notebook. An those scripts should evolve via version control, which is the basis of evolving any good and correct software. And the scripts for producing images for a paper should be good and correct software.
The point is that it makes sense to mix english prose + code to e.g. produce tables or graphs, even if most of the heavy lifting is done separately in code files.
Yep, so what you say makes sense. Isn't it sometimes a bit overly prescriptive to assume all collaborators use Rmarkdown? (Perhaps not! I used to work in biology and statistics and R was very ubiquitous.)
I discovered "How To Design Programs" somewhere late in my first year of using R. Like most beginning R coders with nominal experience in other languages, I wrote a lot of monolithic scripts in a very imperative style. HtDP gave me a mental framework for decomposing larger problems into bite-sized chunks. The lispy roots of R lent itself particularly well to the model of thinking presented in that book.
Ever since then, I've pined for the graphing calculator parts in a more modern Scheme. When ggplot and then the tidyverse (neé hadleyverse) came on the scene, I was even more convinced that Scheme, especially Racket, was the ideal future for data science. If R could support a large ecosystem like tidyverse, just imagine what the metaprogramming facilities of Racket could do!
But I think those graphing calculator parts are hard to reproduce. Attempts to clone ggplot2 fall short year after year, because most other languages don't have grid graphics to build on top of. R is a deep ecosystem on "an OK scheme," which is damned hard to beat.
Aside: my first year with R, was in an urban planning masters program and I was terrified of my first big kid statistics course (taught in SPSS). I decided I'd give myself bonus work by learning R. While it was absurd to be doing my stats homework in SPSS, then R, then reviewing HtDP on top of the rest of my course load, I did ace that stats course. :-)
My first reaction, is "why not on a modern Scheme as opposed to Common Lisp" and in so thinking, I have demonstrated exactly why no lisp / scheme has ever achieved critical mass :-)
This hits home for me. We are just starting to use R for risk modeling where I work. R, more than any language I've ever used, makes me appreciate "worse is better". From a theoretical "aesthetic" perspective R is a mess. Yet for data processing all those theoretical concerns don't matter. It just works.
It's honestly kind of humbling that something so theoretically messy can be so practically coherent. It makes me question my assumptions about simplicity.
That said I still run into trouble with package deprecations. I was trying to install the optmatch package (deprecated but still used by causal inference packages) and had a really tough time getting it to compile on macOS.
When I think back to the era you're describing what I recall was people winging around hacky scripts being the norm regardless of their environment. While still not something I'd think of as software engineering best practices, what I see now is less Wild West.
I think R 3.0 introduced namespaces which fixed a lot of the really crazy stuff.
Also, I was writing Sweave in 2010 for my thesis, and I definitely wasn't alone.
Really there wasn't a lot of thought put into the language. We figure that if it ends up being a total failure, we can just pivot.
library(tidyverse)
library(scales)
download.file(url = "https://api.coronavirus.data.gov.uk/v2/data?areaType=overview&metric=covidOccupiedMVBeds&metric=newAdmissions&metric=newCasesBySpecimenDate&metric=newDeaths28DaysByDeathDate&metric=newPeopleReceivingFirstDose&format=csv", destfile = "./data.csv", method = "wget")
read_csv("./data.csv") %>%
pivot_longer(names_to = "Data", cols = c(newCasesBySpecimenDate,
covidOccupiedMVBeds,
newAdmissions,
newDeaths28DaysByDeathDate)) %>%
mutate(Data = factor(Data)) %>%
mutate(Data = recode_factor(Data, newCasesBySpecimenDate = "New Cases",
newAdmissions = "Admissions",
newDeaths28DaysByDeathDate = "Deaths",
covidOccupiedMVBeds = "Ventilated")) %>%
ggplot(aes(y = value, x = date, colour = Data))+
geom_point(size = 1, colour = "gray", alpha = 0.6)+
geom_smooth(type = "LOESS", span = 0.1)+
labs(y = "Daily rate", x = "Date", colour = "UK COVID-19")+
scale_x_date(date_breaks = "months", date_labels = "%b-%y")+
scale_y_log10(labels = comma(10 ^ (0:5),
accuracy = 1),
breaks = 10 ^ (0:5))+
theme(axis.text.x = element_text(angle = 45, hjust = 1))Unfortunately, it’s also a lot of work. In this case, you’ve posted an intermediate stage artifact from R. If one of the many Python programmers reading this want to produce a comparable artifact they need to understand or run that code. That alone reduces your likelihood of getting any substantial replies.
Maybe add a link to an image of the resulting plot?
Yes, I love these things and very curious to see what hackernews comes up with. Project Euler was an eye opener for how things could be optimised in different languages.
import pandas as pd
import matplotlib.pyplot as plt
import matplotlib.dates as mdates
import seaborn as sn
df = (pd.read_csv("/tmp/overview_2022-03-21.csv") # i just used curl beforehand
.assign(date=lambda x: pd.to_datetime(x["date"]))
.set_index("date")
.melt(value_vars=[
"newCasesBySpecimenDate",
"covidOccupiedMVBeds",
"newAdmissions",
"newDeaths28DaysByDeathDate"],
var_name="Data", ignore_index=False)
.assign(Data=lambda x: x["Data"].replace({
"newCasesBySpecimenDate": "New Cases",
"newAdmissions": "Admissions",
"newDeaths28DaysByDeathDate": "Deaths",
"covidOccupiedMVBeds": "Ventilated"
}))
)
ax = sn.scatterplot(data=df, x=df.index, y=df["value"], hue="Data")
ax.set(xlabel="Date", ylabel="Daily rate", yscale="log")
ax.xaxis.set_major_formatter(mdates.DateFormatter("%b"))
plt.show()
I spend 2 minutes on the pandas part and 20 minutes on the plotting part, which really says it all. Seaborn's support for smoothing is really bad and doesn't play nicely with datetimes for some reason, so if I wanted smoothing I'd need to do it myself. And the other stuff I left out requires going into matplotlib's documentation which I don't want to spend time on.pandas is as good or better than R's dataframe manipulation, but R's plotting tools are best in class. I hate all the python plotting libraries.
You don't need the curl -- read_csv works with URLs directly.
The lambda can be replaced by passing parse_dates=[“date”]
I've given up discussing this. My personal opinion is that MPL is the assembly vs ggplots2's Python.
Yes, maybe there are some things that are doable in MPL but not in ggplot2, but never (and I mean - never -) has any one of my colleagues found a single example that I couldnt recreate in ggplot2 with a more readable code that resulted in a better looking plot.
MPL has this Latex like "once you get the hang of it, you will never use anything else in life for anything" and OOP's "If its good code then it is OOP and if its not, then you're not doing correct OOP" mythisque hanging over it. When confronted with bad MPL code that results in bad plots, its always one of those two.
ggplot2 is the best plotting package out there and imo one of the best "end-user" packages in any language. Also, Hadley is a saint.
import CSV
using Chain: @chain
using DataFrames
import Downloads
using Gadfly
using Dates
@chain begin
Downloads.download(
"https://api.coronavirus.data.gov.uk/v2/data?areaType=overview&metric=covidOccupiedMVBeds&metric=newAdmissions&metric=newCasesBySpecimenDate&metric=newDeaths28DaysByDeathDate&metric=newPeopleReceivingFirstDose&format=csv",
)
CSV.File
DataFrame
stack(
[:newCasesBySpecimenDate, :covidOccupiedMVBeds, :newAdmissions, :newDeaths28DaysByDeathDate];
variable_name = :Data,
)
transform(
:Data =>
(
x -> replace(
x,
"newCasesBySpecimenDate" => "NewCases",
"newAdmissions" => "Admissions",
"newDeaths28DaysByDeathDate" => "Deaths",
"covidOccupiedMVBeds" => "Ventilated",
)
) => :Data,
)
subset(:value => ByRow(!ismissing)) # Can't plot Geom.smooth with missings
plot(
_,
x = :date,
y = :value,
colour = :Data,
layer(Geom.smooth(method = :loess, smoothing = 0.1)),
layer(Geom.point),
Scale.y_log10(),
Guide.xlabel("Date"),
Guide.ylabel("Daily rate"),
Guide.xlabel("Angle"),
Guide.colorkey(title = "UK COVID-19"),
)
endIt revealed to me that there's a buglet in `forcats::last()` (https://github.com/tidyverse/forcats/issues/303) and made me wonder if `pivot_longer()` should be able to rename the columns as you pivot them (https://github.com/tidyverse/tidyr/issues/1338)
Renaming factors is one of those things that always seems a bit awkward. I think I've used several methods. Passing a list of named vectors into `levels(x)` allows a many-to-one mapping but was quite dangerous. I've used revalue and mapvalues from plyr. fct_recode is new to me. But yes, renaming while reshaping could be quite convenient. Just looking at fct_recode now, it looks really nice. Seems to support many-to-one and being able to pass it a name vector is very convenient.
Learnt so much today, and this is even before trying out the python and julia examples! Many thanks for this and all your work in R!
The thing about R, for me and many others, is that it's very much an everyday grind language. Especially with Rstudio, its natural domain is as one of "notebook" languages like python, julia, matlab, and mathematica but with a more clear focus towards the tasks of data-analysis. I just tell the BI-tool people that R is excel on 'roids.
R frustrates me a lot, however. But I think the frustration comes out of the fact that when I am using R and get stuck, I am always in the middle of doing something that I need to get done and I don't feel like diving into a long "vignette". Moreover, the documentation is usually too terse and generalized for me to just understand it immediately. Even though I've been using R for years (albeit in fits and starts rather than continuously), there are things about it that I've just never picked up-- I just DON'T KNOW (or care) what F S3 and S4 mean. Unlike the OP, who clearly knows more R than myself, I grit my teeth when I am looking at docs and see the "..." in the arg list.
I suspect that this is part of the heritage from R's beginnings. I once tried to read John Chamber's book but found the presentation complete ass-backwards and impractical for my immediate needs. The Tidyverse has been great, it's far more consistent and ggplot is a kick-ass tool to have in your box. The drawback is that it makes Base-R seem really alien and if you want to be good at R, you have to know more than just the Tidyverse, IMHO.
- It has a bunch of different types of classes, and they all behave differently. Debugging isn't awful, but it's harder than it should be. Also, the documentation isn't clear about which classes to use.
- A lot of the workhorse functions suffer from parameter glut. Despite having different kinds of classes, almost all functions expect plain vectors. Packages like survival show how objects make it easier to read code, reuse data, and validate data. Without the base packages doing it more, everyone's chosen their own systems. The community's been gravitating to organizing "objects" as rows in tables (i.e. tidy).
- The way a function uses an argument might surprisingly change based on other arguments given (e.g., `binom.test`). And then the documentation won't have examples for the different use cases.
- Most users don't have the time or desire to become better R programmers. They have other work to do. For my own work, I write packages with custom classes, functions, and template documents. For collaboration, I keep things very plain and rarely go beyond dplyr; very often, the script goes between two steps executed in a GUI software.
To me, it's the tooling around it. Everything is done in R-studio and it's focus is to generate statistical documents.
The result is a sub optimum solution. It lacks good tooling around installing and running R programs. R programs don't import, they include. It doesn't make it more readable. R-studio is very Emacs like in the sense that it just lacks a decent editor. Due to R-studio being the default, there's not much support for other editors.
But RStudio is amazing... Easily best environment I've used for any programming language (well, except Pharo).
I can even pop up an interactive R command prompt session and do whatever I want in it, even quick ggplot2 graphs. Help shows up just as you would expect, and plots pop up in new windows. RStudio is much less advanced than people think it is, it's really just managing R's windows for you and doing generic IDE work. R is doing all the heavy lifting.
If you're wanting a more "import" like thing you might want to look into making R packages instead of scripts. You don't have to submit them to CRAN, and you can execute them pretty simply on the R command line as well.
That being said, it's really not an OOP or software development tool. It's definitely geared toward data science, but you can automate that very well for generating graphs automatically and reporting for whatever reason needed.
I'd also look into the "knitr" package, which is what all of the Rmarkdown is based around. So for instance, most of my Makefiles are based around a simple command like:
R -e "library(knitr); knit2html('index.Rmd')"
Then I just code using VIM on index.Rmd. You can probably set this up however you like with the R command line.For interactive it literally is just typing "R" in the command prompt. For help, things like "?ggplot", "??knitr", or whatever, so you can open multiple interactive sessions like you were using IPython or something. When you print a plot, it just pops up in a new window.
You can also use "R" to just execute R raw if you are trying to do it without Rmarkdown. I just prefer the HTML output. Pretty sure all the RStudio RMarkdown stuff just calls knitr as well.
The output looks the same as anything on RPubs (and there is a way to publish to RPubs, I used to have to do that at one point), random one from the first page:
In R they default to just vomiting some internal exception often with no context, and I can count the times I have encountered helpful or even seemingly deliberately constructed error messaging in single-place base five. Even kernel development is in some sense better because at least there you are in context and the layers are traversable.
Seemingly one of the primary skills of an R programmer is serving as an informal database of “what the fuck does this mean?” when the issue is something trivially detectable like passing a list vs a vector.
For me it's a pragmatic workhorse tool that I use and aside from frequently getting frustrated with the task at hand, it has never failed me in the end.
I think R is much like a handheld power tool. I have no interest in diving deep into the workings of the tool because when I need to use a drill, for instance, I just need to drill holes and anything else is an annoying distraction (I realize that sounds bad!).
I've also worked with Mathematica and JMP in the past. They're very capable, but not as good at general-purpose data-wrangling as R is today (given Rstudio, knitr, shiny, all the specialized libraries, and most especially the tidyverse).
I found numpy, scipy, pandas, and plotly docs to be quite clear and extensive. The only docs I have found to be confusing are matplotlib's and the Python standard library's. Not sure what packages you are referring to?
Also, the built-in types are documented in one page, going from boolean to sequence types and even type annotations. Every Ctrl + F gives me 20 different results, which is annoying as hell.
It's probably possible at least in theory to structure docs so you have a terse section followed by a verbose section for each thing, but I've yet to develop the discipline or the competence to pull that off remotely regularly.
Maybe in another 15 years.
I should have been more specific, the ... frustration for me comes up mostly in ggplot, Which usually directs you to layer(). Which gets parameter string documentation like:
* geom - The geometric object to use display the data
* stat - The statistical transformation to use on the data for this layer, as a string.
These are two hugely important parameters, with really big concepts and abstractions under them, but the documentation is of the style "foobar(): this is a function that foos the bar", documentation that restates the information in the name, but with more words, and no insight on where to go next.
So now a person is two pages deep into documentation, and it's actually circular documentation because layer() has a ... argument that gets passed back to what? The function documentation that you came from? For a newcomer it's a completely twisty series of passages, and as an experienced user who reaches for ggplot before any other tool, it's confusing.
The other function based confusion is that the list of aesthetics is not connected quite well enough to the mapping argument from aes(). What aesthetic values the function understands is probably one of the most important things about looking up the function. But reading the parameter documentation, it's not clear that there's an entire section below that describes that crucial material, far further down the page. And on a long long page it's easy to accidentally skip over that section when skimming.
(These are the sorts of frustration I have with typical Python documentation, btw, so maybe my brain is just different from typical engineers)
All that said, I find the documentation to be saying a lot more than it did in the past, and it sounds like it has been continually improving.
One of the best compliments I can think of is that with ggplot, easy things are easy and hard things are possible. But I haven’t been able to figure out how to fully work the system.
(Thanks for all of the work!)
And then there is an entire ggplot2 book (there are many, but this one was written by Hadley): https://ggplot2-book.org/
And yet.. it somehow works. It makes data analysis and statistical modelling a pleasure. It somehow gives off a sense of lightness, and makes it easy to investigate and explore. I would guess I am genuinely 2x as productive in R as I would be in Python on similar tasks.
I know it's not a "proper" language, but I think that, maybe, not everything has to be exactly like "proper" software engineering?
The beauty of R is that you can write one line of code and use some hot-off-the-PhD-thesis cutting-edge-just-published-in-J.-Stat.-Soft-chunk of statistical analysis in your totally different, completely whacky problem, and it's fast, and (by and large) works.
Of course, that's its biggest problem as well. Scientifically, it will quite happily give you a 150 mm howitzer to aim at your foot, assuming you know best.
I think you mean "poorly-documented-cobbled-together-under-deadlines-never-to-be-maintained by someone who has no idea of software principles". Very few labs have a dedicated software engineer to actually turn this software into a usable/hackable tool let alone maintain it.
I've been working with Python for the last year and appreciate how much it helps with general IT problems, but I would still stick to R for statistical/data analysis.
This seems highly unlikely, based on my 20+ years with R. Yes, using wrong data structures/algorithms can lead to slow code, but switching languages won't fix this.
rprof and microbenchmark are your friends if you really need to optimize your code.
and (as in python, and as several others have pointed out), if you have something especially challenging, write it in C/C++/fortran instead, and link it to R.
I came here with sleeves rolled up to defend the language, but was pleasantly surprised to find it was already being done much better than I could have.
It's interesting to see how R elicits such a reaction to some programmers. I think it's frequently misunderstood, and R needs to be used in a particular way to allow it to fly.
When I've tried to recreate analyses in Python or Julia, they have nowhere near the fluency of R. It isn't possible to know this if you're messing around with if statements and other procedural methods of achieving things which are better suited to other languages, but rather when crunching data for analysis and graphically visualising the results.
I also understand that it's due to R's lisp-y-ness that allows us to have tidyverse in the first place.
Question for Hadley - there have been a couple of projects to fuse the speed of data.table and tidyverse. What do you think of this aim and are you tempted to change tidyverse to get to the speeds of data.table, or would that require too much of a fundamental change?
I truly, genuinely dislike the language. I think it's very productive, and I appreciate that Matlab costs an arm and a leg (and god help you once you start paying for some of the nicer packages on top) - but Matlab has spoiled me immensely on the language front.
To me, Matlab feels like a language that was designed with an intent to appeal to folks with some understanding of traditional procedural programming, but nudged into treating matrices as first class citizens.
R feels like a language that was built for people who were using excel, and have never written a line of code in their life - it's riddled with completely unintuitive, frustrating, intentionally obtuse operators and terms for things that have perfectly fine definitions in normal programming.
The difference is that I have 20+ years of programming experience (including quite a bit of functional programming) that I can easily port over to Matlab, and which becomes literal baggage trying to use R. The end result is that I will use R, but I basically always walk away frustrated and infuriated, even when the problem is solved.
It doesn't go well at all with a procedural method.
I don't think so. Most people who come to R after years of Excel find it just as alien as you do.
I also recall my pushback was along the lines of "who on earth would want that". Yeah, it's a good thing I'm not the person coming up with these things :)
In the intervening time I've become a large advocate for the pattern of chained operators. So I'd imagine I'd enjoy piping in R. And if that means I'm emulating a common Excel workflow, that's fine. I won't have the childish response of "ew, Excel" :)
Where are you getting that from? To start with the pipe operator has been independently reinvented multiple times in R, and neither ‘magrittr’ nor ‘dplyr’ were the first to introduce the pipe operator into R. And (at least when I was exposed to it), the pipe operator had nothing whatsoever to do with Excel. Instead, it was an attempt to introduce the composability concepts from the UNIX shell and Haskell composition into R.
EDIT: I found the conversation in question but it involved deleted tweets. And those deleted tweets are the one that reference the package name. Sigh. It was just after the release of magrittr and several months after dplyr
I have no idea where you get that impression, most Excel power users I have met take a long time to understand how to use the pipe operator in R.
The S language predates the first release of Excel by 11 years.
> and which becomes literal baggage trying to use R
I've had the opposite experience. My experience was that having a broad array of programming experience made it easier to pick up the weirder corners of R. It became more likely that I'd seen *something* similar to that construct in the past. The converse has also been true. Seeing all the weird corners in R has made it easier to pick up new concepts in other languages & paradigms as it's been more likely I've seen *something* similar from R.
An important factor not often mentioned is that I think R really helps individual developers/very small teams to be productive.
The the main reason that statisticians love it is that the libraries useful to them are much better in R than elsewhere (though Python keeps encroaching in that turf, and "real developers" dislike Python a lot less than they do R).
The main reasons that developers hate it is that it is very unlike almost all other languages that they're used to. This is very valid since outside the narrow domain of statistics, there's probably nothing that R does better than other languages. So for a dev who occasionally dabbles with R by necessity, the otherness serves nothing but frustration.
Still, I wonder how much criticism there is against R as a programming language, that is not some variation on this works very differently from other languages. IMHO the sub-setting syntax, and countless x-apply variations are big warts. I'm not a big fan of Tidyverse, and even less of the schism between base- and Tidy-R. I read some seemingly fundamental criticism about R's deficient scoping rules, but I'm not nearly knowledgeable enough to judge their merits.
I guess it doesn't help that almost nobody learned R as their first computer (as opposed to statistics) language. Personally, I learned C, Matlab, Python/numpy, SQL, R in that order. R does seem to be quirkier than all the others, except maybe SQL. But I don't dislike working in R any more than working in any other language.
I rarely found an important difference between the languages besides having to transpose some matrices here and there.
Having used both in a professional setting - and coming in with a fair bit of programming experience - Matlab is generally a pleasure to use. It's different where it needs to be in order to treat matrices as first class citizens, but otherwise you can apply many of the same intuitions and paradigms that you would in any other language.
R on the other hand... R is a fucking disaster of inconsistency. I find myself incredibly frustrated attempting to do simple and sane things - things that I know are only a line or two in Matlab (or even python) and instead fighting with "which version of the 12 different slight variations of this operator are you attempting to use today!" hell scape.
My strong guess is that if you have no coding experience, and you learn R fairly thoroughly - it will feel very nice. My problem is that for anyone with actual coding experience, it's like being given a keyboard with a qwerty layout, but which is actually using dvorak. All your intuitions are pointlessly wrong - not because they are actually problematic, but because R has decided that the A key is really on the other side of the fucking keyboard.
They each have a few strengths over the other, but generally speaking, I much prefer the language consistency of Matlab.
My general experience is - industrial shops will be using Matlab. Almost all of my Matlab work was aerospace related (think sigint/radar/signal processing/modeling).
R is more popular in education environments - but I strongly suspect that's just because it's free. Post-grads don't have much lab funding to work with at the best of times, and Matlab with an associated set of plugins/libraries specific for your task can easily run 30k a seat.
Personally - I find it pretty telling that most places with money choose Matlab. Doesn't inherently make it better, but it does mean Matlab is getting used in places where mistakes are expensive, and there's a focus and consistency to the tooling that I just think is desperately lacking in R.
Aside from two statisticians I had as professors, I am yet to meet someone with deep understanding of statistics who doesn't speak R as first language ...
I found it way easier to grasp the meaning of statistics by playing with R than by reading the maths.
I thought that was real Scottmen.
Because real Scottsmen prefer:
table.loc[(table.column > 2) | (table.column2 < 3)].reset_index(drop=True)
to
table[column > 2 & column2 < 3, ]
and everyone knows this!
not to mention, if you aren't managing 100 virtual environments and 100 conda environments (with different syntax for requirements), you aren't a real scottsman!
table.query('column > 2 or column2 < 3')
If you want. I'm not sure why you're dropping the index there.But that's just why it's useful - R is great when you are an expert, but becoming an expert takes years. The perspective of new users is really important. (I've been using R almost 20 years, have written several packages, and still feel like an amateur. Indeed, I'd never heard of `**` as an alias for `^` until today; nor `sequence`, which apparently has always been in base; and I still can't remember what `sweep` does.)
I thought some of these arguments were better than others. True that base R regex is confusing and messy (and that stringi/stringr are improvements). False that allowing string concatenation with `+` would be a good idea. That's just a footgun waiting to go off, given that R also is weakly typed. Expecting `nchar(1000)` to magically work seems naïve. `<<-` (roughly, global assignment) is an ugly necessity and a code smell, not a cool language feature.
An awful lot of these problems are fixed, or try to be fixed, in the tidyverse. Not using tidyverse is a bit unusual because most beginners nowadays, I think, start with the tidyverse more than with base R.
For me the worst part of R is simply it fails silently. This is really deadly, especially when you are producing scientific results. There are so many places where R will plug gamely on after you have done something deeply inappropriate. Given how badly scientists code, one has to worry.
I don't agree that "R won’t change" is the base problem. It's not so simple. R is used for science. I like very much that my code from 2008 will probably still work if someone wants to replicate my results. I appreciate the R-core team's work in making this true. There are genuine trade-offs here.
If you want emotional relief, it's worth following https://twitter.com/whydoesR.
Maybe Julia is the way forward? Or is R "worse is better"?
Julia is well worth learning, if you do computationally-expensive work. It is kind of a pain to use interactively, though. I use both R and Julia in my research. Think of Julia as the new Fortran, though, not the new R.
You can inspect the assembly code for anything you're working on, and that can be quite helpful at times; see e.g. https://youtu.be/wU6c8CDRXJE?t=3887.
I think the reason why quite a few high-performance people (I mean in the science community -- I don't know much about other communities) are excited about Julia is simply that well-respected experts are also excited. An example is the Julia implementation of the MIT GCM (general circulation model) for the ocean; for similar projects, see https://github.com/CliMA.
Programmer effort is also a factor in scientific computation. If the system can do some of your work for you, so much the better; see e.g. https://www.youtube.com/watch?v=rZS2LGiurKY for a lecture that touches upon how the framework of Julia eases the burden of machine-learning tasks.
As I say, though, I do not have deep experience with Julia. I've rewritten one of my numerical models in Julia and the speed is about the same as before, but my code is much shorter and easier to understand. I would not burn up 6 months translating a complex code, but nor would I start a 6-month coding project in Fortran anymore.
I see problems when people take an imperative approach to solving numerical problems, and something like Python is better suited to that. Also, R isn't really set up to work with matrices like Matlab/Julia are.
R is for data manipulation. 90% of what I do in R is manipulate dataframes or matrices and then run machinelearningmodel(mydataframe) or ggplot(mydataframe). And for this it is incredibly efficient. You can rightly argue that some elements of the language are quirky but that's missing the point.
> Asked over 100 Stack Overflow R questions.
As a tangent I find a hundred questions asked the first year for a very mature language is a lot.
In a way, you could argue that the entire tidyverse is a huge effort of a band-aid.
So for all the irritating design choices and idiosyncracies, R is still a network of islands that work incredibly well for people, as long as they don't ever go to sea.
I agree with the author's sentiment - I love a lot of what R has, but there is a lot of small madnesses.
There are so many unique PL ideas in R (may not actually be unique but certainly unique among common languages today)
- first class environments - named, default parameters and even the ... parameter which encourages the pattern of hierarchical library functions - there's one large customizable main workhorse function, and many wrapper functions that specify some defaults or add some behavior, but all the underlying customizations are exposed through ... - copy on write as a default - ability to choose evaluation strategy
But I also wonder how many of these cool ideas would actually work well in a saner language
The boost you get from the slightly better expressiveness of R over something like Julia or Python is not worth the headaches you'll run into down the road in trying to maintain whatever you wrote 6 months later, or God forbid, trying to integrate your code into someone else's work.
R was my first language and in hindsight that was a HUGE mistake. So much of the R code out there is horribly written, and even when it isn't you still have to deal with all of the issues the author here points out. If you pick up R as your first language, you will end up picking up all sorts of bad habits;
R is fine if you're working solo and you don't plan on maintaining or reusing or reusing your code. For everything else, R is garbage. It took me a year or more to undo all of the bad habits I picked up learning R.
I don't agree with the "worse is better" comparison in the comments here. "Worse is Better" was meant to refer to the idea of "Don't make the perfect the enemy of the good", among other things. It was not meant to be used as a justification for poor design. If anything, python for data analysis fits the "worse is better" philosophy much better than R. It's not as well optimized for data work compared to R, but it's much simpler, more consistent, less error prone, and it plays well with others.
In some research fields (e.g. scientific fields that use R) the ground rules are that the code needs to be understandable and it needs to be clear that the libraries involved were used correctly. That's basically it. Even hardcoded directories are common. Good development practices are not widely understood to be important and in general many people are just starting to get the hang of version control and might not use it at all.
If R enables you to solve a statistical problem you have right now and it does this in a way that is better or more comprehensible for the people who use it, that means it has a niche. As someone with software development experience in a bunch of other languages, I agree with you that R is full of weird warts, but let's not forget that there are areas where its value is still obvious.
Citation: my partner works in a scientific field where R is predominant.
For sure. As much frustration as I had with R, at the time it was an enormous improvement over the stuff that came before it. And its emergence and success led to other languages improving their data and analytics capabilities.
I understand it's frustrating trying to use a language you don't understand. And instead of reading the language manual you go on rambling.
"" is an empty string (almost) as you know it from other languages.
character(0) is an empty vector of type character (i.e. a vector with no elements). This vector doesn't even contain an empty string.
R is a vectorized language. You almost always deal with vectors. "" actually is a character(1), a character vector of length 1. Once you understand this, there is a chance for you to enjoy R.
I'd agree with this assessment. If you start doing R and it feels weird to you then -- in my opinion -- you're probably in the wrong place. Meanwhile, for the cognoscenti -- the researcher, the statistician -- R behaves just as you'd expect. That is the draw -- a language developed around statistics.
R is not a great computing environment for computer science. E.g. writing iterative algorithms. Almost everything worth a damn in R is written in C++ and then FFId in. Those who do not want to use C++ can write their algorithms in Python or Julia -- and they often do. Arguably the defacto for computing oriented machine learning is Python, not R.
And even if the author is the problem, I wouldn't accuse them of not reading enough.
> Python has two types of empty string, array('u',) and ""
you'd probably conclude that hasn't really understood what he read.
This OP complaint seems like weird nitpicking about R. Many languages have different empty/null-types for different variable-types. Also, don't get me started on "nulls" in C, C-strings, C++ strings, or memory allocation.
All languages are complicated.
R's data types are one of its most alien part, and that's why I think if you're coming from another language, chapter 20 of Hadley Wickham's book[1] is the most important one.
In that sense, it's similar to the problem many languages have with NULL, but on steroids: you can have NULL, NA, character(0) (or anythingelse(0)), or '' as your null result and each of them are tested for in different ways.
Obviously this won't be a problem for the various battle-tested standard libraries, but a lot of my work in R at least is assembling somewhat-novel analysis pipelines based on quite new statistics code.
With respect to dealing with return values, you can circumvent some pain points by using identical(), isTRUE(), isFALSE() in if conditions instead of, e.g., `==` which many people use because this is what they know from other languages. The assertive package is also nice.
As others have mentioned, just use tidyverse. I picked it up 4 years ago, and last week I went back to the code I wrote then.
I was productive in minutes. I could read the code, modify it, and easily test it in the REPL. The docs for dplyr are good.
ggplot2 is still awesome and the docs are good there too. ggplot2 is the fastest way to figure out what you want and make a pretty plot.
(However one thing that still annoys me is that R moves faster than Debian. So it's possible to do install.packages() in R, and it will break telling you your Debian R interpreter is too old. There is no easy solution for this, just a bunch of workarounds)
-----
OK, sure you can call it a polished turd, and to some degree that's true. But a polished turd is better than just using ... a turd!
The error messages in R are not quite as good as Python, but I wouldn't call it a problem. I'm able to localize the source of an error, even when using tidyverse.
My article comparing tidyverse to some other solutions:
What Is a Data Frame? (In Python, R, and SQL) http://www.oilshell.org/blog/2018/11/30.html
----
But would I recommend learning it to anyone else? Absolutely not. We can do so much better.
I would recommend with the caveat that it's one of the hardest languages I've had to learn. However that is partly because it changes how you think. But if you have a certain type of problem then you have to change how you think, or you'll never get it done. Data analysis is surprisingly laborious even for people who have say written compilers and such.
>can't remember the last time I saw a project someone did in R get very much traction anywhere...the only time people talk about R on the internet is to discuss the language itself which is definitely frustrating
There is a lot of R deployed in industry, even in silicon valley, but you have to be in-the-know. R gets plenty of use in statarb & model checking in finance - speaking from personal experience at GS & BofA/ML. My one non-trivial project at Twitter involved working with this team building a model & I remarked - hey this can be done rather easily if you use this library in R - and the teamlead says, yeah that's how we're doing it! But I thought we are a Scala shop, I said. So he says, yeah but imagine building that entire library in Scala from scratch, it'll take forever! So I enquired how he gets it done - you basically spin up a socket server & the jvm sends R commands plus data as payload over the socket, the server runs R and returns the result of the model back as a string, boom done! I said it was kinda janky & he says - I won't tell if you don't ! So that's R for you - it gets the job done & its fast & somewhat messy, but it is used everywhere, yet people won't openly admit to it because its a 30 year old language & we all want to be using the latest & greatest tool.
I now work at a news startup with a few million users, & all of the news personalization is done in R. So when these millions of viewers watch TV, the piece of code that decides which news clip should be shown ahead of which other news clip & which clip comes after - all of that is decided by a block of R code that I wrote. ~ 300 lines of R, uses quanteda, tidytext & parallel under the hood. Pretty much everything I do involves mcmapply, which parallelizes your compute & uses as many cores as you specify. But that's sort of the thing with R - you have to know which functions/libs to use & which ones to avoid. Just switching from tm to quanteda got us a 200% bump in perf. Switching sapply's to mcmapply was another winner. These things aren't documented cleanly - you have to keep up with cran, experiment & see what works best for you.
I can't remember the last time I saw a project someone did in R, or a tutorial on how to do something in R, get very much traction anywhere. It seems the only time people talk about R on the internet is to discuss the language itself (which is definitely frustrating), and it's getting old. Even this awesome, comprehensive document, which I would usually be foaming at the mouth to read, has me going "meh". I'm tired of the subject.
Dark-matter statisticians, I guess?
well, you know, I'm not very active in C++ any more, and I haven't seen an article in over a decade on C++ which received any traction at all. So I guess C++ isn't getting any traction any more either.
Maybe in your field, I work in bioinformatics - before R, perl was widely used as a high-level language.
> Regarding keeping versions straight, all past versions of packages in the CRAN repository are kept on CRAN...
This is woefully inadequate if you need to replicate somebody else's environment. Nobody should think manually guessing and then typing in each package version and hoping they're compatible is a viable option. Not to mention even if you specify an older version of a package it doesn't pull in compatible dependencies, it just pulls in the latest version. There's renv but it's not reached widespread use.
> Regarding tidyverse dependencies you can reduce the number of packages you load by not using library(tidyverse) and instead load the specific packages you need. This will result in fewer packages being loaded
We're talking about replicating other people's work. We don't have any control over their code, and R users are largely ignorant of best-software practices.
Don't forget that I also mentioned the checkpoint package in my post. You only need to know the date for that, not the version of each of the packages.
In your last paragraph I think you are referring more to software development practices than what is available through R. Simply using R or any language doesn't guarantee this.
1. pointing out that, like every other language, base R has idiosyncrasies
2. how use of R is more complex when you’re largely ignorant of the tidyverse, which is crucial for the vast majority of tissue today’s use of R
3. frustration because you’re using a language/ecosystem, that’s targeted for a few specific uses, as a general purpose programming language
This.
I'm interested in non-flamewar non-religious reasons that the tidyverse is bad. He does give some. I think his complaints about inconsistency and a moving target have some validity. However, the price of not using tidyverse is (roughly) paid in the rest of the article. I would definitely not use R without it.
Read his Section 5 on the tidyverse... and see how absolutely minimal his complaints in that section are. E.g. to "purrr" his objections are "largely philosophical"... but he's complaining in the previous section about the annoyance of writing lambdas (which purrr makes even easier).
Yes, R has a big community and there's a lot of quirks in individual packages, especially less-used ones. Yes, there are packages presenting unified interfaces to other quirky outputs (e.g., broom). The necessity of this is not good. The existence of it is good.
HN readers - do you have an "up and coming" language that you think has better structured the fundamentals from R, that you hope will someday have enough capabilities you can use it instead of R? I've tried Julia, which is beautiful but the startup/compilation times were difficult to get over. Is it reasonable to hope Julia will be good for interactive usage someday? Is it already? Are there other candidates in this area?
It's a pleasure to use, though!
If you don’t understand the GG, then ggplot will seem opaque, and no goodness of documentation will suffice.
I don’t mean to blame the user. Perhaps the ggplot documentation could improve by reinforcing the need to understand that or referencing it more frequently?
Most of his examples of WTF's are from base-R. And he's definitely not wrong, as many of these have bitten me a bunch over the years.
> I'm interested in non-flamewar non-religious reasons that the tidyverse is bad.
For the very reason that it's great to use, it's a nightmare to develop with. NSE is super handy as a user, but it's an absolute nightmare to build new functions on top of (dplyr specifically). Like, I now know 2-3 different ways in which quoting/substituting etc can be done for the tidyverse, and I've had to maintain code using them a bunch of times.
It's incredibly annoying, and every time I do it I need to look up Hadley's new approach to NSE (don't get me wrong, I adore using the tidyverse, but I absolutely despise programming with it).
Hope is the operative word here!
I'm writing a language to compete in this area. It's called Mech and I'll be releasing the first beta in October. You can think of it like Matlab + Excel. It's very fast, has default-parallel semantics for operators and functions like Matlab, reactive dataflow like Excel, and supports full interactive coding with no startup/compilation latency issues. It's meant for robots, but I've also designed it to be a better Matlab, and I think it should take on R handily. Fair warning, it's public alpha now so error messages are sparse and the happy path is narrow.
Going to answer with a question: Why is tidyverse == R considered true?
I use ggplot frequently, but for data manipulation data.table is orders of magnitude more powerful. And more stable.
For example, in 4.5.1:
Selecting and deleting at the same time doesn’t work either. For example, data[c(-1, 5)] is an error.
What would it mean for that to work? He seems to acknowledge that "selecting and deleting at the same time" doesn't make sense in 4.11.1 Can you guess what data[-1:5] returns? I can’t either, so don’t ever try it. If you must know, it’s actually an error.
Also in 4.11.1: The : operator is absolutely lovely… until it screws you. The solution is to prefer the seq() functions to using : [....] As I’ve said, seq() and its related functions usually fix this issue.
Maybe the "related functions" fix some issues but seq(a,b) is not different from a:bIn 4.11:
Now what do you think names(foo) <- names(bar) does? Seriously, can you guess? I can think of roughly four realistic guesses. Is it even valid syntax?
How is that surprising? Can the author also think of four realistic guesses about the effect of A[1,2] <- B[3,4] for example?In 4.13:
The index in a for loop uses the same environment as its caller, so loops like for(i in 1:10) will overwrite any variable called i in the parent environment and set it to 10 when the loop finishes. [...] This sounds awful, but I’ve never encountered it in practice.
Is it awful? The same happens in other languages like Python or C if I'm not mistaken. The plot() function has some strange defaults. For example, you need to have a plot before you can plot points [...]
I have no idea what that means. You can plot points using plot() without having a plot beforehand.Edited to add: In 4.5.3:
The $ operator is another case of R quietly changing your data structures.
Is it unexpected that when we extract an element from a data structure we get a different kind of data structure? Is A[1,1] another example of silently changing one data structure (matrix) to another (number)?It annoys the shit out of me in python. I much prefer perl's
for my $x (@array) { ... }
or ES6's for (let x of array) { ... }
Note that given python and ruby only do function level scoping rather than block level, I can -understand- why they work the way they do even if it annoys me. R already has the necessary granular scoping to do the (IMNSHO) sensible thing so it seems like a pointless wart.My -guess- would be that if it was intentional, it came about in R because for loops are rare enough that you want to know the last index more often than you don't because if you don't care about the index presumably you'd've written something else.
(also you could argue the ES6 version would be better written using 'const', but I've lisped sufficiently my fingers invariably generate 'let' when left to their own devices - caveat emptor)
grepl(“es”, “test”)
I searched "R string functions", saw "grep" and wondered if there was something I was missing in the author's challenge.
Is it because it's using regular expressions they don't consider it the correct answer or is it because they aren't as familiar with regular expressions as some other people are, I wonder?
First, there's the constant inelegance/clutter/inefficiency of having to cast into and out of arrays and lists, even when doing basic list comprehension. R, Julia, and Matlab are all vector-based languages (I think), so you avoid having this casting as much.
Secondly, having a native vector type means you don't have to worry about the performance penalty of operating directly on arrays if an existing prebuilt method exists. Since the efficiency of Numpy comes from calling it's underlying C library, you're forced to memorize and use prebuilt Numpy functions rather then just use the more obvious and elegant array manipulations. For example rather than calculating the the cumulative sum like this:
cumsum = reduce(lambda a, b: a+b, [1,2,3,4])
We have do this:
cumsum = np.cumsum([1,2,3,4])
(There are better examples, but this is all I can think of right now).And once you add something like Pytorch tensors on top of this, we now have an additionally layer of casting/redunduncy/memorization of prebuilt functions!
Non-uniformity imposes a burden that will be too much to bear, unless the system offers particular advantages. The fact that several systems co-exist is proof that the advantage-burden balance is favourable in each case.
There is no need to converge on a single tool. Carpenters need both saws and hammers.
In practical applications, language syntax is just part of the story. One must also consider the issue of available libraries. One thing that really stands out with R is its immense collection of well-vetted and well-documented packages. Python and Matlab -- the two main alternatives in my discipline -- fall far behind R in this respect. If there's a journal article on a new statistical technique, then there's a pretty good chance of a package written by same author. And, if that package is on CRAN (the repository for such things) then it has undergone quite rigorous testing on several types of computer, with several versions of R.
My diety! someone was complaining about inconsistent syntax but doesn't recognize inconsistent dependencies?
[1..10]
|> Seq.filter (fun x -> x % 2 = 0)
|> Seq.map (fun x -> x * x * x)Then -- I once made a meme to that Oliver Stone Vietnam movie that said "This is my copy of Stata. Without me it is useless. Without it, I am useless. I must cherish it as I cherish my life..." (In the original said of a rifle.) I was good with Stata, fast and precise and never "ugh, okay, let's open a quickie notebook... there goes my morning..."
At this point, I see R users (typically PhD students and post-docs) doing "science" in R by playing with parameters to functions in poorly-understood packages and publishing papers on which parameters are "best" for data generated from some specialty source.
A very common situation for me is to be pulled in only after a package has been created with some vague hope of fixing performance problems (which R, Rcpp, and RcppParallel make fun to do for me, but I have some C++ background for scientific computing, ymmv). It is extremely common to find that these packages contain fundamental logic errors that probably should invalidate the (already published) results but never got caught because the code ran without actually failing. I guess I'm complaining that people are using buggy packages to write more buggy packages and it just bothers me.
Library-driven development is just how the world works these days. And it should! But I'm not confident that the R bioinformatics world has the kind of guardrails I would prefer to see. I mean, I am reasonably confident tensorflow is functionally correct. Any R package that pulls in too many other R packages to begin with is probably not.
As for the language itself - I guess it is ok. I have some lisp in my background and a fair amount of love for non-traditional array languages. But I don't see much R code that seems to stick to the R "standard library" rather than pulling in a million packages to do anything . . .
That's a lot of bioinformatics, and not specific to R. It is a huge issue with anything vaguely pushbutton in the bioinformatics domain.
People teach the tidyverse to new r users. It makes them think that it’s standard practice to pull in lots of unnecessary but possibly convenient packages. Simple string manipulation should not require an extra package like stringr, but for many users it does. Often, they were taught this way.
Found R when I was looking for something free alternative to Matlab and chanced upon R in 2010/11.
Now a days R is my goto scripting language anytime when I just want to get to the results and don't care about reproducibility.
I also use Shiny as alternative to multiuser scenarios involving spreadsheets since I when work in a financial firm where excel and VBAs still dominate most of the front office functions.
Sure Python could be good tool but once you become fluent with R ecosystem, moving to python just feels too much that can be done with few lines of R code.
For me the conciseness of data.table and ability to cook up shiny web apps with very few lines of code is biggest pull.
The following works and it looks quite natural:
c(list(1, 2, 3), list(LETTERS[1:5]))
Examples:
#learning : a data.frame is a list. x = df with 10 rows, 21 columns, say. as.list(x) gives you a list with 21 elements, one per column
#learning : Above = getting a row and it's previous row using .I() in data.table
The R garbage collector is imperfect in the following (not so) subtle way: it does not move objects (i.e., it does not compact memory) because of the way it interacts with C libraries. (Some other languages/implementations suffer from this too, but others, despite also having to interact with C, manage to have a compacting generational GC which does not suffer from this problem).
golang's collector isn't compacting either - though it uses per-size-class arenas for allocation so you don't end up with fragmentation bloat to nearly the same extent. Part of me wonders if simply building R against jemalloc would get a decent chunk of the same advantages.
I found it pretty interesting that the alternative to R is Haskell for general CLI tools! Seeing some open issues in a popular tool for dealing with ancient DNA (aDNA) about making invalid states impossible within the type system made me genuinely laugh out loud in amazement. I didn't expect that level of technical knowledge within the world of archaeology.
(and if it turns out they already know, -I- don't, and that sounds pretty cool and fun to read up on :)
Cheers!
Fwiw, it’s mostly a tool for parsing and processing this ‘Jan no’ data schema https://poseidon-framework.github.io/#/janno_details
I supplemented the theory parts of my other courses with some of these [2] R books about using the methods instead of deriving and proving properties about them.
There are also some R studio cheat sheets [3].
[2] https://www.routledge.com/Chapman--HallCRC-The-R-Series/book...
`All languages are wrong, some are useful.`
And R is one of them.
R is amazing for data analysis. Also, RStudio is a much more efficient solution for iteratively exploring data than Jupyter. Don't make fun of a screwdriver for not being a hammer.
You can dramatically simplify your life by using lists with lapply and related functions. I teach students with no previous programming experience to do some things that would otherwise be far too complex, but I also have to write a helper function to convert the output into a usable form for further analysis like plotting.
I approached it like you described. I wouldn't want to do a really complex REST API in it, but as a wrapper for calculations, we've got a repeatable pattern to run them in a cost effective manner.
I'm not angered, I'm more wondering about the usefulness of the arbitrary line you've drawn in the sand, and even the shape of the line.
RShiney lets you build interactive webpages with advanced GUIs; if <webdev stack> counts as a programming language, why not R?
Many people like python because it lets you script things, and you can even make your script executable with a shebang at the start (!# /bin/python) -- and while true, that isn't built into R, you can run R scripts programatically (> Rscript myfile.R), or make this executable by putting it in a standard shell script.
It's just so close to being as good, but not quite
Comparing R with others as merely a language is close to meaningless. You have to take the whole ecosystems into account.
LMFAO. I only read the above line and the bit about how tidyverse nearly fixes all the badness that is r. Max lol.
That info could help contextualize the entire piece.
But guess what: it's free.
Which one? I've switched most of my work over to Julia, but I'd much rather use R or Python than Stata or SPSS.
Teaching them to do statistical analysis in Julia will involve you spending 80% of your time teaching them Julia and maybe 20% of your time teaching them statistical analysis.
This works until they run into a use case that doesn't involve running various forms of regression analysis on panel data.
In the parent comment's case, I could imagine that there's an expectation that someone doing neuroscience research will eventually have to expand beyond what's possible in SPSS. In this case, it may make sense to go through the effort of teaching them how to program in Python or Julia.
I find using Python and/or R to be very helpful for teaching, since you can implement the stats procedures using primitives (prob. calculations), so you get some experience with how things work.
Sure it requires some "coding" but nothing harder than using a calculator, so I think it's worth learning.
Julia is a bit more involved (need to learn something about data types), but still would be manageable.
From time to time, I still have to use SPSS. Again and again I'm flabberghasted how bad this overpriced piece of software is.
Academic research, so more econometrics/data science work, but I have some experience with application programming that I've managed to leverage.
> Most people I know who are either statisticians or scientists first and programmers only reluctantly love Stata and SPSS
They are good to the extent that a lot of published - social science - research uses terms and methods that assume you're using one of the two. Of course, having an ok point and click interface also helps.
However, data access, aggregation, and cleaning are easily ninety percent of what's involved in even basic econometric(y) research. It is orders of magnitude easier to do all of this programmatically in R. Once you start working with larger datasets, or once performance becomes an issue, you pretty much have to transition to Python, Julia, or something similar by default.
You almost literally can't come up empty on CRAN.
I rarely reset index -- perhaps its a difference in familiarity? (I use R but it isn't my background, perhaps there is a forced R pattern that is a general antipattern for indexes?)
Recommended watching (32:00 onwards): https://www.youtube.com/watch?v=mWtfZaT7iSc
You will be hard pressed to use pandas without .loc and resetting your index.
What is wrong with .loc, in your view? Genuine question. I used to dislike it but I've been using pandas for a while and I've gotten comfortable with it, and I've forgotten the reasons I used to dislike it.
I often chuckle when people complain about R in production and how it isn't a good general purpose programming language, my experience has been the polar opposite. You can write bad code in any language, and R is no exception, but R allows you to write so much less code and R-core is truly exceptional at backwards compatibility. Our approach to R is basically:
- Don't have a lot of dependencies, and when you do have dependencies, make sure they themselves don't have a lot of dependencies. While we do use shiny as mentioned above, our core models are very dependency light and shiny is just a basic front end.
- data.table (which was designed by quants) is a zero-dependency package that is by far the best tabular data manipulation package that has ever been created since the dawn of time. We generally work on an EC2 instance running linux with a ton of memory. In the < .01% of cases where a dataset doesn't fit in memory (e.g. tick data), we do initial parsing with awk if file based or SQL if DB based and then work in R.
- Check/coerce argument types and lengths on function input to catch and avoid all the quirky edge cases that drive people nuts - it's so easy!
- I hate OOP and I love that R doesn't encourage it. Mutable state, especially for non-software engineers, is the devil. Don't get me wrong, OOP has its place, but the fact that R encourages functional programming is one of the best things about it. The slight inefficiency this produces is almost never a problem.
- R is not slow at all when used correctly. Additionally, the C API is a joy to use when necessary.
- Stick to the base types: vectors, matrices, lists, environments and data.tables (only exception). The fact that you can name, and then use names to index all of the above is stunningly powerful. The only "objects" we really create are lightweight extensions of lists with an S3 print method.
- We have an internal version of renv/packrat that creates a plain text "dependency file" for projects and we pin package versions in docker containers. RConnect doesn't use docker right now, but they do have a versioning system that works quite well in my experience.
I definitely wouldn't want to build something like a company website in R, but I also wouldn't want to build that in C either. R definitely has it's place a server-side language even outside it's assumed domain of statistics.
Haters gonna hate, but joke is on them.