A Quick Introduction to R
github.com
github.com
One of my previous jobs basically turned into an in-house R consultant for a department in a pharmaceutical company, and I caught so many bugs when investigating some other issue which meant the results people were reporting were completely wrong. A really common one is multiplying 2 vectors of unequal length where broadcasting shouldn't be possible and it just recycles the shorter vector - but hey, it ran without error and there's an output so many researchers don't notice.
Not to mention trying to handle errors is pretty miserable, if you want to catch a specific error you have to match the error string, unfortunately the error message changes depending on the locale the R session is running in.
For repl driven development or academic code or exercises it is excellent.
Use Rstudio
Include tidyverse
Turn warnings into errorsHadley's Advanced R is a great reference for getting down to those fundamentals.
I think exceptions where base-R is necessary can be taught as they arise.
I use R professionally for biostatistics and I can't remember the last time I had to use the base syntax because something couldn't be done with the tidyverse approach.
If you learned data.table however, it's better to just stay in data.table. Nothing in the tidyverse can touch the efficiency of data table.
I think it is important to use tidyverse because of the many quirks, surprises, and inconsistencies in base R. It would be helpful if others share their reasoning, or at least point to their favorite blog explanation, so that beginners can understand the problems they will face.
Unfortunately 5 minutes of Googleing failed a to produce a reference for me --- the start of some advanced R book that begins by asking "do you need to read this?" and showing examples whose results are predicted incorrectly by most people. Perhaps another user can provide the info.
* Reference to good HN thread: https://news.ycombinator.com/item?id=20362626
* Particularly pointed notes on base R problems: https://news.ycombinator.com/item?id=20363806
However, some data are better suited to be represented in the form of matrices. Putting matrix-like data in a data.frame is silly, since performance will suffer and you would have to convert it back and forth for many matrix-friendly operations like PCA, tSNE, etc. The creator of data.table shares this opinion [1]. And similar opinions are generally given by people who are familiar with problems that fall outside the data.frame model [2].
[1]: https://twitter.com/mattdowle/status/1037949621773844480?lan...
Most researchers are not programmers and don't care about programming. It's a tool to get the job done and I think you'd run into similar problems with other languages.
At the very least it could warn me. I just tried it in rust and that will error out if you try to divide two ints and store the result in a float which is fine by me.
On a more serious note, I agree that R being too charitable in interpreting things (seemingly without warning) seems to be a problem. You'll have to do some debugging to make sure it actually does what you intended it to do. I've only dabbled in it a bit though.
It's more natural. You never count from zero with real life objects.
In the real world we start counting from 1. CS people cannot stop complaining about it but it makes sense in languages used for mathematics and statistics. Zero-indexing is not very relevant if you don’t care about memory layout.
May I recommend you this fabulous short essay by Dijkstra: https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
It has nothing to do with memory layouts.
It is taken very seriously, though. This “issue” comes up very often when some people come and lecture others about how stupid the language they use is.
> May I recommend you this fabulous short essay by Dijkstra
That essay is not fabulous, it is obnoxious. I know you either love or hate Dijkstra and he enjoyed being a contrarian, but he’s unconvincing. The only point that surfaces during arguments on 0-indexing is iterating over 1..N-1 instead of 0..N. That’s basically what he wrote himself. This could have been solved with just a bit of syntax if it were really a problem, and it remains largely because C did it that way to simplify pointer arithmetics. It does not change the fact that for the vast majority of people, the first element in a list is, well, first.
The proper way of handling this is to allow for arbitrary indices, because you will always find contexts where a different scheme makes sense (e.g. iterating from -10 to 10 is sometimes natural, and would otherwise require some index gymnastics). Insisting that one narrow view is the correct one is just annoying.
> It is taken very seriously, though.
And those who do take it terribly seriously deserve being poked at ;)
Yes, sure, as long as you recognize that as a very subjective determination.
From the statistician's non-programmer POV the syntax of R or some other language are similarly opaque. Learning one vs. another will present similar investments in time. From their perspective, R does not make things more difficult, and the fact that it's more of the lingua franca within the field has it's own benefits.
The people I see complain about R are usually people that learned a different general purpose language first and find that when work requires data analysis they much prefer the GPL for working through the non-analytical portions if their work. (Especially with python where pandas and numpy have made less specialized tasks much easier)
t.test(x, y = NULL, alternative = c("two.sided", "less", "greater"), mu = 0, paired = FALSE, var.equal = FALSE, conf.level = 0.95, …)
A statistician opens the vignette and already knows what all of these variables represent mathematically, and can begin producing analysis immediately.
Heck, my background before using R was python and SPSS and I still prefer R for precisely the example you gave: fine-grained control built in as above, specifying how to handle missing values etc.
I end up using python for large scale data prep.
Unless the thing that makes the language difficult is your expecations. In that case, offering you an alternative mental model that helps you make better decisions when using the language does get you closer to solving your problem.
The beautiful it is to be used interactively, it really takes a lot of practice to write reliable code that doesn't abort with some error now and then.
As someone who mostly writes not-R, my own R irritation comes from a handful of things:
- The dot character "." has no semantic meaning in identifiers. It's just a valid character for names. Looking at function names like "is.numeric" really messes with my reading comprehension.
- Ambiguously, "." also separates identifiers of objects in one of R's type systems from method calls. In some cases, `foo(bar)` and `bar.foo()` are equivalent. But only in some cases.
- Even better, a popular R library defines a function `.()` (i.e., its name is just a single period character), whose job is to expose a surprising quote/unquote expression evaluation semantics.
- This is not to mention the special meaning of "." in formula literals, which are fairly ubiquitous in R.
- Different authors use different naming conventions. Base prefers "as.numeric," Tidyverse might have "to_factor," another library might prefer camel case.
- Finally, R has a surprisingly extensive syntax, exercised by different libraries to different extents, and a correspondingly rich semantics, with "types," "modes," multiple class systems, "expression" objects, immediate and lazy evaluation, expression quoting and unquoting, metaprogramming, and homoiconicity. It is a zoo of a language.
In my experience, the specific things R does well, python does it in a clunkier way.
Statistical software written by statisticians in academia, bioconductor, and quick prototyping is still much faster in R than in python.
My use case is to prototype in R, then move to python if things become more production rather than exploratory.
Once I grokked that R became my default language for anything analytics.
They are just incredibly intuitive and easy to use. ggplot2 has fundamentally influenced how I think about plotting.
With my limited experience, I have never seen anything like it.
The amount of consideration and careful design behind tidyverse APIs (tidyr, ggplot, dplyr) really astounds me. I've never felt the need to actually memorize any of them but they come to me so naturally whenever I type "library(tidyverse)". Very few DSLs, libraries or APIs have ever made me feel this way, and certainly NOT Python and the mess that pandas/matplotlib/scikit is. Even more impressive that he managed to build such a consistent layer atop the hack that is base R.
Note that I've nothing against base R. It really appeals to the hacker in me and it certainly has a ton of cool features (a condition system, multiple function evaluation forms - in what other language are `if`, `while`, `repeat` and even parentheses `(` and the BLOCK STATEMENT `{` all implemented as functions?) but damn if it isn't a mess of corner cases and gotchas.
EDIT: reference https://news.ycombinator.com/item?id=15869039
I use R directly from the terminal quite a bit for any small jobs, like calculations, purely due to the <1000ms boot time.
This has a few advantages, major being that you can run any language with a dynamic REPL this way, without changing your setup. Or, you can even have two files, written in two different languages, open side by side with a corresponding REPLs running beneath each of them. The downside of course is that you miss on auto-completion and other integrations like that. These are not impossible, but you would have to torture your Vim setup quite a bit in order to implement them.
However, you do indeed get autocompletion and many IDE amenities with the language server protocol. Naturally it’s not at the same level as RStudio. But one tool to play with any language is a very nice thing.
If you're like me R is a godsend. You'll also love the tonnes of free packages. You can't get wrong with R if you appreciate simplicity and intuitiveness.
Anyhow, explaining the difference at that part of the tutorial is not easy, so I chose to omit it for now. But might introduce it later, along with "<<-" and "->>", probably after describing closures.
[1]: https://github.com/cran/diptest/blob/master/R/dipTest.R#L37
https://github.com/karoliskoncevicius/tutorial_r_introductio...
Gotta say this is very elegant.
Command line arguments are available as:
args <- commandArgs(trailingOnly=TRUE)
And there are three getopt()-like packages: getopt, optparse, and argparse.
If you have a project root with the folders code, data, etc and are running a project on /path/root/code, you can then just call data_dir <- here::here("data") for the data folder, as the here package uses several always to find the root of a project (e.g., looking for a .git folder).
Takes a weekend to work through the book and you get a statistics refresher as a bonus.
1. It's not zero-indexed (even though most numerical languages aren't)
2. Loops are slow (though if you're looping in R you're probably doing it wrong)
3. It's inconsistent
4. The syntax is weird.
But people don't talk about the somewhat beautiful functional ability of the language to wrangle data almost magically. Its basis in lisp allows for the tidyverse and data.table to exist[1], and ggplot is a formidable analysis/plotting platform that Python doesn't come close to.
"its value is substituted, unless env is .GlobalEnv in which case the symbol is left unchanged."
dim and dims :-)
R could do more here. I really like R.
I tried dims but: Error in dims(iris) : could not find function "dims"
I do find the occasional oddity. I've noticed more very useful messages/warnings (particularly in common tidyverse functions) recently, so I think they help.
To be fair, these quirks are generally very uncommon in day to day use.
https://rdatatable.gitlab.io/data.table/reference/substitute...
2*0 = 1 2*1 = 2 2*-1 = 1/2
When 1 raised to any power equals 1, does the power matter at all? Even if it's unknown, the answer is 1.
1^NA_complex_ # NA
But: 1^NA_real_ # 1Which is an interesting detail in R that should be mentioned anyway, the difference between NA and NaN. Anyone used to languages which just NaN may confuse NA for that non-value.
https://www.r-bloggers.com/2012/08/difference-between-na-and...
https://cran.r-project.org/doc/manuals/r-release/R-lang.html... (lots of details how NA and NaN are handled)
Except - 1^NaN also is 1... now that IMO is wrong. But you can try the same in your browser's JS console and you will get 1 as a result too, so R is not the only one.
There are several NA values in R - NA_integer_, NA_real_, NA_complex_ and NA_character_, and the results will be different if you use some of them. NA_character_ and NA_complex_ will produce errors (different ones).
Also R's cleverness with NA's is not so consistent. For example:
median(c(1,1,1,NA))
Should return 1, since no matter what value is behind NA the median is still 1. But it returns NA.