The R learning curve
datagrad.blogspot.com
datagrad.blogspot.com
There are different ways in which a language can be difficult to learn. Haskell has a steep learning curve because it has a lot of unusual concepts (for those who come from an imperative background) and requires a shift in the way we think about code. J is also difficult to learn because the syntax is crazy, and again, the language is very different from what we're used to.
I found R difficult to learn because it is seems inconsistent to me. Now one reason might be that I picked up bits of R here and there without taking the time to learn the syntax from a book (unlike what I did for other languages), but it took me quite some time to be able to make sense of the different structures (vector, list, matrix, dataframe), their differences, and in particular how functions operate differently on these structures. I also have a deep hatred for the system of attributes (why would anyone want to give attributes to a vector...), and find the indexing system (especially for lists) to be nonsensical. In fact, I think that lists themselves are terrible to work with.
The general impression that I've had learning R is the language is not coherent and systematic in its design in the way other languages are (I know python, c, common lisp, ...), and I find myself spending a lot of time in the interpreter simply trying out things because I'm not sure that it will return what I want, in the format that I want (which never happens to me in any other language).
Now don't get me wrong, I use R daily and it's a useful tool. It has lots of great libraries, including the fantastic ggplot2, and the remarkable Rcpp (best C++ interface I've ever seen.. ok I haven't seen any other, but this one is really great). But learning it was no fun, and if the statistics community decided to move to a cleaner language, I'd definitely be running ahead..
> R began as an experiment in trying to use the methods of Lisp implementors to build a small testbed which could be used to trial some ideas on how a statistical environment might be built. Early on, the decision was made to use an S-like syntax. Once that decision was made, the move toward being more and more like S has been irresistible.
Since I find it basically impossible to remember how 'eval', 'quote', 'substitute', etc. work in R, I suspect that the Lispers are onto something when they say that the lack of syntax in Lisp is an important feature.
[1] http://www.stat.auckland.ac.nz/%7Eihaka/downloads/Interface9...
If it makes you feel any better, I've read several R books and they normally present it in an inconsistent manner as well. I've written programs in some 30 languages over the last 13 years, and I've yet to encounter one that's as difficult to pick up as R. Materials that purport to teach R are, in my experience, almost always presented as a set of recipes for very specific statistical methods. That approach stands in stark contrast from traditional programming introductions that attempt to teach general concepts rather than specific how-tos.
I agree, though, that despite the learning curve, R is rather useful.
One thing I would find very useful is a case study on how to design more complex packages. In particular, I would love to have an executive summary of the inner-working of ggplot2 and/or ddply. In particular, how can I manage to pass formulas as arguments to my functions, how can I achieve the pseudo DSL effect of ggplot2, etc... I know, I could read the source on github, but ggplot2 is pretty big, so a summary would be helpful.
Me and the team I work with are still finding out weird completely illogical errors that effect the language but I think I've somehow started loving it despite all it's flaws and I did actually semi-enjoy the process of figuring out all those little wtf issues.
For those starting out with it reading this book + r inferno helps a lot with understanding the underlying weirdness and inconstancies of the language: http://www.amazon.co.uk/The-Art-Programming-Statistical-Soft...
Survival guide
http://www.win-vector.com/blog/2009/09/survive-r/
Tutorials
http://www.statmethods.net/index.html
http://heather.cs.ucdavis.edu/~matloff/r.old.html
http://en.wikibooks.org/wiki/R_Programming
http://cran.r-project.org/doc/contrib/Paradis-rdebuts_en.pdf
Docs
http://cran.r-project.org/manuals.html
StackOverflow
http://stackoverflow.com/questions/tagged/r?sort=votes&p...
Cheat sheeets
http://cran.r-project.org/doc/contrib/Short-refcard.pdf
R Journal http://journal.r-project.org/
R News (predecessor to R Journal) http://cran.r-project.org/doc/Rnews/
Rseek search engine http://rseek.org/
R could use something like the Python ecosystem intro, explaining where to find stuff, not sure if any of the tutorials are as canonical as for instance Dive Into Python or Learn Python The Hard Way.
It only became popular recently; the expectation used to be that you'd learn R while enrolled in a statistics class. The canonical references are books: Modern Applied Statistics with S, R Graphics, etc. I assume that will change soon.
It can be hard to find the 'right way' to do something.
Hadley Wickham held the ceremony.
Data.Table is the glue that holds together our tenuous marriage.
cats[color="brown",summary(tooth_length),by=list(breed,age)]. This line of code will efficiently grab all the brown cats in your data and summarize their tooth length by mean, median and other quartiles then break it out by the cats breed and age.
Every time I use a language I can't lapply I despise the language and all those involved.
R is hugely inconsistent and terribly inefficient. But its also the most irresistible blend of lispy and object oriented for when you actually need to get data analysis done.
1) Check out http://www.r-bloggers.com/. The site gives you a good sense of the state of R and Tal (who runs the site) has done a great job of promoting R and encouraging the community.
2) Pick one graphics package and stick with it. The standard functions are sufficient but ggplot2 has a more finished appearance.
3) Review available libraries. If you want to do something, someone else likely already has and has posted a package to CRAN.
4) One way for SQL developers to limit the need of learning the idiosyncrasies of the language is to use the sqldf package to manipulate data frames and use ggplot2 (which generally takes data frames as arguments) to display charts.
> I found that regular expressions ... are a great way to isolate the string that one is looking for...
Well I never!
I'm not criticising the guy who wrote this, but I don't know why this has been voted up on Hacker News.
Certainly wasn't the best course I've done on Coursera, although it did "force" me to learn R, so in that sense it achieved its aim.
Even after 2 years of involvement with it, I regularly meet problems that frustrate me for hours, because they are hard to express in R's functional style. Even when I figure out the final solution, it often performs very poorly. In between these frustrations R behaves almost magically. It manipulates huge data sets with an ease and simplicity that defies logic. I've eventually come to the conclusion that an important aspect of R is knowing when NOT to use it. If your task is inherently stateful, involves random access to lots of sparse relational data that is associated in complex ways and with complicated logic - R is going to make your life hell. Write a quick script to pre-process your data and get things into more like a straight data table form and then proceed with R to analyse and visualize it. This is my experience, anyway!
plyr ggplot2 reshape
Learn those 3 and use them for everything you do in R.
data.table is much faster. It's an extension to data.frames that adds some additional constrains/rules that allows for much faster operations including aggregating, subsetting, and merging data.
Cleaning data, however, does not have a steep learning curve or high difficulty level in R-- it has a steep learning curve and high difficulty level period. Implementing good procedures for data munging is 80% of the job.
Have you tried using Google Refine[1]? It's an excellent tool for cleaning datasets and a useful addition to your workflow. I was miffed I didn't know about it when I had to collect and analyze data from surveys awhile ago. I was using Python for cleanup (Google Refine supports Jython).
The problem with Google-Refine is that it is very hard to recreate precisely the same steps to get from raw to processed data.
The documentation, for example. Often very abstract. Not to forget, the dificulty of the subjacent statistics for a lot of people. Some good Books can help with this.
The wow factor of R is very hight in any case. There is a lot of things currently that I can do only with R, thus... keep hammering _and_ take a look to as many R code as you can
It's a bit like JavaScript; if you treat JS like "Java script" or "C without types", you'll suffer. Exploit its true nature and it will flourish.
Some differences in ease of use:
[1] Construct a matrix
MATLAB: y = [1 1; 2 2; 6 6];
R: y <- matrix(c(1, 2, 6, 1, 2, 6), 3)
[2] Insert a new row r = [3 3];
MATLAB: x = [y; r];
R: x <- rbind(y, c(r));
Which is intuitive and concise? ;)
mean(iris(iris(:,5)==2,2))
is more intuitive than mean(iris[iris$Species=="versicolor","Sepal.Width"])?
Or that constructing matrix by mapping f on 1:10 like M=[]
for not_i=1:10
M[:,not_i]=f(not_i)
is more concise that sapply(1:10,f)?
(I know there is arrayfun, but I have never seen it used except in wow-MATLAB-is-functional blog posts) M = repmat(f(1:10), n, 1);
where 'n' gives the number of rows you want in 'M' and 'f' is written in proper "Matlab" style (i.e. behaves reasonably when given an array as input). Or, to be more in the spirit of linear algebra, you could write: M = ones(n,1) * f(1:10);
And, if you only want one row-wise copy, you could (succinctly) write: M = f(1:10);
Or, as you suggested, you could write something like: M = repmat(arrayfun(@(x)f(x), 1:10), n, 1);
Or, getting more silly, and using the handy bsxfun, you could write: M = bsxfun(@times, ones(n,10), f(1:10));
If you don't feel like implementing 'f' so as to permit array inputs, you could modify this to: M = bsxfun(@times, ones(n,10), arrayfun(@(x)f(x), 1:10));
Anyhow, Matlab is very productive if you can effectively wield its powerful built-ins.The problem you've solved has an equally simple implementations in R; f(1:10) for a single copy, matrix(f(1:10),10,n) for n columns, matrix(f(1:10),n,10,byrow=T) for n rows, etc.
The confusion is from the fact that people think that matrices (or data frames) are R's base types -- they are not, only vectors and lists are. Matrix is just something with dim attribute, data frame is a list of equal-length elements with a proper class.
I'm curious about about what you mean by this. How would a single package fix that? And does "fix" mean to make matrices easier or to make dataframes more broadly effective?
R style DataFrames was raison d'etre of Pandas in the first place. (http://pandas.pydata.org/#why-not-r)
Since then, R has gained a huge number of statistics packages. Python has fewer of them.
Hadley's book is really good and a solid treatment (from the author of so many clear, powerful package it's not a surprise), but there have been a lot of changes to ggplot2 since 2009. I have personally found ggplot2 to be so powerful because it lends itself very well to actually learning through using Cookbook-style examples.
also ggplot, while more limited is way easier and faster than d3 for doing the specific things that ggplot is good at. But for statistics, that's usually what you want