The R language, for programmers
johndcook.com
johndcook.com
If you're learning R, learn to use dplyr for data manipulation and ggplot2 for plotting. Both will save you a lot of time.
Edit: I had mentioned data.table as being better than dplyr for performance reasons, but it has a unique learning curve and isn't really good for beginners
I had been learning data.table, but I really like dplyr's % operator and the compositional functions better. I think I'm going to make the move to dplyr.
data.table is not just fast, but is also more memory efficient - we want to highlight in the vignettes as well.
And timings are quite relevant even on 10 million rows: https://gist.github.com/arunsrinivasan/db6e1ce05227f120a2c9
For much larger data, check the project page: https://github.com/Rdatatable/data.table/wiki/Benchmarks-%3A...
Like a lot of good tools, it seems that ggplot2 makes the common case extremely easy, but it can be a struggle to make it work in uncommon situations.
> s2 <- data.frame(x=c(1,2,3,4), y=c(6,7,6,7), series=rep("s2", 4))
> s <- rbind(s1, s2)
> qplot(x, y, data=s, geom=c('point', 'line'), color=series)
Have you looked at Julia at all? I'm only mildly familiar, but it looks super promising and I'm curious if the syntax there seems more normal or predictable for an experienced dev.
Julia looks cool; I think the syntax is meant to look familiar to people who've used Matlab or Octave extensively. I don't do tons of scientific computing, but Julia is on my list of tools to learn.
There are something things which were grafted on to S/R later on. For instance the S4 OOP system was added much later (in S4) and was based on Dylan's OOP system. So there's a case where the GP is correct that it was taking structure from another PL but it's also one which is fairly foreign to a lot of OOP devs
Argument matching is really amazing and useful for prototyping. No doubt there's a penalty, but it's exactly the type of power that's needed to build expressive and useful reusable components with rapidly changing designs. And pattern matching like that really helps with the REPL because it allows far faster exploration with fewer keystrokes. Best programming practice in library code would be to have things more fully fleshed out however.
R is based on S, which itself was purely designed as a statistical programming language.
- pass by value only means code tends to end up as monolithic functions
- very slow in loops so lot contorting to move things to matrix operations
- they just last year got a version out that starts support for vectors and matrices with > 2^31 -1 elements which limits larger data applications.
I find the plotting with ggplot and statistical functionality to be second to none though.
I've actually found R works very well as a functional language with very lean functions. It's perhaps worth noting that R doesn't copy a dataframe in a function call if you don't modify it, which is a very common use-case for me. (I'm not sure if this extends to other datatypes)
> very slow in loops so lot contorting to move things to matrix operations
This is a fair criticism, I think more modern languages like Julia will win out here. That said, R has huge library support, I've often found there are compiled versions for a lot of what I want to do.
> they just last year got a version out that starts support for vectors and matrices with > 2^31 -1 elements which limits larger data applications
Again, a fair criticism. I've never considered R a "big data" tool, my workflow is usually a funnel where each step involves reducing data size by 1-3 orders of magnitude. For example, I may have 1 PB of transactional data, aggregate it in Hadoop to 20 TB of daily aggregated data, run a query that filters and aggregates it further, and then run my analysis in R on final data. In the end I may end up with 20 GB of data, which R can very easily handle.
Also note that loops are slow enough that it is really worth learning the *apply() functions in R to avoid iterating over collections. For a relatively in depth explanation check out Hadley Wickham's book http://adv-r.had.co.nz/Functionals.html
Yes, but aren't they native loops underneath? I've seen it said both ways, that * apply is faster than R loops and that *apply isn't faster than R loops. Would be nice if someone could definitively answer the question and back it up with some stats! :)
EDIT: Thanks chuckcode, sibling post to this, I stand corrected :)
[1] http://stackoverflow.com/questions/2275896/is-rs-apply-famil...
As a Python user, who resorts to R in case of need, the power of R is not in the language, but statistical community & packages.
Typically this is done for manipulating datasets. You might have a data frame with columns Width and Height, and so you want to be able to call do.stuff(Width, Height, data=foo), and have Width and Height automatically taken from within foo. But sometimes it crops up in unexpected places.
What do you love about pandas, is it performance, syntax, access to other python modules? If performance, take a look at R's data.table package: almost any manipulation can be done by reference.
For me, R was my first language, and then I learned Python, and beautiful things like list comprehensions, and it just clicks with my brain a bit more.
In pandas, a group by operation is beautiful
dataset2=dataset1.groupby([stuff], as_index=False).mean()
Same with pivot table...
When I did this with dplyr2 my work flow would be a few more steps, creating the "summarise" object and so forth and so on, which seems like more steps.
Then Sage (sagemath.org) just dazzles me. It's a grand integrated environment using Python with lots of math/stat software built in (including NumPy and R) and lots more optional (including Matlab). You can just go see it and try it at cloud.sagemath.com. If you like it you can continue to use it there or you can download it - it's free open-source software.
Variable names are a hodgepodge of unhelpful single letter abbreviations theSecondArgumentToTheFunction; functions alternate between camelCase, dots, and underscores; and any form of architecture seems at best an afterthought. It seems like the base language encourages this, or at least does nothing to prevent it. It's commonplace to pick on Perl, but the overall quality of popular packages seems considerably lower on CRAN than CPAN. Perhaps this is because Perl is so conscious of its reputation at this point that the remaining programmers take great pain to write clear code?
I feel like R is currently in the stage where Perl and PHP were as the internet was just when the internet started to explode. The first-to-market CGI scripts and libraries, often written by domain expert non-programmers, became the default choices which the rest of the infrastructure was built on. At some point, the weight became too great for the shoddy[1] construction, and most people moved on to languages with better attention to maintainability and foundational detail (Python, Ruby).
Those who remained with the language evolved it in similar directions, by replacing the earlier libraries with better designed ones and by setting a higher standard for community norms. I'm not sure about PHP, but contrary to reputation, modern Perl is often a really clean and consistent language. Julia seems to be playing a parallel role for R, although the new-found strength of Python in the data analysis space complicates the analogy.
But I wonder: is R undergoing (or about to undergo) a similar renaissance? Are there already examples of "Modern R" out there to serve as templates for the future direction of the language? Or is R happy where it is?
[1] Did you know that 'shoddy' was originally a legitimate but low grade of wool, and wasn't necessarily pejorative?
For me what happened was that my thoughts on appropriate naming, structure, etc has evolved over the 6 (I think?) years of the package's existence but I simply haven't had the time to make the wholesale changes necessary. It's on my todo list, but frankly things like "fix actual bugs" have been sitting on that list for a very long time as well.
In general though, I've always found that most packages are pretty crappy and not just for code quality. With a relatively small amount of exceptions what I found over the years was that if you needed to do something it was almost always better to write something yourself than shoehorn someone else's junk into your system. There's an exception w/ Bioconductor, particularly the packages created and maintained by the core devs.
And on your point about the renaissance, yes I believe that has been happening, largely driven by Hadley Wickham.
Anyways, a good place to start would be the Hadleyverse: https://github.com/hadley
One could do a lot worse than following his lead.
'Badass' statistic packages but R always felt a bit 'hacked together'. With Julia on the other hand, I get the impression that there are developers in charge which have a deep understanding about programming languages and computer science. It's (too) early times for Julia but I wouldn't be surprised if in two years many users will (partly) switch.
Running an R script on a server to process data isn't efficient, but does that mean you have to roll your own stats package if you want to have a Java (for example) back-end?
But for proposes of a TL;DR, I believe that a qualitative description of the situation should suffice.
I'll try again:
TL;DR: The language R lacks the quality without a name.
Or:
Jeeze, now I know why none of my previous attempts to learn a little R were fruitful.
Or:
The language R seems to have been developed in isolation and thus it fails to adhere to any particular convention--it is its own beast. Further, it sometimes lacks self-referential integrity and coherence.
Or:
haha cf. PHP or Excel.
:)
I don't see any reason of R including OOP into its design and sometimes it just creates confusions.
There's the S3 OOP system, the S4 OOP system, reference classes and at least one add on package on CRAN which does something different.
So which one are you complaining about? :)
I always thought that R is fairly unmatched in both breadth and depth for statistical work and general data analysis, with the Python stack in second (e.g., numpy+scipy+pandas+...).