Ggplot2 version 2.0.0 released by Hadley Wickham
blog.rstudio.org
blog.rstudio.org
And this is why I gave up on R for other things than ad hoc data exploration (don't get me wrong; it's killer for that). When I finally learned there were no less than four system of OO buried in base R, and still people were unhappy and inventing more, I realised that as far as well structured programming goes, R is kind of beyond help at this point.
From what I can tell, my biggest disappointment with ggplot2 is not addressed - the fact that ggplot evaluates all references not found in the data frame to be plotted in the global namespace. The result is that you can't actually make reusable, composable plots in a reasonable sane way. It sends the whole illusion of composability that ggplot()'s '+' syntax suggests might be there right down the tubes. My plots were all more reusuable before I ported them to ggplot (ie, in base graphics).
Apologies for the whining, and congrats to Hadley!
edit: that being said, ggplot is still better than anything on the python side of things.
1) Software and related services will grow immensely over the next decade because of the "cloud". Some of the growth will be same old, same old application logic software, but most of it will be data analysis software.
2) The language of choice for data analysis software is and will be R due to the biggest software companies (both #1 and #2, among many others) standing firmly behind it along with universities incorporating it in its curriculum, ditching SPSS and SAS.
Very often, to produce specialized plots, I have to send data to the canvas in chunks by performing pre-processing myself. ggplot really doesn't work in this scenario. Combined with the general slowness, it forces me to use alternatives quite frequently.
It's a bummer, really, because I'd like my plots to have a consistent visual style, and doing that across different plotting packages is an issue.
I very often resort to gnuplot when it comes to huge datasets and/or incremental plotting. The same is true also in python (matplotlib is also very slow, independently of the backend). But at least, if you use seaborn (https://github.com/mwaskom/seaborn), you can easily intermix the easiness of plotting through a DataFrame or just supply data arrays.
ggplot is really awesome for what it does, but 1) the syntax doesn't really please me (feels just plainly forced onto the wrong context 2) doesn't scale, which forces me to use alternatives too frequently 3) trying to customize the plot style beyond a few minor tweaks is pure hell.
I took to liking the ggplot syntax immediately. Specifically what do you find odd/forced about it?
The "problem" is that ggplot also takes care of the transformation/reduction step for you.
For example, a KDE plot can source potentially a limitless amount of data while still generating a very simple plot. Likewise for most smoothers.
However, if I have to produce the kde/smoothed line myself, I lose almost all advantages of using ggplot (I have to manually calculate the visual density, scaling and attaching labels is another PITA).
On top of that, as other have said, ggplot really struggles already with thousands of entries. A simple 5x5 faceted scatterplot with ~10k points might take seconds to render on recent hardware. When I plot data interactively for exploration, I might do this hundreds of times a day. I lose all the convenience just in the time wasted for rendering.
This is a regular concern in designing a programming language. Give coders too much freedom and they invent 20 ways to do the same thing, none really clearly better than the other. Restrict the freedom and there's more like 3-4 ways to do something, and usually one is better then the rest.
As a new user to R, I must say that the single hardest thing has been learning all the different ways to do something. Each blog/stackoverflow/etc... article I read seems to propose 5 different ways to do each thing. At the end of it, I'm left not really knowing how things work best.
If you're committed to using R, read John Chamber's book and also look for answers on the old "R mailing list" (a place so mean, it makes stackoverflow look like mister roger's neighborhood).
There's also a legitimate learning challenge: if I want to figure out which OO system is best for most projects, I have to try them all out. And then the payoff for translating to a common system is small.
That said, for new projects I pretty much use S3 and R6 exclusively. (And it is legitimate to have two systems because they are so different and R6 is mostly needed for internal mutability)
To put the scope of the changes in the 2.0.0 release in perspective, I suspect that the simple tutorial is now broken, let alone the code for my more intricate charts. I'm not upset about it through, since all the changes, especially breaking ones, are well-reasoned and well-documented. Hadley did a great job of explaining everything.
I'll have to spend some time diving into ggproto.
I also used the order parameter on a few visualizations, so I'm unsure how to order a stacked bar chart without it. I'll give it another look and see what I find and file as appropriate.
Error in add_ggplot(e1, e2, e2name) : could not find function "is.coord"
Error in (function (el, elname) : "panel.ontop" is not a valid theme element name...
Error in layer(mapping = structure(list(x = element$Line, y = 0, xend = element$Line, : unused arguments (arrow = NULL, lineend = "butt", na.rm = FALSE, colour = "grey50", linetype = 2)
Error : Unknown parameters: guides
Error in FUN(X[[i]], ...) : attempt to apply non-function
Error in coord_tern() : could not find function "coord"
I just can't keep up with this - most of these errors are with objects I am not using (such as 'panel.ontop' when using theme(legend.position="bottom"). For those of us who chose (made the mistake?) of using ggplot2 with larger projects such as shiny, large updates like this which render all our past code mute presents a dilemma - do we keep re-learning this wheel or just stop updating altogether? The latter seems to be the only feasible course. Every plot implementation I had on my website was broken by this update.
I don't mean to be mean, ggplot2 is a wonderful package, but this update presented all headaches.
Otherwise I'd highly recommend using packrat or similar so you can choose on a project by project basis when you want to upgrade selected packages.
[1] https://github.com/hadley/ggplot2/releases/tag/v2.0.0
[2] https://cran.r-project.org/web/packages/ggplot2/vignettes/ex...
The R community really needs to grasp version control head on. There are some great tools, and using git or similar works just fine, but their central packaging system only knows about the latest version of everything.
I would really like to see R develop functionality akin to Maven or SBT, such that an R developer can explicitly specify the exact versions of all dependencies, which will then be installed at the first run.