R: the good parts
hackerretreat.com
hackerretreat.com
$ R
data <- read.csv(file='some_file', header=T, sep=',')
model <- lm(Y ~ COL1 + COL2 + COL3, data=data)
and if you want to use glm -- logistic regression, etc -- it's a trivial change: model <- glm(Y ~ COL1 + COL2 + COL3, family=binomial, data=data)
It really allows people to do quite powerful statistical analyses very simply. Note that they built a dsl for specifying regression equations -- and you don't have to bother with bullshit like quoting column column names; quoting requirements are often hard to explain to new computer users.R's other key feature is it includes a full sql-like data manipulation language to manipulate tabular data; it's so good that every other language that does stats copied it. If df is a dataframe, I can issue predicates on the rows before the comma and columns after the comma, eg
df[ df$col1 < 3 & df$col2 > 7, 'col4']
that takes my dataframe and subsets it so -- row predicates before the comma -- col1 is less than 3 and col2 is greater than 7 -- and column predicates after the comma -- just returns a new dataframe from the subset with col4 in it. It's incredibly powerful and fast.<?
$paragraph = 'hello world';
echo '<p>'.$paragraph.'</p>';
?>
It makes a basic tasks super easy, which allows people who don't know what they're doing to make mistakes. It's just a different philosophy that has some negatives.
R code usually isn't.
R says: 'everything is a vector, and vectors can have missing values'. This is profound. It was only recently that other matrix-oriented language extensions (say panda) got missing values, even though they are meat-and-potatoes for data analysis.
https://www.labkey.org/wiki/home/LabKey%20Server%20Documenta...
(Pardon the extremely old screenshots.)
Agree with the missing values in native datatypes point, although many libraries for other languages had workarounds for missing values that allowed you to get similar results even back in the day, way before python and panda.
I think the general consensus is that for every thing S and R got right, they got a bunch of other stuff not-so-right, but it was worth the pain to work around them because the good stuff was so very good.
R existed in one form or another for a (relatively) small, specialized audience for a long, long time. This included the period where transitioning from S to R included placing a very high priority on maintaining code from S, even being able to replicate the bugs from S. This made a lot of sense from the perspective of a (relatively) small group of statisticians making a data analysis language for their own uses.
Over the last 10 years, R has become massively more popular, and many, many more people are using it outside of its "base" audience. Not surprisingly, many of these new users are discovering that R was not designed with their personal use cases in mind.
Given these constraints, R has done a remarkably good job of adjusting (where it can) to accommodate the diversification in its user base. But you can only go so far without just re-writing it from scratch, and the interests of R Core will always bend towards themselves and their, well, "core" audience.
Our entire back-end is built in R, mostly within the Hadly-verse, and we use Shiny [2] as our web framework.
Our team works a bit differently than most, I suspect; our data-scientists build features and analysis directly in R, and then add the functionality to our Shiny Platform. Our "real devs" are all server + DB, or Android guys. This has created a great development system where all of the "cool findings" and 'awesome visualizations' are immediately implemented in our system, and made available for our clients!
[1] www.Gastrograph.com [2] http://shiny.rstudio.com/
---- EDIT ---
Edited to add; R is a great language and is 100% suitable for production systems. Its older than Python(!), and, with some experience, can be made in to high performance code.
You can call R functions on a server from Android and return the results in a list or array - makes for some really cool internal API magic between our devs and data-scientists :)
I'm rather fond of R as a language, and hop between it and Python as my preferred tools of choice for a given task. I think the package ecosystem is it's biggest plus - for statistical work, Python might have a package to do something, R almost certainly will.
I always check in details with newbies what they are trying to do before I mention that name (it used to be surprisingly hard to come across it randomly) because once I have, it’s a rabbit hole -- and they generally have weird ideas, that need more several single-dimension graphs, easily done with hist() and plot(); however, anything a little subtle benefits so much from that flexibility.
I still don’t understand why bucketed log-scale for histogram and properly typed percentage scales (i.e. “10%” and not “0.1”) are so hard to do, but I love impressing the one guy who tried by showing those casually.
Tweaking plots for presentation is obviously more challenging, but I think it's a challenge with every graphics system.
The end of this article[1] details how to export your plots as SVG files for use in Inkscape or any other vector graphics editor.
[1] http://www.noamross.net/blog/2013/11/20/formatting-plots-for...
http://docs.julialang.org/en/latest/manual/functions/#keywor...
But now that I'm doing analysis on hundreds of thousands of rows, doing aggregation takes awhile. This article convinced me to give those packages a try. If data.table and plyr aggregate functions are indeed paralellizeable, that's a big deal, especially when implementing bootstrap resampling.
I started with Pandas and learned R. I find that R is just better and if R isn't right then Julia or Closure will do the work.
The tools in R are just better and more varied.
On the other hand, I'm taking a Bayesian course now, and am thinking that I could probably do the whole works in Python with little effort. That said, I'm not sure doing things in Python would actually buy me anything over R. If I were to do something other than R it would probably be something like Julia or Clojure, but that would also be more for my amusement rather than for any practical reason.
My current m.o. is to use data.frames as needed and plyr if I need to do any serious manipulation (which means that every time I use plyr, I need to read the docs). There's a lot of benefit to picking one direction and sticking with it...
plyr is wonderful for complex manipulations, but really slow as you start aggregating by a column with many (1000+) distinct values.
The base functions are super-fast (they're written in C), but give you output that you need to manipulate into more useful formats. Additionally, there's some frustrating subtle differences in syntax between them that makes for some annoyances.
These are very simple functions and I generally try to use them before moving on to plyr. The are easy to set up, understand, and faster. Obviously they cannot do everything, though, which is why there is plyr.
That's the option I use, daily.
a code example:
> df <- data.frame(id=c('a','a','a', 'b', 'b'), vals=rnorm(5,10))
> df
id vals
1 a 10.86507
2 a 10.71303
3 a 11.15321
4 b 10.78187
5 b 10.80042
> # calculate a mean on vals grouped by id
> tapply(df$vals, df$id, mean)
a b
10.91044 10.79114
> # similarly a median -- both mean and median are built in functions
> tapply(df$vals, df$id, median)
a b
10.86507 10.79114
>
> # now let's build our own function; I'm going to build a function that drops outliers
> f <- function(xs){ qs <- quantile(xs, probs=c(0.025, 0.975)); mean( xs[xs >= qs[1] & xs <= qs[2]])}
> tapply(df$vals, df$id, f)
a b
10.86507 NaN
>
> # well, this is just a demo and we ran out of bs, so it got NaN, but you see the idea
>
> # and plyr
> library(plyr)
> f2 <- function(dfs){ qs <- quantile(dfs$vals, probs=c(0.025, 0.975)); mean(dfs[ dfs$val >= qs[1] & dfs$val <= qs[2], 'vals'])}
> ddply(df, .(id), f2)
id V1
1 a 10.86507
2 b NaN
plyr relaxes that limitation, that is, X can be a data frame itself -- which brings the huge benefit that your group by logic can operate over more than one column. The function signature changes, but that's the basic innovation. The first letter indicates what X is, and the second letter indicates the output. Thus ddply runs an enhanced tapply over a data frame input (first d) and collects the output into a data frame (the second d). It also offers a bunch of nice enhancements; it's really a solid bit of work.data.table, otoh, removes some of the speed problems with the built in dataframes. It offers keys/indices for quick lookup.
I haven't spent much time with dplyr, but I think it does a couple things: (1) move the plyr code from R to c for performance reasons; it allows you to run plyr operations (with all code written in R) then translates most of that to something that can run in sql against remote dbs (for the obvious reason that plyr/group by produce summary stats, and those can be orders of magnitude smaller than the source data, so pulling that all into R to immediately discard most of it sucks).
One of the cooler features in dplyr is the '%.%' operator which allows you to chain operations. So you can write something like this in dplyr:
Batting %.%
group_by(playerID) %.%
summarise(total = sum(G)) %.%
arrange(desc(total)) %.%
head(5)
which is very readable. That example stolen shamelessly from [1] ;)[1] -- http://blog.rstudio.org/2014/01/17/introducing-dplyr/
I had never put it together that the a main functional difference between tapply and plyr is that you can group by multiple columns.
One thing that I've had to use plyr for: dfrm has three columns: user_id, vals, date
I want to know the average val for the first day of each user_id, second day, and so on. (In other words, average of vals across users relativized to the rank of the date for that user).
This is when plyr is awesome: dfrm <- ddply(dfrm,"user_id",transform,dateRank = order(date))
That one line split my dataframe by user_id, performed a function over every sub-dfrm (ranking by date), then combined them back together into one dfrm, with a new column names dateRank that is appopriate for that {user,date} pair.
Now it's easy to: summaryTable <- data.frame(mean=numeric(), median=numeric, n=numeric()) summaryTable$mean <- tapply(dfrm$val, dfrm$dateRank, mean) summaryTable$median <- tapply(dfrm$val, dfrm$dateRank, median) summaryTable$n <- tapply(dfrm$val, dfrm$dateRank, length)
And so on.
If, for example, you want to aggregate a table using a list of supplied variable names, like if you wanted to abstract out some aggregation code, you need to descend into horrible quote/substitute/etc. hackery to make it work.