My current m.o. is to use data.frames as needed and plyr if I need to do any serious manipulation (which means that every time I use plyr, I need to read the docs). There's a lot of benefit to picking one direction and sticking with it...
My current m.o. is to use data.frames as needed and plyr if I need to do any serious manipulation (which means that every time I use plyr, I need to read the docs). There's a lot of benefit to picking one direction and sticking with it...
a code example:
> df <- data.frame(id=c('a','a','a', 'b', 'b'), vals=rnorm(5,10))
> df
id vals
1 a 10.86507
2 a 10.71303
3 a 11.15321
4 b 10.78187
5 b 10.80042
> # calculate a mean on vals grouped by id
> tapply(df$vals, df$id, mean)
a b
10.91044 10.79114
> # similarly a median -- both mean and median are built in functions
> tapply(df$vals, df$id, median)
a b
10.86507 10.79114
>
> # now let's build our own function; I'm going to build a function that drops outliers
> f <- function(xs){ qs <- quantile(xs, probs=c(0.025, 0.975)); mean( xs[xs >= qs[1] & xs <= qs[2]])}
> tapply(df$vals, df$id, f)
a b
10.86507 NaN
>
> # well, this is just a demo and we ran out of bs, so it got NaN, but you see the idea
>
> # and plyr
> library(plyr)
> f2 <- function(dfs){ qs <- quantile(dfs$vals, probs=c(0.025, 0.975)); mean(dfs[ dfs$val >= qs[1] & dfs$val <= qs[2], 'vals'])}
> ddply(df, .(id), f2)
id V1
1 a 10.86507
2 b NaN
plyr relaxes that limitation, that is, X can be a data frame itself -- which brings the huge benefit that your group by logic can operate over more than one column. The function signature changes, but that's the basic innovation. The first letter indicates what X is, and the second letter indicates the output. Thus ddply runs an enhanced tapply over a data frame input (first d) and collects the output into a data frame (the second d). It also offers a bunch of nice enhancements; it's really a solid bit of work.data.table, otoh, removes some of the speed problems with the built in dataframes. It offers keys/indices for quick lookup.
I haven't spent much time with dplyr, but I think it does a couple things: (1) move the plyr code from R to c for performance reasons; it allows you to run plyr operations (with all code written in R) then translates most of that to something that can run in sql against remote dbs (for the obvious reason that plyr/group by produce summary stats, and those can be orders of magnitude smaller than the source data, so pulling that all into R to immediately discard most of it sucks).
One of the cooler features in dplyr is the '%.%' operator which allows you to chain operations. So you can write something like this in dplyr:
Batting %.%
group_by(playerID) %.%
summarise(total = sum(G)) %.%
arrange(desc(total)) %.%
head(5)
which is very readable. That example stolen shamelessly from [1] ;)[1] -- http://blog.rstudio.org/2014/01/17/introducing-dplyr/
I had never put it together that the a main functional difference between tapply and plyr is that you can group by multiple columns.
One thing that I've had to use plyr for: dfrm has three columns: user_id, vals, date
I want to know the average val for the first day of each user_id, second day, and so on. (In other words, average of vals across users relativized to the rank of the date for that user).
This is when plyr is awesome: dfrm <- ddply(dfrm,"user_id",transform,dateRank = order(date))
That one line split my dataframe by user_id, performed a function over every sub-dfrm (ranking by date), then combined them back together into one dfrm, with a new column names dateRank that is appopriate for that {user,date} pair.
Now it's easy to: summaryTable <- data.frame(mean=numeric(), median=numeric, n=numeric()) summaryTable$mean <- tapply(dfrm$val, dfrm$dateRank, mean) summaryTable$median <- tapply(dfrm$val, dfrm$dateRank, median) summaryTable$n <- tapply(dfrm$val, dfrm$dateRank, length)
And so on.
That's the option I use, daily.
These are very simple functions and I generally try to use them before moving on to plyr. The are easy to set up, understand, and faster. Obviously they cannot do everything, though, which is why there is plyr.
plyr is wonderful for complex manipulations, but really slow as you start aggregating by a column with many (1000+) distinct values.
The base functions are super-fast (they're written in C), but give you output that you need to manipulate into more useful formats. Additionally, there's some frustrating subtle differences in syntax between them that makes for some annoyances.