plyr, while having awesome syntax, is really slow for anything beyond even the most modest dataset. Try using it with 50k rows, not to mention true "big data".
data.table is much faster. It's an extension to data.frames that adds some additional constrains/rules that allows for much faster operations including aggregating, subsetting, and merging data.
Cleaning data, however, does not have a steep learning curve or high difficulty level in R-- it has a steep learning curve and high difficulty level period. Implementing good procedures for data munging is 80% of the job.